StepFun releases five-model StepAudio 3 family
China’s StepFun releases five StepAudio 3 models spanning realtime speech, recognition, synthesis, audio generation, and music through one platform.
Chinese AI company StepFun has released the StepAudio 3 family, a five-model lineup covering realtime conversation, automatic speech recognition, text-to-speech, general audio generation, and music creation. The models are available through StepFun’s open platform, making the release broader than a research preview or a single voice model.
One family, five audio jobs
The lineup includes StepAudio 3 Realtime, StepAudio 3 ASR, StepAudio 3 TTS, StepAudio 3 Gen, and StepAudio 3 Music. StepFun positions the realtime model around full-duplex interaction: it can handle interruptions while continuing to interpret the conversation, reason about the next action, and respond. The separate ASR and TTS models target production pipelines that need control over recognition and voice output independently.
StepAudio 3 Gen is designed to combine several types of sound in one generation framework, including speech, voice design, vocals, sound effects, music, and mixed audio scenes. Its technical report describes a unified architecture trained through progressive pretraining, multitask instruction training, and supervised fine-tuning. That breadth gives developers a route from a spoken prompt to a more complete audio asset rather than requiring a separate service for every media type.
Why the packaging matters
The release arrives as competition in AI audio becomes more fragmented. Major labs are improving realtime dialogue, while specialist providers focus on voice cloning, music, or post-production. StepFun’s choice to ship five connected models through one platform is an attempt to turn a collection of capabilities into an integrated audio stack, particularly for Chinese-language and Asia-based developers.
StepFun says several models rank first globally on Artificial Analysis evaluations, but those claims should be read as benchmark-specific rather than a universal judgment. The more consequential fact is availability: all five models are exposed as usable services at launch. That can shorten the path from prototype to a multimodal product, but it also leaves unresolved questions around voice consent, copyright for music and training data, watermarking, and cross-border access. StepAudio 3 is important because it treats speech, sound design, and music as one product surface; its durability will depend on governance and creator workflow, not leaderboard position alone.