#StepAudio
StepFun’s StepAudio 3 treats audio as a timed situation: conversations that keep moving, scenes on one timeline, and songs planned before synthesis. Several models remain in preview.

sonicfield.org/stepaudio-3-...
StepAudio 3 turns a model family into an audio production stack - Sonic Field - Operative system for sound and listening culture
StepAudio 3 is a five-model audio family whose most interesting advances concern continuity: conversations that keep moving, scenes assembled on one timeline, and songs planned before they are rendere...
sonicfield.org
September 22, 2026 at 5:18 PM
StepFun just dropped StepAudio 2.5 Realtime, and mobile app raters put its AI‑powered sound generation to the test. Curious how it stacks up? Dive into the details and hear the future of generative audio. #StepAudio #RealtimeAI #GenerativeAI

🔗 aidailypost.com/news/stepfun...
May 24, 2026 at 11:33 PM
7. StepAudio 3 Realtime uses Think-While-Speaking for Low-Latency Reasoning LINK
8. IBM Research Explores Agent Task Consistency LINK
September 16, 2026 at 1:00 PM
September 15, 2026 at 6:44 AM
Chengli Feng, Zhiyue Wu, Jiahao Song, Zheqi Dai, Boyang Wang, Ruibin Yuan, Junming Gong, Wenxiao Zhao, Jing Guo, Gang Yu, Xiangyu Zhang, Xuerui Yang, Chao Yan: StepAudio 3 Music Technical Report https://arxiv.org/abs/2609.16034 https://arxiv.org/pdf/2609.16034 https://arxiv.org/html/2609.16034
September 16, 2026 at 6:45 AM
阶跃星辰全新StepAudio 2.5 Realtime实时语音大模型正式发布上线,开发者可直接接入使用。

模型主打副语言感知,能识别语调、语速、停顿、轻笑叹息等细节,精准捕捉用户情绪并自适应调整回复语气。支持高度人设自定义,依托万级原生人设生成百万特征矩阵,角色稳定性拉满。

对话智商情商双升级,可日常闲聊也能胜任面试等专业场景。实测用户体验评分80.41,超越GPT、Gemini同类实时语音模型,国产实时语音AI再迎实力突破。
May 9, 2026 at 10:58 AM
September 14, 2026 at 6:44 AM
[2/2] character settings, and double advancement in EQ and IQ. Currently, StepAudio 2.5 Realtime is fully online.
May 8, 2026 at 2:30 PM
[1/2] 2026-05-08 22:21:21 - [StepAudio 2.5 Realtime, a large real-time voice model released by Step Star] On May 8, StepAudio 2.5 Realtime, a new generation of real-time voice large model, was officially released by Step Star, which realizes real-life depth perception, customizable
May 8, 2026 at 2:30 PM
[1/3] 2026-04-24 12:55:44 - [StepAudio 2.5 ASR is launched today] Today, StepAudio 2.5 ASR is officially released as a new generation automatic speech recognition model. The core breakthrough of this model is the combination of speed and accuracy. It is the first to introduce large
April 24, 2026 at 5:10 AM
[1/2] 2026-04-16 14:51:24 - [The new generation speech generation model StepAudio 2.5 TTS is released in a step] The new generation speech generation model StepAudio 2.5 TTS is released in a step. According to reports, the model has been upgraded around capabilities such as global
April 16, 2026 at 11:05 AM
[1/2] 2026-04-16 14:51:24 - [The new generation speech generation model StepAudio 2.5 TTS is released in a step] The new generation speech generation model StepAudio 2.5 TTS is released in a step. According to reports, the model has been upgraded around capabilities such as global
April 16, 2026 at 11:05 AM
winbuzzer.com/2026/05/25/s...

Chinese AI lab StepFun has launched StepAudio 2.5 Realtime, an end-to-end live voice model for audio input and output in one system, withpersona control and paralinguistic understanding.

#AI #StepAudio25Realtime #StepFun #VoiceAI #AIModels #AIVoices #AIAudio
StepFun Launches StepAudio 2.5 Realtime Live Voice AI Model
StepFun has launched StepAudio 2.5 for live voice AI, pitching persona control while leaving open questions about training-data consent and copyright.
winbuzzer.com
May 25, 2026 at 7:48 PM
StepFun Unveils "StepAudio 2.5 Realtime," Promising End-to-End Voice Synthesis

StepFun launches StepAudio 2.5 Realtime, an end-to-end voice model for roleplaying. Learn about its features and how developers can use GELab-...

https://newsletter.tf/stepfun-stepaudio-2-5-realtime-voice-model-release/
June 22, 2026 at 7:15 AM
StepFun released StepAudio 2.5 Realtime for instant voice generation in roleplaying apps.

https://newsletter.tf/stepfun-stepaudio-2-5-realtime-voice-model-release/
StepFun's StepAudio 2.5 Realtime Voice Model Released for Roleplaying
StepFun launches StepAudio 2.5 Realtime, an end-to-end voice model for roleplaying. Learn about its features and how developers can use GELab-Zero-4B-preview.
newsletter.tf
June 22, 2026 at 7:09 AM
Today's Voice AI News that actually matters:

LiveKit + Smallest AI just nuked voice lag—real-time STT & TTS under 100ms, private & scalable. Wispr Flow goes big in India. OpenLess brings local voice input. StepFun StepAudio shakes up TTS.

Blog: https://voiceai.guru/livekit-tts-breakthrough-wisp...
May 11, 2026 at 6:05 PM
Today's Voice AI News that actually matters:

LiveKit + Smallest AI just nuked voice lag—real-time STT & TTS under 100ms, private & scalable. Wispr Flow goes big in India. OpenLess brings local voice input. StepFun StepAudio shakes up TTS.

Blog: https://voiceai.guru/livekit-tts-breakthrough-wisp...
May 11, 2026 at 6:05 PM
May 25, 2026 at 6:45 AM
Model StepAudio 2.5 Realtime od StepFun pokonał systemy OpenAI i Google w testach interpretacji emocji. Chiński startup wykorzystał 10 tysięcy bazowych wzorców mowy, by stworzyć AI, która słyszy ironię i nigdy nie wychodzi z roli.
Głos, który czuje. Szanghajski StepFun ogrywa OpenAI i Google w emocjonalnym teście
Model StepAudio 2.5 Realtime od StepFun pokonał systemy OpenAI i Google w testach interpretacji emocji. Chiński startup wykorzystał 10 tysięcy bazowych wzorców mowy, by stworzyć AI, która słyszy ironię i nigdy nie wychodzi z roli.
aisight.pl
May 30, 2026 at 2:33 PM
StepFun introduces StepAudio 2.5 TTS, targeting booming AI content market

StepAudio 2.5 TTS is designed to let ordinary users easily become voice directors using natural language.

cntechpost.com/2026/04/16/s...
StepFun introduces StepAudio 2.5 TTS, targeting booming AI content market - CnTechPost
StepAudio 2.5 TTS is designed to let ordinary users easily become voice directors using natural language.
cntechpost.com
April 16, 2026 at 9:15 AM
StepFun's StepAudio 2.5 Realtime delivers end-to-end voice AI with roleplay RLHF and paralinguistic understanding for lifelik... 🔗 https://www.marktechpost.com/2026/05/24/stepfun-releases-stepaudio-2-5-realtime-an-end-to-end-voice-model-with-roleplay-specific-rlhf-and-paralinguistic-comprehension/
May 25, 2026 at 5:01 PM