#SpeechModels
New paper "MultiTalk" extends the Moshi paradigm to long, multi-party, bilingual conversations, releasing 57.6k hours of speech data and codec-frame-level full-duplex modeling for English and Chinese. Aims to address…

#OpenSourceAI #SpeechModels #Moshi #FullDuplex
https://arxiv.org/abs/2609.36903
September 30, 2026 at 2:01 PM
🔥 Alibaba's Qwen drops Qwen-Audio-3.1! New models for speech recognition, synthesis, interaction & *creation*! Plus, major price cuts across the board: TTS down 70%, Realtime 85%, and ASR a whopping 95%! Get ready for powerful, affordable audio AI. #QwenAudio #AI #SpeechModels
Alibaba Launches Qwen-Audio-3.1 and Slashes Speech Model Prices by 95%
Alibaba Cloud’s AI division, Tongyi Qianwen, has announced the official release of the Qwen-Audio-3.1 series, marking a significant evolution in large speech model (LSM) technology. This major update introduces five new …
cryptofox.news
September 23, 2026 at 7:14 AM
Researchers found visual grounding boosts alignment between spoken and written representations, mainly via stronger encoding of word identity. Speech‑based encoders stayed primarily phonetic. https://getnews.me/visual-grounding-impacts-speech-and-text-language-models/ #visualgrounding #speechmodels
September 22, 2025 at 12:54 PM
in the year of our lord 2026, Apple is still forced to use caseless-enums as namespaces in their own public frameworks

since Swift doesn't apparently have anything better

#wwdc

developer.apple.com/documentatio...
SpeechModels | Apple Developer Documentation
Namespace for methods related to model management.
developer.apple.com
June 18, 2025 at 6:09 PM