#SpeechTech
September 29, 2026 at 3:48 PM
🎙️ Can spontaneous speech improve voice pathology detection?

Moving beyond sustained vowels & read speech, our study uses CNNs on MFCC features from natural, spontaneous speech—capturing real-world acoustic nuances & reaching ~92% eval ACC.

link.springer.com/article/10.1...

#SpeechTech #HealthAI
Voice pathology detection on spontaneous speech data using deep learning models - International Journal of Speech Technology
Speech problems are a common issue that affects people everywhere and can affect the quality of their lives. The human speech production system involves various components. Dysfunction of any of these...
link.springer.com
September 25, 2026 at 8:05 AM
Computer-generated voices are getting hard to tell from human ones. What still gives them away? A study of Venezuelan Spanish measured the melody of statements and questions in both, moment by moment.
doi.org/10.25189/267...
#linguistics #speechtech
July 27, 2026 at 5:03 PM
🔍 Dive into the world of Large Vocabulary Speech Recognition! 🤖 Learn to tackle tricky out-of-vocabulary (OOV) terms effectively. Your accuracy depends on term density, not just count! 💡 #SpeechTech
Large Vocabulary Speech Recognition Demystified
An enthusiastic promotion of our latest blog post discussing the complexities of large vocabulary speech recognition and the importance of managing out-of-vocabulary terms.
deepgram.com
April 16, 2026 at 2:15 PM
Why does transcription accuracy matter? A Forrester study found a 15% reduction in misrouted calls can lead to $4.7M in benefits! Dive into our comparison of Deepgram, Google, & AssemblyAI! #SpeechTech
Deepgram vs AssemblyAI vs Whisper: Which Speech-to-Text API Is Best for Developers in 2026?
Promoting the blog post comparing Deepgram, Google, and AssemblyAI on transcription accuracy and its financial impact, highlighting the importance of choosing the right STT provider.
deepgram.com
April 14, 2026 at 5:11 AM
🚀 Choose wisely! ElevenLabs and Deepgram are neck and neck in STT capabilities! From accuracy to deployment models, our latest blog breaks down what you need to know! #SpeechTech #Transcription
ElevenLabs Transcription vs. Deepgram: Which STT API Handles Production?
Promoting a comparison of ElevenLabs and Deepgram STT APIs, highlighting critical decision factors for users in need of reliable transcription solutions.
deepgram.com
April 1, 2026 at 6:32 PM
🚀 Unlock the power of ElevenLabs' Scribe for speech-to-text! Discover how to optimize your transcription while navigating unique constraints. Essential insights await! 🌟 #SpeechTech
Does ElevenLabs Do Speech-to-Text?
An enthusiastic tweet promoting the blog post about ElevenLabs' speech-to-text capabilities, emphasizing its optimization and constraints.
deepgram.com
March 16, 2026 at 10:39 AM
Česká televize spustila nepřetržité skryté titulkování programu ČT24
Česká televize spustila nepřetržité skryté titulkování programu ČT24
Česká televize rozšiřuje přístupnost svého vysílání a na zpravodajském kanálu ČT24 spouští nepřetržité skryté titulkování. Služba je dostupná divákům v pozemním, satelitním i internetovém vysílání.  Nové technologické řešení vychází z dlouhodobé spolupráce se společností SpeechTech a využívá nástroje umělé inteligence. Ty automaticky převádějí mluvené slovo do textové podoby, přičemž systém prošel v loňském roce testovacím provozem během nočního vysílání. Vývojáři se v této fázi zaměřili na zpřesnění identifikace mluvčích, přehlednější členění textu do dvou řádků a minimalizaci časové prodlevy mezi zvukovou stopou a zobrazeným textem. Podle generálního ředitele České televize Hynka Chudárka je nepřetržité titulkování konkrétním krokem k tomu, aby se diváci se sluchovým postižením mohli spolehnout na stejné informace jako ostatní. Přístupnost vysílání označil za samozřejmou součást veřejné služby, nikoliv za nadstandard. Vedoucí služeb pro diváky se smyslovým postižením Petra Kolodějová k tomu dodala, že kromě technologických nástrojů hraje klíčovou roli v udržení kvality a přesnosti informací práce redakčního týmu ČT24. Aktuálně je služba funkční u živého vysílání, v archivu i u zpětného přehrávání. Jedinou výjimkou zůstává živé streamování přes hybridní vysílání HbbTV, kde stávající technologie zatím plnou kompatibilitu s tímto typem titulků neumožňuje. Kromě skrytých titulků Česká televize nadále poskytuje na svých kanálech také audiopopis pro nevidomé a vybrané pořady tlumočené do českého znakového jazyka. Zavedením této novinky instituce překračuje zákonné limity, které jí ukládají povinnost opatřit skrytými titulky minimálně 70 % vysílaných pořadů. V roce 2025 dosáhl podíl titulkovaných pořadů v rámci celé skupiny ČT hodnoty 86,45 %. Vedení televize v současné době prověřuje možnosti, jak automaticky generované titulkování v budoucnu rozšířit i na své další programové okruhy. Pořady tlumočené do českého znakového jazyka loni tvořily 6,95 % vysílání České televize a pořadů s audiopopisem bylo 20,9 %.
www.lupa.cz
February 19, 2026 at 1:40 PM
🔊 Connaissez-vous Voxtral Transcribe 2 (Mistral AI) ? Deux modèles : batch (Voxtral Mini) et Realtime ⚡ <200 ms, open-weights, déployable on‑prem. Supporte 🇫🇷 🇬🇧 🇪🇸 🇩🇪, prix 0,003–0,006$/min. #SpeechTech https://bit.ly/4agUty2
Mistral AI lance sa nouvelle génération de modèles de transcription vocale | LeMagIT
Voxtral Transcribe 2 est la nouvelle famille de modèles de reconnaissance vocale de Mistral. L’offre se décline en une version batch et une version temps réel, publiée en open weights sous licence Apa...
bit.ly
February 7, 2026 at 4:09 PM
Users reported mixed accuracy for the Mandarin tone correction tool, especially at conversational speeds and with tone transformations. This highlights the significant challenge in building models that capture the nuanced reality of spoken Mandarin. #SpeechTech 2/6
January 31, 2026 at 5:00 PM
The tech industry loves billion-dollar voices, but what about the other 7,000 languages?

Excited to keynote the Winter School of Innovation tmrw at Uni of Warsaw. We’ll be tracing the innovation pathway that enables #SpeechTech w/ minimal data.

szkolazimowa.szkolydoktorskie.uw.edu.pl/en/
Front page - Doktorska szkoła zimowa
szkolazimowa.szkolydoktorskie.uw.edu.pl
December 14, 2025 at 1:36 PM
Heureux d'accueillir le prof. Renauld Govain dans le cadre du projet #ANR CREAM sur les langues créoles !
Le #LLL a remis un #corpus unique: 1400 heures de données orales en #kreyòl, transcrites et alignées automatiquement au caractère près 🎧✨
@univorleans.bsky.social
#créolehaïtien #speechtech #NLP
December 8, 2025 at 11:09 AM
🚀 I can’t resist adding these two ‘Realistic AI’ voices to PNL Reader. Hey Swedes, how does this sound to you?

👉 github.com/pnlpal/pnl-r...

#TTS #VoiceAI #PNLReader #Sverige #svenska #Sweden #SpeechTech #webdev #devlog #buildinpublic #indiedev

www.youtube.com/watch?v=7nV0...
Realistic AI voices on PNL Reader
YouTube video by Programming N' Language
www.youtube.com
November 15, 2025 at 3:10 PM
Our pick of the week by @zhihangxie.bsky.social: "#Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in #SpeechLLMs" by Dingdong Wang, Junan Li, Mingyu Cui, et al. (#EMNLP2025)

aclanthology.org/2025.emnlp-m...

#SLU #SpeechTech
November 12, 2025 at 2:43 PM
Our #PickOfTheWeek by @beomseok-lee.bsky.social: "Can Speech LLMs Think while Listening?" by Yi-Jen Shih, @rdesh26.bsky.social, Chunyang Wu, Wei Zhou, SK Bong, Yashesh Gaur, Jay Mahadeokar, Ozlem Kalinli, Mike Seltzer (2025).

#Speech #SpeechLLM #LLM #SpeechTech #AI
Can we make Speech LLMs actually think as they listen? 👂💭
This fascinating work applies CoT inspired by human “thinking while listening”, training models to find the inflection point when reasoning starts.
📄 arxiv.org/abs/2510.07497
Can Speech LLMs Think while Listening?
Recent advances in speech large language models (speech LLMs) have enabled seamless spoken interactions, but these systems still struggle with complex reasoning tasks. Previously, chain-of-thought (Co...
arxiv.org
October 29, 2025 at 1:30 PM
Marco Gaido and Roldano Cattoni presenting our SimulStream Demo at the DI Center Demo Day at FBK!

The open-source tool, which is going to be released soon, natively supports any speech-to-text #HuggingFace models! 🤖

#SpeechTech #Translation
October 10, 2025 at 8:39 AM
A systematic review of ASR research (Jan 2020–Jul 2025) found 71 studies covering 74 datasets for 111 African languages and about 11,200 hours of speech. Read more: https://getnews.me/asr-review-for-african-low-resource-languages-highlights-gaps-and-paths/ #asr #africanlanguages #speechtech
October 3, 2025 at 3:57 AM
A Wasserstein GAN separated the four Mandarin tones without labeled data, forming distinct clusters; training on male speech tokens consistently encoded tone. Read more: https://getnews.me/unsupervised-cnn-learns-mandarin-tonal-categories-without-labels/ #speechtech #unsupervisedlearning
September 25, 2025 at 3:12 AM
Comparing VibeVoice to ElevenLabs, Chatterbox, & Kokoro, the consensus is it's promising but has room to grow. The impressive multilingual capability, especially English/Mandarin, highlights its unique strength. #SpeechTech 3/6
September 4, 2025 at 1:00 PM
Our pick of the week by @zhihangxie.bsky.social: "SimulMEGA: MoE Routers are Advanced Policy Makers for Simultaneous Speech Translation" by Chenyang Le, Bing Han, Jinshun Li, Songyong Chen, and Yanmin Qian (2025)

#Speech #Simultaneous #Translation #MOE #SpeechTech
🚀 SimulMEGA: MoE Routers as advanced policy makers for Simultaneous Speech Translation 🎧🌍
Mixture-of-Experts routing → smarter decisions on when & how to translate, balancing latency vs quality in real-time speech. Paper link at arxiv.org/pdf/2509.012...
arxiv.org
September 3, 2025 at 10:54 AM