Moving beyond sustained vowels & read speech, our study uses CNNs on MFCC features from natural, spontaneous speech—capturing real-world acoustic nuances & reaching ~92% eval ACC.
link.springer.com/article/10.1...
#SpeechTech #HealthAI
Moving beyond sustained vowels & read speech, our study uses CNNs on MFCC features from natural, spontaneous speech—capturing real-world acoustic nuances & reaching ~92% eval ACC.
link.springer.com/article/10.1...
#SpeechTech #HealthAI
doi.org/10.25189/267...
#linguistics #speechtech
doi.org/10.25189/267...
#linguistics #speechtech
Microsoft Launches ‘Mini’ GPT Voice Models in Azure Foundry to Cut Latency and Cost
#AI #Microsoft #Azure #GenAI #OpenAI #VoiceAI #CloudComputing #SpeechTech #EnterpriseAI
Microsoft Launches ‘Mini’ GPT Voice Models in Azure Foundry to Cut Latency and Cost
#AI #Microsoft #Azure #GenAI #OpenAI #VoiceAI #CloudComputing #SpeechTech #EnterpriseAI
Excited to keynote the Winter School of Innovation tmrw at Uni of Warsaw. We’ll be tracing the innovation pathway that enables #SpeechTech w/ minimal data.
szkolazimowa.szkolydoktorskie.uw.edu.pl/en/
Excited to keynote the Winter School of Innovation tmrw at Uni of Warsaw. We’ll be tracing the innovation pathway that enables #SpeechTech w/ minimal data.
szkolazimowa.szkolydoktorskie.uw.edu.pl/en/
Le #LLL a remis un #corpus unique: 1400 heures de données orales en #kreyòl, transcrites et alignées automatiquement au caractère près 🎧✨
@univorleans.bsky.social
#créolehaïtien #speechtech #NLP
Le #LLL a remis un #corpus unique: 1400 heures de données orales en #kreyòl, transcrites et alignées automatiquement au caractère près 🎧✨
@univorleans.bsky.social
#créolehaïtien #speechtech #NLP
#FarFieldRecognition #VoiceRecognition #SpeechTech #SmartSpeakers #AIInteraction
👉 github.com/pnlpal/pnl-r...
#TTS #VoiceAI #PNLReader #Sverige #svenska #Sweden #SpeechTech #webdev #devlog #buildinpublic #indiedev
www.youtube.com/watch?v=7nV0...
👉 github.com/pnlpal/pnl-r...
#TTS #VoiceAI #PNLReader #Sverige #svenska #Sweden #SpeechTech #webdev #devlog #buildinpublic #indiedev
www.youtube.com/watch?v=7nV0...
aclanthology.org/2025.emnlp-m...
#SLU #SpeechTech
aclanthology.org/2025.emnlp-m...
#SLU #SpeechTech
#Speech #SpeechLLM #LLM #SpeechTech #AI
This fascinating work applies CoT inspired by human “thinking while listening”, training models to find the inflection point when reasoning starts.
📄 arxiv.org/abs/2510.07497
#Speech #SpeechLLM #LLM #SpeechTech #AI
The open-source tool, which is going to be released soon, natively supports any speech-to-text #HuggingFace models! 🤖
#SpeechTech #Translation
The open-source tool, which is going to be released soon, natively supports any speech-to-text #HuggingFace models! 🤖
#SpeechTech #Translation
#Speech #Simultaneous #Translation #MOE #SpeechTech
Mixture-of-Experts routing → smarter decisions on when & how to translate, balancing latency vs quality in real-time speech. Paper link at arxiv.org/pdf/2509.012...
#Speech #Simultaneous #Translation #MOE #SpeechTech