#speechRecognition
Berkeley's HuPER models phonetic perception as adaptive inference, hitting state-of-the-art English error rates with just 100 hours of training and zero-shot transfer to 95 languages. Code and models are open-sourced on GitHub.

#OpenSourceAI #NLP #SpeechRecognition
https://arxiv.org/abs/2602.01634
September 29, 2026 at 12:01 AM
Gemini 3.8 Live removes the blank face from AI customer service #CustomerService #SpeechRecognition #ArtificialIntelligence
Gemini 3.8 Live removes the blank face from AI customer service
Google has made generally available an AI system that holds spoken conversations, generates a synchronized face in real time, and switches across 97 languages without breaking stride. The face carries
www.martincid.com
September 25, 2026 at 3:15 PM
“In a room with normal background #noise ..my #SpeechRecognition without #HearingAids is roughly 40–50%. That means I miss about half of what’s being said and have to guess the rest from context.”: buff.ly/hiTT9jz

by @donnietownstudio.bsky.social
#conversation #deaf #disability #disabled #spoonie
All Of Me
Cartoon CJ sits in a dimly lit room infront of a desk, watching a lady sign via a laptop screen. Living with multiple illnesses, disabilities, and daily challenges has taught me a great deal—not ju…
thecrslife.com
September 22, 2026 at 8:30 PM
Simplify your work with smart workflows, ambient voice technology, and medical speech recognition on the go.

🔎 Search Lexacom on the App Store and Google Play

#nhs #digitaltransformation #workflows #productivity #ambientAI #speechrecognition
September 21, 2026 at 2:33 PM
I think so. I see the `SpeechRecognition` type in the list. You can see the full list of added types here: github.com/philipwalton...
modern-web-types/report.md at main · philipwalton/modern-web-types
TypeScript types for new web platform APIs that aren't yet in lib.dom. - philipwalton/modern-web-types
github.com
September 15, 2026 at 6:50 PM
Simplified, intuitive, and packed with new features, download our new app.

Intelligent workflows, ambient voice technology, and speech recognition on the go.

🔎 Search Lexacom on the App Store and Google Play.

#nhs #digitaltransformation #workflows #productivity #ambientAI #speechrecognition
September 14, 2026 at 9:31 AM
Why Speech Recognition Misses Human Context: Dr. Sunday David Ubur’s Affective Architecture

Dr. Sunday Ubur's research explores emotion-aware AI captions that preserve tone, urgency, and context for deaf and hard-of-hearing users.

Telegram AI Digest
#ai #news #speechrecognition
Why Speech Recognition Misses Human Context: Dr. Sunday David Ubur’s Affective Architecture
Dr. Sunday Ubur's research explores emotion-aware AI captions that preserve tone, urgency, and context for deaf and hard-of-hearing users.
hackernoon.com
September 12, 2026 at 1:25 AM
Почему распознавание речи упускает человеческий контекст: Аффективная архитектура доктора Сандея Дэвида Убура

Telegram ИИ Дайджест
#ai #news #speechrecognition
Why Speech Recognition Misses Human Context: Dr. Sunday David Ubur’s Affective Architecture
hackernoon.com
September 12, 2026 at 1:15 AM
Xiaomi open-sources Xiaomi-CocktailASR-1, an industrial-grade speaker diarization LLM! Solves the "cocktail party problem," precisely identifying and transcribing target speakers in multi-speaker audio using a reference voiceprint. Big leap for speech tech! #AI #OpenSource #SpeechRecognition
September 11, 2026 at 6:23 AM
Dr. Sunday Ubur's research explores emotion-aware AI captions that preserve tone, urgency, and context for deaf and hard-of-hearing users. #speechrecognition
Why Speech Recognition Misses Human Context: Dr. Sunday David Ubur’s Affective Architecture
hackernoon.com
September 11, 2026 at 3:52 AM
aeneas

aeneas is a Python/C library and a set of tools to automagically synchronize audio and text (aka forced alignment)

https://github.com/readbeyond/aeneas

#GeneralPurposeOpenSourceTools #SpeechRecognition #CivicTech
September 8, 2026 at 4:52 PM
What do 152,416 words in a week mean to a GP?

For Dr Rafay, it means using Lexacom across all consultations, referrals and admin - saving an hour a day.

➡️ Read the full case study and start your free trial at lexacom.co.uk/faster

#nhs #generalpractice #speechrecognition #healthtech #productivity
September 7, 2026 at 3:57 PM
Heidi fine-tuned and deployed medical speech recognition for clinical documentation, using synthetic multilingual data, distributed GPU training, and scalable AWS serving. The platform supports 2.4M+ consultations weekly. #HealthcareAI #SpeechRecognition #ClinicalDocumentation
Fine-tuning NVIDIA Nemotron Speech ASR on Amazon EC2 for domain adaptation | Amazon Web Services
In this post, we explore how to fine-tune a leaderboard-topping, NVIDIA Nemotron Speech Automatic Speech Recognition (ASR) model; Parakeet TDT 0.6B V2. Using synthetic speech data to achieve superior transcription results for specialised applications, we'll walk through an end-to-end workflow that combines AWS infrastructure with the following popular open-source frameworks.
aws.amazon.com
September 7, 2026 at 8:12 AM
Microsoft’s 10-person team built a speech model that undercuts OpenAI, Google and ElevenLabs on every benchmark — at $0.10 an hour #MaiTranscribe #SpeechRecognition #Google
Microsoft’s 10-person team built a speech model that undercuts OpenAI, Google and ElevenLabs on every benchmark — at $0.10 an hour
Microsoft AI's MAI-Transcribe-2 launched September 3 at $0.10 per audio hour — a 72% drop from its predecessor and below every major competitor. Built by a 10-person team, it ranks first on the FLEURS
www.martincid.com
September 5, 2026 at 7:00 AM
Meta’s new Muse Voice Transcribe sets a new standard. Get real-time, multilingual transcription and advanced speaker identification even in complex group settings—all via an efficient API. #Meta #AI #Transcription #SpeechRecognition #TechNews https://ai.dappcrypto.org/r/ea
September 4, 2026 at 7:28 AM
Microsoft’s new MAI‑Transcribe‑2 slashes cost and cuts transcription time in half—outpacing OpenAI’s frontier models on call‑center audio and multilingual support. Curious how fast and cheap AI can get? Dive in. #MAITranscribe2 #SpeechRecognition #MicrosoftAI

🔗 aidailypost.com/news/microso...
September 3, 2026 at 2:16 PM
🆕 More documentation, without more hours in the day.

Take a look at our latest case study and see how Lexacom changes working days for GPs.

➡️ The full case study is on our website > resources. Link in bio.

#nhs #speechrecognition #productivity #generalpractice #healthtech
August 28, 2026 at 3:04 PM
https://developer.mozilla.org/en-US/docs/Web/API/SpeechRecognition/processLocally

https://webaudio.github.io/web-speech-api/#dom-speechrecognition-processlocally

> **`processLocally` attribute, of type `boolean`**
> This attribute, when set to true, indicates a requirement that the speech […]
Original post on tech.lgbt
tech.lgbt
August 26, 2026 at 9:06 PM
the `SpeechRecognition` web speech api is like the EXACT thing I need for this..
August 26, 2026 at 8:58 PM
See what's new in Lexacom Echo:

- Edit and format text
- Anchor your speech to any window
- Increased personalisation
-️ Dictate in multiple languages
-️ Instantly summarise dictations

➡️ Read more news on our website > Resources. Link in bio.

#speechrecognition #workflow #digitaltransformation
August 21, 2026 at 10:52 AM
Microsoft has confirmed that Dragon Medical Practice Edition will stop working on 14/08/2026.

Switch to Lexacom Echo for fast, accurate medical speech recognition with powerful new capabilities.

Better still, Lexacom Echo is half the price of DMPE.

#speechrecognition #nhs #healthtech
August 11, 2026 at 3:27 PM
Thank you.

the above phrase the #speechrecognition process so often hallucinates

I've not properly configured/tuned it yet

but/though/_and_ I'm full of gratitude, or I aim to be anyway!
August 9, 2026 at 3:34 PM
Microsoft has confirmed that Dragon Medical Practice Edition will stop working on 14/08/2026.

Switch to Lexacom Echo for fast, accurate medical speech recognition with powerful new capabilities.

Better still, Lexacom Echo is half the price of DMPE.

#speechrecognition #nhs #healthtech
August 6, 2026 at 2:52 PM
Whisper can accurately convert Swiss German speech into Standard German in a zero-shot setting, with strong human ratings despite occasional hallucinations. #Whisper #ASR #SwissGerman #AI #SpeechRecognition arxiv.org/html/2404.19...
Does Whisper Understand Swiss German? An Automatic, Qualitative and Human Evaluation
Whisper is a state-of-the-art automatic speech recognition (ASR) model Radford et al. (2022). Although Swiss German dialects are allegedly not part of Whisper’s training data, preliminary experiments…
arxiv.org
August 4, 2026 at 6:52 PM