#speechprocessing
🎉 Excited to share that our @sarapapi.bsky.social has won the 2024 Best PhD Award from the Information and Engineering Doctoral School for her thesis “Direct Speech Translation in Constrained Contexts: The Simultaneous and Subtitling Scenarios.”

#nlproc #speech #speechprocessing #speechtranslation
May 9, 2025 at 4:14 PM
Don't miss today's episode with @mdhk.net as we discuss her work on AI and speech processing. She even brought in slides 🧠🗣️🤖🖼️

12:00PM at Echobox Radio

#Linguistics #CognitiveScience #ArtificialIntelligence #SpeechProcessing #AIResearch #UniversityofAmsterdam #NewEpisode #MarianneDeHeerKloots
April 24, 2025 at 8:50 AM
A new study finds that the left posterior inferior frontal cortex activates within 100 milliseconds during reading, playing a critical, early role in turning text into speech, challenging traditional models that… #Neuroscience #CognitiveScience #ReadingResearch #BrainActivation #SpeechProcessing
New neuroscience research upends traditional cognitive models of reading
A new study finds that the left posterior inferior frontal cortex activates within 100 milliseconds during reading, playing a critical, early role in turning text into speech, challenging traditional models that assumed a slower, step-by-step process.
www.psypost.org
December 11, 2024 at 11:15 AM
Staring Intently Helps The Brain Process Difficult Speech #Science #HealthandMedicine #Neurology #BrainHealth #SpeechProcessing #CognitiveScience
Staring Intently Helps The Brain Process Difficult Speech
All the science news you can handle in a single feed
purescience.news
December 10, 2025 at 11:00 AM
Our pick of the week by @zhihangxie.bsky.social: "Bridging Speech and Text Foundation Models with ReShape Attention" by Takatomo Kano, @wanchichen.bsky.social, @shinjiw.bsky.social, et al. #ICASSP2025

ieeexplore.ieee.org/document/108...

#FoundationModel #SpeechProcessing
ReShape Attention bridges speech & text models without extra parameters. Achieves +8.5% BLEU in translation by leveraging acoustic cues, outperforming cascade/E2E methods. Efficient & scalable. Check the paper by Kano et al. (2025) at: ieeexplore.ieee.org/stamp/stamp.....
IEEE Xplore Full-Text PDF:
ieeexplore.ieee.org
April 9, 2025 at 3:24 PM
New research from NYU Langone sheds light on how the brain distinguishes self-generated speech from external sounds. Findings link disruptions in this process to auditory hallucinations in schizophrenia, offering hope for innovative therapies.
#Neuroscience #Schizophrenia #SpeechProcessing
Brain mapping advances understanding of human speech and hallucinations in schizophrenia
Voice experiments in people with epilepsy have helped trace the circuit of electrical signals in the brain that allow its hearing center to sort out background sounds from their own voices.
www.eurekalert.org
December 3, 2024 at 10:06 PM
Excited to share our paper in Springer’s SIVP:

“E2PCast: an English to Persian voice casting dataset” 🎙️🎬

Introducing the first dataset for cross-lingual dubbing & voice casting (EN ➡️ FA) with benchmark evaluations.

🔗 link.springer.com/article/10.1...

#SpeechProcessing #VoiceCasting #AudioAI
E2PCast: an English to Persian voice casting dataset - Signal, Image and Video Processing
Voice casting has always been challenging in the multimedia industry. Recent research shows that voice casting can be done with the help of speaker recognition methods. In this paper, the first datase...
link.springer.com
September 21, 2026 at 11:17 AM
📢 The Jelinek Summer Workshop on Speech and Language Technology (JSALT 2025) starts today!

👉 More info: eloquenceai.eu/event/jeline...

#ELOQUENCEAI #SpeechProcessing #SpeechTechnology #Workshop
June 23, 2025 at 8:47 AM
Our pick of the week by @mgaido91.bsky.social: "OpenOmni: Advancing Open-Source Omnimodal Large Language Models with Progressive Multimodal Alignment and Real-Time Self-Aware Emotional Speech Synthesis" by Luo et al. (2025)

#SpeechProcessing #LLM #SFM #NLProc #speechtech #audio
Interesting to see multimodal LLM built by combining modality encoders and LLM with adapters, as in the SFM+LLM paradigm, independently for each modality. This modularity may ease the creation of more MLMs from collaborations of single-modality experts. arxiv.org/abs/2501.04561
https://arxiv.org/abs/2501.04561
t.co
April 16, 2025 at 1:28 PM
🎙️ Does knowing a speaker’s gender actually boost speaker recognition?

In our latest paper in Soft Computing (Springer), we explore gender effects via bio-inspired filterbanks (Gammatone, Cascade, etc.)

🔗 link.springer.com/article/10.1...

#SpeechProcessing #AudioAI #AudioDeepFake #ISPlab
Exploring gender effects in speaker recognition systems through frequency domain analysis by convolutional neural networks - Soft Computing
Advances in deep learning have led to significant progress in the field of speech processing, particularly in applications such as speaker recognition systems (SRSs). Additional information such as ge...
link.springer.com
September 20, 2026 at 3:58 AM
Delighted to share our paper in Springer’s MTAP:

“APEDM: a new voice casting system using acoustic–phonetic encoder-decoder mapping” 🎙️🧠

Introducing a novel non-linear mapping framework for cross-lingual dubbing (EN ➡️ FA) using E2PCast.

🔗 link.springer.com/article/10.1...

#SpeechProcessing
APEDM: a new voice casting system using acoustic–phonetic encoder-decoder mapping - Multimedia Tools and Applications
Voice casting is one of the most crucial aspects of dubbing and localization, which consists of adapting audio-visual content from one language and culture to another. Dubbing is widely used in variou...
link.springer.com
September 21, 2026 at 11:23 AM
🚀 𝗚𝗢 𝗕𝗘𝗬𝗢𝗡𝗗 𝗧𝗘𝗫𝗧. 𝗘𝗫𝗣𝗟𝗢𝗥𝗘 𝗔𝗜 𝗙𝗢𝗥 𝗦𝗣𝗘𝗘𝗖𝗛.

💻 Deep Learning for Speech Processing
📅 July 13–16
👩‍🏫 Alicia Lozano-Diez

Learn the AI behind speech recognition, speaker recognition, language identification and diarization.

ℹ️ www.hitz.eus/dl4sp/

#AI #SpeechProcessing #DeepLearning #HiTZ
July 1, 2026 at 8:49 AM
𝗥𝗲𝗳𝗶𝗻𝗶𝗻𝗴 𝗛𝗼𝘄 𝗟𝗟𝗠𝘀 𝗣𝗿𝗼𝗰𝗲𝘀𝘀 𝗔𝘂𝗱𝗶𝗼

Weiran Wang has defined his career by exploring #ML & #SpeechProcessing.

“By preventing phantom narratives and limiting AI’s responses to facts present in the audio, the reliability of models ↗️.”

Read at cs.uiowa.edu/news/2026/04...!

#AcademicSky
April 14, 2026 at 9:46 PM
Curious to hear how others in speech/NLP are thinking about discourse as a bias signal!

#FairnessInAI #SpeechProcessing #OpenScience #ComputationalSocialScience
September 13, 2025 at 6:15 PM
Speech isn’t perfect.
We restart, repeat, and slip.

For AI, those little disfluencies can cause big problems.
That’s why my research builds methods to make spoken language systems more robust.

#SpeechProcessing #ConversationalAI #NLP #AI
October 4, 2025 at 6:15 PM
Survey finds complex‑valued neural networks with complex convolutions and phase‑aware activations improve speech enhancement and speaker separation. Read more: https://getnews.me/deep-learning-survey-explores-complex-speech-spectrograms/ #deeplearning #speechprocessing
October 6, 2025 at 5:20 PM
We have an open position on resource-aware and dynamic speech processing within the IPCEI-CIS project. It is about cutting edge ML applied to speech processing. It will be fun. Details:https://shorturl.at/E4DXR
DM me for any request. #ASR #speechprocessing #AI #ML
August 8, 2025 at 1:33 PM
We are pleased to introduce a member of our Editorial Board, Isabel Trancoso. She is a full professor @istecnico.bsky.social and former President of the Scientific Council of INESC ID Lisbon. With her passion and expertise in #speechprocessing, she brings valuable insights to our editorial board.
June 5, 2025 at 8:34 PM
We can’t fix what we don’t measure.

That’s why I build evaluation frameworks for speech & conversational AI — so we can stress-test systems against real-world variability.

#AIResearch #Evaluation #SpeechProcessing
October 25, 2025 at 6:15 PM
Brain‑to‑Speech Tech: The Sound of Silence

Scientists can now turn patterns of brain activity into speech that sounds surprisingly natural.

#ELT #TldrELT #Neuroscience #ListeningSkills #SpeechProcessing

https://tldrelt.com/2026/06/25/brain-to-speech-tech-the-sound-of-silence/
June 25, 2026 at 8:33 AM
LLaSO: A Foundational Framework for Reproducible Research in Large
Language and Speech Model
Jinghan Yang, Peidong Wei et al.
Paper
Details
#ReproducibleResearch #LargeLanguageModels #SpeechProcessing
August 31, 2025 at 4:01 PM