Whisper for Arabic–English speech with Indian accent
<p>Whisper <a href="https://huggingface.co/datasets/John6666/forum3/blob/main/whisper_ar_en_1.md">might be stuck in the worst possible situation for this model</a>…?</p>
<hr />
<h2><a class="anchor" href="https://discuss.huggingface.co#p-250597-why-this-setting-is-hard-for-vanilla-whisper-1" name="p-250597-why-this-setting-is-hard-for-vanilla-whisper-1"></a>Why this setting is hard for “vanilla” Whisper</h2>
<h3><a class="anchor" href="https://discuss.huggingface.co#p-250597-code-switching-breaks-the-models-strongest-assumptions-2" name="p-250597-code-switching-breaks-the-models-strongest-assumptions-2"></a>Code-switching breaks the model’s strongest assumptions</h3>
<p>Whisper-style models are trained to produce <strong>one coherent transcript</strong> from a window of audio. In code-switch speech, the model must decide (often multiple times per second) whether the next token should come from Arabic script or Latin script, while also handling shared phonetics and loanwords. When the evidence is weak (fast speech, noise, accent), the decoder tends to “commit” to one language and then <strong>keep sampling from that language’s token distribution</strong>, which can spill across the true switch boundary.</p>
<h3><a class="anchor" href="https://discuss.huggingface.co#p-250597-indian-accented-speech-increases-phonetic-ambiguity-3" name="p-250597-indian-accented-speech-increases-phonetic-ambiguity-3"></a>Indian-accented speech increases phonetic ambiguity</h3>
<p>Accents affect:</p>
<ul>
<li>vowel/consonant realizations,</li>
<li>stress timing,</li>
<li>coarticulation patterns,</li>
<li>and segment durations.</li>
</ul>
<p>For short, noisy messages, these shifts are enough to push the model into “low-evidence” decoding where it starts relying more on its language model prior than the acoustic signal. That’s when you see substitutions, omissions, or fluent but wrong text.</p>
<h3><a class="anchor" href="https://discuss.huggingface.co#p-250597-auto-language-detection-is-fragile-on-shortnoisy-audio-4" name="p-250597-auto-language-detection-is-fragile-on-shortnoisy-audio-4"></a>Auto language detection is fragile on short/noisy audio</h3>
<p>In <code>faster-whisper</code>, language detection is performed using the <strong>first ~30 seconds</strong> if you don’t set <code>language=...</code> explicitly. That is a known source of wrong-language outputs if the beginning includes silence/noise or code-switching. (<a href="https://github.com/SYSTRAN/faster-whisper/blob/master/faster_whisper/transcribe.py" title="faster-whisper/faster_whisper/transcribe.py at master">GitHub</a>)<br />
This interacts badly with your setting: once the model “picks” the wrong language early, later chunks are biased.</p>
<h3><a class="anchor" href="https://discuss.huggingface.co#p-250597-repetitionhallucination-loops-are-a-known-failure-mode-with-silencegaps-5" name="p-250597-repetitionhallucination-loops-are-a-known-failure-mode-with-silencegaps-5"></a>Repetition/hallucination loops are a known failure mode with silence/gaps</h3>
<p>Two community-validated mitigations for “stuck repeating / hallucinating after a gap” are:</p>
<ul>
<li>split audio with VAD,</li>
<li>set <code>condition_on_previous_text=False</code>. (<a href="https://github.com/openai/whisper/discussions/679" title="A possible solution to Whisper hallucination #679">GitHub</a>)<br />
This matters for voice messages because they often have pauses and trailing non-speech.</li>
</ul>
<hr />
<h2><a class="anchor" href="https://discuss.huggingface.co#p-250597-practical-pipeline-changes-that-usually-move-the-needle-6" name="p-250597-practical-pipeline-changes-that-usually-move-the-needle-6"></a>Practical pipeline changes that usually move the needle</h2>
<h3><a class="anchor" href="https://discuss.huggingface.co#p-250597-h-1-make-segmentation-your-primary-quality-lever-7" name="p-250597-h-1-make-segmentation-your-primary-quality-lever-7"></a>1) Make segmentation your primary quality lever</h3>
<p>For short messages, segmentation quality often dominates model size.</p>
<p><strong>Target behavior</strong>: feed the decoder 1–8s windows that are “mostly speech”, with small padding and minimal trailing non-speech.</p>
<p><strong>Recommended segmentation recipe</strong></p>
<ul>
<li>VAD to find speech islands</li>
<li>add padding (e.g., 150–300ms)</li>
<li>add overlap (e.g., 100–250ms) to protect word boundaries</li>
<li><strong>explicit tail trimming</strong> after VAD (energy/RMS-based) to remove long quiet endings that trigger hallucinations</li>
<li>cap maximum segment length (e.g., 8–12s); long segments increase drift and LID errors</li>
</ul>
<p>This aligns with common “hallucination fix” guidance: VAD slicing plus disabling conditioning reduces loops when the model can’t find evidence in the current window. (<a href="https://github.com/openai/whisper/discussions/679" title="A possible solution to Whisper hallucination #679">GitHub</a>)</p>
<h3><a class="anchor" href="https://discuss.huggingface.co#p-250597-h-2-default-to-guardrails-on-for-production-decoding-8" name="p-250597-h-2-default-to-guardrails-on-for-production-decoding-8"></a>2) Default to “guardrails ON” for production decoding</h3>
<p>A simple but important rule: don’t disable the thresholds unless you’re deliberately building a repro.</p>
<p>When thresholds are enabled, Whisper-style decoding has mechanisms to suppress output during no-speech/low-confidence regions (implemented in the reference transcribe logic). (<a href="https://github.com/openai/whisper/discussions/29" title="Stops working after long gap with no speech? · openai whisper · Discussion #29 · GitHub">GitHub</a>)</p>
<p><strong>Practical defaults for voice messages</strong></p>
<ul>
<li><code>condition_on_previous_text=False</code> (prevents “carryover text” into gaps) (<a href="https://github.com/openai/whisper/discussions/29" title="Stops working after long gap with no speech? · openai whisper · Discussion #29 · GitHub">GitHub</a>)</li>
<li>keep <code>no_speech_threshold</code>, <code>log_prob_threshold</code>, <code>compression_ratio_threshold</code> enabled (don’t set them to <code>None</code>)</li>
<li>use <code>temperature=0.0</code> for determinism while tuning; add temperature fallback only if needed</li>
</ul>
<h3><a class="anchor" href="https://discuss.huggingface.co#p-250597-h-3-constrain-language-behavior-by-design-dont-rely-on-auto-lid-9" name="p-250597-h-3-constrain-language-behavior-by-design-dont-rely-on-auto-lid-9"></a>3) Constrain language behavior by design (don’t rely on auto-LID)</h3>
<p>Auto-LID being computed on the first ~30s is a known limitation; multiple issues report wrong-language outputs under auto-detection. (<a href="https://github.com/SYSTRAN/faster-whisper/issues/265" title="Improve Language detection #265 - SYSTRAN/faster- ...">GitHub</a>)<br />
There’s also an open request for “limit detection to a subset of languages,” which does not exist as a first-class feature in <code>faster-whisper</code> today. (<a href="https://github.com/SYSTRAN/faster-whisper/issues/1164" title="[Feature Request] Constrain Available Languages when ...">GitHub</a>)</p>
<p><strong>Workarounds that actually help</strong></p>
<ul>
<li>
<p>If you know it’s always Arabic+English, use a <strong>two-pass strategy</strong>:</p>
<ol>
<li>
<p>attempt <code>language="ar"</code> decode</p>
</li>
<li>
<p>attempt <code>language="en"</code> decode</p>
</li>
<li>
<p>pick the better result using a small heuristic:</p>
<ul>
<li>script sanity (Arabic chars ratio vs Latin ratio),</li>
<li>repetition score,</li>
<li>average logprob proxy (if available),</li>
<li>“text produced in low-energy region” penalty.</li>
</ul>
</li>
</ol>
</li>
</ul>
<p>This directly addresses “unstable language ID outputs” without needing new model features.</p>
<h3><a class="anchor" href="https://discuss.huggingface.co#p-250597-h-4-post-processing-that-is-code-switch-aware-10" name="p-250597-h-4-post-processing-that-is-code-switch-aware-10"></a>4) Post-processing that is code-switch aware</h3>
<p>Avoid “English-only cleanup” or “Arabic-only cleanup”; mixed script requires a mixed strategy.</p>
<p><strong>Low-risk post-processing ideas</strong></p>
<ul>
<li>
<p><strong>script-aware normalization</strong></p>
<ul>
<li>normalize Arabic punctuation variants (e.g., Arabic comma/Latin comma)</li>
<li>normalize tatweel and repeated diacritics only if you see them</li>
</ul>
</li>
<li>
<p><strong>repetition filters</strong></p>
<ul>
<li>detect repeated bigrams/trigrams over a threshold and either truncate or mark as suspect</li>
</ul>
</li>
<li>
<p><strong>segment-level confidence flags</strong></p>
<ul>
<li>
<p>mark segments suspicious if:</p>
<ul>
<li>very long text produced while energy is low,</li>
<li>script doesn’t match forced language pass,</li>
<li>high repetition compression-like behavior</li>
</ul>
</li>
</ul>
</li>
</ul>
<h3><a class="anchor" href="https://discuss.huggingface.co#p-250597-h-5-if-whisper-still-struggles-consider-an-alternate-base-model-as-a-reference-11" name="p-250597-h-5-if-whisper-still-struggles-consider-an-alternate-base-model-as-a-reference-11"></a>5) If Whisper still struggles: consider an alternate base model as a reference</h3>
<p>Two candidates worth testing as “sanity checks”:</p>
<ul>
<li>
<p>Meta SeamlessM4T v2: supports Arabic variants (e.g., Modern Standard Arabic, Egyptian, Moroccan) in its published supported language list, and is explicitly evaluated for ASR tasks. (<a href="https://huggingface.co/facebook/seamless-m4t-v2-large" title="facebook/seamless-m4t-v2-large">Hugging Face</a>)<br />
<em>Use case</em>: as a comparison point or fallback for Arabic-heavy segments (not necessarily best at code-switching out of the box).</p>
</li>
<li>
<p>NVIDIA Canary v2: strong multilingual ASR for its supported languages, but public materials emphasize European language coverage; Arabic support is inconsistent across deployments per community reports. (<a href="https://huggingface.co/nvidia/canary-1b-v2" title="nvidia/canary-1b-v2">Hugging Face</a>)<br />
<em>Use case</em>: less compelling if Arabic is core.</p>
</li>
</ul>
<hr />
<h2><a class="anchor" href="https://discuss.huggingface.co#p-250597-existing-fine-tuned-models-you-can-start-from-12" name="p-250597-existing-fine-tuned-models-you-can-start-from-12"></a>Existing fine-tuned models you can start from</h2>
<p>These are not a perfect match (Arabic↔English + Indian accent), but they’re useful starting points.</p>
<h3><a class="anchor" href="https://discuss.huggingface.co#p-250597-arabicenglish-code-switch-whisper-models-13" name="p-250597-arabicenglish-code-switch-whisper-models-13"></a>Arabic–English code-switch Whisper models</h3>
<ul>
<li>
<p><code>MohamedRashad/Arabic-Whisper-CodeSwitching-Edition</code><br />
Fine-tuned on an Arabic-English code-switch dataset; explicitly intended for Arabic speech with embedded English words. License shown as GPL-3.0 (often problematic for commercial use). (<a href="https://huggingface.co/MohamedRashad/Arabic-Whisper-CodeSwitching-Edition" title="MohamedRashad/Arabic-Whisper-CodeSwitching-Edition · Hugging Face">Hugging Face</a>)</p>
</li>
<li>
<p><code>azeem23/whisper-small-codeswitching-ArabicEnglish</code><br />
A smaller Whisper variant fine-tuned for Arabic-English code-switching, based on the same dataset. (<a href="https://huggingface.co/azeem23/whisper-small-codeswitching-ArabicEnglish" title="azeem23/whisper-small-codeswitching-ArabicEnglish">Hugging Face</a>)</p>
</li>
</ul>
<h3><a class="anchor" href="https://discuss.huggingface.co#p-250597-indian-accent-english-whisper-model-14" name="p-250597-indian-accent-english-whisper-model-14"></a>Indian-accent English Whisper model</h3>
<ul>
<li><code>Tejveer12/Indian-Accent-English-Whisper-Finetuned</code><br />
Fine-tuned on the Indian-accent English dataset (<code>WillHeld/india_accent_cv</code>). (<a href="https://huggingface.co/Tejveer12/Indian-Accent-English-Whisper-Finetuned" title="Tejveer12/Indian-Accent-English-Whisper-Finetuned · Hugging Face">Hugging Face</a>)<br />
The model repository indicates an MIT license in its metadata/commit history. (<a href="https://huggingface.co/Tejveer12/Indian-Accent-English-Whisper-Finetuned/commit/26b6d18da8db69db8290513077c1b57727b43181" title="Training in progress, step 7000 · Tejveer12/Indian-Accent- ...">Hugging Face</a>)</li>
</ul>
<p><strong>How to use these in practice</strong></p>
<ul>
<li>Use the Indian-accent model as an <strong>English-pass decoder</strong> for English-dominant segments.</li>
<li>Use a code-switch model as the <strong>Arabic-pass decoder</strong> (especially for Arabic-dominant segments with English insertions).</li>
<li>Or: use these as initialization targets for your own adapter fine-tune (next section).</li>
</ul>
<hr />
<h2><a class="anchor" href="https://discuss.huggingface.co#p-250597-fine-tuning-whisper-for-your-exact-data-practical-recipe-15" name="p-250597-fine-tuning-whisper-for-your-exact-data-practical-recipe-15"></a>Fine-tuning Whisper for your exact data (practical recipe)</h2>
<h3><a class="anchor" href="https://discuss.huggingface.co#p-250597-h-1-use-adapter-style-fine-tuning-lora-first-16" name="p-250597-h-1-use-adapter-style-fine-tuning-lora-first-16"></a>1) Use adapter-style fine-tuning (LoRA) first</h3>
<p>Full fine-tuning of large Whisper checkpoints is expensive and easy to overfit. For accent + code-switch adaptation, LoRA usually gets you most of the gain with lower risk.</p>
<p>The Hugging Face PEFT guide shows an int8 + LoRA training approach for Whisper ASR specifically. (<a href="https://huggingface.co/docs/peft/v0.6.0/en/task_guides/int8-asr" title="int8 training for automatic speech recognition">Hugging Face</a>)</p>
<p><strong>Why LoRA helps here</strong></p>
<ul>
<li>You’re adapting pronunciation + boundary behavior, not learning a new language.</li>
<li>You want to preserve general robustness while nudging the model toward your accent and code-switch distribution.</li>
</ul>
<h3><a class="anchor" href="https://discuss.huggingface.co#p-250597-h-2-build-a-training-mix-that-matches-your-deployment-distribution-17" name="p-250597-h-2-build-a-training-mix-that-matches-your-deployment-distribution-17"></a>2) Build a training mix that matches your deployment distribution</h3>
<p>Aim for three buckets:</p>
<ol>
<li><strong>In-domain</strong>: your actual voice messages (even 10–50 hours helps if transcripts are consistent)</li>
<li><strong>Indian-accent English</strong>: augment English segments with accent data (e.g., <code>WillHeld/india_accent_cv</code>) (<a href="https://huggingface.co/datasets/WillHeld/india_accent_cv" title="WillHeld/india_accent_cv · Datasets at Hugging Face">Hugging Face</a>)</li>
<li><strong>Arabic–English code-switch</strong>: add code-switch examples (e.g., MohamedRashad dataset/models; also consider Mixat for methodology even if dialect differs) (<a href="https://huggingface.co/MohamedRashad/Arabic-Whisper-CodeSwitching-Edition" title="MohamedRashad/Arabic-Whisper-CodeSwitching-Edition · Hugging Face">Hugging Face</a>)</li>
</ol>
<p>If you lack real Arabic↔English code-switch hours, synthetic code-switch generation is an active research direction (phrase-level mixing) and can be used to bootstrap. (<a href="https://www.isca-archive.org/interspeech_2025/nguyen25_interspeech.pdf" title="Can we train ASR systems on Code-switch without real ...">isca-archive.org</a>)</p>
<h3><a class="anchor" href="https://discuss.huggingface.co#p-250597-h-3-keep-transcript-conventions-strict-and-stable-18" name="p-250597-h-3-keep-transcript-conventions-strict-and-stable-18"></a>3) Keep transcript conventions strict and stable</h3>
<p>For code-switch, consistency matters more than perfection:</p>
<ul>
<li>keep Arabic in Arabic script and English in Latin script</li>
<li>avoid random transliterations</li>
<li>normalize punctuation and casing rules consistently across the dataset</li>
</ul>
<h3><a class="anchor" href="https://discuss.huggingface.co#p-250597-h-4-training-choices-that-matter-most-for-your-case-19" name="p-250597-h-4-training-choices-that-matter-most-for-your-case-19"></a>4) Training choices that matter most for your case</h3>
<ul>
<li>
<p>Start from a multilingual checkpoint (e.g., Whisper small/medium/large-v3 depending on budget)</p>
</li>
<li>
<p>Use <code>task="transcribe"</code> (not translate)</p>
</li>
<li>
<p>Ensure audio is standardized to 16kHz mono</p>
</li>
<li>
<p>Filter or downweight:</p>
<ul>
<li>clips with extremely low SNR,</li>
<li>clips with unreliable transcripts,</li>
<li>clips with long non-speech tails (or trim them)</li>
</ul>
</li>
</ul>
<h3><a class="anchor" href="https://discuss.huggingface.co#p-250597-h-5-evaluation-dont-rely-on-one-wer-number-20" name="p-250597-h-5-evaluation-dont-rely-on-one-wer-number-20"></a>5) Evaluation: don’t rely on one WER number</h3>
<p>Use at least:</p>
<ul>
<li>overall WER</li>
<li><strong>English-only WER on English spans</strong></li>
<li><strong>Arabic-only WER/CER on Arabic spans</strong></li>
<li>a “switch-boundary” check (simple proxy): count how often the script flips in the right neighborhood of known switch points (even a heuristic boundary test catches regressions quickly)</li>
</ul>
<hr />
<h2><a class="anchor" href="https://discuss.huggingface.co#p-250597-high-quality-references-to-follow-end-to-end-21" name="p-250597-high-quality-references-to-follow-end-to-end-21"></a>High-quality references to follow end-to-end</h2>
<ul>
<li>Hugging Face blog: “Fine-Tune Whisper For Multilingual ASR with Transformers” (step-by-step). (<a href="https://huggingface.co/blog/fine-tune-whisper" title="Fine-Tune Whisper For Multilingual ASR with 🤗 Transformers">Hugging Face</a>)</li>
<li>PEFT int8 + LoRA ASR guide for Whisper (T4-friendly training approach). (<a href="https://huggingface.co/docs/peft/v0.6.0/en/task_guides/int8-asr" title="int8 training for automatic speech recognition">Hugging Face</a>)</li>
<li>Whisper hallucination mitigation discussion: VAD slicing + <code>condition_on_previous_text=False</code>. (<a href="https://github.com/openai/whisper/discussions/679" title="A possible solution to Whisper hallucination #679">GitHub</a>)</li>
<li>Code-switch dataset methodology reference: Mixat paper (how they build and analyze code-mixed Arabic/English speech). (<a href="https://aclanthology.org/2024.sigul-1.26/" title="Mixat: A Data Set of Bilingual Emirati-English Speech">ACL Anthology</a>)</li>
<li><code>faster-whisper</code> language detection limitations and wrong-language reports. (<a href="https://github.com/SYSTRAN/faster-whisper/issues/265" title="Improve Language detection #265 - SYSTRAN/faster- ...">GitHub</a>)</li>
</ul>
<hr />
<h2><a class="anchor" href="https://discuss.huggingface.co#p-250597-a-concrete-starting-plan-for-your-production-pipeline-22" name="p-250597-a-concrete-starting-plan-for-your-production-pipeline-22"></a>A concrete “starting plan” for your production pipeline</h2>
<ol>
<li>
<p><strong>Segment aggressively</strong> (VAD + pad + overlap + explicit tail trim) before decoding. (<a href="https://github.com/openai/whisper/discussions/679" title="A possible solution to Whisper hallucination #679">GitHub</a>)</p>
</li>
<li>
<p><strong>Decode with guardrails on</strong> and <code>condition_on_previous_text=False</code> by default for voice messages. (<a href="https://github.com/openai/whisper/discussions/679" title="A possible solution to Whisper hallucination #679">GitHub</a>)</p>
</li>
<li>
<p><strong>Two-pass language strategy</strong> per segment:</p>
<ul>
<li>run forced Arabic decode, forced English decode</li>
<li>choose output by script sanity + repetition penalty (+ score proxy if available)</li>
</ul>
</li>
<li>
<p><strong>Fallback policy</strong>: if output is suspicious (wrong script, repetition, text in low energy), re-decode with stricter thresholds and/or shorter segment.</p>
</li>
<li>
<p><strong>Fine-tune via LoRA</strong> using your in-domain audio + Indian-accent English + Arabic-English code-switch data. (<a href="https://huggingface.co/docs/peft/v0.6.0/en/task_guides/int8-asr" title="int8 training for automatic speech recognition">Hugging Face</a>)</p>
</li>
</ol>