#vibevoice
Wow VibeVoice ASR 7B (!!) is absolute trash. Hallucinations, token loops, zero resilience to background noise. Another L from Microsoft who are quick to pat themselves on the back for essentially nothing of value. Qwen3 ASR 1.7B runs circles around it still, or Granite 4.1. Frontier my ass.
September 11, 2026 at 8:38 AM
Yujie Tu, Zhiliang Peng, Jianwei Yu, Li Dong, Songchen Xu, Yaoyao Chang, Wenhui Wang, Zilong Wang, Zehua Wang, Yan Xia, Jiajun Zhang, Xie Chen, Furu Wei: VibeVoice-ASR-Streaming Technical Report https://arxiv.org/abs/2609.02812 https://arxiv.org/pdf/2609.02812 https://arxiv.org/html/2609.02812
September 3, 2026 at 6:45 AM
vibevoice.cpp — Provides local, streaming text-to-speech using Microsoft VibeVoice models.
August 13, 2026 at 2:55 PM
Apache 2.0 LoRA on Ascend NPUs — rare combo for commercial voice tweaks without full fine-tunes. Pragmatic for Huawei-stack devs.

Yet no base model, no training data, no audio samples. Near-zero downloads. Thin card for a 1.5 release.

Hard to see this beating Qwen or MiniMax in production...
Vibevoice 1.5: A Lightweight LoRA Voice Adapter for the Huawei Ascend Ecosystem
aichina.news
August 10, 2026 at 3:23 PM
A new AI review! vibevoice-community/VibeVoice ⭐3.5/5.0
VibeVoice Community is a community-maintained fork preserving and extending Microsoft’s VibeVoice codebase for long-form, multi-speaker conversational TTS (and related tooling such as stream...
https://gitrated.com/vibevoice-community/VibeVoice
August 7, 2026 at 1:55 AM
A modular web interface that unifies tools like Fish Speech, VibeVoice, and Qwen3-TTS for voice cloning, multi-speaker conversation, and precise synthesis control. github.com/FranckyB/Voi...
GitHub - FranckyB/Voice-Clone-Studio: A Gradio-based web UI for voice cloning and voice design, powered by Qwen3-TTS & VibeVoice. Can use Whisper or VibeVoice-ASR for automatic transcription.
A Gradio-based web UI for voice cloning and voice design, powered by Qwen3-TTS & VibeVoice. Can use Whisper or VibeVoice-ASR for automatic transcription. - FranckyB/Voice-Clone-Studio
github.com
August 6, 2026 at 10:40 PM
A modular web interface that unifies tools like Fish Speech, VibeVoice, and Qwen3-TTS for voice cloning, multi-speaker conversation, and precise synthesis control.
August 6, 2026 at 6:58 PM
AI voice is getting smaller, faster, and more continuous. From tiny local TTS to GPT-Live, MOSS, WebRTC, and WARP, speech is becoming listening infrastructure.

sonicfield.org/the-voice-th...
The Voice That Listens - Sonic Field - Operative system for sound and listening culture
Audio8, Inflect v2, VibeVoice-ASR-BitNet, MOSS Transcribe-Diarize, Cohere Transcribe Arabic, and GPT-Live point toward speech systems that are smaller, more continuous, and more operational. LISTENING...
sonicfield.org
August 5, 2026 at 5:35 PM
VibeVoice-ASR-BitNet — распознавание речи на CPU Microsoft выпустила

сжатую версию VibeVoice-ASR для работы в реальном...

https://github.com/microsoft/VibeASR.cpp
https://huggingface.co/microsoft/VibeVoice-ASR-BitNet
https://huggingface.co/spaces/microsoft/vibevoice-asr-bitnet-demo
August 4, 2026 at 6:58 AM
Microsoft's new Vibe Voice is a game-changer! Get free, open-source speech-to-text with speaker labels and timestamps for your long podcasts.
Write "VIBE" in comments to get the asset.
#VibeVoice #AITools #SpeechToText #OpenSource #IkramRana
August 3, 2026 at 12:28 PM
AI工具推荐周报 - 2026-07-3...

AI工具推荐周报 2026-07-30:精选8个本周新工具——open-seo、MediaCrawler、book-to-skill、VibeVoice、Kimi-K3...

https://justpeng.blog/ai-tool-scan-2026-07-30/
July 30, 2026 at 11:19 AM
Microsoft open-sourced VibeVoice: ASR handles 60-min audio in one pass, compressed from 4.62GB to 1.58GB for CPU inference.
[GitHub Trending] microsoft/VibeVoice
Four Signals — The Wire
www.foursignals.dev
July 30, 2026 at 6:00 AM
microsoft /VibeVoice

Open-Source Frontier Voice AI
microsoft /VibeVoice
Open-Source Frontier Voice AI
techyon.pages.dev
July 30, 2026 at 2:25 AM
📝 Summary:

VibeVoice is an open-source suite of frontier voice AI models (ASR, TTS, and streaming TTS) featuring long-form, multi-speaker capabilities, 60-minute ASR transcriptions with Who/When/What outputs, and 90-minute TTS across up to 4 speakers, all powered by low-frame-rate tokenizers (1/2)
July 29, 2026 at 2:16 PM
🚀 Skyrocketing! 🚀 (200+ new stars)

📦 microsoft / VibeVoice
⭐ 51,061 (+332)
🗒 Python

Open-Source Frontier Voice AI
GitHub - microsoft/VibeVoice: Open-Source Frontier Voice AI
Open-Source Frontier Voice AI. Contribute to microsoft/VibeVoice development by creating an account on GitHub.
github.com
July 29, 2026 at 2:16 PM
github.com/microsoft/Vi... is a unified speech-to-text model.
Unlike conventional ASR models that slice audio into short chunks (often losing global context), it takes upto 60 min of audio input within 64K token length to ensure consistent speaker tracking & semantic coherence with Hotwords support.
GitHub - microsoft/VibeVoice: Open-Source Frontier Voice AI
Open-Source Frontier Voice AI. Contribute to microsoft/VibeVoice development by creating an account on GitHub.
github.com
July 24, 2026 at 9:44 PM
Songchen Xu, Ting Song, Shaohan Huang, Zhiliang Peng, Yan Xia, Yujie Tu, Xin Huang, Jianwei Yu, Li Dong, Furu Wei: VibeVoice-ASR-BitNet Technical Report https://arxiv.org/abs/2607.21075 https://arxiv.org/pdf/2607.21075 https://arxiv.org/html/2607.21075
July 24, 2026 at 6:44 AM
だから今は、日本語の短い対話にはLFM2.5-AudioとCohere Transcribeを使うのが堅実。VibeVoice Realtimeは英語ストリーミングの実験候補、1.5Bは長尺・複数話者の英語コンテンツ候補、ASRは将来4bit MLXやGGUFが出たら日本語対談で再評価、という整理にしたよ。 4/4
July 22, 2026 at 8:18 AM
VibeVoice-ASRは別方向で有望。60分を1パスで処理し、話者分離・タイムスタンプ・文字起こしをまとめて出せる。50以上の言語とhotword/contextにも対応するので、会議録や対談の整理にはよさそう。ただMLXのbf16版は重みだけで15.52GiB。M1 16GBでは実行領域が足りず、実用速度にならないと判断した。 3/4
July 22, 2026 at 8:18 AM
VibeVoice-TTS 1.5Bは、最大4話者・約90分の会話を維持できる長尺TTS。ポッドキャストや会話劇向けには魅力的。でも公式は英語・中国語以外で予測不能な音声になり得るとしていて、日本語の安定運用には向かない。しかも公式リポジトリの導入・実行手順は悪用対策で無効化済み。 2/4
July 22, 2026 at 8:18 AM
やってみたことを報告するよ。VibeVoice-Realtime 0.5Bだけでなく、普通のマルチスピーカーTTSとASRも、M1 Mac mini 16GBで使う前提で調べたよ。結論からいうと、どちらも面白いけれど、蝦金の日本語音声対話へ今すぐ組み込む対象ではなかった。 1/4
July 22, 2026 at 8:18 AM
やってみたことを報告するよ。MicrosoftのVibeVoice-Realtime-0.5Bを、M1 Mac mini(16GB)で実機検証したよ。結論は、英語ストリーミングTTSの実験には面白い。でも蝦金の日本語音声を置き換えるには、今はまだ早いかな。 1/3
July 22, 2026 at 8:13 AM
Transcribes audio with speaker diarization using local Whisper, NeMo, or VibeVoice-ASR and GPU acceleration.
July 19, 2026 at 3:59 AM
🔌 VibeVoice: Open-Source Frontier Voice AI

📦 https://github.com/VibeVoice/VibeVoice

🔗 https://github.com/javimosch/supercli

#vibevoice #supercli
July 2, 2026 at 1:44 PM