Yet no base model, no training data, no audio samples. Near-zero downloads. Thin card for a 1.5 release.
Hard to see this beating Qwen or MiniMax in production...
Yet no base model, no training data, no audio samples. Near-zero downloads. Thin card for a 1.5 release.
Hard to see this beating Qwen or MiniMax in production...
VibeVoice Community is a community-maintained fork preserving and extending Microsoft’s VibeVoice codebase for long-form, multi-speaker conversational TTS (and related tooling such as stream...
https://gitrated.com/vibevoice-community/VibeVoice
VibeVoice Community is a community-maintained fork preserving and extending Microsoft’s VibeVoice codebase for long-form, multi-speaker conversational TTS (and related tooling such as stream...
https://gitrated.com/vibevoice-community/VibeVoice
sonicfield.org/the-voice-th...
sonicfield.org/the-voice-th...
сжатую версию VibeVoice-ASR для работы в реальном...
https://github.com/microsoft/VibeASR.cpp
https://huggingface.co/microsoft/VibeVoice-ASR-BitNet
https://huggingface.co/spaces/microsoft/vibevoice-asr-bitnet-demo
сжатую версию VibeVoice-ASR для работы в реальном...
https://github.com/microsoft/VibeASR.cpp
https://huggingface.co/microsoft/VibeVoice-ASR-BitNet
https://huggingface.co/spaces/microsoft/vibevoice-asr-bitnet-demo
Write "VIBE" in comments to get the asset.
#VibeVoice #AITools #SpeechToText #OpenSource #IkramRana
Write "VIBE" in comments to get the asset.
#VibeVoice #AITools #SpeechToText #OpenSource #IkramRana
AI工具推荐周报 2026-07-30:精选8个本周新工具——open-seo、MediaCrawler、book-to-skill、VibeVoice、Kimi-K3...
https://justpeng.blog/ai-tool-scan-2026-07-30/
AI工具推荐周报 2026-07-30:精选8个本周新工具——open-seo、MediaCrawler、book-to-skill、VibeVoice、Kimi-K3...
https://justpeng.blog/ai-tool-scan-2026-07-30/
Open-Source Frontier Voice AI
Open-Source Frontier Voice AI
VibeVoice is an open-source suite of frontier voice AI models (ASR, TTS, and streaming TTS) featuring long-form, multi-speaker capabilities, 60-minute ASR transcriptions with Who/When/What outputs, and 90-minute TTS across up to 4 speakers, all powered by low-frame-rate tokenizers (1/2)
VibeVoice is an open-source suite of frontier voice AI models (ASR, TTS, and streaming TTS) featuring long-form, multi-speaker capabilities, 60-minute ASR transcriptions with Who/When/What outputs, and 90-minute TTS across up to 4 speakers, all powered by low-frame-rate tokenizers (1/2)
📦 microsoft / VibeVoice
⭐ 51,061 (+332)
🗒 Python
Open-Source Frontier Voice AI
Unlike conventional ASR models that slice audio into short chunks (often losing global context), it takes upto 60 min of audio input within 64K token length to ensure consistent speaker tracking & semantic coherence with Hotwords support.
Unlike conventional ASR models that slice audio into short chunks (often losing global context), it takes upto 60 min of audio input within 64K token length to ensure consistent speaker tracking & semantic coherence with Hotwords support.
📦 https://github.com/VibeVoice/VibeVoice
🔗 https://github.com/javimosch/supercli
#vibevoice #supercli
📦 https://github.com/VibeVoice/VibeVoice
🔗 https://github.com/javimosch/supercli
#vibevoice #supercli