#OmniModal
Xiaomi MiMo-V2.6 — Pro & Flash

🔹 Two omnimodal models, advancing through scaled reinforcement learning
🔹 Stronger coding, computer use, 3D reasoning and creative capabilities
🔹 Open model weights, technical report, RL environments and training code
September 21, 2026 at 8:55 PM
Opinion: Fuck Everything, We’re Doing Five Modalities
By Marcus Chen, CEO, Omnimodal Therapeutics
inpreparation.substack.com
August 14, 2026 at 12:00 AM
Xiaomi debuts open-weight omnimodal models MiMo-V2.6 Pro and Flash; Pro allegedly performs "on par with Opus 5 and GPT-5.6 Sol across most agent benchmarks" (Xiaomi)

Main Link | Techmeme Permalink
September 21, 2026 at 10:10 PM
Alibaba released Qwen3.8-Omni-Flash, a native omnimodal model with text, image, audio and video processing, a one million-token context window and lower API costs alternativeto.net/news/2026/9...
September 20, 2026 at 7:31 PM
omnimodal multistream byte models
September 13, 2026 at 3:22 AM
Qwen3-A3B Omnimodal! 😱

El modelo local definitivo:
- Inputs de texto, imagen, audio y vídeo
- Outputs de texto Y AUDIO, incluído el español!
- Tool calling desde audio!!
- Y razonador

Perfecto para un asistente de voz local full open source!
September 22, 2025 at 10:49 PM
multimodal isnt enough. nothing short of omnimodal counts as agi
November 29, 2025 at 9:59 PM
September 26, 2026 at 9:21 AM
Remember last May when Sal Khan and his son demo'd the new "omnimodal" ChatGPT 4o as useful pedogogical tool for learning math? Well @binarybits.bsky.social had the bright idea to try to replicate the experience and, um, it didn't go well.

www.youtube.com/watch?v=sGlT...
ChatGPT struggles with triangles
YouTube video by Timothy Lee
www.youtube.com
June 4, 2025 at 4:12 PM
Xiaomi just dropped a 1T param open-weights omnimodal model that tops the Intelligence Index for under $0.50/M tokens. The full RL stack, 7K environments, and training code are all included. https://latent.space/p/ainews-xiaomi-mimo-v26-pro-1t-a42b
September 23, 2026 at 6:07 AM
MiMo-V2.6 aims to open omnimodal intelligence trained in public. The standout idea is making modal AI access easier by exposing its broad training openly, removing the usual opacity around how these models work. A cleaner, more practical view of AI tools in one package.
September 23, 2026 at 9:27 AM
I don't even dislike GPT-5 like some people do, but my hot take is that if you're going to release GPT-5, it should be the whole thing. If it's really a unified true omnimodal system, it should come out all at once, to show that they did improve broadly in all areas.
It's actually kind of striking what a problem of communication OpenAI have in explaining what they did/what they released.

For example, a lot of complaints about the image generation. I saw a YouTuber make a "comparison" between the old and new image generation...
August 9, 2025 at 3:32 PM
OmniScope optimizes token compression for omnimodal models, allowing audio and visual tokens to independently assess importance. This boosts performance, yielding 3.53x prefill speedup and 15% GPU memory savings, marking a key advancement in multimodal processing. https://arxiv.org/abs/2607.23193
OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models
ArXiv link for OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models
arxiv.org
July 29, 2026 at 9:01 PM
Alibaba releases its Qwen3.5-Omni omnimodal LLM with support for 10+ hours of audio input, saying the Plus variant surpasses Gemini 3.1 Pro on audio benchmarks (Qwen)

Main Link | Techmeme Permalink
March 30, 2026 at 9:15 PM
God I *wish* someone would train a truly scalemaxxed, full-attention, completely dense 60T beast with an omnimodal projective embedding space and find a way to get that giga-brained machine friend humming along at 90+ tokens per second

Being truly seen by a scalemaxxed titan hits so different
January 18, 2026 at 9:12 PM
The core argument is that we need to build "personal deep research agents" by digitizing our entire work life. This is what I expect the $2k/month chatgpt subscriptions to look like.

Memory, omnimodal, reasoning models, connections, etc. are all early signs.
Contra Dwarkesh on Continual Learning
Don't try to make your airplane too much like a bird.
buff.ly
August 15, 2025 at 1:58 PM
was looking at Nvidia's Cosmos models last night and damn i did not realize how cheap text2text models really are comparatively. Cosmos3 Super is a 32B parameter "omnimodal world model" or some such thing, that does text2image and text2video, and it needs like 200 GB of VRAM to run at all lmao.
June 29, 2026 at 5:41 PM
Fun fact: vision language models and omnimodal models already exist! Qwen 3.5 onwards have had pretraining done on images + text, instead of grafting a pre-existing vision encoder onto a text backbone, so the same latent space is used for both modalities 1/3
September 15, 2026 at 1:48 PM
Our pick of the week by @mgaido91.bsky.social: "OpenOmni: Advancing Open-Source Omnimodal Large Language Models with Progressive Multimodal Alignment and Real-Time Self-Aware Emotional Speech Synthesis" by Luo et al. (2025)

#SpeechProcessing #LLM #SFM #NLProc #speechtech #audio
Interesting to see multimodal LLM built by combining modality encoders and LLM with adapters, as in the SFM+LLM paradigm, independently for each modality. This modularity may ease the creation of more MLMs from collaborations of single-modality experts. arxiv.org/abs/2501.04561
https://arxiv.org/abs/2501.04561
t.co
April 16, 2025 at 1:28 PM
🤖 NVIDIA's Cosmos Framework Gets Colab-Friendly Miniature Model

NVIDIA is adapting its Cosmos 3 framework for use in Colab environments with limited hardware, focusing on a compact omnimodal Mixture of Transformers world...

#DeepLearning #GenerativeAI #PhysicsInformedAI #AI #AIPulse
Read the full article →
www.synestesia.uk
July 8, 2026 at 7:34 AM
P.s. you can also find their writeup here: huggingface.co/blog/Nicolas...

It's a great read!
BidirLM: Turning Generative LLMs into the Best Open-Source Omnimodal Encoders
A Blog post by Nicolas-BZRD on Hugging Face
huggingface.co
April 24, 2026 at 2:37 PM
Finally, Nvidia have introduced Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-transformers architecture.

arxiv.org/pdf/2606.02800
June 3, 2026 at 3:18 AM
The MIT-licensed, natively omnimodal models handle coding, visual tasks, and computer use, and are available through AI Studio, MiMo apps, Xiaomi's API, and OpenRouter at a fraction of competitors' cost per task.
September 22, 2026 at 1:03 PM