#MixtureOfExperts
Une astuce de #Deepseek est le #MixtureOfExperts : pour chaque prédiction, seulement 10^10 des 10^11 paramètres sont utilisés.

Voilà qui réduit les coûts d'un ordre de grandeur (non seulement à l'entraînement, avec 3M GPU-heures, mais à l'inférence aussi !).
January 28, 2025 at 2:40 PM
I'm not a coder—you had me beat at [Hello World]. But with OpenAI, Claude, and Qwen as my dev team, I'm building delta : kitsune anyway. 18K+ lines of code and growing. I'm the architect. They're the execution. #AI #AIforGood #MixtureOfExperts #LocalAI

deltakitsune.medium.com/midweek-i-am...
Midweek: I Am Not a Coder
Let’s be clear — I understand general code logic, but I am not a coder. Never claimed to be.
deltakitsune.medium.com
November 6, 2025 at 1:47 AM
LingBot-Video: A New Open-Source MoE Model for Embodied Video Generation

https://pneumetron.com/news/ai_research/lingbot-video-open-source-moe-embodied-video-generation-b466e9

#videogeneration #mixtureofexperts #embodiedintelligence #opensource
July 14, 2026 at 2:49 AM
Tencent just dropped Hy3, a 21‑billion‑parameter MoE LLM with 256K context and sparse activation—matching the big players while staying open‑source. Curious how it stacks up? Dive in! #TencentHy3 #MixtureOfExperts #21BParameters

🔗 aidailypost.com/news/tencent...
July 6, 2026 at 6:35 PM
It was a pleasuring joining IBM's Model of Experts podcast for the first time. This podcast typically covers some of the latest developments in the world of AI; however, we decided to inject a bit of quantum computing in the conversation. Check it out: ibm.biz/BdGwRD

#MixtureOfExperts
- YouTube
Enjoy the videos and music you love, upload original content, and share it all with friends, family, and the world on YouTube.
ibm.biz
March 7, 2025 at 2:16 PM
July 6, 2026 at 1:02 PM
A huge thank you to #LeonardHackel for such a thought-provoking and forward-looking presentation! 🙌

#FoundationModels #AI #EarthObservation #Geospatial #MixtureOfExperts
February 27, 2026 at 10:35 AM
winbuzzer.com
February 14, 2026 at 3:44 PM
#Term: #MixtureOfBlockAttention (#Moba) – #ArtificialIntelligence - https://with.ga/y4ww8
"Mixture of Block Attention (MoBA) is an efficient, #SparseAttention mechanism for #Transformer models that applies the routing logic of #MixtureOfExperts (#Moe) to sequence blocks instead of...
September 11, 2026 at 9:36 AM
#Term: #MixtureOfBlockAttention (#Moba) – #ArtificialIntelligence - https://with.ga/y4ww8
"Mixture of Block Attention (MoBA) is an efficient, #SparseAttention mechanism for #Transformer models that applies the routing logic of #MixtureOfExperts (#Moe) to sequence blocks instead of...
September 11, 2026 at 9:30 AM
L'article scientifique n'est pas encore sorti, mais dans un blogpost ils parlent d'optimisation à plusieurs niveaux :
- Efficacité algorithmique (avec quantification des poids, MixtureOfExperts, décodage spéculatif, compilation...)
- Hardware & datacenters optimisés
- Gestion des modèles en idle
August 25, 2025 at 8:46 AM
Episode 36 of the IBM #MixtureOfExperts with Kate Soule, Chris Hay, and me, along with the best host in the world, @timhwang.bsky.social! We talk o3, DeepSeek-V3, what it means to be an author, and switching between different models with different personalities.
www.youtube.com/watch?v=QzER...
OpenAI o3, DeepSeek-V3, and the Brundage/Marcus AI bet
YouTube video by IBM Technology
www.youtube.com
January 3, 2025 at 7:23 PM
Arcee AI just poured half its VC into an open-reasoning model—using a 4-of-256 mixture-of-experts that fires per token. Could this reshape LLM efficiency? Dive into the details. #ArceeAI #MixtureOfExperts #OpenReasoning

🔗 aidailypost.com/news/arcee-a...
April 12, 2026 at 9:18 AM
🔥 Alibaba Qwen3-Next: 10x effizienter, 90% weeniger Trainingskosten!

▶️ Entdecke Hybrid-MoE nun
▶️ Aktiviere 262K Kontext!
▶️ Starte SGLang Turbo nun

#ai #ki #artificialintelligence #qwen3next #alibaba #llms #mixtureofexperts

🔥 Jetzt KLICKEN & KOMMENTIEREN! 💭

kinews24.de/qwen3-next-a...
September 12, 2025 at 11:40 AM
A fascinating paper on Kimi K3, a 2.8T-parameter MoE model with a 1M-token context, native vision, and novel attention designs - open weights, top-tier results, and real innovation. See link below. #LLMs #OpenSourceAI #MixtureOfExperts #AIresearch
https://arxiv.org/abs/2607.24653
August 3, 2026 at 6:29 AM
Alibaba just dropped Qwen3.8‑Max, a 2.4 trillion‑parameter Mixture‑of‑Experts LLM that dwarfs the 27B sibling. Could this be the next OpenAI challenger? Dive into the specs and what it means for the open‑source race. #Qwen3.8Max #MixtureOfExperts #2.4TParameters

🔗 aidailypost.com/news/alibaba...
August 3, 2026 at 8:43 AM
xAI releases the base model weights and network architecture of Grok-1, a 314B parameter Mixture-of-Experts model, under the Apache 2.0 license. Explore the possibilities of this open-source advancement in AI. #xAI #Grok1 #OpenSource #MixtureOfExperts #AIModel
March 17, 2024 at 11:50 PM
🚀 Kimi AI just dropped its AgentENV stack—open‑source, Firecracker microVMs, Mixture‑of‑Experts, and distributed agentic RL for Kimi K3. Ready to spin up massive agent farms? Dive in and see how Moonshot AI is reshaping training pipelines. #AgentENV #MixtureOfExperts #FirecrackerVMs

🔗
July 28, 2026 at 6:34 AM
Alibaba has introduced QwQ-Max-Preview, a new AI reasoning model designed to challenge OpenAI and DeepSeek #AI #Alibaba #QwQMaxPreview #QwenChat #GenAI #MixtureOfExperts #China
Alibaba Unveils QwQ-Max-Preview to Compete with OpenAI and DeepSeek - WinBuzzer
QwQ-Max-Preview is leveraging Mixture-of-Experts (MoE) architecture and comaptibility with the OpenAI API.
buff.ly
February 25, 2025 at 8:31 PM
winbuzzer.com/2026/07/17/m...

Moonshot AI has launched its 2.8-trillion-parameter Kimi K3 model, but a high hallucination rate might tempers its frontier-model pitch.

#AI #MoonshotAI #KimiK3 #AIModels #MultimodalAI #MixtureOfExperts #AIReasoningModels #AIBenchmarks #ChinaAI
Moonshot AI Unveils 2.8T-Parameter Kimi K3 AI Model
Moonshot AI has launched its 2.8-trillion-parameter Kimi K3 model, but a high hallucination rate might tempers its frontier-model pitch.
winbuzzer.com
July 17, 2026 at 9:42 AM
Cracking the JAX MoE training puzzle: variable expert token counts, routing tricks, and the NVIDIA Transformer Engine boost. See how DeepSeek‑V3 and GB200 handle conditional computation. Dive in for the nitty‑gritty! #JAX #MixtureOfExperts #ConditionalComputation

🔗 aidailypost.com/news/jax-moe...
September 14, 2026 at 4:54 PM