#Mamba2
Another great hnet deep dive by @main_horse (on the other network). If you ever care about making tokenizerless happen and gigantic mamba2 layer. main-horse.github.io/hnet/eng-1gpu/
September 3, 2025 at 12:39 PM
Бесплатная модель Nemotron Ultra от NVIDIA: где и как применить

Под Nemotron Ultra сейчас имеется в виду NVIDIA-Nemotron-3-Ultra-550B-A55B: 550 млрд параме…

Подробнее: https://gruzdevv.ru/stati/nemotron-ultra-nvidia-besplatnyj-dostup/

#NVIDIA #NemotronUltra #OpenSourceLLM #LLM #Mamba2 #IT #SaaS
September 26, 2026 at 6:00 PM
Cartesia.ai distills LLaMa into a state space model with interesting results

The distilled version seems to consistently outperform on reasoning tasks, even though it follows a Mamba2 architecture. So, the "reasoning" capabilities seem to emerge from the underlying architecture.
February 22, 2025 at 3:36 AM
wait you have BOTH gated deltanet AND mamba2? dang… that’s sick. no kidding it’s a hybrid
March 5, 2026 at 9:32 PM
Blog: Bamba: Inference-Efficient Hybrid Mamba2 Model ( huggingface.co/blog/bamba )
Models: huggingface.co/collections/...
Bamba: Inference-Efficient Hybrid Mamba2 Model
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
huggingface.co
December 18, 2024 at 10:15 PM
Nemotron 3

A new hybrid mamba2/attention LLM from NVIDIA that beats Qwen3-30B-A3B (same size & shape)

Notes:
* 1M context, with incredible recall past 256K
* New open datasets
* 10 open source RL environments

Overall this is a huge win for neolabs

huggingface.co/nvidia/NVIDI...
December 16, 2025 at 1:15 PM
Bamba-9B "hybrid Mamba2" model delivers significant improvements in throughput and latency, enhancing real-time application performance.

Benchmarking with vLLM against Llama 3.1 8B for long contexts shows:
🔹 2.5x throughput improvement
🔹 2x lower latency

Repo: github.com/foundation-m...
GitHub - foundation-model-stack/bamba: Train, tune, and infer Bamba model
Train, tune, and infer Bamba model. Contribute to foundation-model-stack/bamba development by creating an account on GitHub.
github.com
December 18, 2024 at 10:00 PM
New Mamba‑3 slashes state size in half while keeping Mamba‑2 perplexity, adds ~4% LM gain and cuts latency. Curious how the architecture pulls this off? Dive into the details! #Mamba3 #Mamba2 #InferenceLatency

🔗 aidailypost.com/news/mamba3-...
March 17, 2026 at 11:45 PM
a dedicated accelerator on FPGA with hardware-algorithm co-design to promote the deployment efficiency of Mamba2. Specifically, we successfully achieve 8-bit quantization for linear layers through Hadamard transformation to eliminate outliers. [3/7 of https://arxiv.org/abs/2505.18975v1]
May 27, 2025 at 5:54 AM
Check out our latest work on post-training quantization for Mamba2 models! #LLM #SSM #Quantization
We’re excited to pre-release our latest work: Quamba2
🔧 Supports W4A8 / W4A16 / W4AX / W8A8 for Mamba1 and Mamba2
🚀 Achieves 4× memory reduction and 3× generation speedup
⚡️ Enables 8B model inference on Orin Nano 8G at 13 tokens/sec
🔥 Outperforms W4A8KV4 Llama3-8B in both speed and quality
April 5, 2025 at 11:59 PM
We’re excited to pre-release our latest work: Quamba2
🔧 Supports W4A8 / W4A16 / W4AX / W8A8 for Mamba1 and Mamba2
🚀 Achieves 4× memory reduction and 3× generation speedup
⚡️ Enables 8B model inference on Orin Nano 8G at 13 tokens/sec
🔥 Outperforms W4A8KV4 Llama3-8B in both speed and quality
April 5, 2025 at 5:27 PM
Songlin Yang, Jan Kautz, Ali Hatamizadeh
Gated Delta Networks: Improving Mamba2 with Delta Rule
https://arxiv.org/abs/2412.06464
December 10, 2024 at 6:52 AM
Not only transformers! Meta has released the second version of Mamba (well, almost) with super-efficient inference. The quality still falls short of transformers, but it's still very good.
huggingface.co/blog/bamba
Bamba: Inference-Efficient Hybrid Mamba2 Model
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
huggingface.co
December 18, 2024 at 9:32 PM
PlantCAD2: A Long-Context DNA Language Model for Cross-Species Functional Annotation in Angiosperms
www.biorxiv.org/content/10.1...
March 10, 2026 at 3:47 PM
Meet NVIDIA's Nemotron 3.5 Lightning – a lean, open‑weight reasoning model built for AI agents and agentic workflows. Think faster coding assistants and smarter instruction systems. Dive into the specs! #Nemotron3.5Lightning #AIagent #Mamba2

🔗 aidailypost.com/news/nvidia-...
August 14, 2026 at 11:25 AM
Zyphra released Zamba2-VL, an open-source vision-language model family that uses a Mamba2-transformer hybrid architecture to target lower-latency multimodal inference for documents, OCR, counting and edge AI tasks.

The release covers three model sizes: 1.2B, 2.7B and 7B parameters.

#AI
Zyphra’s Zamba2-VL Tests Hybrid AI For Faster Vision-Language Models
Zyphra released Zamba2-VL, an open-source vision-language model family that uses a Mamba2-transformer hybrid architecture to target lower-latency multimodal inference for documents, OCR, counting and edge AI tasks.
stechtimes.com
September 9, 2026 at 3:00 PM
NVIDIA Nemotron 3.5 Lightning 30B-A3B: Hybrid MoE (Mamba2/Transformer) mit 3B aktiven Parametern. NVFP4-Training und Multi-Token Prediction (MTP) ermöglichen native spekulative Dekodierung. 1M Kontextfenster, OpenMDW-Lizenz.
nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-Base-BF16
Official Hugging Face namespace: nvidia; Model ID: nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-Base-BF16; Pipeline: text-generation; Downloads: 336; Tags: transformers, safetensors, nemotron_h, text-generation, nvidia, pytorch, nemotron-3…
huggingface.co
August 11, 2026 at 12:40 PM
Nemotron-Labs-3-Puzzle-75B-A9B-FP8 nutzt LatentMoE, Mamba2 und Multi-Token Prediction. Auf 8×B200 steigt der Durchsatz um 2×, die H100-Konkurrenz bei 1M-Token-Kontext auf 8 Requests. Lizenz: OpenMDW-1.1. AIME25: 89.4, MMLU-Pro: 82.0.
nvidia/NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-FP8
Official Hugging Face namespace: nvidia; Model ID: nvidia/NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-FP8; Pipeline: text-generation; Downloads: 43; Tags: transformers, safetensors, nemotron_h_puzzle, text-generation, nvidia, pytorch, nemotron-3…
huggingface.co
July 7, 2026 at 10:45 AM
NVIDIA AI Brings Nemotron-3-Nano-30B to NVFP4 with Quantization Aware Distillation (QAD) for Efficient Reasoning Inference

NVIDIA has released Nemotron-Nano-3-30B-A3B-NVFP4, a production checkpoint that runs a 30B parameter reasoning model in 4 bit NVFP4 format while keeping accuracy close to its…
NVIDIA AI Brings Nemotron-3-Nano-30B to NVFP4 with Quantization Aware Distillation (QAD) for Efficient Reasoning Inference
NVIDIA has released Nemotron-Nano-3-30B-A3B-NVFP4, a production checkpoint that runs a 30B parameter reasoning model in 4 bit NVFP4 format while keeping accuracy close to its BF16 baseline. The model combines a hybrid Mamba2 Transformer Mixture of Experts architecture with a Quantization Aware Distillation (QAD) recipe designed specifically for NVFP4 deployment. Overall, it is an ultra-efficient NVFP4 precision version of Nemotron-3-Nano that delivers up to 4x higher throughput on Blackwell B200.
nexttech-news.com
February 2, 2026 at 7:47 AM
TII Abu-Dhabi Released Falcon H1R-7B: A New Reasoning Model Outperforming Others in Math and Coding with only 7B Params with 256k Context Window

Technology Innovation Institute (TII), Abu Dhabi, has released Falcon-H1R-7B, a 7B parameter reasoning specialized model that matches or exceeds many 14B…
TII Abu-Dhabi Released Falcon H1R-7B: A New Reasoning Model Outperforming Others in Math and Coding with only 7B Params with 256k Context Window
Technology Innovation Institute (TII), Abu Dhabi, has released Falcon-H1R-7B, a 7B parameter reasoning specialized model that matches or exceeds many 14B to 47B reasoning models in math, code and general benchmarks, while staying compact and efficient. It builds on Falcon H1 7B Base and is available on Hugging Face under the Falcon-H1R collection. Falcon-H1R-7B is interesting because it combines 3 design choices in 1 system, a hybrid Transformer along with Mamba2 backbone, a very long context that reaches 256k tokens in standard vLLM deployments, and a training recipe that mixes supervised long form reasoning with reinforcement learning using GRPO.
nexttech-news.com
January 7, 2026 at 12:46 PM
2512.17351
言語モデルのアーキテクチャの違いを理解することは困難であり、特に学術規模の事前学習(例:13億パラメータ、1000億トークン)では、結果がノイズやランダム性に支配されることが多い。この課題を克服するため、我々は中核的なモデル能力を分離・評価する制御された合成事前学習タスクを導入する。この枠組...
January 8, 2026 at 12:11 AM
Gavin Tao, Yinuo Wang, Jinzhao Zhou
Can SSD-Mamba2 Unlock Reinforcement Learning for End-to-End Motion Control?
https://arxiv.org/abs/2509.07593
September 10, 2025 at 5:55 AM