Под Nemotron Ultra сейчас имеется в виду NVIDIA-Nemotron-3-Ultra-550B-A55B: 550 млрд параме…
Подробнее: https://gruzdevv.ru/stati/nemotron-ultra-nvidia-besplatnyj-dostup/
#NVIDIA #NemotronUltra #OpenSourceLLM #LLM #Mamba2 #IT #SaaS
Под Nemotron Ultra сейчас имеется в виду NVIDIA-Nemotron-3-Ultra-550B-A55B: 550 млрд параме…
Подробнее: https://gruzdevv.ru/stati/nemotron-ultra-nvidia-besplatnyj-dostup/
#NVIDIA #NemotronUltra #OpenSourceLLM #LLM #Mamba2 #IT #SaaS
The distilled version seems to consistently outperform on reasoning tasks, even though it follows a Mamba2 architecture. So, the "reasoning" capabilities seem to emerge from the underlying architecture.
The distilled version seems to consistently outperform on reasoning tasks, even though it follows a Mamba2 architecture. So, the "reasoning" capabilities seem to emerge from the underlying architecture.
Models: huggingface.co/collections/...
Models: huggingface.co/collections/...
A new hybrid mamba2/attention LLM from NVIDIA that beats Qwen3-30B-A3B (same size & shape)
Notes:
* 1M context, with incredible recall past 256K
* New open datasets
* 10 open source RL environments
Overall this is a huge win for neolabs
huggingface.co/nvidia/NVIDI...
A new hybrid mamba2/attention LLM from NVIDIA that beats Qwen3-30B-A3B (same size & shape)
Notes:
* 1M context, with incredible recall past 256K
* New open datasets
* 10 open source RL environments
Overall this is a huge win for neolabs
huggingface.co/nvidia/NVIDI...
https://gigazine.net/news/20241015-zamba2-7b-released/
https://gigazine.net/news/20241015-zamba2-7b-released/
Benchmarking with vLLM against Llama 3.1 8B for long contexts shows:
🔹 2.5x throughput improvement
🔹 2x lower latency
Repo: github.com/foundation-m...
Benchmarking with vLLM against Llama 3.1 8B for long contexts shows:
🔹 2.5x throughput improvement
🔹 2x lower latency
Repo: github.com/foundation-m...
🔗 aidailypost.com/news/mamba3-...
🔗 aidailypost.com/news/mamba3-...
🔧 Supports W4A8 / W4A16 / W4AX / W8A8 for Mamba1 and Mamba2
🚀 Achieves 4× memory reduction and 3× generation speedup
⚡️ Enables 8B model inference on Orin Nano 8G at 13 tokens/sec
🔥 Outperforms W4A8KV4 Llama3-8B in both speed and quality
🔧 Supports W4A8 / W4A16 / W4AX / W8A8 for Mamba1 and Mamba2
🚀 Achieves 4× memory reduction and 3× generation speedup
⚡️ Enables 8B model inference on Orin Nano 8G at 13 tokens/sec
🔥 Outperforms W4A8KV4 Llama3-8B in both speed and quality
🔧 Supports W4A8 / W4A16 / W4AX / W8A8 for Mamba1 and Mamba2
🚀 Achieves 4× memory reduction and 3× generation speedup
⚡️ Enables 8B model inference on Orin Nano 8G at 13 tokens/sec
🔥 Outperforms W4A8KV4 Llama3-8B in both speed and quality
Gated Delta Networks: Improving Mamba2 with Delta Rule
https://arxiv.org/abs/2412.06464
Gated Delta Networks: Improving Mamba2 with Delta Rule
https://arxiv.org/abs/2412.06464
huggingface.co/blog/bamba
huggingface.co/blog/bamba
www.biorxiv.org/content/10.1...
www.biorxiv.org/content/10.1...
🔗 aidailypost.com/news/nvidia-...
🔗 aidailypost.com/news/nvidia-...
The release covers three model sizes: 1.2B, 2.7B and 7B parameters.
#AI
The release covers three model sizes: 1.2B, 2.7B and 7B parameters.
#AI
NVIDIA has released Nemotron-Nano-3-30B-A3B-NVFP4, a production checkpoint that runs a 30B parameter reasoning model in 4 bit NVFP4 format while keeping accuracy close to its…
NVIDIA has released Nemotron-Nano-3-30B-A3B-NVFP4, a production checkpoint that runs a 30B parameter reasoning model in 4 bit NVFP4 format while keeping accuracy close to its…
Technology Innovation Institute (TII), Abu Dhabi, has released Falcon-H1R-7B, a 7B parameter reasoning specialized model that matches or exceeds many 14B…
Technology Innovation Institute (TII), Abu Dhabi, has released Falcon-H1R-7B, a 7B parameter reasoning specialized model that matches or exceeds many 14B…
言語モデルのアーキテクチャの違いを理解することは困難であり、特に学術規模の事前学習(例:13億パラメータ、1000億トークン)では、結果がノイズやランダム性に支配されることが多い。この課題を克服するため、我々は中核的なモデル能力を分離・評価する制御された合成事前学習タスクを導入する。この枠組...
言語モデルのアーキテクチャの違いを理解することは困難であり、特に学術規模の事前学習(例:13億パラメータ、1000億トークン)では、結果がノイズやランダム性に支配されることが多い。この課題を克服するため、我々は中核的なモデル能力を分離・評価する制御された合成事前学習タスクを導入する。この枠組...
Can SSD-Mamba2 Unlock Reinforcement Learning for End-to-End Motion Control?
https://arxiv.org/abs/2509.07593
Can SSD-Mamba2 Unlock Reinforcement Learning for End-to-End Motion Control?
https://arxiv.org/abs/2509.07593