The model weights and technical report.
Model weights: huggingface.co/moonshotai/K...
Tech report: github.com/MoonshotAI/K...
Tech blog: kimi.com/blog/kimi-k3
The model weights and technical report.
Model weights: huggingface.co/moonshotai/K...
Tech report: github.com/MoonshotAI/K...
Tech blog: kimi.com/blog/kimi-k3
They found that Muon optimizer can be scaled up using the follow techniques:
• Adding weight decay
• Carefully adjusting the per-parameter update scale
📚 Code: github.com/MoonshotAI/M...
🤗 Model: huggingface.co/moonshotai
📜 Paper: github.com/MoonshotAI/M...
They found that Muon optimizer can be scaled up using the follow techniques:
• Adding weight decay
• Carefully adjusting the per-parameter update scale
📚 Code: github.com/MoonshotAI/M...
🤗 Model: huggingface.co/moonshotai
📜 Paper: github.com/MoonshotAI/M...
- better front end coding
- context increase 128K->256K
huggingface.co/moonshotai/K...
- better front end coding
- context increase 128K->256K
huggingface.co/moonshotai/K...
Paper: github.com/MoonshotAI/K...
Repo: github.com/MoonshotAI/K...
Model: huggingface.co/moonshotai/K...
Paper: github.com/MoonshotAI/K...
Repo: github.com/MoonshotAI/K...
Model: huggingface.co/moonshotai/K...
Replaces traditional fixed, uniform attention with learned, input-dependent depth-wise attention
Much smarter attention, very little extra computational cost
github.com/MoonshotAI/A...
Replaces traditional fixed, uniform attention with learned, input-dependent depth-wise attention
Much smarter attention, very little extra computational cost
github.com/MoonshotAI/A...
A new 16B model
The Muon optimizer is 2x more data efficient than AdamE, but only for matrix parameters
note: this is a big deal
huggingface.co/moonshotai
A new 16B model
The Muon optimizer is 2x more data efficient than AdamE, but only for matrix parameters
note: this is a big deal
huggingface.co/moonshotai
huggingface.co/moonshotai/K...
✨ 1T MoE / 32B active / 256K context
✨ Agent Swarm: 300 sub-agents × 4,000 steps
✨ Modified MIT
huggingface.co/moonshotai/K...
✨ 1T MoE / 32B active / 256K context
✨ Agent Swarm: 300 sub-agents × 4,000 steps
✨ Modified MIT
The other labs keep a lot more of their methodology proprietary, but I have to assume SOTA is SOTA, and they are all doing similar things.
The other labs keep a lot more of their methodology proprietary, but I have to assume SOTA is SOTA, and they are all doing similar things.
A high-performance CUTLASS-based implementation of Kimi Delta Attention kernels. Achieves 1.72×–2.22× prefill speedup over the flash-linear-attention baseline on H20, and works as a drop-in backend for flash-linear-attention.
github.com/MoonshotAI/F...
A high-performance CUTLASS-based implementation of Kimi Delta Attention kernels. Achieves 1.72×–2.22× prefill speedup over the flash-linear-attention baseline on H20, and works as a drop-in backend for flash-linear-attention.
github.com/MoonshotAI/F...
huggingface.co/moonshotai/K...
github.com/MoonshotAI/K...
huggingface.co/moonshotai/K...
github.com/MoonshotAI/K...
Moonshot AI's Kimi Infra team dropped K2 Vendor Verifier where you can visually see the difference in tool call accuracy across providers on OpenRouter. TogetherAI looks really bad.
github.com/MoonshotAI/K...
Moonshot AI's Kimi Infra team dropped K2 Vendor Verifier where you can visually see the difference in tool call accuracy across providers on OpenRouter. TogetherAI looks really bad.
github.com/MoonshotAI/K...
Moonshot AI, of kimi.ai fame validated Muon’s scalability (github.com/MoonshotAI/M...), and now
Essential AI is stating that it expands Pareto frontier over AdamW on the compute-time tradeoff (arxiv.org/abs/2505.02222)
Moonshot AI, of kimi.ai fame validated Muon’s scalability (github.com/MoonshotAI/M...), and now
Essential AI is stating that it expands Pareto frontier over AdamW on the compute-time tradeoff (arxiv.org/abs/2505.02222)
Moonlight: 3B/16B MoE model trained with Muon on 5.7T tokens, advancing the Pareto frontier with better performance at fewer FLOPs.
huggingface.co/moonshotai
Moonlight: 3B/16B MoE model trained with Muon on 5.7T tokens, advancing the Pareto frontier with better performance at fewer FLOPs.
huggingface.co/moonshotai
diffuse.one/p/d1-007
diffuse.one/p/d1-007