DeepSeek-R1 (Preview) Results. The model performs in the vicinity of o1-Medium providing SOTA reasoning performance on LiveCodeBench.
DeepSeek-R1 (Preview) Results. The model performs in the vicinity of o1-Medium providing SOTA reasoning performance on LiveCodeBench.
• 78.4% SWE-bench Verified
• 66.1% LiveCodeBench v6 — 99.1% easy · 86.7% medium, full 442-problem window
• 262K native context
• 78.4% SWE-bench Verified
• 66.1% LiveCodeBench v6 — 99.1% easy · 86.7% medium, full 442-problem window
• 262K native context
huggingface.co/collections/...
✨Pro (1.6T/49B active) & Flash (284B/13B active)
✨1M token context
✨LiveCodeBench 93.5 🤯 beats Gemini 3.1 Pro & Opus 4.6
✨MIT licensed
huggingface.co/collections/...
✨Pro (1.6T/49B active) & Flash (284B/13B active)
✨1M token context
✨LiveCodeBench 93.5 🤯 beats Gemini 3.1 Pro & Opus 4.6
✨MIT licensed
* Open weights, apache 2
* 32B
* beats o3-mini
* for TTC they train an extra head as a reward model to do binary classification
hf: huggingface.co/MetaStoneTec...
paper: arxiv.org/abs/2507.01951
* Open weights, apache 2
* 32B
* beats o3-mini
* for TTC they train an extra head as a reward model to do binary classification
hf: huggingface.co/MetaStoneTec...
paper: arxiv.org/abs/2507.01951
https://olud.ai/leaderboard.html
#LLM #Benchmarks #OpenSource #AI
https://olud.ai/leaderboard.html
#LLM #Benchmarks #OpenSource #AI
GPQA: 54.4%
MMLU-Pro: 72.5%
Humanity's Last Exam: 3.7%
LiveCodeBench: 38.5%
Measured independently, not self-reported →https://olud.ai/leaderboard.html
#LLM #Benchmarks #OpenSource #AI
GPQA: 54.4%
MMLU-Pro: 72.5%
Humanity's Last Exam: 3.7%
LiveCodeBench: 38.5%
Measured independently, not self-reported →https://olud.ai/leaderboard.html
#LLM #Benchmarks #OpenSource #AI
The flagship model, Qwen2.5-Coder-32B-Instruct, reaches 92.7 in HumanEval, 90.2 in MBPP, 31.4 in LiveCodeBench, 73.7 in Aider, 85.1 in Spider, and 68.9 in CodeArena!
The flagship model, Qwen2.5-Coder-32B-Instruct, reaches 92.7 in HumanEval, 90.2 in MBPP, 31.4 in LiveCodeBench, 73.7 in Aider, 85.1 in Spider, and 68.9 in CodeArena!
It achieves:
- AIME24: 69.0
- AIME25: 53.6
- LiveCodeBench: 44.4
It achieves:
- AIME24: 69.0
- AIME25: 53.6
- LiveCodeBench: 44.4
https://olud.ai/leaderboard.html
#LLM #Benchmarks #OpenSource #AI
https://olud.ai/leaderboard.html
#LLM #Benchmarks #OpenSource #AI
- SOTA short-CoT performance, outperforming GPT-4o and Claude Sonnet 3.5 on 📐AIME, 📐MATH-500, 💻 LiveCodeBench by a large margin (up to +550%)
- Long-CoT performance matches o1 across multiple modalities (👀MathVista, 📐AIME, 💻Codeforces, etc)
Demo: kimi.ai
- SOTA short-CoT performance, outperforming GPT-4o and Claude Sonnet 3.5 on 📐AIME, 📐MATH-500, 💻 LiveCodeBench by a large margin (up to +550%)
- Long-CoT performance matches o1 across multiple modalities (👀MathVista, 📐AIME, 💻Codeforces, etc)
Demo: kimi.ai
uuuh, guys this isn’t a boring model
This crushes all the agentic benchmarks, even beating out the already-impressive qwen3-30b-a4b
It’s hanging with some already impressive mid-sized models at only 4B
uuuh, guys this isn’t a boring model
This crushes all the agentic benchmarks, even beating out the already-impressive qwen3-30b-a4b
It’s hanging with some already impressive mid-sized models at only 4B
huggingface.co/deepseek-ai/...
Upgrades include:
✨ MATH-500: 74.8% → 82.8%
✨ LiveCodebench: 29.2% → 34.38%
✨ Writing & reasoning improved on internal tests.
✨ Enhanced file upload & webpage summarization UX
huggingface.co/deepseek-ai/...
Upgrades include:
✨ MATH-500: 74.8% → 82.8%
✨ LiveCodebench: 29.2% → 34.38%
✨ Writing & reasoning improved on internal tests.
✨ Enhanced file upload & webpage summarization UX
- AIME 2025: 50.2 → 67.4
- LiveCodeBench v5: 53.1 → 61.1
- LiveCodeBench v6: 47.9 → 54.9
- Codeforces ELO: 2024
Paper (includes both training recipe and implementation details): arxiv.org/abs/2505.16400
Model: huggingface.co/nvidia/AceRe...
- AIME 2025: 50.2 → 67.4
- LiveCodeBench v5: 53.1 → 61.1
- LiveCodeBench v6: 47.9 → 54.9
- Codeforces ELO: 2024
Paper (includes both training recipe and implementation details): arxiv.org/abs/2505.16400
Model: huggingface.co/nvidia/AceRe...
Happy to see some European action in the usable model space.
Mistral blog post: mistral.ai/news/mistral-3
Happy to see some European action in the usable model space.
Mistral blog post: mistral.ai/news/mistral-3
GPQA: 33.1%
MMLU-Pro: 39.7%
Humanity's Last Exam: 6.6%
LiveCodeBench: 9.3%
Measured independently, not self-reported →https://opensourceai.tech/leaderboard.html
#LLM #Benchmarks #OpenSource #AI
GPQA: 33.1%
MMLU-Pro: 39.7%
Humanity's Last Exam: 6.6%
LiveCodeBench: 9.3%
Measured independently, not self-reported →https://opensourceai.tech/leaderboard.html
#LLM #Benchmarks #OpenSource #AI
this has been one of my favorite local models, and now we get an even better version!
better instruction following, tool use & coding. Nice small MoE!
huggingface.co/Qwen/Qwen3-3...
this has been one of my favorite local models, and now we get an even better version!
better instruction following, tool use & coding. Nice small MoE!
huggingface.co/Qwen/Qwen3-3...
the constraint didn't limit reasoning — it stopped the model drowning in its own scratchpad. structure as liberation
https://andthattoo.dev/blog/structured_cot
the constraint didn't limit reasoning — it stopped the model drowning in its own scratchpad. structure as liberation
https://andthattoo.dev/blog/structured_cot
Deepseek V3 on livecodebench (highest non-reasoning model)
Deepseek V3 on livecodebench (highest non-reasoning model)
www.youtube.com/live/f5nPfp3...
www.youtube.com/live/f5nPfp3...