#LiveCodeBench
DeepSeek-R1 is coming soon.

DeepSeek-R1 (Preview) Results. The model performs in the vicinity of o1-Medium providing SOTA reasoning performance on LiveCodeBench.
January 17, 2025 at 7:31 PM
I want LiveCodeBench but showing how many tokens/€s/KWhs it takes to get to 100%. Not interested in scores less than 100%.
April 4, 2026 at 4:15 PM
Bad Theory Labs' BTL-4 35B and Macaw 2.7B a frontier agentic reasoning model and an on-device Mac agent. (open-weight)

• 78.4% SWE-bench Verified
• 66.1% LiveCodeBench v6 — 99.1% easy · 86.7% medium, full 442-problem window
• 262K native context
August 7, 2026 at 4:38 PM
DeepSeek V4 🐳🔥 makes 1M token context & SOTA Agentic/coding more accessible!!

huggingface.co/collections/...

✨Pro (1.6T/49B active) & Flash (284B/13B active)
✨1M token context
✨LiveCodeBench 93.5 🤯 beats Gemini 3.1 Pro & Opus 4.6
✨MIT licensed
April 24, 2026 at 6:48 AM
XBai-o4: a new supermodel

* Open weights, apache 2
* 32B
* beats o3-mini
* for TTC they train an extra head as a reward model to do binary classification

hf: huggingface.co/MetaStoneTec...
paper: arxiv.org/abs/2507.01951
August 2, 2025 at 9:53 PM
📊 Llama 3.1 Tulu3 405B’s independent benchmark run: GPQA 51.6%, MMLU-Pro 71.6%, HLE 3.3%, LiveCodeBench 29.1%. The gap between reasoning and coding tells the real story — see how it stacks up on the full leaderboard.

https://olud.ai/leaderboard.html

#LLM #Benchmarks #OpenSource #AI
September 21, 2026 at 12:00 AM
📊 Solar Pro 2 (Preview) (Non-reasoning) — the actual numbers

GPQA: 54.4%
MMLU-Pro: 72.5%
Humanity's Last Exam: 3.7%
LiveCodeBench: 38.5%

Measured independently, not self-reported →https://olud.ai/leaderboard.html

#LLM #Benchmarks #OpenSource #AI
September 22, 2026 at 10:00 PM
Alibaba's Qwen2.5-Coder-32B-Instruct

The flagship model, Qwen2.5-Coder-32B-Instruct, reaches 92.7 in HumanEval, 90.2 in MBPP, 31.4 in LiveCodeBench, 73.7 in Aider, 85.1 in Spider, and 68.9 in CodeArena!
November 11, 2024 at 6:31 PM
Qwen2.5-Max, a large MoE LLM pretrained on massive data and post-trained with curated SFT and RLHF recipes. It achieves competitive performance against the top-tier models, and outcompetes DeepSeek V3 in benchmarks like Arena Hard, LiveBench, LiveCodeBench, GPQA-Diamond.
January 28, 2025 at 4:03 PM
Nvidia's AceMath-RL-Nemotron-7B, an open math model trained with reinforcement learning from the SFT-only checkpoint: Deepseek-R1-Distilled-Qwen-7B.

It achieves:
- AIME24: 69.0
- AIME25: 53.6
- LiveCodeBench: 44.4
April 25, 2025 at 1:38 AM
Try: "Qwen3 4B (Non-reasoning) posts 39.8% GPQA, 58.6% MMLU-Pro, 3.3% Humanity's Last Exam, 23.3% LiveCodeBench. These aren't vendor claims — they're measured

https://olud.ai/leaderboard.html

#LLM #Benchmarks #OpenSource #AI
September 25, 2026 at 8:00 PM
Kimi k1.5 --- an o1-level multi-modal model

- SOTA short-CoT performance, outperforming GPT-4o and Claude Sonnet 3.5 on 📐AIME, 📐MATH-500, 💻 LiveCodeBench by a large margin (up to +550%)
- Long-CoT performance matches o1 across multiple modalities (👀MathVista, 📐AIME, 💻Codeforces, etc)

Demo: kimi.ai
January 20, 2025 at 4:49 PM
Qwen3-4B Instruct & Thinking

uuuh, guys this isn’t a boring model

This crushes all the agentic benchmarks, even beating out the already-impressive qwen3-30b-a4b

It’s hanging with some already impressive mid-sized models at only 4B
August 6, 2025 at 5:24 PM
DeepSeek-V2.5-1210 🔥 the updated version of DeepSeek-V2.5 just released!
huggingface.co/deepseek-ai/...
Upgrades include:
✨ MATH-500: 74.8% → 82.8%
✨ LiveCodebench: 29.2% → 34.38%
✨ Writing & reasoning improved on internal tests.
✨ Enhanced file upload & webpage summarization UX
deepseek-ai/DeepSeek-V2.5-1210 · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
huggingface.co
December 10, 2024 at 10:02 AM
- AIME 2024: 69.7 → 78.6
- AIME 2025: 50.2 → 67.4
- LiveCodeBench v5: 53.1 → 61.1
- LiveCodeBench v6: 47.9 → 54.9
- Codeforces ELO: 2024

Paper (includes both training recipe and implementation details): arxiv.org/abs/2505.16400
Model: huggingface.co/nvidia/AceRe...
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Despite recent progress in large-scale reinforcement learning (RL) for reasoning, the training recipe for building high-performing reasoning models remains elusive. Key implementation details of front...
arxiv.org
May 23, 2025 at 7:01 AM
Qwen3-Max: Just Scale It

it’s now safe to say Qwen is a frontier lab

qwen.ai/blog?id=2413...
September 23, 2025 at 10:42 PM
Mistral dropped ministral 3B, 8B, 14B models and the big one - a seemingly deepseek shaped Mistral large 3, 675B moe brick. All apache 2!

Happy to see some European action in the usable model space.

Mistral blog post: mistral.ai/news/mistral-3
December 2, 2025 at 7:19 PM
📊 DBRX Instruct — the actual numbers

GPQA: 33.1%
MMLU-Pro: 39.7%
Humanity's Last Exam: 6.6%
LiveCodeBench: 9.3%

Measured independently, not self-reported →https://opensourceai.tech/leaderboard.html

#LLM #Benchmarks #OpenSource #AI
July 21, 2026 at 7:00 AM
The biggest misses in terms of percentage points have been LiveCodeBench Pro Hard (actual: 53.8% in May 2026; forecast: 12%-14% by the end of 2026) and FrontierMath Tiers 1-3 unadjusted (actual: 40.7% as of late 2025; forecast: around 30%-31%). Superforecasters were more skeptical than exper
September 24, 2026 at 8:26 AM
yesssss! a small update to Qwen3-30B-A3B

this has been one of my favorite local models, and now we get an even better version!

better instruction following, tool use & coding. Nice small MoE!

huggingface.co/Qwen/Qwen3-3...
July 29, 2025 at 4:53 PM
constrain a reasoning model's think block to a 5-line grammar. thinking tokens drop 43x. performance goes UP 14pp on LiveCodeBench.

the constraint didn't limit reasoning — it stopped the model drowning in its own scratchpad. structure as liberation
https://andthattoo.dev/blog/structured_cot
April 26, 2026 at 3:11 PM
Microsoft releases Phi4-reasoning, a 14B distilled from o3-mini

huggingface.co/microsoft/Ph...
May 3, 2025 at 11:45 AM
@rohanpaul_ai https://x.com/rohanpaul_ai/status/1872636577291366645 #x-rohanpaul_ai

Deepseek V3 on livecodebench (highest non-reasoning model)
December 27, 2024 at 2:00 PM