#LiveCodeBench
Try: "Qwen3 4B (Non-reasoning) posts 39.8% GPQA, 58.6% MMLU-Pro, 3.3% Humanity's Last Exam, 23.3% LiveCodeBench. These aren't vendor claims — they're measured

https://olud.ai/leaderboard.html

#LLM #Benchmarks #OpenSource #AI
September 25, 2026 at 8:00 PM
The biggest misses in terms of percentage points have been LiveCodeBench Pro Hard (actual: 53.8% in May 2026; forecast: 12%-14% by the end of 2026) and FrontierMath Tiers 1-3 unadjusted (actual: 40.7% as of late 2025; forecast: around 30%-31%). Superforecasters were more skeptical than exper
September 24, 2026 at 8:26 AM
📊 Solar Pro 2 (Preview) (Non-reasoning) — the actual numbers

GPQA: 54.4%
MMLU-Pro: 72.5%
Humanity's Last Exam: 3.7%
LiveCodeBench: 38.5%

Measured independently, not self-reported →https://olud.ai/leaderboard.html

#LLM #Benchmarks #OpenSource #AI
September 22, 2026 at 10:00 PM
📊 Llama 3.1 Tulu3 405B’s independent benchmark run: GPQA 51.6%, MMLU-Pro 71.6%, HLE 3.3%, LiveCodeBench 29.1%. The gap between reasoning and coding tells the real story — see how it stacks up on the full leaderboard.

https://olud.ai/leaderboard.html

#LLM #Benchmarks #OpenSource #AI
September 21, 2026 at 12:00 AM
Llama 4 Maverick is 13x larger than Gemma 4 31B yet scores just 43.4% on LiveCodeBench v6 against Gemma's 80%. Bigger MoE did not buy better code. With Apache 2.0 on Gemma and Qwen, Meta's license-locked default is over. Which open model do you actually reach for?
September 17, 2026 at 10:14 PM
UkisAI cut Qwen3.8-27B's median GPQA thinking tokens 58% in a five-seed BF16 same-stack test, while the score barely moved. LiveCodeBench rose, but AIME fell. Less overthinking wasn't a free win.
September 16, 2026 at 8:39 AM
FYI: GLM-5.3-Flash aka "Ox Alpha" is out

320B MoE, 18B active

Blog post -> z.ai/blog/glm-5.3...
August 26, 2026 at 2:39 PM
Don't Claim Benchmark-Oriented Optimization Improves General Coding Capability -- Diverse Evaluation Is Required
tech_blogs_arxiv | Author: Egor Shibaev, Vera Kudrevskaia, Timur Galimzyanov, Mikhail Evtikhiev, Ana Terna, Rastislav Rabatin, Timur Kudashev, Timofey Bryksin, Arina Puchkova, Patrik Bar…
Don't Claim Benchmark-Oriented Optimization Improves General Coding Capability -- Diverse Evaluation Is Required
Post-training papers, model cards, and blog posts often treat scores on a small set of coding benchmarks (e.g., SWE-bench and LiveCodeBench) as evidence of broad coding capability, both for research artifacts and user-facing systems. We argue that optimization for these benchmarks leads to measuring
arxiv.org
August 17, 2026 at 4:13 AM
Quantising a 235B MoE to W8A8 with 22B active is a genuine single-node sweet spot—the memory cut is real, and Ascend support widens the runway beyond Nvidia shops.

But fidelity rests on a sparse eval footprint: no MMLU-Pro or LiveCodeBench, and the 2507 tag smells like a community re-quant, not...
Qwen3-235B-A22B MoE Model Optimised for High-Performance Inference
aichina.news
August 16, 2026 at 1:50 PM
📊 Llama 3.2 Instruct 90B (Vision) — the actual numbers

GPQA: 43.2%
MMLU-Pro: 67.1%
Humanity's Last Exam: 4.5%
LiveCodeBench: 21.4%

Measured independently, not self-reported →https://olud.ai/leaderboard.html

#LLM #Benchmarks #OpenSource #AI
August 15, 2026 at 1:00 PM
Qwen3.8-27B-FP8: a 27B dense VL model that hits 84.3 OSWorld-Verified, 90.3 LiveCodeBench v6, 79.0 QwenSWEBench. Fits on 1-2 GPUs, 1M context via Qwen Cloud coming. Self-hostable sweet spot. https://huggingface.co/Qwen/Qwen3.8-27B-FP8
August 15, 2026 at 5:02 AM
NVIDIA-Nemotron-Labs-Teacher-STEM: 550B Parameter (55B aktiv), LatentMoE mit Mamba-2, MoE und Multi-Token Prediction, bis 1M Tokens Kontext. Laut Model Card auf GPQA, MMLU-Pro und LiveCodeBench v6 auf Niveau von DeepSeek V4 Pro. Läuft ab 4xB200/GB200 oder 8xH100, Lizenz OpenMDW-1.1.
nvidia/NVIDIA-Nemotron-Labs-Teacher-STEM
Official Hugging Face namespace: nvidia; Model ID: nvidia/NVIDIA-Nemotron-Labs-Teacher-STEM; Pipeline: text-generation; Downloads: 0; Tags: transformers, safetensors, nemotron_h, text-generation, nvidia, pytorch, nemotron-3, latent-moe
huggingface.co
August 14, 2026 at 9:41 PM
Compare models by intelligence, not just speed

Every curated model carries an IQ badge sourced from the Artificial Analysis Intelligence Inde, a composite of MMLU-Pro, GPQA Diamond, LiveCodeBench, AIME, and more.
August 14, 2026 at 8:50 PM
@LuminaBench

🚨Qwen3.8 27B benchmarks are ridiculous

How is a 27B model even putting up numbers like this

• SWE bench Pro: 61.7
• QwenSWEBench: 79.0
• CoWorkBench: 70.7
• LiveCodeBench: 90.3

It’s genuinely Opus 4.6 level while being small enough...

https://x.com/i/web/status/2088283230495551590
August 14, 2026 at 6:55 PM
Viernes de releases!

GLM 5.3, de tamaño intermedio (700B), mirando de tú a tú a los gigantes como Fable, Sol 5.6 o Kimi K3.

Y el más esperado para los local-AI-enjoyers: Qwen 3.8 27B, compitiendo con Opus 4.6 y GPT 5.5!!! El modelo más potente de hace 6 meses, ahora en el PC de tu casa*
August 14, 2026 at 6:34 PM
Zhipu’s rumination-trained 32B posts strong zero-shot GPQA and LiveCodeBench scores — enough to press Qwen and DeepSeek at this size.

But these are vendor-run benchmarks, with no independent eval yet; the edge over GLM-4.5-Air is unverified beyond Modelers.

Hard to see this displacing...
Zhipu AI Refines Reasoning with New GLM-Z1-Rumination-32B Model
aichina.news
August 8, 2026 at 9:57 AM
Bad Theory Labs' BTL-4 35B and Macaw 2.7B a frontier agentic reasoning model and an on-device Mac agent. (open-weight)

• 78.4% SWE-bench Verified
• 66.1% LiveCodeBench v6 — 99.1% easy · 86.7% medium, full 442-problem window
• 262K native context
August 7, 2026 at 4:38 PM
Do Code Language Models Use Tests? A Behavioral and Representational Study of Test-Driven Code Generation

Yunhao Liang et al.

#arXiv #cs.SE
Do Code Language Models Use Tests? A Behavioral and Representational Study of Test-Driven Code Generation
Public tests are widely used to guide large language model code generation, but whether models treat them as executable specifications or merely as extra prompt context remains unclear. We study test-driven code generation on HumanEval+, MBPP+, and recent LiveCodeBench tasks using Qwen2.5-Coder-7B …
arxiv.org
July 31, 2026 at 12:05 AM
📊 DBRX Instruct — the actual numbers

GPQA: 33.1%
MMLU-Pro: 39.7%
Humanity's Last Exam: 6.6%
LiveCodeBench: 9.3%

Measured independently, not self-reported →https://opensourceai.tech/leaderboard.html

#LLM #Benchmarks #OpenSource #AI
July 21, 2026 at 7:00 AM
AlgoBench: Benchmarking Algorithmic Adaptation in Code Generation
High pass rates on established programming benchmarks such as HumanEval and LiveCodeBench do not always show whether a model can reason about algorithms. Many fixed benchmarks eventually become part of the public training ecosystem through released problem statements, editorials, and generated solutions, allowing later models to improve partly by exposure rather than by stronger algorithmic ability. We introduce ALGOBENCH, a framework that automatically builds novel algorithmic problems from known competitive-programming problems through structured constraint-shifting transformations. Each accepted ALGOBENCH variant is traceable to a source problem, but must make the original reference algorithm fail. Beyond pass@$k$, we introduce complexity-aware metrics -- including OPTT, OPTS, TRAPRATE, GAPT, and CONSENS -- to test whether a solution is not only functionally correct but also asymptotically suitable for the generated problem. Experiments across multiple LLMs and prompting strategies show that performance drops sharply on ALGOBENCH variants, retrieval can increase reuse of the old algorithm, and many correct-looking solutions fail to meet the required complexity. Error analysis shows that failures are mainly algorithmic rather than implementation-level, suggesting that ALGOBENCH evaluates adaptation beyond functional correctness.
arxiv.org
July 2, 2026 at 4:35 AM
📈 Agent optimizer: improves your agents from their past runs.

It proposes deltas and tests them by replaying only the affected parts of the run, with 95% KV-cache reuse.

Outperforms the current best meta-optimizer (MetaHarness) by 27% on LiveCodeBench, and is 2x as fast.
July 1, 2026 at 5:04 PM
4. They also created a new benchmark called LiveCodeBench-Pro-Dafny.

It contains 250 competition-style programming problems translated into Dafny with proper formal specifications.
July 1, 2026 at 1:36 PM
What to do about it

Contamination-resistant benchmarks (LiveBench, LiveCodeBench) update monthly. The International AI Safety Report 2026 (a hundred experts) states it plainly: capabilities are outpacing safety measures.
June 30, 2026 at 1:54 PM