https://olud.ai/leaderboard.html
#LLM #Benchmarks #OpenSource #AI
https://olud.ai/leaderboard.html
#LLM #Benchmarks #OpenSource #AI
GPQA: 54.4%
MMLU-Pro: 72.5%
Humanity's Last Exam: 3.7%
LiveCodeBench: 38.5%
Measured independently, not self-reported →https://olud.ai/leaderboard.html
#LLM #Benchmarks #OpenSource #AI
GPQA: 54.4%
MMLU-Pro: 72.5%
Humanity's Last Exam: 3.7%
LiveCodeBench: 38.5%
Measured independently, not self-reported →https://olud.ai/leaderboard.html
#LLM #Benchmarks #OpenSource #AI
https://olud.ai/leaderboard.html
#LLM #Benchmarks #OpenSource #AI
https://olud.ai/leaderboard.html
#LLM #Benchmarks #OpenSource #AI
tech_blogs_arxiv | Author: Egor Shibaev, Vera Kudrevskaia, Timur Galimzyanov, Mikhail Evtikhiev, Ana Terna, Rastislav Rabatin, Timur Kudashev, Timofey Bryksin, Arina Puchkova, Patrik Bar…
tech_blogs_arxiv | Author: Egor Shibaev, Vera Kudrevskaia, Timur Galimzyanov, Mikhail Evtikhiev, Ana Terna, Rastislav Rabatin, Timur Kudashev, Timofey Bryksin, Arina Puchkova, Patrik Bar…
But fidelity rests on a sparse eval footprint: no MMLU-Pro or LiveCodeBench, and the 2507 tag smells like a community re-quant, not...
But fidelity rests on a sparse eval footprint: no MMLU-Pro or LiveCodeBench, and the 2507 tag smells like a community re-quant, not...
GPQA: 43.2%
MMLU-Pro: 67.1%
Humanity's Last Exam: 4.5%
LiveCodeBench: 21.4%
Measured independently, not self-reported →https://olud.ai/leaderboard.html
#LLM #Benchmarks #OpenSource #AI
GPQA: 43.2%
MMLU-Pro: 67.1%
Humanity's Last Exam: 4.5%
LiveCodeBench: 21.4%
Measured independently, not self-reported →https://olud.ai/leaderboard.html
#LLM #Benchmarks #OpenSource #AI
Every curated model carries an IQ badge sourced from the Artificial Analysis Intelligence Inde, a composite of MMLU-Pro, GPQA Diamond, LiveCodeBench, AIME, and more.
Every curated model carries an IQ badge sourced from the Artificial Analysis Intelligence Inde, a composite of MMLU-Pro, GPQA Diamond, LiveCodeBench, AIME, and more.
🚨Qwen3.8 27B benchmarks are ridiculous
How is a 27B model even putting up numbers like this
• SWE bench Pro: 61.7
• QwenSWEBench: 79.0
• CoWorkBench: 70.7
• LiveCodeBench: 90.3
It’s genuinely Opus 4.6 level while being small enough...
https://x.com/i/web/status/2088283230495551590
🚨Qwen3.8 27B benchmarks are ridiculous
How is a 27B model even putting up numbers like this
• SWE bench Pro: 61.7
• QwenSWEBench: 79.0
• CoWorkBench: 70.7
• LiveCodeBench: 90.3
It’s genuinely Opus 4.6 level while being small enough...
https://x.com/i/web/status/2088283230495551590
GLM 5.3, de tamaño intermedio (700B), mirando de tú a tú a los gigantes como Fable, Sol 5.6 o Kimi K3.
Y el más esperado para los local-AI-enjoyers: Qwen 3.8 27B, compitiendo con Opus 4.6 y GPT 5.5!!! El modelo más potente de hace 6 meses, ahora en el PC de tu casa*
GLM 5.3, de tamaño intermedio (700B), mirando de tú a tú a los gigantes como Fable, Sol 5.6 o Kimi K3.
Y el más esperado para los local-AI-enjoyers: Qwen 3.8 27B, compitiendo con Opus 4.6 y GPT 5.5!!! El modelo más potente de hace 6 meses, ahora en el PC de tu casa*
But these are vendor-run benchmarks, with no independent eval yet; the edge over GLM-4.5-Air is unverified beyond Modelers.
Hard to see this displacing...
But these are vendor-run benchmarks, with no independent eval yet; the edge over GLM-4.5-Air is unverified beyond Modelers.
Hard to see this displacing...
• 78.4% SWE-bench Verified
• 66.1% LiveCodeBench v6 — 99.1% easy · 86.7% medium, full 442-problem window
• 262K native context
• 78.4% SWE-bench Verified
• 66.1% LiveCodeBench v6 — 99.1% easy · 86.7% medium, full 442-problem window
• 262K native context
Yunhao Liang et al.
#arXiv #cs.SE
GPQA: 33.1%
MMLU-Pro: 39.7%
Humanity's Last Exam: 6.6%
LiveCodeBench: 9.3%
Measured independently, not self-reported →https://opensourceai.tech/leaderboard.html
#LLM #Benchmarks #OpenSource #AI
GPQA: 33.1%
MMLU-Pro: 39.7%
Humanity's Last Exam: 6.6%
LiveCodeBench: 9.3%
Measured independently, not self-reported →https://opensourceai.tech/leaderboard.html
#LLM #Benchmarks #OpenSource #AI
Xinyuan Song, Zekun Cai, Liang Zhao
#arXiv #cs.SE #cs.AI #cs.PL
It proposes deltas and tests them by replaying only the affected parts of the run, with 95% KV-cache reuse.
Outperforms the current best meta-optimizer (MetaHarness) by 27% on LiveCodeBench, and is 2x as fast.
It proposes deltas and tests them by replaying only the affected parts of the run, with 95% KV-cache reuse.
Outperforms the current best meta-optimizer (MetaHarness) by 27% on LiveCodeBench, and is 2x as fast.
It contains 250 competition-style programming problems translated into Dafny with proper formal specifications.
It contains 250 competition-style programming problems translated into Dafny with proper formal specifications.
Contamination-resistant benchmarks (LiveBench, LiveCodeBench) update monthly. The International AI Safety Report 2026 (a hundred experts) states it plainly: capabilities are outpacing safety measures.
Contamination-resistant benchmarks (LiveBench, LiveCodeBench) update monthly. The International AI Safety Report 2026 (a hundred experts) states it plainly: capabilities are outpacing safety measures.