#HumanEval
"With 65.2% accuracy on HumanEval for code generation, DeepSeek V3 sets a new standard for AI performance."

at google, the job title for a software engineer who writes 35% incorrect code is "former software engineer"
March 25, 2025 at 10:46 PM
Salesforce released the open-weight of CoDA-1.7B: a text diffusion coding model that outputs tokens bidirectionally in parallel.

⚡️ Faster inference, 1.7B rivaling 7B.
📊 54.3% HumanEval | 47.6% HumanEval+ | 55.4% EvalPlus

Model: huggingface.co/Salesforce/C...
Report: github.com/SalesforceAI...
Salesforce/CoDA-v0-Instruct · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
huggingface.co
October 5, 2025 at 2:18 PM
The quality of Qwen2.5-Code-32B-Instruct remains virtually unaffected by quantization.

"HumanEval benchmark of EXL2 quants of popular local LLMs (2.5 through 8.0 bpw covered)"

Link: www.reddit.com/r/LocalLLaMA...
November 17, 2024 at 11:16 PM
Meanwhile, further independent evaluation confirms SYNTH performs better than Nanochat/Fineweb across nearly all benchmarks.

(Humaneval to be expected: we did not out any code inside yet)
Outside of HumanEval, the model pretrained on the SYNTH dataset outperforms the standard nanochat on every task. I report results for both the standard chat template (Harmony) as well as the Qwen3 Chat template used in the SYNTH dataset.
February 1, 2026 at 10:50 PM
Announcing CheatGPT, a revolutionary model that achieves SoTA on HumanEval! It's incredibly sample-efficient – just ONE training sample – and *tiny*, fitting on your Casio wristwatch!
April 1, 2025 at 7:19 PM
Verified hand-written Lean solutions to HumanEval. ~ Markus Himmel. github.com/TwoFX/human-... #ITP #LeanProver
GitHub - TwoFX/human-eval-lean: Hand-written verified Lean solutions for the HumanEval benchmark
Hand-written verified Lean solutions for the HumanEval benchmark - TwoFX/human-eval-lean
github.com
April 30, 2025 at 11:11 AM
Alibaba's Qwen2.5-Coder-32B-Instruct

The flagship model, Qwen2.5-Coder-32B-Instruct, reaches 92.7 in HumanEval, 90.2 in MBPP, 31.4 in LiveCodeBench, 73.7 in Aider, 85.1 in Spider, and 68.9 in CodeArena!
November 11, 2024 at 6:31 PM
On the the most commonly-used benchmarks is MMLU, which multiple choice about knowledge across lots of fields. Others include writing programs in Python (HumanEval) or solving reasoning puzzles involving recognizing patterns (ARC-AGI), or even taking standardized tests like the LSAT.
January 27, 2025 at 4:53 PM
Looking for evidence of data leakage in the #HumanEval code generation #LLM benchmark? Check out our analysis of a subset of HumanEval tasks vs comparable tasks on #ChatGPT, #Claude and #Llama. @riddhimore.bsky.social
March 7, 2025 at 3:22 PM
🆕 GPT-4o vs Claude 3.5 Sonnet: HumanEval Pass@1 Gap

#LLM #AI #CodeGeneration
https://tildalice.io/gpt-4o-vs-claude-3-5-sonnet-humaneval-benchmark/
June 1, 2026 at 6:04 PM
Speaking of LLM’s and svelte 5, @khromov.se has built this to help

github.com/khromov/svel...

@sveltesociety.dev
GitHub - khromov/svelte-bench: An LLM benchmark for Svelte 5 based on the HumanEval methodology
An LLM benchmark for Svelte 5 based on the HumanEval methodology - khromov/svelte-bench
github.com
May 9, 2025 at 2:46 PM
Watch and learn!

Let's observe Qwen2.5-coder:0.5b on OpenAI HumanEval.

`pip install observers`

And start collecting your data on the @huggingface.bsky.social Hub.

Dataset: huggingface.co/datasets/dav...
Library: github.com/cfahlgren1/o...
November 21, 2024 at 1:46 PM
Ask five frontier labs for their HumanEval scores and you’ll get nearly identical numbers. That’s not differentiation. It’s benchmark saturation.

See how labs are building fresh, expert-reviewed evaluations from production data: imerit.ai/resources/bl...

#LLMEvaluation #AIBenchmarks #GenAI
September 25, 2026 at 3:36 PM
Grok vs DeepSeek

Grok = live data + trends + contextual analysis
DeepSeek = coding + structured reasoning + efficiency

Different strengths. Different workflows.

Explore the full comparison: aicomparison.ai/grok-vs-deep...

#GrokVsDeepSeek #Grok #DeepSeek #AI #AIAutomation
Grok vs DeepSeek 2026: Full AI Comparison & Benchmarks
Compare Grok and DeepSeek on AIME math scores, HumanEval coding accuracy, API pricing, and real-time data access before choosing an AI model in 2026.
aicomparison.ai
September 24, 2026 at 5:58 AM
Our paper "Addressing #DataLeakage in #HumanEval Using Combinatorial Test Design" has been accepted for publication at the Int’l Conference on Software Testing, Verification & Validation (#ICST2025), Short Papers, Vision & Emerging Results.
co-author: Riddhi More
arxiv.org/abs/2412.01526
#benchmark
January 30, 2025 at 5:01 PM
💡 For code tasks like MBPP and HumanEval, how you measure changes everything.
✅ Using continuous metrics (probability of correct answers) gives much better decision accuracy.
❌ Accuracy alone? No better than random 😬
April 15, 2025 at 1:02 PM
Outside of HumanEval, the model pretrained on the SYNTH dataset outperforms the standard nanochat on every task. I report results for both the standard chat template (Harmony) as well as the Qwen3 Chat template used in the SYNTH dataset.
February 1, 2026 at 9:55 PM
3/3

Current features:

• Python coder-reviewer loops
• HumanEval integration
• Custom prompts
• Configurable model combinations
• PDF reports
June 9, 2026 at 8:24 AM
Qwen 3.8 27B spent 32,000 tokens and 14 minutes thinking about one HumanEval problem, then returned an empty response. It never reached an answer. That single problem was 7.2% of everything it generated across all 164.

#LocalAI
Why Qwen 3.8 27B Feels Slow: Reasoning Tokens Measured
Qwen 3.8 27B generates at full speed on a 3090 and still crawls. All 164 HumanEval problems measured on the chat path: 92.8% of output is thinking.
insiderllm.com
August 17, 2026 at 11:00 PM
AI benchmarks are the new vanity metrics.

"We scored 94.2% on HumanEval" tells you nothing about whether the model can debug a race condition in production at 3 AM.

Show me the benchmark for "handles ambiguity without hallucinating confidence."
March 29, 2026 at 7:11 AM
Style, Not Self: Surface Cues Explain Zero-Shot Code Attribution by Large Language Models

Ehsan Barkhordar, Surendrabikram Thapa

#arXiv #cs.AI #cs.CL #cs.SE
Style, Not Self: Surface Cues Explain Zero-Shot Code Attribution by Large Language Models
If a language model can recognize code it wrote, it may favor that code as a judge, and instances of one model monitoring each other could collude. We test this zero-shot on current commercial models. Five LLMs generate solutions to MBPP, HumanEval, and DS-1000, seven more to MBPP, and models act a…
arxiv.org
September 25, 2026 at 4:18 PM
A fun, clever idea from @upiter.bsky.social : treat code generation as a sequential editing problem -- this gives you loads of training data from synthetically editing existing code

And it works! Higher performance on HumanEval, MBPP, and CodeContests across small LMs like Gemma-2, Phi-3, Llama 3.1
Our paper showing that LMs benefit from human-like abstractions for code synthesis was accepted to ICLR! 🇸🇬

We show that order matters in code gen. -- casting code synthesis as a sequential edit problem by preprocessing examples in SFT data improves LM test-time scaling laws
February 13, 2025 at 3:42 PM
trynimbus.org/nimbus-model... #buildinginpublic #buildinpublic my little 6gb 9B contribution to the open source community. Hoping for a solid MLX run so Mac Neo users can deploy it comfortably. Smaller and smarter…smaller and smarter… huggingface.co/Nimbus-Labs/...
Nimbus Models — Small models. Serious work.
Nimbus-9B v2.1 reaches 89.0% HumanEval and 87.3% MBPP with native thinking. See all four verified EvalPlus scores.
trynimbus.org
July 21, 2026 at 10:36 AM
3.1 70B vs 3.3 70B:

Code Generation
> HumanEval: 80.5% → 88.4% (+7.9%)
> MBPP EvalPlus: 86.0% → 87.6% (+1.6%)

Steerability
> IFEval: 87.5% → 92.1% (+4.6%)

Reasoning & Math
> GPQA Diamond (CoT): 48.0% → 50.5% (+2.5%)
> MATH (CoT): 68.0% → 77.0% (+9%)
December 6, 2024 at 6:20 PM
'Rivals Claude Opus' needs context. Benchmark parity on MMLU and HumanEval doesn't tell you much about agentic reliability in production. At 744B MoE the inference costs also shift the math on 'open-source is cheaper.'
March 12, 2026 at 6:32 PM