#Humaneval
"With 65.2% accuracy on HumanEval for code generation, DeepSeek V3 sets a new standard for AI performance."

at google, the job title for a software engineer who writes 35% incorrect code is "former software engineer"
March 25, 2025 at 10:46 PM
Salesforce released the open-weight of CoDA-1.7B: a text diffusion coding model that outputs tokens bidirectionally in parallel.

⚡️ Faster inference, 1.7B rivaling 7B.
📊 54.3% HumanEval | 47.6% HumanEval+ | 55.4% EvalPlus

Model: huggingface.co/Salesforce/C...
Report: github.com/SalesforceAI...
Salesforce/CoDA-v0-Instruct · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
huggingface.co
October 5, 2025 at 2:18 PM
The quality of Qwen2.5-Code-32B-Instruct remains virtually unaffected by quantization.

"HumanEval benchmark of EXL2 quants of popular local LLMs (2.5 through 8.0 bpw covered)"

Link: www.reddit.com/r/LocalLLaMA...
November 17, 2024 at 11:16 PM
Meanwhile, further independent evaluation confirms SYNTH performs better than Nanochat/Fineweb across nearly all benchmarks.

(Humaneval to be expected: we did not out any code inside yet)
Outside of HumanEval, the model pretrained on the SYNTH dataset outperforms the standard nanochat on every task. I report results for both the standard chat template (Harmony) as well as the Qwen3 Chat template used in the SYNTH dataset.
February 1, 2026 at 10:50 PM
Announcing CheatGPT, a revolutionary model that achieves SoTA on HumanEval! It's incredibly sample-efficient – just ONE training sample – and *tiny*, fitting on your Casio wristwatch!
April 1, 2025 at 7:19 PM
Alibaba's Qwen2.5-Coder-32B-Instruct

The flagship model, Qwen2.5-Coder-32B-Instruct, reaches 92.7 in HumanEval, 90.2 in MBPP, 31.4 in LiveCodeBench, 73.7 in Aider, 85.1 in Spider, and 68.9 in CodeArena!
November 11, 2024 at 6:31 PM
Verified hand-written Lean solutions to HumanEval. ~ Markus Himmel. github.com/TwoFX/human-... #ITP #LeanProver
GitHub - TwoFX/human-eval-lean: Hand-written verified Lean solutions for the HumanEval benchmark
Hand-written verified Lean solutions for the HumanEval benchmark - TwoFX/human-eval-lean
github.com
April 30, 2025 at 11:11 AM
On the the most commonly-used benchmarks is MMLU, which multiple choice about knowledge across lots of fields. Others include writing programs in Python (HumanEval) or solving reasoning puzzles involving recognizing patterns (ARC-AGI), or even taking standardized tests like the LSAT.
January 27, 2025 at 4:53 PM
Looking for evidence of data leakage in the #HumanEval code generation #LLM benchmark? Check out our analysis of a subset of HumanEval tasks vs comparable tasks on #ChatGPT, #Claude and #Llama. @riddhimore.bsky.social
March 7, 2025 at 3:22 PM
Grok vs DeepSeek

Grok = live data + trends + contextual analysis
DeepSeek = coding + structured reasoning + efficiency

Different strengths. Different workflows.

Explore the full comparison: aicomparison.ai/grok-vs-deep...

#GrokVsDeepSeek #Grok #DeepSeek #AI #AIAutomation
Grok vs DeepSeek 2026: Full AI Comparison & Benchmarks
Compare Grok and DeepSeek on AIME math scores, HumanEval coding accuracy, API pricing, and real-time data access before choosing an AI model in 2026.
aicomparison.ai
September 24, 2026 at 5:58 AM
Ask five frontier labs for their HumanEval scores and you’ll get nearly identical numbers. That’s not differentiation. It’s benchmark saturation.

See how labs are building fresh, expert-reviewed evaluations from production data: imerit.ai/resources/bl...

#LLMEvaluation #AIBenchmarks #GenAI
September 25, 2026 at 3:36 PM
🆕 GPT-4o vs Claude 3.5 Sonnet: HumanEval Pass@1 Gap

#LLM #AI #CodeGeneration
https://tildalice.io/gpt-4o-vs-claude-3-5-sonnet-humaneval-benchmark/
June 1, 2026 at 6:04 PM
Speaking of LLM’s and svelte 5, @khromov.se has built this to help

github.com/khromov/svel...

@sveltesociety.dev
GitHub - khromov/svelte-bench: An LLM benchmark for Svelte 5 based on the HumanEval methodology
An LLM benchmark for Svelte 5 based on the HumanEval methodology - khromov/svelte-bench
github.com
May 9, 2025 at 2:46 PM
Flash-dLLM speeds up diffusion LLM inference by tackling a hidden bottleneck: GPU memory I/O. With fused KV-cache kernels and self draft-and-verify decoding, it reaches 11× speedup on HumanEval over prior methods.

http://arxiv.org/abs/2609.26796v1
September 23, 2026 at 6:15 PM
Watch and learn!

Let's observe Qwen2.5-coder:0.5b on OpenAI HumanEval.

`pip install observers`

And start collecting your data on the @huggingface.bsky.social Hub.

Dataset: huggingface.co/datasets/dav...
Library: github.com/cfahlgren1/o...
November 21, 2024 at 1:46 PM
Sub-3-bit quantization survives short answers but collapses on long reasoning. Distilling on the model's own quantized decoding path lifts BF16 retention from 35% to 70% on MATH-500 and 66% to 91% on HumanEval.
Low-bit cost shows up in long outputs.
September 23, 2026 at 4:30 PM
Style, Not Self: Surface Cues Explain Zero-Shot Code Attribution by Large Language Models

Ehsan Barkhordar, Surendrabikram Thapa

#arXiv #cs.AI #cs.CL #cs.SE
Style, Not Self: Surface Cues Explain Zero-Shot Code Attribution by Large Language Models
If a language model can recognize code it wrote, it may favor that code as a judge, and instances of one model monitoring each other could collude. We test this zero-shot on current commercial models. Five LLMs generate solutions to MBPP, HumanEval, and DS-1000, seven more to MBPP, and models act a…
arxiv.org
September 25, 2026 at 4:18 PM
Our paper "Addressing #DataLeakage in #HumanEval Using Combinatorial Test Design" has been accepted for publication at the Int’l Conference on Software Testing, Verification & Validation (#ICST2025), Short Papers, Vision & Emerging Results.
co-author: Riddhi More
arxiv.org/abs/2412.01526
#benchmark
January 30, 2025 at 5:01 PM
💡 For code tasks like MBPP and HumanEval, how you measure changes everything.
✅ Using continuous metrics (probability of correct answers) gives much better decision accuracy.
❌ Accuracy alone? No better than random 😬
April 15, 2025 at 1:02 PM
Outside of HumanEval, the model pretrained on the SYNTH dataset outperforms the standard nanochat on every task. I report results for both the standard chat template (Harmony) as well as the Qwen3 Chat template used in the SYNTH dataset.
February 1, 2026 at 9:55 PM
3/3

Current features:

• Python coder-reviewer loops
• HumanEval integration
• Custom prompts
• Configurable model combinations
• PDF reports
June 9, 2026 at 8:24 AM
AI benchmarks are the new vanity metrics.

"We scored 94.2% on HumanEval" tells you nothing about whether the model can debug a race condition in production at 3 AM.

Show me the benchmark for "handles ambiguity without hallucinating confidence."
March 29, 2026 at 7:11 AM
Qwen 3.8 27B spent 32,000 tokens and 14 minutes thinking about one HumanEval problem, then returned an empty response. It never reached an answer. That single problem was 7.2% of everything it generated across all 164.

#LocalAI
Why Qwen 3.8 27B Feels Slow: Reasoning Tokens Measured
Qwen 3.8 27B generates at full speed on a 3090 and still crawls. All 164 HumanEval problems measured on the chat path: 92.8% of output is thinking.
insiderllm.com
August 17, 2026 at 11:00 PM
A fun, clever idea from @upiter.bsky.social : treat code generation as a sequential editing problem -- this gives you loads of training data from synthetically editing existing code

And it works! Higher performance on HumanEval, MBPP, and CodeContests across small LMs like Gemma-2, Phi-3, Llama 3.1
Our paper showing that LMs benefit from human-like abstractions for code synthesis was accepted to ICLR! 🇸🇬

We show that order matters in code gen. -- casting code synthesis as a sequential edit problem by preprocessing examples in SFT data improves LM test-time scaling laws
February 13, 2025 at 3:42 PM
trynimbus.org/nimbus-model... #buildinginpublic #buildinpublic my little 6gb 9B contribution to the open source community. Hoping for a solid MLX run so Mac Neo users can deploy it comfortably. Smaller and smarter…smaller and smarter… huggingface.co/Nimbus-Labs/...
Nimbus Models — Small models. Serious work.
Nimbus-9B v2.1 reaches 89.0% HumanEval and 87.3% MBPP with native thinking. See all four verified EvalPlus scores.
trynimbus.org
July 21, 2026 at 10:36 AM