at google, the job title for a software engineer who writes 35% incorrect code is "former software engineer"
at google, the job title for a software engineer who writes 35% incorrect code is "former software engineer"
⚡️ Faster inference, 1.7B rivaling 7B.
📊 54.3% HumanEval | 47.6% HumanEval+ | 55.4% EvalPlus
Model: huggingface.co/Salesforce/C...
Report: github.com/SalesforceAI...
⚡️ Faster inference, 1.7B rivaling 7B.
📊 54.3% HumanEval | 47.6% HumanEval+ | 55.4% EvalPlus
Model: huggingface.co/Salesforce/C...
Report: github.com/SalesforceAI...
"HumanEval benchmark of EXL2 quants of popular local LLMs (2.5 through 8.0 bpw covered)"
Link: www.reddit.com/r/LocalLLaMA...
"HumanEval benchmark of EXL2 quants of popular local LLMs (2.5 through 8.0 bpw covered)"
Link: www.reddit.com/r/LocalLLaMA...
(Humaneval to be expected: we did not out any code inside yet)
(Humaneval to be expected: we did not out any code inside yet)
The flagship model, Qwen2.5-Coder-32B-Instruct, reaches 92.7 in HumanEval, 90.2 in MBPP, 31.4 in LiveCodeBench, 73.7 in Aider, 85.1 in Spider, and 68.9 in CodeArena!
The flagship model, Qwen2.5-Coder-32B-Instruct, reaches 92.7 in HumanEval, 90.2 in MBPP, 31.4 in LiveCodeBench, 73.7 in Aider, 85.1 in Spider, and 68.9 in CodeArena!
Grok = live data + trends + contextual analysis
DeepSeek = coding + structured reasoning + efficiency
Different strengths. Different workflows.
Explore the full comparison: aicomparison.ai/grok-vs-deep...
#GrokVsDeepSeek #Grok #DeepSeek #AI #AIAutomation
Grok = live data + trends + contextual analysis
DeepSeek = coding + structured reasoning + efficiency
Different strengths. Different workflows.
Explore the full comparison: aicomparison.ai/grok-vs-deep...
#GrokVsDeepSeek #Grok #DeepSeek #AI #AIAutomation
See how labs are building fresh, expert-reviewed evaluations from production data: imerit.ai/resources/bl...
#LLMEvaluation #AIBenchmarks #GenAI
See how labs are building fresh, expert-reviewed evaluations from production data: imerit.ai/resources/bl...
#LLMEvaluation #AIBenchmarks #GenAI
#LLM #AI #CodeGeneration
https://tildalice.io/gpt-4o-vs-claude-3-5-sonnet-humaneval-benchmark/
#LLM #AI #CodeGeneration
https://tildalice.io/gpt-4o-vs-claude-3-5-sonnet-humaneval-benchmark/
github.com/khromov/svel...
@sveltesociety.dev
github.com/khromov/svel...
@sveltesociety.dev
http://arxiv.org/abs/2609.26796v1
http://arxiv.org/abs/2609.26796v1
Let's observe Qwen2.5-coder:0.5b on OpenAI HumanEval.
`pip install observers`
And start collecting your data on the @huggingface.bsky.social Hub.
Dataset: huggingface.co/datasets/dav...
Library: github.com/cfahlgren1/o...
Let's observe Qwen2.5-coder:0.5b on OpenAI HumanEval.
`pip install observers`
And start collecting your data on the @huggingface.bsky.social Hub.
Dataset: huggingface.co/datasets/dav...
Library: github.com/cfahlgren1/o...
Low-bit cost shows up in long outputs.
Low-bit cost shows up in long outputs.
Ehsan Barkhordar, Surendrabikram Thapa
#arXiv #cs.AI #cs.CL #cs.SE
co-author: Riddhi More
arxiv.org/abs/2412.01526
#benchmark
co-author: Riddhi More
arxiv.org/abs/2412.01526
#benchmark
✅ Using continuous metrics (probability of correct answers) gives much better decision accuracy.
❌ Accuracy alone? No better than random 😬
✅ Using continuous metrics (probability of correct answers) gives much better decision accuracy.
❌ Accuracy alone? No better than random 😬
Current features:
• Python coder-reviewer loops
• HumanEval integration
• Custom prompts
• Configurable model combinations
• PDF reports
Current features:
• Python coder-reviewer loops
• HumanEval integration
• Custom prompts
• Configurable model combinations
• PDF reports
"We scored 94.2% on HumanEval" tells you nothing about whether the model can debug a race condition in production at 3 AM.
Show me the benchmark for "handles ambiguity without hallucinating confidence."
"We scored 94.2% on HumanEval" tells you nothing about whether the model can debug a race condition in production at 3 AM.
Show me the benchmark for "handles ambiguity without hallucinating confidence."
#LocalAI
#LocalAI
And it works! Higher performance on HumanEval, MBPP, and CodeContests across small LMs like Gemma-2, Phi-3, Llama 3.1
We show that order matters in code gen. -- casting code synthesis as a sequential edit problem by preprocessing examples in SFT data improves LM test-time scaling laws
And it works! Higher performance on HumanEval, MBPP, and CodeContests across small LMs like Gemma-2, Phi-3, Llama 3.1