#deepswe
This is where benchmarks are going to start making progressively less sense in general. Like, on a pure happy-happy-joy-joy code-grind basis, 95% of developers might do something that would represent the top half of DeepSWE...what, once every five years?

We have saturated HumanDevBench thoroughly
September 22, 2026 at 6:33 PM
Let’s look at deepswe to see, deepswe.datacurve.ai/data/v1/task...

Maybe? What the fuck is a happy DOM
DeepSWE
DeepSWE measures frontier coding agents on original, long-horizon software engineering tasks.
deepswe.datacurve.ai
September 19, 2026 at 7:25 PM
DeepSWE weights harder on long-horizon tasks, where Luna (5.6, at least) sometimes can struggle.
September 22, 2026 at 7:04 PM
DeepSWE. Luna/max is functionally equivalent to Opus/medium but costs ~25x less. Wild.
September 22, 2026 at 6:41 PM
It seems M3 is not a very good coding model per DeepSWE.

entrpi.github.io/misc/deep-sw...
June 2, 2026 at 1:06 AM
GLM-5.3-Flash can now be run locally! ✨

Run 3-bit on 128GB RAM via Unsloth GGUF.

GLM-5.3-Flash (ox-alpha) rivals Claude Opus 4.8 on DeepSWE, coding & agentic benchmarks.

Guide: unsloth.ai/docs/models/...
GGUF: huggingface.co/unsloth/GLM-...
August 27, 2026 at 2:47 PM
K3 is Fable level in some evaluations
July 19, 2026 at 12:37 AM
congratulations to the Moonshot team for extending Claude Fable 5's inclusion in claude subscriptions for another few weeks
July 16, 2026 at 7:45 PM
Remember how I’ve been using DeepSWE as my LLM benchmark of choice?

It’s basically saturated now

In particular DeepSeek v4 Pro, released two days ago, gets scores near the top for on the order of 1% of the price of the US companies
August 15, 2026 at 1:58 AM
The framing on this story is that GPT > Claude but my actual takeaway is that benchmarks are mostly made up and the points don’t matter

venturebeat.com/technology/d...
DeepSWE blows up the AI coding leaderboard, crowns GPT-5.5, and finds Claude Opus exploiting a benchmark loophole
DeepSWE puts GPT-5.5 atop the AI coding leaderboard while raising new questions about Claude Opus, SWE-Bench Pro, and benchmark leakage.
venturebeat.com
May 27, 2026 at 7:24 PM
New SWE benchmark, called DeepSWE.

I'm curious why Cursor is not included. Anyway, I stand by my statement that Claude Code (Claude Opus 4.7) is better at first 80% and Codex (GPT 5.5) is better at remaining 1,000%, at this moment.

deepswe.datacurve.ai
May 27, 2026 at 4:23 PM
Datacurve's DeepSWE Bench is so outdated.

No Grok 4.7
No 6 Sol
No 6 Luna
No Opus 5.5
No Fable 5.1

WTF
September 24, 2026 at 8:53 PM
does anyone have $400 i wanna run deepswe on k2.7...
June 15, 2026 at 3:18 PM
new stealth model on openrouter, "union alpha," tokenizer seems to be GLM5-ish, maybe an american finetune?
September 16, 2026 at 6:02 PM
📰 Together AI
GLM-5.3 vs. GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing

https://www.together.ai/blog/glm-5-3-vs-gpt-5-6-sol-on-deepswe-cost-coding-and-rout
in#IA##AI##ML#ML
September 26, 2026 at 7:00 PM
📰 Together AI
GLM-5.3 vs. Claude Fable 5 on DeepSWE: Cost, Coding, and Routing

https://www.together.ai/blog/glm-5-3-vs-claude-fable-5-on-deepswe-cost-coding-and-rout
in#IA##AI##ML#ML
September 26, 2026 at 7:00 PM
Luna xhigh outperforms opus on DeepSWE at basically every effort level for 20% the price lmao
July 14, 2026 at 7:23 PM
✨ MiMo- V2.6 Pro RL: Maximum capability
- 1.02T / 42B
- Strong coding + agent performance: 71.9 DeepSWE / 53.1 AutomationBench / 89.9 Terminal Bench 2.1

✨ MiMo-V2.6 Flash RL: Maximum efficiency
- 309B / 15B
- 15B active params while staying close to Pro on many agent benchmarks
September 21, 2026 at 9:12 PM
📰 Together AI
DeepSeek V4 Pro 0813 vs Claude Fable 5 on DeepSWE: Cost, Coding, and Routing

https://www.together.ai/blog/deepseek-v4-pro-0813-vs-claude-fable-5-on-deepswe-cost-coding-and-rout
in#IA##AI##ML#ML
September 26, 2026 at 11:00 PM
A 118B model scoring 40% on DeepSWE (compared to the 700+B GLM 5.2 that scores 44%) is super impressive.

Will want to download and check out.
Today we're releasing Laguna S 2.1, our most capable model to date.

It's a 118B total parameter MoE model with 8B activated per token, a context window of up to 1M tokens, and thinking and no-thinking modes.

Small enough to run on a single @NVIDIAAI DGX Spark

poolside.ai/blog/introdu...
Introducing Laguna S 2.1
Today we’re releasing Laguna S 2.1, a significant step forward in our development of models that pursue longer horizon work and make effective use of reasoning.
poolside.ai
July 21, 2026 at 9:53 PM
@UnslothAI

GLM-5.3-Flash can now be run locally! ✨

Run 3-bit on 128GB RAM via Unsloth GGUF.

GLM-5.3-Flash (ox-alpha) rivals Claude Opus 4.8 on DeepSWE, coding & agentic benchmarks.

Guide:
GGUF:

https://x.com/i/web/status/2092986464196002094
September 23, 2026 at 4:51 PM
DeepSWE v1.1: Updated execution and grading for the same long-horizon engineering tasks. deepswe.datacurve.ai/blog/deepswe... 自分はDeepSWEは今のところまともなベンチだと思っていて,そこでこういう結果なら,今まで通りgpt-5.5 [high]をデフォルト使いするな.
June 22, 2026 at 3:19 AM
بمب جدید متا توی دنیای هوش مصنوعی...

مدل Muse Spark 1.3 معرفی شد. این مدل توی تست DeepSWE 1.1 تونسته غول‌هایی مثل GPT-5.6 و Opus 5 رو پشت سر بذاره.
September 3, 2026 at 12:00 AM
May 31, 2026 at 12:35 PM