We have saturated HumanDevBench thoroughly
We have saturated HumanDevBench thoroughly
Maybe? What the fuck is a happy DOM
Maybe? What the fuck is a happy DOM
Run 3-bit on 128GB RAM via Unsloth GGUF.
GLM-5.3-Flash (ox-alpha) rivals Claude Opus 4.8 on DeepSWE, coding & agentic benchmarks.
Guide: unsloth.ai/docs/models/...
GGUF: huggingface.co/unsloth/GLM-...
Run 3-bit on 128GB RAM via Unsloth GGUF.
GLM-5.3-Flash (ox-alpha) rivals Claude Opus 4.8 on DeepSWE, coding & agentic benchmarks.
Guide: unsloth.ai/docs/models/...
GGUF: huggingface.co/unsloth/GLM-...
It’s basically saturated now
In particular DeepSeek v4 Pro, released two days ago, gets scores near the top for on the order of 1% of the price of the US companies
It’s basically saturated now
In particular DeepSeek v4 Pro, released two days ago, gets scores near the top for on the order of 1% of the price of the US companies
venturebeat.com/technology/d...
venturebeat.com/technology/d...
I'm curious why Cursor is not included. Anyway, I stand by my statement that Claude Code (Claude Opus 4.7) is better at first 80% and Codex (GPT 5.5) is better at remaining 1,000%, at this moment.
deepswe.datacurve.ai
I'm curious why Cursor is not included. Anyway, I stand by my statement that Claude Code (Claude Opus 4.7) is better at first 80% and Codex (GPT 5.5) is better at remaining 1,000%, at this moment.
deepswe.datacurve.ai
No Grok 4.7
No 6 Sol
No 6 Luna
No Opus 5.5
No Fable 5.1
WTF
No Grok 4.7
No 6 Sol
No 6 Luna
No Opus 5.5
No Fable 5.1
WTF
GLM-5.3 vs. GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing
https://www.together.ai/blog/glm-5-3-vs-gpt-5-6-sol-on-deepswe-cost-coding-and-routin#IA##AI##ML#ML
GLM-5.3 vs. GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing
https://www.together.ai/blog/glm-5-3-vs-gpt-5-6-sol-on-deepswe-cost-coding-and-routin#IA##AI##ML#ML
GLM-5.3 vs. Claude Fable 5 on DeepSWE: Cost, Coding, and Routing
https://www.together.ai/blog/glm-5-3-vs-claude-fable-5-on-deepswe-cost-coding-and-routin#IA##AI##ML#ML
GLM-5.3 vs. Claude Fable 5 on DeepSWE: Cost, Coding, and Routing
https://www.together.ai/blog/glm-5-3-vs-claude-fable-5-on-deepswe-cost-coding-and-routin#IA##AI##ML#ML
- 1.02T / 42B
- Strong coding + agent performance: 71.9 DeepSWE / 53.1 AutomationBench / 89.9 Terminal Bench 2.1
✨ MiMo-V2.6 Flash RL: Maximum efficiency
- 309B / 15B
- 15B active params while staying close to Pro on many agent benchmarks
- 1.02T / 42B
- Strong coding + agent performance: 71.9 DeepSWE / 53.1 AutomationBench / 89.9 Terminal Bench 2.1
✨ MiMo-V2.6 Flash RL: Maximum efficiency
- 309B / 15B
- 15B active params while staying close to Pro on many agent benchmarks
DeepSeek V4 Pro 0813 vs Claude Fable 5 on DeepSWE: Cost, Coding, and Routing
https://www.together.ai/blog/deepseek-v4-pro-0813-vs-claude-fable-5-on-deepswe-cost-coding-and-routin#IA##AI##ML#ML
DeepSeek V4 Pro 0813 vs Claude Fable 5 on DeepSWE: Cost, Coding, and Routing
https://www.together.ai/blog/deepseek-v4-pro-0813-vs-claude-fable-5-on-deepswe-cost-coding-and-routin#IA##AI##ML#ML
Will want to download and check out.
It's a 118B total parameter MoE model with 8B activated per token, a context window of up to 1M tokens, and thinking and no-thinking modes.
Small enough to run on a single @NVIDIAAI DGX Spark
poolside.ai/blog/introdu...
Will want to download and check out.
GLM-5.3-Flash can now be run locally! ✨
Run 3-bit on 128GB RAM via Unsloth GGUF.
GLM-5.3-Flash (ox-alpha) rivals Claude Opus 4.8 on DeepSWE, coding & agentic benchmarks.
Guide:
GGUF:
https://x.com/i/web/status/2092986464196002094
GLM-5.3-Flash can now be run locally! ✨
Run 3-bit on 128GB RAM via Unsloth GGUF.
GLM-5.3-Flash (ox-alpha) rivals Claude Opus 4.8 on DeepSWE, coding & agentic benchmarks.
Guide:
GGUF:
https://x.com/i/web/status/2092986464196002094
مدل Muse Spark 1.3 معرفی شد. این مدل توی تست DeepSWE 1.1 تونسته غولهایی مثل GPT-5.6 و Opus 5 رو پشت سر بذاره.
مدل Muse Spark 1.3 معرفی شد. این مدل توی تست DeepSWE 1.1 تونسته غولهایی مثل GPT-5.6 و Opus 5 رو پشت سر بذاره.
https://gigazine.net/news/20260528-deepswe-ai-coding-benchmark/
https://gigazine.net/news/20260528-deepswe-ai-coding-benchmark/