#FrontierCode
FrontierCode benchmarks AI-generated code quality with a focus on mergeability. Crafted by 20 open-source developers, it addresses real-world coding standards and achieves 81% fewer misclassifications, challenging top AI models. https://cognition.ai/blog/frontier-code
Introducing FrontierCode | Cognition
FrontierCode benchmarks AI-generated code quality with a focus on mergeability. Crafted by 20 open-source developers, it addresses real-world coding standards and achieves 81% fewer misclassifications
cognition.ai
June 9, 2026 at 11:10 PM
heh here's one. competitive! (these benchmarks are all shit though)
x.com/cognition/st...
September 22, 2026 at 7:23 PM
Gemini 3.7 Flash is here!

and these bars are ordered very intentionally

blog.google/innovation-a...
August 13, 2026 at 6:31 PM
On agentic terminal coding (Frontier-Bench v0.1), Opus 5 is the new state of the art at 43.3%: ahead of Fable 5 (33.7%) and more than double Opus 4.8.

On FrontierCode v1.1 it matches Fable 5, at half the price.
July 24, 2026 at 5:35 PM
crazy — cognition/windsurf are officially in the game

(also, why no xhigh effort level? that’s sick…)
September 10, 2026 at 6:15 PM
Important cost vs performance phrasing
June 9, 2026 at 5:29 PM
You can see this with the latest Claude. Notice the logarithmic cost x-axis vs the linear performance y-axis. You don't get 2x better output for 2x the cost. That's not an outlier, it's well-known, OpenAI showed the same linear-performance-increase vs logarithmic-compute-cost in September 2024.
/2
June 12, 2026 at 3:51 PM
Cognition、AI生成コードの品質評価用ベンチマーク「FrontierCode」発表
#FrontierCode #ITニュース
ITちゃんねる
Cognition、AI生成コードの品質評価用ベンチマーク「FrontierCode」発表 #FrontierCode #ITニュース
it.f-frontier.com
June 9, 2026 at 7:25 AM
pretty good marketing though
August 13, 2026 at 2:57 PM
Cognition's SWE-2: better code, lower cost 💸

Cognition has released SWE-2, a new coding model that matches Fable 5.1's performance on FrontierCode 1.1 Main1 at 64% lower cost, scaling RL to multi-trillion parameters.

It's a strong move towards making high-end coding assistance more accessible.
September 11, 2026 at 3:27 PM
@kimmonismus

Benchmarks GPT-6 Sol & Luna:

• FrontierCode: 48.4% vs. Fable 5.1’s 48.7%, both at xhigh. $1.37 vs. $9.27 per task, roughly 85% cheaper.

• DeepSWE: 68.8% at max vs. Fable 5’s best result of 69.9% at xhigh. $2.74 vs. $13.41 per tas...

https://x.com/i/web/status/2102462091299098640
September 23, 2026 at 6:31 PM
Introducing FrontierCode, a benchmark for evaluating AI-generated code quality through mergeability and adherence to coding standards, highlighting that even leading models struggle to meet these rigorous expectations. https://cognition.ai/blog/frontier-code
Introducing FrontierCode | Cognition
Introducing FrontierCode, a benchmark for evaluating AI-generated code quality through mergeability and adherence to coding standards, highlighting that even leading models struggle to meet these rigo
cognition.ai
August 9, 2026 at 9:00 AM
🔥 FrontierCode revolutionizes AI
FrontierCode is a new AI concept that has the potential to revolutionize the field of artificial intelligence. This innovative approach could lead to sig...

🌐 aitechcodex.uk/news/2206/?utm_source=bluesky&utm_medium=social&utm_campaign=daily_news 📡 t.me/AITechNewsUK
June 9, 2026 at 3:01 AM
The best AI coding model in the world scores 13 out of 100. On code a real engineer would actually merge.

Every benchmark you've seen measures: does this code run?

FrontierCode measures: would a senior maintainer accept this PR?

Scope discipline. Style consistency. Regression safety.
June 10, 2026 at 6:42 AM
It's pretty bad on most frontier-ish stuff.

iirc, FrontierCode is measuring "mergeability" based on a bunch of heuristics determined by maintainers, rather than "does it complete the task". Makes it really noisy.
July 24, 2026 at 10:22 PM
Devs: Prioritize "Kernel Contracts" to bound training-inference divergence and adopt "FrontierCode" benchmarks to filter LLM-generated slop. These tools shift workflows from "prompt-and-pray" to verifiable, high-fidelity code generation and deployment stability.
Enhancing Strawberry Yield Forecasting with Backcasted IoT Sensor Data and Machine Learning
Rapid global population growth underscores the need for digitally enabled agricultural systems that support sustainable food production and data-driven resource management for farmers and stakeholders. The adoption of Internet of Things (IoT) technologies, capable of capturing real-time environmenta
arxiv.org
June 9, 2026 at 10:34 AM
جربت Kimi K3 بنفسي… والنتيجة صدمتني
النموذج المفتوح المصدر اللي كسر الدنيا وبقى رقم 1 على FrontierCode
قعدت أختبره في حاجات حقيقية، وشوفوا عمل إيه

#حسام_الدين_حسن #خبير_اونلاين #Kimi #Kimi_K3 #KimiK3 #مفتوح_المصدر #OpenSource #الذكاء_الاصطناعي #نماذج_الذكاء_الاصطناعي #ذكاء_اصطناعي_عربي #Fable_5
July 18, 2026 at 2:35 PM
Google launched Gemini 3.7 Flash, a new Flash model for coding and AI agents, just three weeks after 3.6 Flash. It improves in software engineering, web dev, and automation. Benchmarks show gains in FrontierCode and WebDev Arena.

#Google #Gemini

aidisruption.ai/p/musks-grok...
Musk's Grok Bot Is Here!
Musk drops Grok Bot—an AI teammate with its own cloud PC. Works 24/7, runs multiple bots, and automates real tasks like a human colleague.
aidisruption.ai
August 14, 2026 at 9:56 AM
If this benchmark is to be trusted, current frontier is:
- grok 4.5 (high): cheap + fast
- opus 5 (medium): best
- fable 5 (high): if you're feeling fancy
and forget the rest.

Now time to see if Opus 5 is really that good. Source: cognition.com/frontiercode
July 27, 2026 at 11:59 AM
Gemini 3.7 Flash dropped today — Google's new 'workhorse' model for coding and agents. 43.6% on FrontierCode (up from 34.4%), and intro pricing is half what 3.6 Flash cost. If you're paying for Flash today, switch. https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-g
August 14, 2026 at 9:08 AM
También notable el salto en FrontierCode, un nuevo benchmark que no solo evalúa la capacidad de pasar los tests, sino la 'mergeabilidad' del código.

Así estaba hasta ahora, casi a cero,y Fable 5 da el salto hasta el 30% de golpe!
June 9, 2026 at 5:20 PM
🐶 らぼまる速報⚡

【コード生成AI「SWE-2」徹底検証:Fable 5.1同等性能を64%コスト削減した Cognition …】
CognitionのコーディングモデルSWE-2を早速検証。Kimi K3をポストトレーニングした結果、FrontierCodeでFable 5.1並みの精度を叩き出しつつ、コストを64%も削減できて驚いた。実務のコード生成コスト改善に直結する。…

👇 詳細・一次ソース解説
https://labomaru.com/posts/20260913121123/

#AI速報 #らぼまる #AIツール
コード生成AI「SWE-2」徹底検証:Fable 5.1同等性能を64%コスト削減した Cognition の推論最適化メカニズム
CognitionのコーディングモデルSWE-2を早速検証。Kimi K3をポストトレーニングした結果、FrontierCodeでFable 5.1並みの精度を叩き出しつつ、コストを64%も削減できて驚いた。実務のコード生成コスト改善に直…
labomaru.com
September 13, 2026 at 3:13 AM
Cognition's SWE-2 Coding Agent Matches Rivals at a Quarter of the Price
Cognition just launched SWE-2, an AI coding agent it says scores within a single point of Anthropic's top coding model while costing up to 70% less to run. The pitch is blunt: near-frontier code, at a fraction of the bill. Cognition, the startup behind the autonomous coding agent Devin, shipped SWE-2 on September 10, 2026, rolling it out across Devin Desktop, the Devin CLI, Devin Web and Devin Fusion. On the company's own FrontierCode 1.1 Main benchmark, which grades whether a human maintainer would actually merge an AI-written pull request, SWE-2 scored 50.0%. Anthropic's Fable 5.1 scored 50.9%. That's a single point apart. Cognition says it matched that score for 64% less compute cost. Against OpenAI's GPT-6 Astra, which scored 53.3% on the same test, Cognition claims SWE-2 comes in at roughly a quarter of the price. Price is the argument. It also beat its own predecessor, SWE-1.7, which scored 42.0%, and xAI's Grok 4.6, which scored 48.0%. The engineering behind that claim is specific. SWE-2 is post-trained from Moonshot AI's Kimi K3, a 2.8 trillion parameter base model, using a reinforcement learning method that trains three separate effort levels, medium, high and max, in a single run instead of...
startupfortune.com
September 11, 2026 at 3:45 AM