#FrontierCode
@kimmonismus

Benchmarks GPT-6 Sol & Luna:

• FrontierCode: 48.4% vs. Fable 5.1’s 48.7%, both at xhigh. $1.37 vs. $9.27 per task, roughly 85% cheaper.

• DeepSWE: 68.8% at max vs. Fable 5’s best result of 69.9% at xhigh. $2.74 vs. $13.41 per tas...

https://x.com/i/web/status/2102462091299098640
September 23, 2026 at 6:31 PM
heh here's one. competitive! (these benchmarks are all shit though)
x.com/cognition/st...
September 22, 2026 at 7:23 PM
💻 Cognition lanza SWE-2: un modelo de codificación post-entrenado desde Kimi K3 que iguala a Fable 5.1 en FrontierCode con un 64% menos de coste cibered.com/inteligencia...

#Cognition #SWE2 #IA #InteligenciaArtificial #Programacion #MachineLearning #CiberED
Cognition lanza SWE-2: un modelo de codificación post-entrenado desde Kimi K3 que iguala a Fable 5.1 en FrontierCode con un 64% menos de coste | Novedades IA | CIBERED
Cognition presenta SWE-2, su modelo de codificación más capaz: 50% en FrontierCode, 64% más barato que Fable 5.1, pero solo disponible dentro de Devin.
cibered.com
September 16, 2026 at 4:29 AM
2/ Côté performance, Cognition (Devin) sort SWE-2 : le niveau de Fable 5.1 sur FrontierCode (50,0%), pour 64% moins cher. Base : Kimi K3, le modèle ouvert de Moonshot AI (2 800 milliards de paramètres), affiné par renforcement. https://www.lefilia.fr/article/6157540
September 13, 2026 at 6:40 AM
🐶 らぼまる速報⚡

【コード生成AI「SWE-2」徹底検証:Fable 5.1同等性能を64%コスト削減した Cognition …】
CognitionのコーディングモデルSWE-2を早速検証。Kimi K3をポストトレーニングした結果、FrontierCodeでFable 5.1並みの精度を叩き出しつつ、コストを64%も削減できて驚いた。実務のコード生成コスト改善に直結する。…

👇 詳細・一次ソース解説
https://labomaru.com/posts/20260913121123/

#AI速報 #らぼまる #AIツール
コード生成AI「SWE-2」徹底検証:Fable 5.1同等性能を64%コスト削減した Cognition の推論最適化メカニズム
CognitionのコーディングモデルSWE-2を早速検証。Kimi K3をポストトレーニングした結果、FrontierCodeでFable 5.1並みの精度を叩き出しつつ、コストを64%も削減できて驚いた。実務のコード生成コスト改善に直…
labomaru.com
September 13, 2026 at 3:13 AM
🚀 Cognition just dropped its SWE-2 coding model—matching the rival at 64% lower cost! Powered by Devin, Kimi K3 and post‑training tricks, it’s a game‑changer for RL‑driven dev. Curious? Dive into the details. #SWE2 #CognitionAI #FrontierCode

🔗 aidailypost.com/news/cogniti...
September 13, 2026 at 12:18 AM
Cognition, Kimi K3 기반 코딩 모델 SWE-2 공개: Fable 5.1과 1점 차, 비용은 64% 절감

SWE-2 소개: 코딩 모델의 경쟁 축이 점수에서 비용 대비 성능으로

Cognition이 현지 시간 2026년 9월 10일 자사의 코딩 특화 모델 SWE-2 를 공개했습니다. 핵심 수치는 FrontierCode 1.1 Main에서 50.0%입니다. 같은 벤치마크에서 Claude Fable 5.1이 기록한 50.9%에 1점이 채 못 미치지만, 롤아웃 한 번에 드는 평균 비용은 64% 낮습니다. Cognition은…
Cognition, Kimi K3 기반 코딩 모델 SWE-2 공개: Fable 5.1과 1점 차, 비용은 64% 절감
SWE-2 소개: 코딩 모델의 경쟁 축이 점수에서 비용 대비 성능으로 Cognition이 현지 시간 2026년 9월 10일 자사의 코딩 특화 모델 SWE-2 를 공개했습니다. 핵심 수치는 FrontierCode 1.1 Main에서 50.0%입니다. 같은 벤치마크에서 Claude Fable 5.1이 기록한 50.9%에 1점이 채 못 미치지만, 롤아웃 한 번에 드는 평균 비용은 64% 낮습니다. Cognition은 이 결과를 두고 파레토 프론티어(Pareto Frontier), 즉 더 싸게 만들려면 성능을 포기해야 하는 경계선 자체를 밀어냈다고 표현합니다. C...
discuss.pytorch.kr
September 12, 2026 at 1:31 AM
Cognition's SWE-2 hits 50% on FrontierCode for 64% less cost than Fable 5.1, using RL scale to flatten the cost-performance curve.
Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
Four Signals — The Wire
www.foursignals.dev
September 11, 2026 at 4:00 PM
Cognition's SWE-2: better code, lower cost 💸

Cognition has released SWE-2, a new coding model that matches Fable 5.1's performance on FrontierCode 1.1 Main1 at 64% lower cost, scaling RL to multi-trillion parameters.

It's a strong move towards making high-end coding assistance more accessible.
September 11, 2026 at 3:27 PM
7. Cognition SWE-2 achieves 50% on FrontierCode 1.1 at 64% lower cost LINK
8. T1: 122B Mixture-of-Experts Model for Long-Horizon Tool Use LINK
September 11, 2026 at 1:00 PM
Cognition ships SWE-2, nearing GPT-6 Astra at a quarter of the cost
Cognition's new SWE-2 coding model scores within a point of Fable 5.1 on the FrontierCode 1.1 Main benchmark while costing 64% less, and comes within a few points of GPT-6 Astra at a quarter of its price.
Cognition ships SWE-2, nearing GPT-6 Astra at a quarter of the cost
Cognition's new SWE-2 coding model scores within a point of Fable 5.1 on the FrontierCode 1.1 Main benchmark while costing 64% less, and comes within a few points of GPT-6 Astra at a quarter of its price.
glonce.com
September 11, 2026 at 8:12 AM
💡 Summary:

SWE-2は、SWE-1.7を上回る能力とコスト効率を実現した最新のコーディングモデルで、RLを用いて全ての意思決定レベルを1回の訓練で学習させ、 FrontierCodeのパレート前線を全面的に押し上げた。 FrontierCode 1.1 MainでSWE-2はSWE-1.7やGrok 4.6を上回り、GPT-5.6 SolやFable 5/5.1と同等以上の性能を大幅なコスト削減で達成し、GPT-6 Astraに肉薄する程度のコスト効率を実現している。 さらに、データ拡充・検証強化・リアルタイムのオンラインドラフトモデル統合などの改善を通じて、 (1/2)
September 11, 2026 at 7:44 AM
Cognition dropped SWE-2: 50.0% on FrontierCode 1.1 — within 1 point of Fable 5.1, a quarter the cost of GPT-Astra. Trained with a cost-penalized RL objective so all effort levels ship in one run. https://cognition.com/blog/swe-2
September 11, 2026 at 5:02 AM
Cognition's SWE-2 Coding Agent Matches Rivals at a Quarter of the Price
Cognition just launched SWE-2, an AI coding agent it says scores within a single point of Anthropic's top coding model while costing up to 70% less to run. The pitch is blunt: near-frontier code, at a fraction of the bill. Cognition, the startup behind the autonomous coding agent Devin, shipped SWE-2 on September 10, 2026, rolling it out across Devin Desktop, the Devin CLI, Devin Web and Devin Fusion. On the company's own FrontierCode 1.1 Main benchmark, which grades whether a human maintainer would actually merge an AI-written pull request, SWE-2 scored 50.0%. Anthropic's Fable 5.1 scored 50.9%. That's a single point apart. Cognition says it matched that score for 64% less compute cost. Against OpenAI's GPT-6 Astra, which scored 53.3% on the same test, Cognition claims SWE-2 comes in at roughly a quarter of the price. Price is the argument. It also beat its own predecessor, SWE-1.7, which scored 42.0%, and xAI's Grok 4.6, which scored 48.0%. The engineering behind that claim is specific. SWE-2 is post-trained from Moonshot AI's Kimi K3, a 2.8 trillion parameter base model, using a reinforcement learning method that trains three separate effort levels, medium, high and max, in a single run instead of...
startupfortune.com
September 11, 2026 at 3:45 AM
Cognition's SWE-2 sits within a point of Fable 5.1 on FrontierCode at 64% lower cost, Devin-only. Also: a second mathematician on trusting OpenAI with unpublished math, Raschka on looped transformers, Anthropic's economic scenarios. Synth Daily: https://argosvix.com/en/live/news/2026-09-11
September 10, 2026 at 10:44 PM
Cognition, Fable 5.1·GPT-Astra에 필적하는 코딩 모델 SWE-2 출시

2.8조 매개변수 Kimi K3 를 후속 학습한 SWE-2는 FrontierCode 1.1 Main에서 50.0%를 기록해 Fable 5.1과 0.9%p 차이를 보이면서도 비용은 64% 낮음 추론 노력 수준별로 기반 모델의 비용·성능 곡선 기울기에 맞춘 선형 비용 페널티 를 적용해, 한 ...
Cognition, Fable 5.1·GPT-Astra에 필적하는 코딩 모델 SWE-2 출시
2.8조 매개변수 Kimi K3 를 후속 학습한 SWE-2는 FrontierCode 1.1 Main에서 50.0%를 기록해 Fable 5.1과 0.9%p 차이를 보이면서도 비용은 64% 낮음 추론 노력 수준별로 기반 모델의 비용·성능 곡선 기울기에 맞춘 선형 비용 페널티 를 적용해, 한 ...
news.hada.io
September 10, 2026 at 8:00 PM
crazy — cognition/windsurf are officially in the game

(also, why no xhigh effort level? that’s sick…)
September 10, 2026 at 6:15 PM
FrontierCode Extended って Main + 簡単なタスクだったみたい。

Main を見ると astra-low がよかった。
September 4, 2026 at 9:49 PM
claude系は選択肢なく fable-5.1-low → opus-5-medium(fable 枠使い切ったら)

codex系は 5.6-luna-xhigh/max → 6-astra-medium → 6-astra-high

自分が $200 プラン両方持ってたら
gpt-6-astra-high, fable-5.1-medium のみで使えるだけ使いそう。

FrontierCode なのであくまで実装タスクのみの話。
September 4, 2026 at 9:30 PM
推しベンチマークである FrontierCode の成績
コスパモデルをスコア順で見ると
小: gpt-5.6-luna-xhigh/max, grok-4.6-low
中: grok-4.6-medium, claude-fable-5.1-low
大: gpt-6-astra-high, claude-fable-5.1-medium

意外と grok がいいところを抑えている。
Cursor ユーザーは基本 grok-4.6-low/medium で困ったら fable-5.1-medium という使い方がよさそう。
September 4, 2026 at 9:20 PM
This lines up for me tbh. FrontierCode has consistently been semi inversely aligned with my preferences lol
September 2, 2026 at 8:49 AM
RT @dluzar: Fable 5.1 underperforming Fable 5 from ≥high on FrontierCode.
September 2, 2026 at 8:47 AM
FrontierCode のベンチ結果が Fable 5.1 や Opus 5 だと medium が一番成績いいのは、それ以上にするとスコープ外の仕事もしてしまうから。というのはわかりやすくてよい。
September 2, 2026 at 5:09 AM
RT @cognition: Introducing Fable 5.1 in Devin

Fable-level intelligence is now 54% cheaper – making it even cheaper than Opus – due to a change in caching.

Devin’s Fusion harness is now even smarter and cheaper than before, matching Fable 5.1 on FrontierCode at 47% lower cost. Here's how:
September 2, 2026 at 8:46 AM