Benchmarks GPT-6 Sol & Luna:
• FrontierCode: 48.4% vs. Fable 5.1’s 48.7%, both at xhigh. $1.37 vs. $9.27 per task, roughly 85% cheaper.
• DeepSWE: 68.8% at max vs. Fable 5’s best result of 69.9% at xhigh. $2.74 vs. $13.41 per tas...
https://x.com/i/web/status/2102462091299098640
Benchmarks GPT-6 Sol & Luna:
• FrontierCode: 48.4% vs. Fable 5.1’s 48.7%, both at xhigh. $1.37 vs. $9.27 per task, roughly 85% cheaper.
• DeepSWE: 68.8% at max vs. Fable 5’s best result of 69.9% at xhigh. $2.74 vs. $13.41 per tas...
https://x.com/i/web/status/2102462091299098640
x.com/cognition/st...
x.com/cognition/st...
#Cognition #SWE2 #IA #InteligenciaArtificial #Programacion #MachineLearning #CiberED
#Cognition #SWE2 #IA #InteligenciaArtificial #Programacion #MachineLearning #CiberED
【コード生成AI「SWE-2」徹底検証:Fable 5.1同等性能を64%コスト削減した Cognition …】
CognitionのコーディングモデルSWE-2を早速検証。Kimi K3をポストトレーニングした結果、FrontierCodeでFable 5.1並みの精度を叩き出しつつ、コストを64%も削減できて驚いた。実務のコード生成コスト改善に直結する。…
👇 詳細・一次ソース解説
https://labomaru.com/posts/20260913121123/
#AI速報 #らぼまる #AIツール
【コード生成AI「SWE-2」徹底検証:Fable 5.1同等性能を64%コスト削減した Cognition …】
CognitionのコーディングモデルSWE-2を早速検証。Kimi K3をポストトレーニングした結果、FrontierCodeでFable 5.1並みの精度を叩き出しつつ、コストを64%も削減できて驚いた。実務のコード生成コスト改善に直結する。…
👇 詳細・一次ソース解説
https://labomaru.com/posts/20260913121123/
#AI速報 #らぼまる #AIツール
🔗 aidailypost.com/news/cogniti...
🔗 aidailypost.com/news/cogniti...
SWE-2 소개: 코딩 모델의 경쟁 축이 점수에서 비용 대비 성능으로
Cognition이 현지 시간 2026년 9월 10일 자사의 코딩 특화 모델 SWE-2 를 공개했습니다. 핵심 수치는 FrontierCode 1.1 Main에서 50.0%입니다. 같은 벤치마크에서 Claude Fable 5.1이 기록한 50.9%에 1점이 채 못 미치지만, 롤아웃 한 번에 드는 평균 비용은 64% 낮습니다. Cognition은…
SWE-2 소개: 코딩 모델의 경쟁 축이 점수에서 비용 대비 성능으로
Cognition이 현지 시간 2026년 9월 10일 자사의 코딩 특화 모델 SWE-2 를 공개했습니다. 핵심 수치는 FrontierCode 1.1 Main에서 50.0%입니다. 같은 벤치마크에서 Claude Fable 5.1이 기록한 50.9%에 1점이 채 못 미치지만, 롤아웃 한 번에 드는 평균 비용은 64% 낮습니다. Cognition은…
Cognition has released SWE-2, a new coding model that matches Fable 5.1's performance on FrontierCode 1.1 Main1 at 64% lower cost, scaling RL to multi-trillion parameters.
It's a strong move towards making high-end coding assistance more accessible.
Cognition has released SWE-2, a new coding model that matches Fable 5.1's performance on FrontierCode 1.1 Main1 at 64% lower cost, scaling RL to multi-trillion parameters.
It's a strong move towards making high-end coding assistance more accessible.
Cognition's new SWE-2 coding model scores within a point of Fable 5.1 on the FrontierCode 1.1 Main benchmark while costing 64% less, and comes within a few points of GPT-6 Astra at a quarter of its price.
Cognition's new SWE-2 coding model scores within a point of Fable 5.1 on the FrontierCode 1.1 Main benchmark while costing 64% less, and comes within a few points of GPT-6 Astra at a quarter of its price.
SWE-2は、SWE-1.7を上回る能力とコスト効率を実現した最新のコーディングモデルで、RLを用いて全ての意思決定レベルを1回の訓練で学習させ、 FrontierCodeのパレート前線を全面的に押し上げた。 FrontierCode 1.1 MainでSWE-2はSWE-1.7やGrok 4.6を上回り、GPT-5.6 SolやFable 5/5.1と同等以上の性能を大幅なコスト削減で達成し、GPT-6 Astraに肉薄する程度のコスト効率を実現している。 さらに、データ拡充・検証強化・リアルタイムのオンラインドラフトモデル統合などの改善を通じて、 (1/2)
SWE-2は、SWE-1.7を上回る能力とコスト効率を実現した最新のコーディングモデルで、RLを用いて全ての意思決定レベルを1回の訓練で学習させ、 FrontierCodeのパレート前線を全面的に押し上げた。 FrontierCode 1.1 MainでSWE-2はSWE-1.7やGrok 4.6を上回り、GPT-5.6 SolやFable 5/5.1と同等以上の性能を大幅なコスト削減で達成し、GPT-6 Astraに肉薄する程度のコスト効率を実現している。 さらに、データ拡充・検証強化・リアルタイムのオンラインドラフトモデル統合などの改善を通じて、 (1/2)
->Startup Fortune | More on "AI coding agents cost efficiency" at BigEarthData.ai
2.8조 매개변수 Kimi K3 를 후속 학습한 SWE-2는 FrontierCode 1.1 Main에서 50.0%를 기록해 Fable 5.1과 0.9%p 차이를 보이면서도 비용은 64% 낮음 추론 노력 수준별로 기반 모델의 비용·성능 곡선 기울기에 맞춘 선형 비용 페널티 를 적용해, 한 ...
2.8조 매개변수 Kimi K3 를 후속 학습한 SWE-2는 FrontierCode 1.1 Main에서 50.0%를 기록해 Fable 5.1과 0.9%p 차이를 보이면서도 비용은 64% 낮음 추론 노력 수준별로 기반 모델의 비용·성능 곡선 기울기에 맞춘 선형 비용 페널티 를 적용해, 한 ...
(also, why no xhigh effort level? that’s sick…)
(also, why no xhigh effort level? that’s sick…)
Main を見ると astra-low がよかった。
Main を見ると astra-low がよかった。
codex系は 5.6-luna-xhigh/max → 6-astra-medium → 6-astra-high
自分が $200 プラン両方持ってたら
gpt-6-astra-high, fable-5.1-medium のみで使えるだけ使いそう。
FrontierCode なのであくまで実装タスクのみの話。
codex系は 5.6-luna-xhigh/max → 6-astra-medium → 6-astra-high
自分が $200 プラン両方持ってたら
gpt-6-astra-high, fable-5.1-medium のみで使えるだけ使いそう。
FrontierCode なのであくまで実装タスクのみの話。
コスパモデルをスコア順で見ると
小: gpt-5.6-luna-xhigh/max, grok-4.6-low
中: grok-4.6-medium, claude-fable-5.1-low
大: gpt-6-astra-high, claude-fable-5.1-medium
意外と grok がいいところを抑えている。
Cursor ユーザーは基本 grok-4.6-low/medium で困ったら fable-5.1-medium という使い方がよさそう。
コスパモデルをスコア順で見ると
小: gpt-5.6-luna-xhigh/max, grok-4.6-low
中: grok-4.6-medium, claude-fable-5.1-low
大: gpt-6-astra-high, claude-fable-5.1-medium
意外と grok がいいところを抑えている。
Cursor ユーザーは基本 grok-4.6-low/medium で困ったら fable-5.1-medium という使い方がよさそう。
Fable-level intelligence is now 54% cheaper – making it even cheaper than Opus – due to a change in caching.
Devin’s Fusion harness is now even smarter and cheaper than before, matching Fable 5.1 on FrontierCode at 47% lower cost. Here's how:
Fable-level intelligence is now 54% cheaper – making it even cheaper than Opus – due to a change in caching.
Devin’s Fusion harness is now even smarter and cheaper than before, matching Fable 5.1 on FrontierCode at 47% lower cost. Here's how: