#ArtificialAnalysis
New from ArtificialAnalysis: The price of compute per KG set to surpass the price of cocaine in 2029
September 26, 2026 at 10:28 PM
Chronically refreshing ArtificialAnalysis this morning because I'm addicted to seeing benchmarks for new models. I'm probably not even going to use GLM 5.3 anytime soon. I'm a simple man who likes when numbers go brrr.
August 14, 2026 at 2:15 PM
8. And China still has the smartest models on almost every benchmark, here is the recent ArtificialAnalysis rankings.
January 7, 2026 at 3:27 PM
ArtificialAnalysis claims Xiaomi MiMo 2.6 Pro (Which is a 1TA42B model) is delivering roughly 6 Sol performance at ~5.6 Luna prices, so. There's that.
September 25, 2026 at 8:41 PM
GLM-5.2 is on ArtificialAnalysis now
June 17, 2026 at 11:31 AM
The contempt of AI from insurance adjusters gets at a serious problem with AI: it still constantly makes up stuff and it can't be trusted with data. Here's state of the art LLMs (including just-released Astra & Fable 5.1) still confidently barfing up wrong answers a third of the time.
September 3, 2026 at 11:55 PM
When I say "bullshitting is inherent to LLMs," I don't mean it colloquially, I mean it empirically. Here's the bleeding edge of the frontier (Opus, Sol, Fable, Astra), and the lowest bullshit rate (answering wrongly instead of admitting ignorance) is 45%. artificialanalysis.ai/evaluations/...
September 8, 2026 at 11:14 AM
LLM Video Generation Models ELO score by ArtificialAnalysis

So many great AI Video are available, most of them are on
@replicate.com

It's addictive, I want to try them ALL!
December 17, 2024 at 6:32 PM
GPT-6 Sol, Luna

噂通り発表された
Terraが消えるというのも噂通り

ArtificialAnalysisのベンチマークでは性能はそれほぼ変わっていない
一方、APIでは半額くらいになっていて、コストパフォーマンスがいいモデルとなっている
トークン使用量も大きくは変わっていなくて、APIの値下げは直接感じられるはず

Astraもでて、Terraの立ち位置が微妙になったのだろうが、このように命名法則がすぐに変わっていくと追っている側としてはやりづらい

Claude Opus 5.5も同日に発表されていて、Grok 4.7も先日発表されたし、競争が激化している
September 22, 2026 at 9:33 PM
Oh nice. yup, that is probably the latest flash (0731) which gets a better score on artificialanalysis than Pro.
August 3, 2026 at 11:54 PM
Claude Opus 5.5

ArtificialAnalysisのベンチマークを見ると、Opus 5から性能が大幅に向上し、Fableをも上回る結果を出している
API価格も20%引きとなったが、トークン使用量も増え結局は値段は同じくらいになってしまっている

今まではAPIの価格は上がる一方だったが、トップ2社が価格を下げる方向に進んでいて、コスパも注視しているのが分かる

個人的にはトップモデルが必要なことは多くなく、それなりの性能のモデルが安く使えるなら歓迎する
September 22, 2026 at 9:49 PM
i feel like ArtificialAnalysis is the only benchmark i care about bc they do total cost vs performance comparisons

A10B is going to feel real snappy vs Sonnet
October 27, 2025 at 11:43 AM
cooking a thing, might just turn this into a lil' site with filterable bench scores and visualizations too

numbers for subs from x.com/semianalysis..., cost per task is from artificialanalysis, ant says quotas are still extended but i didn't correct for that to give them the best fighting chance
July 23, 2026 at 9:55 PM
GPT-5.6シリーズが発表されていた

ArtificialAnalysisのベンチマーク結果を見る限り、Solの性能はFableに匹敵と言ってよさそう

でも、個人的には、性能よりも、コスパの方が重要に感じる

Fableと同等のSolは、Fableの3分の1のコストで、ArtificialAnalysisのベンチマークを実行したという

性能も重要だけど、一般ユーザーは、それと同じくらいコスパが重要

最近のモデルは、性能は良くなっているものの、トークン使用量が増え、結局かかる費用が上がるという現象が起こっている
その中で、コスパのことを考えて作られたモデルは重要な意味を持つと思う
July 9, 2026 at 8:15 PM
ArtificialAnalysis Intelligence Index v4.1

It’s always been a composite over lots of benchmarks. They upgraded several, dropped some, added some, and generally skewed more toward agentic tasks

artificialanalysis.ai/methodology/...
June 16, 2026 at 11:57 AM
Faster. Not smarter. Mercury 2.5 runs 770 tokens per second, burns 35M tokens to score 12 on the Artificial Analysis index, where the median model burns 85M to hit 13. One point less smart, the basics dialed.

#artificialanalysis

https://links.madcoolsocial.com/eXhAE2F
September 24, 2026 at 5:10 PM
ArtificialAnalysis doesn't rate Qwen 3.8 Max as high as Kimi K3, on a par with Sonnet 5
August 3, 2026 at 10:14 PM
Introducing Claude Opus 5...

#Anthropic #ArtificialAnalysis #FreeCAD

Read more on Simon Willison's Weblog: https://postreads.co/feed-item/90502/click?source=bluesky
July 31, 2026 at 4:00 AM
oh very interesting, artificialanalysis is claiming 25, but openrouter is claiming 50. I am finding it feels more like 50, but that does still pale in comparison to 150 from the biggies
June 3, 2025 at 4:30 PM
[JP] GPT-6 Sol (Max)の知能・性能・価格を徹底分析!Artificial Analysis最新ベンチマーク公開
[EN] In-Depth Analysis of GPT-6 Sol (Max): Intelligence, Performance, and Prici…

https://ai-minor.com/blog/en/2026-09-23-1790117983886-gpt_6_sol__max__intelligence__performance_and_pric

#GPT-6Sol #ベンチマーク #ArtificialAnalysis #AI #Tech
In-Depth Analysis of GPT-6 Sol (Max): Intelligence, Performance, and Pricing Unveiled! Latest Benchmarks from Artificial Analysis
In-Depth Analysis of GPT-6 Sol (Max): Intelligence, Performance, and Pricing Unveiled! Latest Benchmarks from Artificial Analysis
ai-minor.com
September 23, 2026 at 12:26 AM
Muse Spark 1.3

ArtificialAnalysisのベンチマークを見る限り、ClaudeのFableとOpusの間の性能
価格はClaude系等の同等性能のモデルと比較すると相当安い(xhighですらGemini 3.8 FlashやGPT 5.6 Terra(max)の同じくらい)
また、スピードも速い(GeminiのFlashに匹敵するくらい)
1Mあたりの価格も性能にしては相当安く設定されている

Muse Codeも発表し、本格的にAnthropicやOpenAI2対抗してきた
性能も問題なく、その割に速度が速く価格が安い
トップモデルの競争にMetaが参入してきた
September 2, 2026 at 9:01 PM
Not as much as an uplift in the ArtificialAnalysis Index score for Deepseek V4 Pro 0813 vs V4 Flash 0731 as I was hoping for (53 vs 52). But the coding specific benchmarks look like a stronger improvement than the general purpose ones covered in the index. RL'ed a bit harder on coding?
August 13, 2026 at 12:46 PM
also, if i load a bunch of models up on artificialanalysis, it seems to follow a logarithmic curve on price per task vs. “intelligence,” which is what scaling laws have predicted all along
August 14, 2026 at 1:01 PM
New #ArtificialAnalysis metric!

This is exactly the direction #AI should take.

Users don’t care which #LLM is slightly better at arithmetic, programming, or isolated tasks. The real challenge is #multidisciplinary AI—models that can handle #real-world problems holistically.

x.com/ArtificialAn...
February 14, 2025 at 12:24 PM