#ChatbotArena
Pretty crazy that a bunch of LLMs you can run locally on an iPhone would probably beat the original ChatGPT on ChatBotArena.
November 15, 2024 at 4:51 AM
Very pleased to see Tulu 3 70B more or less tied with Llama 3.1 70B Instruct on style controlled ChatBotArena. The only model anywhere close to that with open code and data for post-training! Lots of stuff people can build on.

Next looking for OLMo 2 numbers.
January 8, 2025 at 5:13 PM
ChatBotArena is far from the first eval to be overfit to. It's becoming underrated. Likely the single most impactful evaluation project since ChatGPT. The labs are the ones releasing these slightly off models.
April 28, 2025 at 3:48 PM
Up to 5/15 top models in ChatBotArena being open weights. Yes, there are caveats as always, but this is THE BEST opportunity people have ever had to build things on open models.

R1, Gemma 3, DeepSeek V3, QwQ, Command A 🫡
www.interconnects.ai/p/gemma-3-ol...
Gemma 3, OLMo 2 32B, and the growing potential of open-source AI
Leading open-weight models and the first open-source model to clearly surpass GPT 3.5 (the very last version).
www.interconnects.ai
March 18, 2025 at 2:47 PM
Zkouším AI chatbot DeepSeek: chat.deepseek.com S novým modelem R1 se hned dostal na přední příčky žebříčku ChatbotArena (viz www.zive.cz/clanky/nejle...), odpovídá rychle a správně. Ale je to Čína se vším, co k tomu patří, takže obavy, co člověk vkládá, by měly být větší než jinde.
DeepSeek
Chat with DeepSeek AI.
chat.deepseek.com
January 25, 2025 at 9:03 AM
J'ai testé Grok via ChatbotArena. Sur beaucoup de tests (c'est vague mais je n'ai pas compté) je n'ai quasiment jamais eu un meilleur résultat via Grok.
Pour faire simple, en variant énormément les prompts, c'est O1 mini et Claude 3.5 qui sortent, et parfois
Mistral, mais Grok, c'est non.
November 28, 2024 at 5:07 PM
So in the latest new way to waste time I bring you... the Chatbot Arena: where two soulless text machines wrestle for your fleeting approval while you sip iced coffee and pretend you're contributing to the future of civilization.
#AI #ChatbotArena #machinelearning

toryhayward.com/index.php/20...
Chatbot Arena - Tory Hayward
So in the latest new way to waste time I bring you... the Chatbot Arena.
toryhayward.com
May 14, 2025 at 7:03 AM
not really. I think I got too excited for this one. ChatBotArena is only one of multiple needed tests
November 17, 2024 at 12:30 AM
GPT-4o-mini changed ChatBotArena
Why people are frustrated with a small, cheap, and maybe weak model scoring so high. And how to understand Llama 3.1’s results on the community's favorite benchmark.
Read it here:
GPT-4o-mini changed ChatBotArena
And how to understand Llama 3.1’s results on the community's favorite benchmark.
www.interconnects.ai
July 31, 2024 at 3:08 PM
ChatbotArena的なシステムでTanuki-8x8Bを始めとする大規模言語モデルの日本語性能を評価する(2024年8月)
https://zenn.dev/matsuolab/articles/95fa297ef12a14
ChatbotArena的なシステムでTanuki-8x8Bを始めとする大規模言語モデルの日本語性能を評価する(2024年8月)
zenn.dev
August 31, 2024 at 4:34 AM
OpenAI’s GPT-4o has taken back the top AI spot from Google’s Gemini-Exp, showcasing creative and contextual enhancements in the Chatbot Arena leaderboard. #chatgpt #ai #openai #google #chatbotarena

winbuzzer.com/2024/11/21/o...
OpenAI Updates GPT-4o to Retake Top Spot from Google in Chatbot Arena - WinBuzzer
OpenAI’s GPT-4o has taken back the top AI spot from Google’s Gemini-Exp, showcasing creative and contextual enhancements in the Chatbot Arena leaderboard.
winbuzzer.com
November 21, 2024 at 10:27 AM
スピードはChatbotArenaで比較したため、正確ではないかもだが、Geminiよりは少し遅かった。
Geminiと比較したが、GPT-4o、4o-mini、Claude 3.5 Sonnet、3.5 Haikuとかと比べると、精度は劣っても圧倒的に安いのである程度は価値があるかもしれない。
それでもGPT系を使わないのならば、Geminiが良さそうな気はする(精度はChatbotArenaでProは4o、Flashはminiよりちょい下、値段は相当安いし)。
December 5, 2024 at 11:30 AM
He also doesn’t take responsibility for failures (gaming ChatbotArena), and overstates impact of features.

Claims 1 billion users of Meta AI without being able to give a real example of use (besides “having hard conversations”).

Vacuous predictions like “people will want to have fun” with tools 🤦‍♂️
April 30, 2025 at 8:07 PM
Wow! I didn't really like Gemma 2, but Gemma 3, released today, is awesome. It comes in four sizes, 1b, 4b, 12b and 27b. It's super fast and except for the 1b version it can even handle images.

The 27B version apparently outperforms both DeepSeek v3 and LLaMA3-405 on the ChatbotArena benchmark […]
Original post on mastodon.social
mastodon.social
March 12, 2025 at 1:10 PM
#ChatbotArena Italia è una piattaforma che ha l'obiettivo di comparare e valutare i Large Language Models sulla lingua italiana. 🤖🇮🇹

Se volete partecipare, basta sottoporre un prompt a due modelli #AI scelti a caso dal sistema e votare la migliore. C'è anche la classifica!

indigo.ai/it/chatbot-a...
indigo.ai | Chatbot Arena Italia
indigo.ai
February 27, 2025 at 6:51 AM
Eine globale Rangliste der KI-Sprachmodelle – und die Möglichkeit, mehrere LLMs blind zu vergleichen: Beides gibt es auf #ChatbotArena. Tipp: Hier war Deepseek zu entdecken, bevor der mediale Hype ausgebrochen ist.
#clickomaniach
Treffen sich zwei Sprachmodelle in einem Boxring … – Clickomania
Bei Chatbot Arena hetzen wir 86 Sprach­mo­del­le auf­einan­der los. Es gibt eine glo­bale Rang­liste und in­te­res­san­te Ein­blicke da­rüber, welche KIs auf dem Auf­stieg sind.
blog.clickomania.ch
February 20, 2025 at 7:40 AM
Elon Musk's xAI just dropped Grok 3, an AI model that outperforms GPT-4o, Gemini, and DeepSeek—leading the Chatbot Arena rankings and redefining the AI race! 🏆🤖

Read more: venturebeat.com/ai/elon-musk... #ArtificialIntelligence #MachineLearning #ElonMusk #Grok3 #ChatbotArena
Elon Musk just released an AI that’s smarter than ChatGPT — here’s why that matters
Elon Musk's xAI launches Grok 3, outperforming ChatGPT and Google Gemini in benchmarks with 200,000 GPUs and advanced reasoning capabilities, intensifying AI competition days after failed OpenAI bid.
venturebeat.com
February 18, 2025 at 8:00 PM
Amazon Novaを使ってみた所感

Nova Pro, Lite, MicroはそれぞれGemini 1.5 Pro-002, Flash-002, Flash-8b-001の1〜2割ほど性能が下がった感じ。少なくとも日本語ではNovaはそれほど良いモデルではない。ChatbotArenaのランキングに乗っていないので、多言語だと分からないが、値段もGemini系統の1〜2割安いくらい(すごく表が読みづらく、間違っているかも)だから、Novaが特別いいというわけではないかな。
December 5, 2024 at 11:27 AM
今日のZennトレンド

大規模言語モデルを開発するにあたっての事前・事後学習の戦略メモー特に合成データについてー
この記事は、大規模言語モデル(LLM)の開発における事前・事後学習の戦略について、特に合成データの重要性を論じています。
著者は、従来の「事後学習は少量で済む」「モデルから知識を取り出すのは容易」といった仮説が、実際には成立せず、LLMに知識やスキルを習得させるには膨大な量の訓練データが必要であると主張しています。
具体的には、合成データを用いた独自の開発経験や、モデルの知識獲得や指示追従能力に関する課題について、複数の仮説と具体的な例を挙げながら解説しています。
大規模言語モデルを開発するにあたっての事前・事後学習の戦略メモー特に合成データについてー
関連URLTanuki-8x8BTanuki-8B大規模言語モデルTanuki-8B, 8x8Bの位置づけや開発指針など全体像フルスクラッチで開発した大規模言語モデルTanuki-8B, 8x8Bの性能についての技術的な詳細Japanese MT-Benchにおける性能の詳細とJasterに関する一部言及ChatbotArena的なシステムでTanuki-8x8Bを始めとする大規模言語モデルの日
zenn.dev
August 31, 2024 at 9:10 AM
🤖 7.7M+ votes rank 385+ models.
🏆 Claude leads text leaderboard.
🎨 Claude & Qwen top multimodal.
💻 Claude, Kimi excel in coding.
#ChatbotArena #AILeaderboard #Claude #MultimodalAI #CodingAI
View in Timelines
August 18, 2026 at 3:08 PM
Grok3がChatbotArenaで1位になってるのでいろいろテストしよ
February 21, 2025 at 4:27 AM
Mô hình AI Trung Quốc lọt Top 10 thế giới về hiệu suất

Thêm một mô hình AI Trung Quốc lọt Top 10 toàn cầu về đánh giá hiệu suất! Qwen2.5-Max của Alibaba Cloud đã vượt trội hơn các mô hình khác như DeepSeek-V3, o1-mini và Claude-3.5-Sonnet để chiếm vị trí trong bảng xếp hạng của Chatbot Arena. Đây…
Mô hình AI Trung Quốc lọt Top 10 thế giới về hiệu suất
Thêm một mô hình AI Trung Quốc lọt Top 10 toàn cầu về đánh giá hiệu suất! Qwen2.5-Max của Alibaba Cloud đã vượt trội hơn các mô hình khác như DeepSeek-V3, o1-mini và Claude-3.5-Sonnet để chiếm vị trí trong bảng xếp hạng của Chatbot Arena. Đây là cột mốc quan trọng cho ngành công nghiệp trí tuệ nhân tạo. #AI #TrungQuốc #Top10 #HiệuSuất #Qwen2.5Max #AlibabaCloud #ChatbotArena…
drtao.vn
February 7, 2025 at 12:32 AM