#sciarena
Introducing SciArena, a platform for benchmarking models across scientific literature tasks. Inspired by Chatbot Arena, SciArena applies a crowdsourced LLM evaluation approach to the scientific domain. 🧵
July 1, 2025 at 3:02 PM
o3, an AI model developed by the creators of ChatGPT, has been ranked the best AI tool for answering science questions in multiple fields

go.nature.com/4lzaWlC
OpenAI's o3 tops new AI league table for answering scientific questions
SciArena uses votes by researchers to evaluate large language models’ responses on technical topics.
go.nature.com
July 12, 2025 at 5:32 PM
A few new challengers enter SciArena—including DeepSeek-V3.2-Exp and Claude Sonnet 4.5 🔬
September 29, 2025 at 7:20 PM
A new model enters SciArena. 👀 Welcome Moonshot AI's Kimi K2! SciArena lets you benchmark models across scientific literature tasks, applying a crowdsourced LLM evaluation approach to the scientific domain.

🧪 Learn more and try SciArena here:
buff.ly/J2kYM0Q
July 17, 2025 at 3:30 PM
SciArena update: our Olmo 3.1 32B Instruct scores 963.6 Elo overall at just $0.17/100 calls—ahead of OpenAI’s GPT-OSS-20B. In Engineering, it hits 1039.2 Elo, only 2.5 behind GPT-OSS-120B—a model ~4× its size. 🧵
January 16, 2026 at 5:57 PM
We built SciArena to test how well AI models handle scientific literature questions, as judged by researchers.

It's retiring July 15, and the results are in: ~1,700 users cast ~3,900 votes.

Here's what they told us. 🧵
July 15, 2026 at 5:06 PM
Today we released SciArena, an open evaluation platform where researchers can compare and vote on foundation models for scientific literature tasks. 👇
July 1, 2025 at 5:15 PM
o3, an AI model developed by the creators of ChatGPT, has been ranked the best AI tool for answering science questions in multiple fields

go.nature.com/44FcZ0i
OpenAI's o3 tops new AI league table for answering scientific questions
SciArena uses votes by researchers to evaluate large language models’ responses on technical topics.
go.nature.com
July 10, 2025 at 9:38 AM
For a deeper look at what SciArena revealed about how scientists evaluate AI-generated answers to Qs about the scientific literature, read our NeurIPS 2025 Spotlight paper—which includes findings beyond the leaderboard.

📄 buff.ly/A9OrLph
SciArena: An Open Evaluation Platform for Non-Verifiable Scientific Literature-Grounded Tasks
We present SciArena, an open and collaborative platform for evaluating foundation models on scientific literature-grounded tasks. Unlike traditional benchmarks for scientific literature understanding…
arxiv.org
July 15, 2026 at 5:07 PM
SciArena leaderboard update! 🔬
We've added new frontier models – including GPT-5.1 and Gemini 3 Pro Preview – to our arena for scientific literature tasks. The new rankings: o3 holds #1, Gemini 3 Pro Preview lands at #2, Claude Opus 4.1 sits at #3, GPT-5 at #4, & GPT-5.1 debuts at #5. 🧵
December 1, 2025 at 8:24 PM
🚨 SciArena update + evaluation of new models including GPT-5! 🚨
With thousands of new votes, new LLMs are reshaping our leaderboard for scientific literature tasks.
o3 still leads—but GPT-5, Claude Opus 4.1, & more are closing the gap.
August 21, 2025 at 7:45 PM
SciArena revolutionizes evaluating foundation models on scientific literature tasks via community-driven voting, amassing over 20,000 votes. With insights on model performance and the new SciArena-Eval benchmark, it enhances automated evaluation methods in science. https://arxiv.org/abs/2507.01001
SciArena: An Open Evaluation Platform for Non-Verifiable Scientific Literature-Grounded Tasks
ArXiv link for SciArena: An Open Evaluation Platform for Non-Verifiable Scientific Literature-Grounded Tasks
arxiv.org
January 24, 2026 at 9:41 PM
The pattern is clear: Olmo 3.1 32B Instruct is exceptionally strong in Engineering and Healthcare with standout efficiency. Compare against other models in SciArena → buff.ly/5M7jj1h & try it here → buff.ly/CNJT5VD
Ai2 SciArena
sciarena.allen.ai
January 16, 2026 at 5:57 PM
Have a tough scientific research question? Submit it, compare citation-grounded model responses, and vote. The leaderboard updates regularly as the community weighs in → sciarena.allen.ai
Ai2 SciArena
sciarena.allen.ai
September 29, 2025 at 7:00 PM
o3, an AI model developed by the creators of ChatGPT, has been ranked the best AI tool for answering science questions in multiple fields, according to a benchmarking platform launched last week by the Allen Institute for Artificial Intelligence
🧪
www.nature.com/articles/d41...
OpenAI's o3 tops new AI league table for answering scientific questions
SciArena uses votes by researchers to evaluate large language models’ responses on technical topics.
www.nature.com
July 11, 2025 at 7:00 AM
SciArena has three key components:
1️⃣ SciArena platform: Where anyone can submit questions, vote on their preferred output, and see how models rank.
2️⃣ Leaderboard: A ranking system based on community votes.
3️⃣ SciArena-Eval: A meta-evaluation benchmark from human preference data.
July 1, 2025 at 5:15 PM
We also introduce SciArena-Eval, the first meta-evaluation benchmark for scientific literature tasks built on collected human preference data. The goal is to understand – and improve – LLM-based evaluations in this area.

Both SciArena and SciArena-Eval are now available.
July 1, 2025 at 3:02 PM
o3 finished on top of the SciArena leaderboard – ahead of Claude Opus 4.1, Gemini 3 Pro Preview, & open-weights models like DeepSeek-R1 – with answers researchers called more detailed + to the point.

Learn more in our updated blog: buff.ly/o1EyhhG
SciArena: A new platform for evaluating foundation models in scientific literature tasks | Ai2
Discover how SciArena is being used to evaluate foundation models’ capabilities in scientific literature tasks through community-driven, literature-grounded, and multi-disciplinary reasoning.
allenai.org
July 15, 2026 at 5:07 PM
SciArena also allowed us to collect high-quality, expert-annotated ground truth on evaluation data.

This is a unique resource & especially important as AI agents are increasingly used to judge the quality of other AI agents, and such evaluations need to be grounded + verified.
July 15, 2026 at 5:07 PM
We've added Grok 4, the latest model from xAI, to our SciArena platform! SciArena allows you to benchmark models across scientific literature tasks, applying a crowdsourced LLM evaluation approach to the scientific domain. 👇
July 14, 2025 at 2:04 PM
Discover SciArena by AllenAI—a community-driven platform that ranks AI models on scientific literature tasks using expert votes and real research papers. Open, transparent, and advancing AI for science! Explore more: sciarena.allen.ai #AI #SciArena #ResearchAI
July 3, 2025 at 5:02 PM
#sciarena is a crowdsourced platform built by the #alleninstituteforai (#ai2, @allenai) for testing #ai tools on scientific research tasks.
https://sciarena.allen.ai/

You can find its latest results in a July 1 blog post. If you use any AI tools for research purposes, note that the SciArena […]
Original post on fediscience.org
fediscience.org
July 3, 2025 at 1:34 PM
✍️ Learn more in our blog: buff.ly/RjOhOU2
🚀 Visit SciArena to cast your votes: buff.ly/PxYFNyu
💾 Download the dataset: buff.ly/OKATySI
💻 Check out the codebase: buff.ly/nsOOo6U
yale-nlp/SciArena · Datasets at Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
huggingface.co
July 1, 2025 at 3:02 PM