go.nature.com/4lzaWlC
go.nature.com/4lzaWlC
🧪 Learn more and try SciArena here:
buff.ly/J2kYM0Q
🧪 Learn more and try SciArena here:
buff.ly/J2kYM0Q
It's retiring July 15, and the results are in: ~1,700 users cast ~3,900 votes.
Here's what they told us. 🧵
It's retiring July 15, and the results are in: ~1,700 users cast ~3,900 votes.
Here's what they told us. 🧵
go.nature.com/44FcZ0i
go.nature.com/44FcZ0i
📄 buff.ly/A9OrLph
📄 buff.ly/A9OrLph
We've added new frontier models – including GPT-5.1 and Gemini 3 Pro Preview – to our arena for scientific literature tasks. The new rankings: o3 holds #1, Gemini 3 Pro Preview lands at #2, Claude Opus 4.1 sits at #3, GPT-5 at #4, & GPT-5.1 debuts at #5. 🧵
We've added new frontier models – including GPT-5.1 and Gemini 3 Pro Preview – to our arena for scientific literature tasks. The new rankings: o3 holds #1, Gemini 3 Pro Preview lands at #2, Claude Opus 4.1 sits at #3, GPT-5 at #4, & GPT-5.1 debuts at #5. 🧵
With thousands of new votes, new LLMs are reshaping our leaderboard for scientific literature tasks.
o3 still leads—but GPT-5, Claude Opus 4.1, & more are closing the gap.
With thousands of new votes, new LLMs are reshaping our leaderboard for scientific literature tasks.
o3 still leads—but GPT-5, Claude Opus 4.1, & more are closing the gap.
#chatbot #chatgpt #claude #deepseek #gemini #intelligenzaartificiale #sciarena
guruhitech.com/sciarena-la-...
#chatbot #chatgpt #claude #deepseek #gemini #intelligenzaartificiale #sciarena
guruhitech.com/sciarena-la-...
🧪
www.nature.com/articles/d41...
🧪
www.nature.com/articles/d41...
1️⃣ SciArena platform: Where anyone can submit questions, vote on their preferred output, and see how models rank.
2️⃣ Leaderboard: A ranking system based on community votes.
3️⃣ SciArena-Eval: A meta-evaluation benchmark from human preference data.
1️⃣ SciArena platform: Where anyone can submit questions, vote on their preferred output, and see how models rank.
2️⃣ Leaderboard: A ranking system based on community votes.
3️⃣ SciArena-Eval: A meta-evaluation benchmark from human preference data.
Both SciArena and SciArena-Eval are now available.
Both SciArena and SciArena-Eval are now available.
Learn more in our updated blog: buff.ly/o1EyhhG
Learn more in our updated blog: buff.ly/o1EyhhG
This is a unique resource & especially important as AI agents are increasingly used to judge the quality of other AI agents, and such evaluations need to be grounded + verified.
This is a unique resource & especially important as AI agents are increasingly used to judge the quality of other AI agents, and such evaluations need to be grounded + verified.
https://sciarena.allen.ai/
You can find its latest results in a July 1 blog post. If you use any AI tools for research purposes, note that the SciArena […]
https://sciarena.allen.ai/
You can find its latest results in a July 1 blog post. If you use any AI tools for research purposes, note that the SciArena […]
🚀 Visit SciArena to cast your votes: buff.ly/PxYFNyu
💾 Download the dataset: buff.ly/OKATySI
💻 Check out the codebase: buff.ly/nsOOo6U
🚀 Visit SciArena to cast your votes: buff.ly/PxYFNyu
💾 Download the dataset: buff.ly/OKATySI
💻 Check out the codebase: buff.ly/nsOOo6U