#BixBench
Introducing BixBench, a benchmark for AI agents in bioinformatics, built with ScienceMachine. We've created 53 scenarios with 296 questions testing AI on computational biology challenges. BixBench includes evaluation metrics and an open-source LLM environment.
March 4, 2025 at 3:15 PM
#arXiv BixBench: a Comprehensive Benchmark for LLM-based Agents in Computational Biology arxiv.org/abs/2503.00096 over 50 real-world scenarios of practical biological data analysis with nearly 300 associated open-answer questions designed to measure the ability of LLM-based agents.
March 4, 2025 at 5:07 PM
single-cell is a fun field. for instance, one of the heavily curated bixbench scenarios is about interpreting the results of a sc analysis and comparing to ground truth. this ground truth is, of course, based on DE analyses with some truly remarkable p-values for n=5
March 6, 2025 at 2:44 AM
💡 How do you judge an autonomous AI data analyst? We revisited BixBench (github.com/FUture-House...), an existing benchmark where AI agents analyse real biomedical datasets, & rescored its answers with Karenina.
September 7, 2026 at 1:23 PM
🚀 OpenAI представила GPT-Rosalind – ИИ-модель для ускорения разработки лекарств! 💊 Цель – сократить исследования с 10-15 лет. Модель показала лучшие результаты на BixBench и LABBench2. Доступна корпоративным клиентам в США. #AI #Pharma #DeSci
Cryptovka
CryptoMarket and Blockchain News
cryptovka.ru
April 18, 2026 at 1:03 PM
Equips Claude Code with 199 bioinformatics skills for RNA-seq, single-cell analysis, and drug discovery. Boosted BixBench scores from 65% to 92%.
July 13, 2026 at 9:16 PM
17% accuracy in the open-answer regime, and no better than random in a multiple-choice setting. By exposing the current limitations of frontier models, we hope BixBench can spur the development of agents capable of [6/7 of https://arxiv.org/abs/2503.00096v1]
March 4, 2025 at 6:18 AM
may be near, but no extensive benchmarks for measuring progress have been introduced to date. We therefore present the Bioinformatics Benchmark (BixBench), a dataset comprising over 50 real-world scenarios of practical [3/7 of https://arxiv.org/abs/2503.00096v1]
March 4, 2025 at 6:18 AM
Mitchener, Laurent, Tenmann, Narayanan, Wellawatte, White, Sani, Rodriques: BixBench: a Comprehensive Benchmark for LLM-based Agents in Computational Biology https://arxiv.org/abs/2503.00096 https://arxiv.org/pdf/2503.00096 https://arxiv.org/html/2503.00096
March 4, 2025 at 6:17 AM
[2025-03-11] 📚 Updates in #ELM

(1) BixBench: a Comprehensive Benchmark for LLM-based Agents in Computational Biology
(2) <a href="https://researchtrend.ai/papers/2503.05788" class="hover:underline text-blue-600 dark:text-sky-400 no-card-link" target="_blank" rel="noopener" data-link="bsky">Emergent Abilities in Large Language Models: A Survey
(3) Emergent Abilities in Large Language Models: A Survey

🔍 More at researchtrend.ai/communities/ELM
March 11, 2025 at 3:51 AM
[2025-03-11] 📚 Updates in #LLMAG

(1) BixBench: a Comprehensive Benchmark for LLM-based Agents in Computational Biology
(2) <a href="https://researchtrend.ai/papers/2503.05944" class="hover:underline text-blue-600 dark:text-sky-400 no-card-link" target="_blank" rel="noopener" data-link="bsky">Enhancing Reasoning with Collaboration and Memory
(3) Enhancing Reasoning with Collaboration and Memory

🔍 More at researchtrend.ai/communities/LLMAG
March 11, 2025 at 3:51 AM
[2025-03-11] 📚 Updates in #LM&MA

(1) BixBench: a Comprehensive Benchmark for LLM-based Agents in Computational Biology
(2) <a href="https://researchtrend.ai/papers/2503.06074" class="hover:underline text-blue-600 dark:text-sky-400 no-card-link" target="_blank" rel="noopener" data-link="bsky">Towards Conversational AI for Disease Management
(3) Towards Conversational AI for Disease Management

🔍 More at researchtrend.ai/communities/LM&MA
March 11, 2025 at 3:51 AM
BixBench: a Comprehensive Benchmark for LLM-based Agents in Computational Biology
Large Language Models (LLMs) and LLM-based agents show great promise in accelerating scientific research. Existing benchmarks for measuring this potential and guiding future development continue to evolve from pure recall and rote knowledge tasks, towards more practical work such as literature review and experimental planning. Bioinformatics is a domain where fully autonomous AI-driven discovery may be near, but no extensive benchmarks for measuring progress have been introduced to date. We therefore present the Bioinformatics Benchmark (BixBench), a dataset comprising over 50 real-world scenarios of practical biological data analysis with nearly 300 associated open-answer questions designed to measure the ability of LLM-based agents to explore biological datasets, perform long, multi-step analytical trajectories, and interpret the nuanced results of those analyses. We evaluate the performance of two frontier LLMs (GPT-4o and Claude 3.5 Sonnet) using a custom agent framework we open source. We find that even the latest frontier models only achieve 17% accuracy in the open-answer regime, and no better than random in a multiple-choice setting. By exposing the current limitations of frontier models, we hope BixBench can spur the development of agents capable of conducting rigorous bioinformatic analysis and accelerate scientific discovery.
arxiv.org
October 10, 2025 at 4:25 AM
BixBench: a Comprehensive Benchmark for LLM-based Agents in Computational Biology
Large Language Models (LLMs) and LLM-based agents show great promise in accelerating scientific research. Existing benchmarks for measuring this potential and guiding future development continue to evolve from pure recall and rote knowledge tasks, towards more practical work such as literature review and experimental planning. Bioinformatics is a domain where fully autonomous AI-driven discovery may be near, but no extensive benchmarks for measuring progress have been introduced to date. We therefore present the Bioinformatics Benchmark (BixBench), a dataset comprising over 50 real-world scenarios of practical biological data analysis with nearly 300 associated open-answer questions designed to measure the ability of LLM-based agents to explore biological datasets, perform long, multi-step analytical trajectories, and interpret the nuanced results of those analyses. We evaluate the performance of two frontier LLMs (GPT-4o and Claude 3.5 Sonnet) using a custom agent framework we open source. We find that even the latest frontier models only achieve 17% accuracy in the open-answer regime, and no better than random in a multiple-choice setting. By exposing the current limitations of frontier models, we hope BixBench can spur the development of agents capable of conducting rigorous bioinformatic analysis and accelerate scientific discovery.
arxiv.org
October 10, 2025 at 4:25 AM
BixBench: a Comprehensive Benchmark for LLM-based Agents in Computational Biology
Large Language Models (LLMs) and LLM-based agents show great promise in accelerating scientific research. Existing benchmarks for measuring this potential and guiding future development continue to evolve from pure recall and rote knowledge tasks, towards more practical work such as literature review and experimental planning. Bioinformatics is a domain where fully autonomous AI-driven discovery may be near, but no extensive benchmarks for measuring progress have been introduced to date. We therefore present the Bioinformatics Benchmark (BixBench), a dataset comprising over 50 real-world scenarios of practical biological data analysis with nearly 300 associated open-answer questions designed to measure the ability of LLM-based agents to explore biological datasets, perform long, multi-step analytical trajectories, and interpret the nuanced results of those analyses. We evaluate the performance of two frontier LLMs (GPT-4o and Claude 3.5 Sonnet) using a custom agent framework we open source. We find that even the latest frontier models only achieve 17% accuracy in the open-answer regime, and no better than random in a multiple-choice setting. By exposing the current limitations of frontier models, we hope BixBench can spur the development of agents capable of conducting rigorous bioinformatic analysis and accelerate scientific discovery.
arxiv.org
March 11, 2025 at 4:22 AM
BixBench: a Comprehensive Benchmark for LLM-based Agents in Computational Biology
Large Language Models (LLMs) and LLM-based agents show great promise in accelerating scientific research. Existing benchmarks for measuring this potential and guiding future development continue to evolve from pure recall and rote knowledge tasks, towards more practical work such as literature review and experimental planning. Bioinformatics is a domain where fully autonomous AI-driven discovery may be near, but no extensive benchmarks for measuring progress have been introduced to date. We therefore present the Bioinformatics Benchmark (BixBench), a dataset comprising over 50 real-world scenarios of practical biological data analysis with nearly 300 associated open-answer questions designed to measure the ability of LLM-based agents to explore biological datasets, perform long, multi-step analytical trajectories, and interpret the nuanced results of those analyses. We evaluate the performance of two frontier LLMs (GPT-4o and Claude 3.5 Sonnet) using a custom agent framework we open source. We find that even the latest frontier models only achieve 17% accuracy in the open-answer regime, and no better than random in a multiple-choice setting. By exposing the current limitations of frontier models, we hope BixBench can spur the development of agents capable of conducting rigorous bioinformatic analysis and accelerate scientific discovery.
arxiv.org
March 4, 2025 at 6:05 AM
2. 科研与安全突破
•学术前沿:
◦发现拉姆齐数新证明,GeneBench遗传学分析超越GPT-5.4,BixBench生物信息学领先;
◦能处理歧义数据、识别混杂因素,相当于专家数天工作量。
•安全防护:
◦新增网络安全/生物学红队测试,200+真实场景验证,号称"最强安全框架"
April 24, 2026 at 12:31 PM
Skill-Augmented Frontier Agents Nearly Saturate BixBench-Verified-50 https://www.biorxiv.org/content/10.64898/2026.04.28.721523v1
May 1, 2026 at 7:47 PM
OpenAI just dropped GPT‑Rosalind, and it's crushing the BixBench benchmark with a record‑breaking score. Curious how this new model stacks up? Dive into the details and see why the AI community is buzzing. #GPTRosalind #BixBench #OpenAIResearch

🔗 aidailypost.com/news/openai-...
April 16, 2026 at 7:35 PM
AI news of the day · 18.05.26
OpenAI GPT-Rosalind: Inside the 0.751 BixBench Score and $2.6B Pharma Bet [2026] - tech-insider.org
OpenAI GPT-Rosalind: Inside the 0.751 BixBench Score and $2.6B Pharma Bet [2026] &n...
Subnewsletter
#ProteinDesign #StructuralBiology #Bioinformatics #ProteinEngineering
OpenAI GPT-Rosalind: Inside the 0.751 BixBench Score and $2.6B Pharma Bet [2026] - tech-insider.org
OpenAI GPT-Rosalind: Inside the 0.751 BixBench Score and $2.6B Pharma Bet [2026] &nbsp;&nbsp; tech-insider.org
news.google.com
May 18, 2026 at 8:00 AM
• Global SOTA on Data Analysis Benchmarks: BixBench 48.78% open-answer, 55.12% multiple-choice + refusal, 64.39% multiple-choice (no refusal) - outperforming systems like Edison Scientific and Kepler.
February 6, 2026 at 3:33 PM
• Global SOTA on Data Analysis Benchmarks: BixBench 48.78% open-answer, 55.12% multiple-choice + refusal, 64.39% multiple-choice (no refusal) - outperforming systems like Edison Scientific and Kepler.
February 6, 2026 at 3:26 PM
Skill-Augmented Frontier Agents Nearly Saturate BixBench-Verified-50 https://www.biorxiv.org/content/10.64898/2026.04.28.721523v1
May 1, 2026 at 7:47 PM