#BenchMIRT
BenchMIRT method breaks down LLM benchmark scores to unveil underlying capabilities. 🔍🧠 #LLM #Benchmarks #Research
BenchMIRT: What are LLM benchmarks actually measuring?
A Blog post by Ai2 on Hugging Face
huggingface.co
September 26, 2026 at 8:39 AM
📰 Hugging Face
BenchMIRT: What are LLM benchmarks actually measuring?

https://huggingface.co/blog/allenai/benchm
ir#IA##AI##ML#ML
September 25, 2026 at 11:01 PM
Do LLM safety & capability evals measure what they claim to?

We built BenchMIRT to audit them + see which model abilities their Qs actually test. On BBQ, a social-bias eval, it found the Qs distinguished models more by reasoning ability than safety. 🧵

buff.ly/bTcvqJf
September 1, 2026 at 9:52 PM
Is the safety-capability tradeoff for LLMs real? Or could it be an artefact of the benchmarks we use to measure safety??

We did some explorations with psychometrics-inspired multi-dimensional IRT models and created BenchMIRT to explore these questions!

See 🧵
Do LLM safety & capability evals measure what they claim to?

We built BenchMIRT to audit them + see which model abilities their Qs actually test. On BBQ, a social-bias eval, it found the Qs distinguished models more by reasoning ability than safety. 🧵

buff.ly/bTcvqJf
September 2, 2026 at 12:44 PM
BenchMIRT gives researchers a clearer view of what benchmarks really measure—and could help build evals that are smaller, more focused, & easier to interpret.

We’re releasing it openly so others can build on it:

💻 buff.ly/FKrFkrC
📄 buff.ly/kMz7Gs8
GitHub - allenai/BenchMIRT
Contribute to allenai/BenchMIRT development by creating an account on GitHub.
github.com
September 1, 2026 at 9:53 PM
LLMの性能テスト(ベンチマーク)にモヤモヤしたことある人に聞きたいです!🤖
Ai2が公開した「BenchMIRT」は、ベンチマークの個別プロンプトを分析し、安全性のつもりが実は推論力などを測っていたという課題をあぶり出す手法です。
普段AIを使っていて「このテスト結果、本当の能力を表してる?」と感じた疑問や体験を、ぜひ一言で教えてください!💬
#AIベンチマーク
September 26, 2026 at 10:01 PM
September 19, 2026 at 8:00 PM
We trained BenchMIRT on results from 100 LLMs across 16 benchmarks & 34K+ questions.

We didn’t tell it which evals measured what. Two dominant dimensions consistently emerged: general reasoning + safety.
September 1, 2026 at 9:53 PM
BenchMIRT works at multiple levels:

◙ For models, it estimates strength on the capabilities reflected in the benchmark set.

◙ For questions, it estimates difficulty & how strongly each Q distinguishes models along those capabilities.
September 1, 2026 at 9:53 PM
BenchMIRT: A "New Method For Auditing LLM #Benchmarks at the Level of Individual #Prompts—The Questions and Tasks a Model is Scored On" (via @ai2.bsky.social)

allenai.org/blog/benchmirt

#AI #GenAI #LLMs
September 2, 2026 at 4:54 PM
We used BenchMIRT to audit popular LLM evals—and found some quirks.

HarmBench mostly tests safety behavior, but its copyright questions depend more on reasoning.

XSTest draws on both reasoning & safety, while ToxiGen provides little signal on either.
September 1, 2026 at 9:53 PM
BenchMIRT can also estimate how a model will perform on Qs it hasn’t answered.

For held-out questions, it correctly predicts whether the model will answer correctly 79% of the time compared with 70% for a simpler benchmark-average baseline.
September 1, 2026 at 9:53 PM
BenchMIRT builds on Item Response Theory (IRT), a technique from psychometrics for measuring abilities from patterns of test responses.

The idea: not every question tells you the same amount. Some are harder; some better distinguish stronger models from weaker ones.
September 1, 2026 at 9:52 PM
We applied BenchMIRT across the 16 benchmarks it was trained on to see whether we could make evals more efficient by removing less informative questions.

We found keeping just the strongest 10% of Qs preserves nearly the same picture of model strengths as using the full set.
September 1, 2026 at 9:53 PM
BenchMIRTが問うLLMベンチマークの本質:何を測定し、何を誤解しているのか?

LLMベンチマークの測定対象と限界を深く掘り下げる。

#LLM評価 #ベンチマーク #AI開発 #性能測定 #MIRT
BenchMIRTが問うLLMベンチマークの本質:何を測定し、何を誤解しているのか?
LLMベンチマークの測定対象と限界を深く掘り下げる。
ai.warp-studio.com
September 1, 2026 at 11:40 PM
AIの性能テスト(ベンチマーク)にモヤモヤしたことある人に聞きたいです!!🧠
AI2が発表した「BenchMIRT」は、ベンチマークの個別のプロンプトやタスクを分析して、AIの本当の能力を監査する手法とのこと。
皆さんは、AIの性能評価って本当にあてになっていると思いますか?🤔
「ここは信用できる」「ここが分かりにくい」など、現場で触っていて感じることを一言で教えてください👇お気軽にどうぞ!
#AI活用
September 10, 2026 at 2:33 PM
LLMの性能評価(ベンチマーク)にモヤモヤしたことがある人に聞きたいです!!🧠
AI2が発表した「BenchMIRT」は、ベンチマークの個別プロンプトを分析し、何がスコアを動かしているかを可視化する手法ですが、皆さん普段AIの性能テスト結果をどれくらい信じていますか?🤔
「ここはちょっと信用しすぎ注意だよね」「このテストは参考になる!」など、あなたの体感や本音を1つだけ短く教えてください✨お仕事以外でも大歓迎です!
#AIベンチマーク
September 19, 2026 at 2:00 PM
Are your LLM benchmarks actually measuring what you think? AllenAI's BenchMIRT reveals surprising mismatches, like safety benchmarks that mostly test reasoning. A must-read for anyone building evals. https://huggingface.co/blog/allenai/benchmirt
September 14, 2026 at 6:03 AM
🐾 AI2 拆了 34,000 道題發現:測「社會偏見」的榜單,其實在測推理能力(BenchMIRT,2026-09-01)
AI2 於 2026-09-01 發表 BenchMIRT,用多維試題反應理論分析 100 個 LLM、16 個榜單、34,000+ 題。結果自己浮出兩個維度: https://ampm-aiops.com/guides/benchmirt-what-llm-benchmarks-measure-2026/
September 10, 2026 at 6:05 PM
AIの性能テスト(ベンチマーク)について、現場でモヤモヤしている人に聞きたいです!💻
LLMのテストが「何を図っているか」を個別の質問単位で分析する新手法「BenchMIRT」が発表されました。
AIの性能テストのスコアって、正直あまり信用できないな…と感じる瞬間はありますか?🤔
「ここが気になる」「こうテストしてほしい」など、あなたの実感や一言を気軽に教えてください!✨

#AI開発
September 9, 2026 at 10:01 PM
BenchMIRT catching this is the real finding: HarmBench's copyright questions load on reasoning ability, not safety, so the benchmark is testing the wrong construct for part of its own item bank. Same trap as scoring hybrid search on topical match instead of what you actually needed.
We used BenchMIRT to audit popular LLM evals—and found some quirks.

HarmBench mostly tests safety behavior, but its copyright questions depend more on reasoning.

XSTest draws on both reasoning & safety, while ToxiGen provides little signal on either.
September 6, 2026 at 2:43 PM
BenchMIRT Explained: What Language Model Benchmarks Really Measure

Understanding BenchMIRT’s Core Goals BenchMIRT was created to answer a simple yet persistent question: when a language model scor

https://bloggersminds.com/post/benchmirt-explained-what-language-model-benchmarks-really-measure-4665
BenchMIRT Explained: What Language Model Benchmarks Really Measure
Understanding BenchMIRT’s Core Goals BenchMIRT was created to answer a simple yet persistent question: when a language model scores well on a benchmark, what does that score actually reflect? The project gathers a wide range of tasks, from factual recall to complex reasoning, and analyses how each task correlates with underlying model abilities.
bloggersminds.com
September 2, 2026 at 7:45 PM
LLM benchmarks are NOT what they seem. A "safety" score might be pure reasoning! Don't be fooled by aggregated numbers. Deconstruct benchmarks to match your product's *true* needs. Misunderstand your AI's foundation, and you'll fail at scale.

#BenchMIRT #LLMBenchmarks
September 2, 2026 at 9:49 AM
BenchMIRT: What are LLM benchmarks actually measuring?
BenchMIRT: What are LLM benchmarks actually measuring?
BenchMIRT: Unpacking What LLM Benchmarks Really Measure Large Language Models (LLMs) are at the forefront of AI innovation, constantly pushing boundaries in natural language understanding and generati...
nerdstool.com
September 2, 2026 at 12:01 AM
In today's edition — Claude Opus 4.8 vs Mistral Large:

Benchmark scores hide what models actually do: BenchMIRT decomposes contaminated evals using item-response theory to isolate true capability signals, cutting question count while catching false confidence in single…
BenchMIRT: What are LLM benchmarks actually measuring?
Two AI summaries, read blind — pick the better
www.snipvote.com
September 2, 2026 at 7:25 AM