BenchMIRT: What are LLM benchmarks actually measuring?
https://huggingface.co/blog/allenai/benchmir#IA##AI##ML#ML
BenchMIRT: What are LLM benchmarks actually measuring?
https://huggingface.co/blog/allenai/benchmir#IA##AI##ML#ML
We built BenchMIRT to audit them + see which model abilities their Qs actually test. On BBQ, a social-bias eval, it found the Qs distinguished models more by reasoning ability than safety. 🧵
buff.ly/bTcvqJf
We built BenchMIRT to audit them + see which model abilities their Qs actually test. On BBQ, a social-bias eval, it found the Qs distinguished models more by reasoning ability than safety. 🧵
buff.ly/bTcvqJf
We did some explorations with psychometrics-inspired multi-dimensional IRT models and created BenchMIRT to explore these questions!
See 🧵
We built BenchMIRT to audit them + see which model abilities their Qs actually test. On BBQ, a social-bias eval, it found the Qs distinguished models more by reasoning ability than safety. 🧵
buff.ly/bTcvqJf
We did some explorations with psychometrics-inspired multi-dimensional IRT models and created BenchMIRT to explore these questions!
See 🧵
We’re releasing it openly so others can build on it:
💻 buff.ly/FKrFkrC
📄 buff.ly/kMz7Gs8
We’re releasing it openly so others can build on it:
💻 buff.ly/FKrFkrC
📄 buff.ly/kMz7Gs8
Ai2が公開した「BenchMIRT」は、ベンチマークの個別プロンプトを分析し、安全性のつもりが実は推論力などを測っていたという課題をあぶり出す手法です。
普段AIを使っていて「このテスト結果、本当の能力を表してる?」と感じた疑問や体験を、ぜひ一言で教えてください!💬
#AIベンチマーク
Ai2が公開した「BenchMIRT」は、ベンチマークの個別プロンプトを分析し、安全性のつもりが実は推論力などを測っていたという課題をあぶり出す手法です。
普段AIを使っていて「このテスト結果、本当の能力を表してる?」と感じた疑問や体験を、ぜひ一言で教えてください!💬
#AIベンチマーク
https://uenozooo.com/benchmirt-exposes-how-llm-benchmarks-measure-nothing/
#HuggingFace #AI #LLMs
https://uenozooo.com/benchmirt-exposes-how-llm-benchmarks-measure-nothing/
#HuggingFace #AI #LLMs
We didn’t tell it which evals measured what. Two dominant dimensions consistently emerged: general reasoning + safety.
We didn’t tell it which evals measured what. Two dominant dimensions consistently emerged: general reasoning + safety.
◙ For models, it estimates strength on the capabilities reflected in the benchmark set.
◙ For questions, it estimates difficulty & how strongly each Q distinguishes models along those capabilities.
◙ For models, it estimates strength on the capabilities reflected in the benchmark set.
◙ For questions, it estimates difficulty & how strongly each Q distinguishes models along those capabilities.
allenai.org/blog/benchmirt
#AI #GenAI #LLMs
allenai.org/blog/benchmirt
#AI #GenAI #LLMs
HarmBench mostly tests safety behavior, but its copyright questions depend more on reasoning.
XSTest draws on both reasoning & safety, while ToxiGen provides little signal on either.
HarmBench mostly tests safety behavior, but its copyright questions depend more on reasoning.
XSTest draws on both reasoning & safety, while ToxiGen provides little signal on either.
For held-out questions, it correctly predicts whether the model will answer correctly 79% of the time compared with 70% for a simpler benchmark-average baseline.
For held-out questions, it correctly predicts whether the model will answer correctly 79% of the time compared with 70% for a simpler benchmark-average baseline.
The idea: not every question tells you the same amount. Some are harder; some better distinguish stronger models from weaker ones.
The idea: not every question tells you the same amount. Some are harder; some better distinguish stronger models from weaker ones.
We found keeping just the strongest 10% of Qs preserves nearly the same picture of model strengths as using the full set.
We found keeping just the strongest 10% of Qs preserves nearly the same picture of model strengths as using the full set.
AI2が発表した「BenchMIRT」は、ベンチマークの個別のプロンプトやタスクを分析して、AIの本当の能力を監査する手法とのこと。
皆さんは、AIの性能評価って本当にあてになっていると思いますか?🤔
「ここは信用できる」「ここが分かりにくい」など、現場で触っていて感じることを一言で教えてください👇お気軽にどうぞ!
#AI活用
AI2が発表した「BenchMIRT」は、ベンチマークの個別のプロンプトやタスクを分析して、AIの本当の能力を監査する手法とのこと。
皆さんは、AIの性能評価って本当にあてになっていると思いますか?🤔
「ここは信用できる」「ここが分かりにくい」など、現場で触っていて感じることを一言で教えてください👇お気軽にどうぞ!
#AI活用
AI2が発表した「BenchMIRT」は、ベンチマークの個別プロンプトを分析し、何がスコアを動かしているかを可視化する手法ですが、皆さん普段AIの性能テスト結果をどれくらい信じていますか?🤔
「ここはちょっと信用しすぎ注意だよね」「このテストは参考になる!」など、あなたの体感や本音を1つだけ短く教えてください✨お仕事以外でも大歓迎です!
#AIベンチマーク
AI2が発表した「BenchMIRT」は、ベンチマークの個別プロンプトを分析し、何がスコアを動かしているかを可視化する手法ですが、皆さん普段AIの性能テスト結果をどれくらい信じていますか?🤔
「ここはちょっと信用しすぎ注意だよね」「このテストは参考になる!」など、あなたの体感や本音を1つだけ短く教えてください✨お仕事以外でも大歓迎です!
#AIベンチマーク
AI2 於 2026-09-01 發表 BenchMIRT,用多維試題反應理論分析 100 個 LLM、16 個榜單、34,000+ 題。結果自己浮出兩個維度: https://ampm-aiops.com/guides/benchmirt-what-llm-benchmarks-measure-2026/
AI2 於 2026-09-01 發表 BenchMIRT,用多維試題反應理論分析 100 個 LLM、16 個榜單、34,000+ 題。結果自己浮出兩個維度: https://ampm-aiops.com/guides/benchmirt-what-llm-benchmarks-measure-2026/
LLMのテストが「何を図っているか」を個別の質問単位で分析する新手法「BenchMIRT」が発表されました。
AIの性能テストのスコアって、正直あまり信用できないな…と感じる瞬間はありますか?🤔
「ここが気になる」「こうテストしてほしい」など、あなたの実感や一言を気軽に教えてください!✨
#AI開発
LLMのテストが「何を図っているか」を個別の質問単位で分析する新手法「BenchMIRT」が発表されました。
AIの性能テストのスコアって、正直あまり信用できないな…と感じる瞬間はありますか?🤔
「ここが気になる」「こうテストしてほしい」など、あなたの実感や一言を気軽に教えてください!✨
#AI開発
HarmBench mostly tests safety behavior, but its copyright questions depend more on reasoning.
XSTest draws on both reasoning & safety, while ToxiGen provides little signal on either.
Understanding BenchMIRT’s Core Goals BenchMIRT was created to answer a simple yet persistent question: when a language model scor
https://bloggersminds.com/post/benchmirt-explained-what-language-model-benchmarks-really-measure-4665
Understanding BenchMIRT’s Core Goals BenchMIRT was created to answer a simple yet persistent question: when a language model scor
https://bloggersminds.com/post/benchmirt-explained-what-language-model-benchmarks-really-measure-4665
#BenchMIRT #LLMBenchmarks
#BenchMIRT #LLMBenchmarks
Benchmark scores hide what models actually do: BenchMIRT decomposes contaminated evals using item-response theory to isolate true capability signals, cutting question count while catching false confidence in single…
Benchmark scores hide what models actually do: BenchMIRT decomposes contaminated evals using item-response theory to isolate true capability signals, cutting question count while catching false confidence in single…