#MMLU
Is MMLU Western-centric? 🤔

As part of a massive cross-institutional collaboration:
🗽Find MMLU is heavily overfit to western culture
🔍 Professional annotation of cultural sensitivity data
🌍 Release improved Global-MMLU 42 languages

📜 Paper: arxiv.org/pdf/2412.03304
📂 Data: hf.co/datasets/Coh...
December 5, 2024 at 4:31 PM
Announcing Global-MMLU - an improved MMLU Open dataset with evaluation coverage across 42 languages.

The result of months of work with the goal of advancing Multilingual LLM evaluation.

Built together with the community and amazing collaborators at Cohere4AI, MILA, MIT, and many more.
December 6, 2024 at 8:59 AM
Jake Tapper deserves this.
February 27, 2026 at 4:22 AM
We're releasing Mistral Small 3!
- 24B params, 81% MMLU
- Latency optimized: 150 tokens/s
- Competitive with Llama-3.3 70B, Qwen-2.5 32B, GPT4o-mini
- Apache 2.0
mistral.ai/news/mistral...
Mistral Small 3
Apache 2.0, 81% MMLU, 150 tokens/s
mistral.ai
January 30, 2025 at 9:17 PM
ooooooooo that’s real bad 😭 you’re all gonna need to update your local copies of MMLU soon
TIGER-Lab/MMLU-Pro · Leading Whitespace Leaks Correct Choice
It seems like certain categories have an oddity that an empty space is the first character for a subset correct answers. This means that a strategy of random guessing + pick-if-starts-with-space wi...
huggingface.co
January 16, 2026 at 5:58 PM
> Evaluating on language
modeling, MMLU, and GSM8K benchmarks, our method reduces cache miss rates by over
50%, with negligible impact on perplexity (0.1%–3%) and downstream task accuracy (<0.1%)

yup just implemented this and our results are even better lol
arxiv.org/pdf/2412.00099 local model understanders might appreciate this :)
arxiv.org
September 25, 2026 at 4:54 AM
The newest Project Voltage illustration, "39!!", compiles every single Hatsune Miku featured in the Pokémon collaboration thus far. Drawn by mmlu!
March 9, 2025 at 10:03 AM
For clarity -- great project, but most of the MMLU errors we found (and fixed) in our MMLU Redux paper (arxiv.org/abs/2406.04127) are also present in this dataset. We also provide a curated version of MMLU, so it's easy to fix 😊
Announcing Global-MMLU - an improved MMLU Open dataset with evaluation coverage across 42 languages.

The result of months of work with the goal of advancing Multilingual LLM evaluation.

Built together with the community and amazing collaborators at Cohere4AI, MILA, MIT, and many more.
December 6, 2024 at 9:26 AM
MMLU-Redux just touched down at #NAACL2025! 🎉
Wish I could be there for our "Are We Done with MMLU?" poster today (9:00-10:30am in Hall 3, Poster Session 7), but visa drama said nope 😅
If anyone's swinging by, give our research some love! Hit me up if you check it out! 👋
May 2, 2025 at 1:00 PM
Our work on fixing and improving MMLU ("Are We Done with MMLU?", arxiv.org/abs/2406.04127; NAACL 2025) is featured on the DeepSeek home page! deepseek.com -- it's always amazing to see real-world applications and the impact of academic research 🚀🚀🚀
January 28, 2025 at 9:39 AM
January 22, 2026 at 5:27 PM
Roboverse: unified simulation + dataset + benchmarking that supports many different robotics simulators including nvidia omniverse. One step closer to robotics getting its mmlu, maybe.
April 8, 2025 at 2:05 AM
Introducing Global-MMLU🌍: A multilingual benchmark featuring MMLU translations in 42 languages crafted with:
✅ Human curation
✅ Extensive metadata
✅ Insights into cultural sensitivity

Proud to have collaborated with Shivalika Singh, @sarahooker.bsky.social and Cohere For AI!
Is MMLU Western-centric? 🤔

As part of a massive cross-institutional collaboration:
🗽Find MMLU is heavily overfit to western culture
🔍 Professional annotation of cultural sensitivity data
🌍 Release improved Global-MMLU 42 languages

📜 Paper: arxiv.org/pdf/2412.03304
📂 Data: hf.co/datasets/Coh...
December 5, 2024 at 7:46 PM
Major AI Large Language Models (LLMs) ranked by capabilities (MMLU)
interactive https://informationisbeautiful.net/visualizations/the-rise-of-generative-ai-large-language-models-llms-like-chatgpt/
November 6, 2024 at 3:55 PM
Garbage in, garbage out -- nice gem for the Italian-speaking folks on this platform 😅 TLDR, in arxiv.org/abs/2406.04127 we found that MMLU contains TONS of errors, and looks like all these seamlessly propagated to this new "Global MMLU" dataset
December 6, 2024 at 1:17 PM
What is MMLU? It’s a benchmark used to evaluate how well AI models handle knowledge and reasoning across a wide range of subjects.

MMLU helps compare AI performance beyond simple language generation.
nyvoraai.github.io/ai-news/what...

#MMLU #LLM #AIModels #ArtificialIntelligence #MachineLearning
August 3, 2026 at 3:00 PM
Along with Baguettotron we release the smallest viable language model to date. Monad, a 56M transformer, trained on the English part of SYNTH with non-random performance on MMLU. Desiging Monad an engineering challenge requiring a custom tiny tokenizer. huggingface.co/PleIAs/Monad
November 10, 2025 at 5:33 PM
Since SYNTH has been designed to train for reasoning capacities, we get actual reasoning signals very early in training. For Baguettotron, we find that MMLU starts to get non-random after less than 10 billion tokens and quickly achieve near-SOTA performance.
November 10, 2025 at 5:32 PM
Translating MMLU is great, but global users of multilingual #LLMs don't care all that much about an LLM's understanding of US Law!

Our new #NLProc work centers multilingual #LLM evaluations toward regional knowledge in 44 languages.
🚀 Introducing INCLUDE 🌍: A multilingual LLM evaluation benchmark spanning 44 languages!

Contains *newly-collected* data, prioritizing *regional knowledge*.
Setting the stage for truly global AI evaluation.
Ready to see how your model measures up?
#AI #Multilingual #LLM #NLProc
December 2, 2024 at 4:26 PM
Clever test of AI reasoning ability adds the option "none of these" to the common MMLU benchmark, forcing the AI to consider options rather than just picking the best

The result is a big drop in accuracy for most models, though Reasoners (o3 & DeepSeek) hold up much better arxiv.org/pdf/2502.12896
February 20, 2025 at 4:34 PM
You would think moral questions are universal,
MMLU only asks about US morals...
better translations and separation by sensitivity
👇👇
Is MMLU Western-centric? 🤔

As part of a massive cross-institutional collaboration:
🗽Find MMLU is heavily overfit to western culture
🔍 Professional annotation of cultural sensitivity data
🌍 Release improved Global-MMLU 42 languages

📜 Paper: arxiv.org/pdf/2412.03304
📂 Data: hf.co/datasets/Coh...
December 5, 2024 at 7:52 PM
I’ll be travelling to London from Wednesday to Friday for an upcoming event and would be very happy to meet up! 🚀
I'd love to chat about my recent works (DeCoRe, MMLU-Redux, etc.). DM me if you’re around! 👋

DeCoRe: arxiv.org/abs/2410.18860
MMLU-Redux: arxiv.org/abs/2406.04127
November 18, 2024 at 1:48 PM
our new Olmo Hybrid model combines attention with linear RNN layers

🍣training efficiency is crazy good. the model reaches same MMLU score as Olmo 3 in 50% of the tokens. also see this in many other tasks

as always: weights, data, ckpts, training code, etc. all fully open
March 5, 2026 at 6:47 PM
I think @natolambert.bsky.social this means Ai2 run club would get us a few more MMLU points?..
November 21, 2024 at 12:41 AM
We created SuperBPE🚀, a *superword* tokenizer that includes tokens spanning multiple words.

When pretraining at 8B scale, SuperBPE models consistently outperform the BPE baseline on 30 downstream tasks (+8% MMLU), while also being 27% more efficient at inference time.🧵
March 21, 2025 at 4:48 PM