#hellaswag
New Pleias paper: "What the HellaSwag?"
HellaSwag is currently on of the most widely LLM benchmarks in the world. We introduce a new critical method to assess the validity of standard LLM evals and show it does not accurately measure common sense reasoning. arxiv.org/abs/2504.07825
April 14, 2025 at 3:44 PM
September 20, 2026 at 2:49 AM
today's grok did not seem like the sort of model that is extremely good at hellaswag, no
July 9, 2025 at 2:53 AM
My apologies ChatGPT, I was unfamiliar with your HellaSwag
November 19, 2024 at 4:18 PM
We are proud to announce that we trained 1.5B, 8B, and 24B generative language models from scratch on 2 to 4 tera-tokens of carefully curated, high-quality data covering French, English and code. We release our models and code under open-source licences. Thread👇
November 12, 2025 at 5:05 PM
Blows my mind that model souping Just Works™️

Same model, same data, train 3-5 times with different seeds, 1-2 extra points on MMLU, Hellaswag, ARC, GSM8k, etc
November 4, 2024 at 9:08 PM
sooo it's the same? #hellaswag #positivity
August 8, 2026 at 6:02 AM
She raised someone w #hellaswag
September 10, 2024 at 3:16 PM
machine learning is neat, but every company has a stupid dumb joke name I hate it. a tool called HellaSwag is awful! be professional you weirdos
July 31, 2025 at 3:32 PM
Neat, I’m only seven years behind!
September 16, 2026 at 2:14 AM
MMLU: Does not measure reasoning ability.

Hellaswag: Does not measure reasoning ability.

ARC: Does not measure reasoning ability.

GPQA: Does not measure reasoning ability.

Did you have others?
December 20, 2024 at 8:11 PM
we don’t talk enough about how stupid machine learning jargon is. like you can’t expect me to take your results on a benchmark called ‘HellaSwag’ seriously?
August 15, 2024 at 9:32 AM
One of the most big AI benchmarks is called HellaSwag. Just when you think this field can't be any more cringe inducing...
December 7, 2023 at 10:04 AM
HellaSwag has a typical messy story: a merge of ActivityNet (ultimately Youtube transcripts) and WikiHow supplemented with synthetic generations by the… first GPT model (not even GPT-2).
April 14, 2025 at 3:45 PM
LLaDA-2.0: the largest text diffusion model ever

- 100B-A6B MoE architecture
- 535 tok/s
- Competitive with Qwen3-30B-A3B

🤔🤔🤔

huggingface.co/inclusionAI/...
December 13, 2025 at 2:33 PM
No, buddy, they don't. MMLU measures "knowledge acquired in pretraining." Hellaswag measures "the ability to predict the ending of a narrative from multiple choice answers." ARC claims to measure "skill acquisition," but based on the examples, I would question their use of the term "skill."
December 20, 2024 at 8:25 PM
I feel pity for this model trained on material only up to 1914, as it struggles to answer HellaSwag questions based on the 21c. It's getting decent validation loss for its size (on pre-1915 text) but has zero chance answering eval questions about smartphones etc. (Sample question in next tweet) +
December 6, 2024 at 8:26 PM
Este paper cuestiona un benchmark central en NLP. HellaSwag ha sido ampliamente utilizado en ránkings y comparativas de modelos. Si los análisis de HellaSwag no son correctos, las decisiones sobre qué modelos son “mejores” podrían estar basadas en evaluaciones defectuosas.

arxiv.org/abs/2504.078...
What the HellaSwag? On the Validity of Common-Sense Reasoning Benchmarks
Common-sense reasoning is a key language model capability because it encapsulates not just specific factual knowledge but rather general language and world understanding. Measuring common-sense reason...
arxiv.org
April 14, 2025 at 4:16 PM
They found that
- DLMs beat AR on same number of tokens
- A 1B DLM trained on just 1B tokens hits 56% HellaSwag & 33% MMLU — no tricks, no cherry-picks.
- No saturation: more epochs = more gains
August 9, 2025 at 11:09 PM
venturebeat.com/ai/ai21-labs...

Lots to look forward too with this new (bona fide!) open source hashtag#GenAI model from AI21 Labs. Because ...Attention isn't all you need (pun intended).
AI21 Labs juices up gen AI transformers with Jamba
According to AI21 Labs Jamba can outperform traditional transformer-based models on generative reasoning tasks as measured by benchmarks such as HellaSwag.
venturebeat.com
March 28, 2024 at 8:24 PM
Concretely, this means many HellaSwag prompts are badly structured, even outright inconsistent and even when they make sense, do not yield to one obvious good answer. Hence, do good results require proper understanding/parsing of the prompt sentence?
April 14, 2025 at 3:45 PM
Yeah but clearly that was all planted evidence by Obama, and the children were all actually transdimensional Democrat demons, can't you people see what a hoax it all was? Azealia Banks told me raping minors is hellaswag so wtf do you people even know 😒
July 29, 2025 at 1:24 AM
🎯 Picking the right benchmark to hillclimb on makes all the difference.
Some tasks are much more sensitive at small scale:
💡 MMLU & ARC Easy give strong signals earlier than HellaSwag
🚫 Others stay hard to predict—even at larger scales
April 15, 2025 at 1:02 PM
HellaSwag and illustrate them with various evaluations using generative language models of different sizes. We argue that this benchmark does not accurately measure common-sense reasoning and, therefore, should not be used for evaluation in its [6/8 of https://arxiv.org/abs/2504.07825v1]
April 11, 2025 at 6:02 AM