HellaSwag is currently on of the most widely LLM benchmarks in the world. We introduce a new critical method to assess the validity of standard LLM evals and show it does not accurately measure common sense reasoning. arxiv.org/abs/2504.07825
HellaSwag is currently on of the most widely LLM benchmarks in the world. We introduce a new critical method to assess the validity of standard LLM evals and show it does not accurately measure common sense reasoning. arxiv.org/abs/2504.07825
Same model, same data, train 3-5 times with different seeds, 1-2 extra points on MMLU, Hellaswag, ARC, GSM8k, etc
Same model, same data, train 3-5 times with different seeds, 1-2 extra points on MMLU, Hellaswag, ARC, GSM8k, etc
Hellaswag: Does not measure reasoning ability.
ARC: Does not measure reasoning ability.
GPQA: Does not measure reasoning ability.
Did you have others?
Hellaswag: Does not measure reasoning ability.
ARC: Does not measure reasoning ability.
GPQA: Does not measure reasoning ability.
Did you have others?
- 100B-A6B MoE architecture
- 535 tok/s
- Competitive with Qwen3-30B-A3B
🤔🤔🤔
huggingface.co/inclusionAI/...
- 100B-A6B MoE architecture
- 535 tok/s
- Competitive with Qwen3-30B-A3B
🤔🤔🤔
huggingface.co/inclusionAI/...
arxiv.org/abs/2504.078...
arxiv.org/abs/2504.078...
- DLMs beat AR on same number of tokens
- A 1B DLM trained on just 1B tokens hits 56% HellaSwag & 33% MMLU — no tricks, no cherry-picks.
- No saturation: more epochs = more gains
- DLMs beat AR on same number of tokens
- A 1B DLM trained on just 1B tokens hits 56% HellaSwag & 33% MMLU — no tricks, no cherry-picks.
- No saturation: more epochs = more gains
Lots to look forward too with this new (bona fide!) open source hashtag#GenAI model from AI21 Labs. Because ...Attention isn't all you need (pun intended).
Lots to look forward too with this new (bona fide!) open source hashtag#GenAI model from AI21 Labs. Because ...Attention isn't all you need (pun intended).
Some tasks are much more sensitive at small scale:
💡 MMLU & ARC Easy give strong signals earlier than HellaSwag
🚫 Others stay hard to predict—even at larger scales
Some tasks are much more sensitive at small scale:
💡 MMLU & ARC Easy give strong signals earlier than HellaSwag
🚫 Others stay hard to predict—even at larger scales