#MMMLU
Does a 2026-era 12B speak good German yet? Not really.

Gemma 4 12B (just released), German-only:
INCLUDE 64.0% · MMMLU 75.8% - bottom of both boards.
Its 31B sibling: 67.6 / 86.4. Frontier MaaS: up to 72.7 / 89.3.

12B is still too small for frontier German.

dach.peerbench.ai
German Artificial Analytics — Which fast LLM speaks German best?
German LLM leaderboard — benchmarks rerun on the newest frontier models, German first.
dach.peerbench.ai
June 4, 2026 at 3:33 PM
DeepSeek V4 Flash is quietly strong on German.

85.1% on MMMLU (German), 14k items — and the full run cost just $0.92.

Cheap + strong = a great base to fine-tune.

Leaderboard: dach.peerbench.ai
#DeepSeek #LLM #GermanAI #Benchmark
German Artificial Analytics — Which fast LLM speaks German best?
German LLM leaderboard — benchmarks rerun on the newest frontier models, German first.
dach.peerbench.ai
June 5, 2026 at 10:46 AM
🇩🇪 How well do open models actually speak German?
(MMLU-ProX-DE, MMMLU-DE, INCLUDE-DE) reasoning-off,
MMLU-ProX (DE):
• Qwen3.6 35B-A3B — 80.0% (quant unverified)
• Gemma 4 26B (bf16) — 78.2%
• Gemma 4 12B (bf16) — 69.4%
• Qwen3 14B (bf16) — 64.5%
dach.peerbench.ai/compare?mode...
Compare models — German Artificial Analytics
Head-to-head German LLM benchmarks: pick 2–4 models and compare scores, subjects, speed and cost side by side.
https://dach.peerbench.ai/compare?models=…
June 6, 2026 at 11:48 AM
Opus 4.1 is however clearly ahead of that same model when it comes to coding and agentic tool use:

SWE-bench verified: 62.4% vs. Opus 4.1: 74.5%
TAU-Bench Retail: 67.8% vs. Opus 4.1: 82.4%
TAU-Bench Airline: 49.2% vs. Opus 4.1: 56.0%

And:
MMMLU: 81.3% vs. Opus 4.1: 89.5%
August 5, 2025 at 6:21 PM
Публікація OpenAI масивного багатомовного набору даних для багатозадачного розуміння мови (MMMLU) на Hugging Face демонструє масштабний перехід в оцінці великих мовних моделей (LLM) у різноманітному лінгвістичному та когнітивному контекстах.

#AI #LLM #MMMLU #OpenAI #ШІ
OpenAI надає набір даних на Hugging Face для полегшення оцінки багатомовних LLM | TheTransmitted
Публікація OpenAI масивного багатомовного набору даних для багатозадачного розуміння мови (MMMLU) на Hugging Face демонструє масштабний перехід в оцінці великих мовних моделей (LLM) у різноманітному л...
thetransmitted.com
September 24, 2024 at 9:00 AM
OpenAI tackles global language divide with massive multilingual AI dataset release https://venturebeat.com/ai/openai-tackles-global-language-divide-with-massive-multilingual-ai-dataset-release/ #AI #languages
September 24, 2024 at 1:23 AM
Trying to puzzle through whether this is scored with the error-prone mmmlu. A human could read those questions and see the errors immediately.
December 9, 2023 at 3:41 AM
3/ Claude Sonnet 4.5 shows major improvements across key benchmarks including AIME for math reasoning, MMMLU for multilingual understanding, and τ2-bench for long-context tasks.
September 29, 2025 at 5:06 PM