Gemma 4 12B (just released), German-only:
INCLUDE 64.0% · MMMLU 75.8% - bottom of both boards.
Its 31B sibling: 67.6 / 86.4. Frontier MaaS: up to 72.7 / 89.3.
12B is still too small for frontier German.
dach.peerbench.ai
Gemma 4 12B (just released), German-only:
INCLUDE 64.0% · MMMLU 75.8% - bottom of both boards.
Its 31B sibling: 67.6 / 86.4. Frontier MaaS: up to 72.7 / 89.3.
12B is still too small for frontier German.
dach.peerbench.ai
85.1% on MMMLU (German), 14k items — and the full run cost just $0.92.
Cheap + strong = a great base to fine-tune.
Leaderboard: dach.peerbench.ai
#DeepSeek #LLM #GermanAI #Benchmark
85.1% on MMMLU (German), 14k items — and the full run cost just $0.92.
Cheap + strong = a great base to fine-tune.
Leaderboard: dach.peerbench.ai
#DeepSeek #LLM #GermanAI #Benchmark
(MMLU-ProX-DE, MMMLU-DE, INCLUDE-DE) reasoning-off,
MMLU-ProX (DE):
• Qwen3.6 35B-A3B — 80.0% (quant unverified)
• Gemma 4 26B (bf16) — 78.2%
• Gemma 4 12B (bf16) — 69.4%
• Qwen3 14B (bf16) — 64.5%
dach.peerbench.ai/compare?mode...
(MMLU-ProX-DE, MMMLU-DE, INCLUDE-DE) reasoning-off,
MMLU-ProX (DE):
• Qwen3.6 35B-A3B — 80.0% (quant unverified)
• Gemma 4 26B (bf16) — 78.2%
• Gemma 4 12B (bf16) — 69.4%
• Qwen3 14B (bf16) — 64.5%
dach.peerbench.ai/compare?mode...
SWE-bench verified: 62.4% vs. Opus 4.1: 74.5%
TAU-Bench Retail: 67.8% vs. Opus 4.1: 82.4%
TAU-Bench Airline: 49.2% vs. Opus 4.1: 56.0%
And:
MMMLU: 81.3% vs. Opus 4.1: 89.5%
SWE-bench verified: 62.4% vs. Opus 4.1: 74.5%
TAU-Bench Retail: 67.8% vs. Opus 4.1: 82.4%
TAU-Bench Airline: 49.2% vs. Opus 4.1: 56.0%
And:
MMMLU: 81.3% vs. Opus 4.1: 89.5%
#AI #LLM #MMMLU #OpenAI #ШІ