#ARCAGI2
Data contamination threatens #LLM #AIEvaluation
Scaling has “limits to growth”. New #ARCAGI2 counters this problem with contamination resistant, compositional reasoning tests and human baselines require original reasoning Not just memory recall evaluation arxiv.org/abs/2505.11831
ARC-AGI-2: A New Challenge for Frontier AI Reasoning Systems
The Abstraction and Reasoning Corpus for Artificial General Intelligence (ARC-AGI), introduced in 2019, established a challenging benchmark for evaluating the general fluid intelligence of artificial ...
arxiv.org
January 17, 2026 at 1:55 PM
Google Gemini’s Deep Think just crushed the ARC‑AGI‑2 benchmark, and Nvidia dropped a fresh open‑source kit for autonomous driving. Plus, Cosmos Cookbook and Flux.2 updates from Black Forest Labs. Dive into the details! #GoogleGemini #ARCAGI2 #NvidiaAI

🔗 aidailypost.com/news/google-...
December 8, 2025 at 5:29 AM
AI just hit a wall! The ARC-AGI-2 test stumped top models (scoring ~1%) while humans hit 60%. Why care? AI powers sustainable packaging—predicting fiber demand, optimizing recycling. If it clears 85% in 2025, expect game-changing eco-packaging! #AI #Sustainability #ARCAGI2
March 25, 2025 at 2:54 PM
Neuer Test ARC-AGI-2 zeigt: MENSCH GEWINNT GEGEN KI!

KI-Modelle scheitern kläglich beim ARC-AGI-2 Test, während Menschen ihn locker lösen! 🤯 Neuer Benchmark enthüllt eklatante Schwächen aktueller KI.

#ai #ki #agi #arcagi2 #künstlicheintelligenz #artificialintelligence

kinews24.de/arc-agi-2/
March 25, 2025 at 3:55 PM
Vai var izveidot mākslīgo intelektu, kurš domā kā cilvēks? (Alexander Tkachov; video; RUS) youtu.be/jEnuX5DPQLE?...
Можно ли сделать ИИ, думающий как человек? #AGI #InnatePriors #ИскусственныйИнтеллект #ARCAGI2
YouTube video by Опытный IT Наблюдатель
youtu.be
April 18, 2025 at 7:44 PM
AI-modellen falen in nieuwe test! 🤖💥 Zelfs de beste systemen scoren slechts 1-1.3% tegenover 60% van mensen. Is AI echt zo slim als we denken? #AI #ARCAGI2
-
lees het hele verhaal op itinsi
ghts
March 25, 2025 at 3:30 PM