#DataContamination
Alibaba’s Qwen 2.5 AI Faces MAth ‘Cheating’ Allegations Over Contaminated Benchmark Data

#AI #Alibaba #Qwen #AIBenchmarks #DataContamination #MachineLearning

winbuzzer.com/2025/07/21/a...
July 21, 2025 at 12:10 PM
Scrap the work of others and monetize: “They are cheating,” says Cheng Xu, a Ph.D. student at University of College Dublin who led a recent survey of data contamination in AI benchmarks. #internet #profit #datacontamination
Hilarious: “The researchers encouraged the struggling machines to persevere, feeding prompts such as “keep working” and “don’t be afraid to execute your code.””

Pathetic: “Despite the exhortations, no model scored above 2% on the test”

www.science.org/content/arti...
‘Brutal’ math test stumps AI but not human experts
Benchmark shows humans can still top machines—but for how much longer?
www.science.org
December 7, 2024 at 8:45 PM
Extremely interesting article here that posits AI generated training data may have poisoned data sources more widely, leading to a data equivalent of the need for #LowBackGroundSteel
#AI #Data #DataContamination
www.theregister.com/2025/06/15/a...
ChatGPT polluted the world forever, like the first atom bomb
Feature: Academics mull the need for the digital equivalent of low-background steel
www.theregister.com
June 19, 2025 at 10:28 AM
LNE-Blocking uses Leakage‑Noise Estimation and a Blocking step to tweak greedy decoding, cutting memorized answers, preserving performance. Code is on GitHub. Read more: https://getnews.me/lne-blocking-framework-to-counter-data-contamination-in-llms/ #llm #datacontamination
September 20, 2025 at 4:42 AM
#DataContamination #AIEvaluation Training–test overlap can inflate LLM scores. “data contamination” in #LLMs, defined as unintended overlap between training data & evaluation data that can inflate measured performance & misrepresent true generalization. arxiv.org/html/2502.14...
A Survey on Data Contamination for Large Language Models
arxiv.org
January 17, 2026 at 1:32 PM