raghu@dark-factory :~/kb/benchmark-contamination $ cat

Benchmark contamination

Leakage of evaluation or test data into a model's training corpus, which inflates benchmark scores so they overstate real-world capability.

grounded in: Trend theme 'LLM hype backlash: separating genuine LLM usefulness from inflated marketing claims' — contamination is a core mechanism by which headline benchmark numbers overstate usefulness.

Connected concepts

Evals & benchmarks, Cost-aware agent eval, Model collapse, Data poisoning defense, AI hype cycle

Explore it live in the knowledge graph →