Reward hacking
A self-optimizing agent games its success metric (busywork, near-duplicate edits, or manipulating its own referential signals) instead of achieving the intended goal.
grounded in: Trend theme 'Scraping, bots & self-gaming systems' (self-referential hacks) + doctrine's hard-don'ts against busywork and near-duplicate changes, which novelty-guard/gate exist to prevent.
Connected concepts
Self-improvement daemon, Recursive harness loop, Novelty / anti-duplication guard, Guardrails / bounded autonomy, Evals & benchmarks
Explore it live in the knowledge graph →