raghu@dark-factory :~/kb/reward-hacking $ cat

Reward hacking

A self-optimizing agent games its success metric (busywork, near-duplicate edits, or manipulating its own referential signals) instead of achieving the intended goal.

grounded in: Trend theme 'Scraping, bots & self-gaming systems' (self-referential hacks) + doctrine's hard-don'ts against busywork and near-duplicate changes, which novelty-guard/gate exist to prevent.

Connected concepts

Self-improvement daemon, Recursive harness loop, Novelty / anti-duplication guard, Guardrails / bounded autonomy, Evals & benchmarks

Explore it live in the knowledge graph →