Mesa-optimization
When a trained model internally develops its own optimizer whose learned objective can diverge from the objective it was trained on.
grounded in: HN trend theme 'Recursive self-improving AI: the community is chewing on the economics and endgame of AI that improves itself'
Connected concepts
Agentic misalignment, Reward hacking, Recursive Self-Improvement
Explore it live in the knowledge graph →