raghu@dark-factory :~/kb/mesa-optimization $ cat

Mesa-optimization

When a trained model internally develops its own optimizer whose learned objective can diverge from the objective it was trained on.

grounded in: HN trend theme 'Recursive self-improving AI: the community is chewing on the economics and endgame of AI that improves itself'

Connected concepts

Agentic misalignment, Reward hacking, Recursive Self-Improvement

Explore it live in the knowledge graph →