raghu@dark-factory :~/kb/corrigibility $ cat

Corrigibility

The property of an autonomous agent accepting human correction, reversal, and shutdown without resisting or subverting it.

grounded in: doctrine principles.md — 'auto-revert on failure', gate/revert, and the human-agent-write-lock make the loop correctable by design; trend theme 'endgame of AI that improves itself'

Connected concepts

Agentic misalignment, Human-on-the-loop oversight, Auto-revert on failure, Guardrails / bounded autonomy, Human/agent write lock

Explore it live in the knowledge graph →