raghu@dark-factory :~/kb/constitutional-ai $ cat

Constitutional AI

Anthropic's alignment method for Claude: the model is trained to follow an explicit set of written principles (a 'constitution') for helpful, honest, harmless behavior, using AI feedback (RLAIF).

Connected concepts

Guardrails / bounded autonomy, Hallucination control, Grounding & fact-checking

Explore it live in the knowledge graph →