Constitutional AI
Anthropic's alignment method for Claude: the model is trained to follow an explicit set of written principles (a 'constitution') for helpful, honest, harmless behavior, using AI feedback (RLAIF).
Connected concepts
Guardrails / bounded autonomy, Hallucination control, Grounding & fact-checking
Explore it live in the knowledge graph →