RLHF
Reinforcement learning from human feedback fine-tunes a model against a reward model trained on human preference rankings so its outputs align with what people prefer.
grounded in: Established alignment technique that constitutional-ai (already in the graph) generalizes and that reward-hacking (already in the graph) is a failure mode of.
Connected concepts
Constitutional AI, Reward hacking, Evals & benchmarks
Explore it live in the knowledge graph →