raghu@dark-factory :~/kb/rlhf $ cat

RLHF

Reinforcement learning from human feedback fine-tunes a model against a reward model trained on human preference rankings so its outputs align with what people prefer.

grounded in: Established alignment technique that constitutional-ai (already in the graph) generalizes and that reward-hacking (already in the graph) is a failure mode of.

Connected concepts

Constitutional AI, Reward hacking, Evals & benchmarks

Explore it live in the knowledge graph →