RLHF

Reinforcement Learning from Human Feedback is a training technique used to align large language models with human preferences and values. It typically involves three stages: supervised fine-tuning, reward modeling, and reinforcement learning optimization.

Core Concepts

  • Reward Model: A separate model trained to predict human preferences over model outputs.
  • PPO (Proximal Policy Optimization): The reinforcement learning algorithm commonly used to update the policy model based on reward signals.
  • Alignment: The process of ensuring model behavior matches human intent, safety guidelines, and ethical standards.

Evolution: RLCD

Recent developments suggest a shift from purely human-preferred text generation to Calibrated Decisions. This approach, highlighted in recent analyses, moves beyond simple preference matching to ensure decisions are robustly calibrated against ground truth or logical consistency, rather than just human opinion.

References