RLHF
Reinforcement Learning from Human Feedback is a training technique used to align large language models with human preferences and values. It typically involves three stages: supervised fine-tuning, reward modeling, and reinforcement learning optimization.
Core Concepts
- Reward Model: A separate model trained to predict human preferences over model outputs.
- PPO (Proximal Policy Optimization): The reinforcement learning algorithm commonly used to update the policy model based on reward signals.
- Alignment: The process of ensuring model behavior matches human intent, safety guidelines, and ethical standards.
Evolution: RLCD
Recent developments suggest a shift from purely human-preferred text generation to Calibrated Decisions. This approach, highlighted in recent analyses, moves beyond simple preference matching to ensure decisions are robustly calibrated against ground truth or logical consistency, rather than just human opinion.
- Key Insight: Human feedback can be noisy or biased; calibrated decision-making aims for more reliable alignment.
- Proposed by: Diogo Almeida, co-inventor of the technique behind ChatGPT.
- Analysis: See Jev: RLCD’s Shift from Human-Preferred Text to Calibrated Decisions for a detailed breakdown of this shift.