Reinforcement Learning from Human Feedback
Reinforcement Learning from Human Feedback (rlhf) is a technique used to align large language models with human preferences. It involves training a model to generate responses that humans rate as high-quality, often using a reward model derived from human annotations.
Core Mechanism
- Supervised Fine-Tuning (SFT): Initial training on high-quality human-written data.
- Reward Modeling: Training a separate model to predict human preferences based on pairs of responses.
- Reinforcement Learning: Optimizing the policy model using the reward model via algorithms like PPO (Proximal Policy Optimization).
Evolution: From RLHF to RLCD
Recent developments indicate a shift away from purely human-preferred text generation toward more calibrated decision-making processes. This evolution addresses limitations in traditional RLHF, such as reward hacking and lack of factual grounding.
- RLCD (Reinforcement Learning from Calibrated Decisions): A proposed framework focusing on calibrated outcomes rather than just human preference scores.
- Key Insight: Diogo Almeida, co-inventor of the technique behind ChatGPT, highlights a fundamental shift in AI training paradigms.
- Objective: Moving beyond “human-preferred text” to ensure models make decisions that are robustly calibrated to truth and utility.
- Implication: This shift challenges the core idea that human preference alone is sufficient for optimal model alignment.
Related Concepts
- Reward Modeling
- Proximal Policy Optimization
- AI Alignment
- large-language-models