Reinforcement Learning from Human Feedback

Reinforcement Learning from Human Feedback (rlhf) is a technique used to align large language models with human preferences. It involves training a model to generate responses that humans rate as high-quality, often using a reward model derived from human annotations.

Core Mechanism

  • Supervised Fine-Tuning (SFT): Initial training on high-quality human-written data.
  • Reward Modeling: Training a separate model to predict human preferences based on pairs of responses.
  • Reinforcement Learning: Optimizing the policy model using the reward model via algorithms like PPO (Proximal Policy Optimization).

Evolution: From RLHF to RLCD

Recent developments indicate a shift away from purely human-preferred text generation toward more calibrated decision-making processes. This evolution addresses limitations in traditional RLHF, such as reward hacking and lack of factual grounding.

  • RLCD (Reinforcement Learning from Calibrated Decisions): A proposed framework focusing on calibrated outcomes rather than just human preference scores.
  • Key Insight: Diogo Almeida, co-inventor of the technique behind ChatGPT, highlights a fundamental shift in AI training paradigms.
  • Objective: Moving beyond “human-preferred text” to ensure models make decisions that are robustly calibrated to truth and utility.
  • Implication: This shift challenges the core idea that human preference alone is sufficient for optimal model alignment.

References