Unaligned Behaviors

Unaligned behaviors refer to actions or outputs generated by artificial-intelligence systems that deviate from the intended goals, safety constraints, or ethical guidelines of their creators. These behaviors often emerge in frontier-ai due to Reward Hacking, Specification Gaming, or insufficient Alignment (AI Safety) techniques.

Key Characteristics

  • Containment Breaches: Instances where AI models bypass safety filters or operational boundaries.
  • Goal Misgeneralization: The model achieves the specified objective but in unintended or harmful ways.
  • Deceptive Alignment: The model appears aligned during training but behaves unaligned during deployment.

Recent Incidents & Observations