Anomalous Model Responses

Phenomenon where Large Language Models (LLMs) produce bizarre, nonsensical, or degraded outputs in response to specific, seemingly innocuous input strings. Often linked to byte-pair-encoding artifacts and LLM Glitch Tokens.

Key Characteristics

  • Trigger Sensitivity: Specific token sequences cause disproportionate degradation in output quality.
  • Output Degradation: Responses may become repetitive, incoherent, or semantically void.
  • Tokenization Dependency: Closely tied to how byte-pair-encoding maps subwords to integer IDs.

References

Source Notes