Attention Heads

Definition

Sub-components of the Multi-Head Attention mechanism within Transformer architectures. They enable the model to simultaneously attend to information from different representation subspaces at different positions.

Mechanism

  • Each head performs independent Scaled Dot-Product Attention using partitioned (Query), (Key), and (Value) projections.
  • The outputs of all heads are concatenated and linearly transformed to produce the final output of the attention layer.

Inference and Performance Context


Backlink: 2026 04 22 LLM Inference Engines Memory Mapping and Performance Optimization

Source Notes