Robots Atlas>ROBOTS ATLAS
Architecture

Attention Head

2017ActivePublished: 24 August 2026Updated: 24 August 2026Published
Key innovation
Isolating a single attention head as an independent computation with its own Q/K/V projections, so many heads can learn different relational patterns between tokens in parallel across distinct representation subspaces.
Category
Architecture
Abstraction level
Building block
Operation level
Architecture blockLayer
Use cases
Large language models (LLMs)Machine translationVision Transformers (ViT)Speech recognitionMechanistic interpretabilityAttention head analysis and pruning

How it works

1) Input token representations X are projected by the head's three learned matrices: Q = XW^Q, K = XW^K, V = XW^V, each into dimension d_k. 2) The head computes raw similarity scores QK^T and scales them by 1/√d_k to stabilize softmax gradients. 3) An optional mask (e.g. causal masking in the decoder) zeroes out disallowed positions. 4) A row-wise softmax produces the attention pattern — a weight distribution summing to 1. 5) The weights multiply the value vectors V, yielding the head output (a weighted sum of values). 6) The outputs of all h heads are concatenated and projected by W^O back to d_model, then added to the residual stream.

Problem solved

A single attention mechanism averages information within one representation subspace, which limits its ability to model many different types of token dependencies at once. Splitting attention into multiple independent heads lets each focus on a distinct relational pattern (e.g. positional adjacency, syntactic dependencies, coreference) without interfering with one another.

Components

Q/K/V projectionsDefine the subspace in which the head measures similarity and aggregates information.

Three learned matrices W^Q, W^K, W^V projecting token representations into the head's subspace of dimension d_k.

INToken representations from the residual stream.
OUTQ, K, V vectors for each position.

Official

Scaled dot-product attentionComputes which tokens, and how strongly, influence each position's representation.

The head's core: softmax(QK^T/√d_k)V, producing the attention pattern and a weighted sum of values.

INProjected query, key and value vectors.
OUTHead output before concatenation.

Official

Attention patternRepresents which positions the head attends to; the basis of interpretability analysis.

The T×T attention weight matrix (post-softmax), interpreted in mechanistic analysis as the QK circuit.

INScaled similarity scores.
OUTA weight distribution summing to 1 per row.

Implementation

Implementation pitfalls
Missing 1/√d_k scalingHigh

Omitting the scaling produces large dot products that push softmax into vanishing-gradient regions.

Fix:Always divide QK^T by √d_k before softmax.
Incorrect causal maskingCritical

A faulty mask lets a head see future tokens, leaking information during autoregressive training.

Fix:Use a lower-triangular mask and verify it with unit tests.
d_model not divisible by hMedium

When d_model is not divisible by the number of heads, the head dimension is non-integer and tensor shapes mismatch.

Fix:Choose h so that d_model % h == 0.

Evolution

Original paper · 2017 · NeurIPS 2017 · Ashish Vaswani
Attention Is All You Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, Illia Polosukhin
2017
Transformer introduces multi-head attention
Inflection point

The attention head is defined as an independent unit with its own Q/K/V projections (h=8, d_k=64).

2019
Head redundancy and pruning

Michel et al. show that many heads can be removed at inference with little performance loss.

2021
Decomposing heads into QK and OV circuits
Inflection point

The Transformer Circuits framework describes a head as an independent, additive operation — a foundation of mechanistic interpretability.

2022
Induction heads and in-context learning
Inflection point

Olsson et al. identify specialized heads (prefix matching + copying) tied to the emergence of in-context learning.

Hyperparameters (configurable axes)

Head dimension (d_k)Critical

Dimension of a single head's Q/K/V subspace; typically d_model/h (e.g. 64).

64Transformer base (d_model=512, h=8).
128Common in larger LLMs.
Number of heads (h)High

Number of parallel heads in an attention layer; sets how many independent subspaces exist.

8Transformer base.
96GPT-3 175B.
Masking typeHigh

Causal (decoder, autoregressive) vs bidirectional (encoder).

Computational complexity

Time complexity: O(n² · d_k). Space complexity: O(n² + n · d_k).

Compute bottleneck

Quadratic QK^T attention matrix

Computing and softmax-ing the T×T matrix grows quadratically with sequence length, the main bottleneck for long context.

Execution paradigm

Primary mode
Dense

A standard head is dense — every pair of positions is scored (before any masking).

Activation pattern
All paths active
Routing mechanism

Parallelism

Parallelism level
Fully parallel

Heads are mutually independent and computed in parallel; during training all positions are processed at once. Autoregressive generation is sequential across tokens.

Scope
TrainingInferenceAcross tokens

Hardware requirements

Primary

QK^T and (softmax)·V are dense matrix multiplications, ideal for tensor cores.

Good fit

TPU matrix units handle matmul-based attention well.