Robots Atlas>ROBOTS ATLAS
Architecture

RoPE

2021ActiveUpdated: 24 August 2026Published
Key innovation
Encodes token position by rotating query and key vectors by an angle proportional to position, so the attention dot product depends on relative token distance and naturally extrapolates to longer sequences.
Category
Architecture
Abstraction level
Primitive
Operation level
LayerArchitecture block
Use cases
Long-context models (LLaMA, Mistral, Gemma)Extrapolation to sequences longer than trainingStandard component in modern LLMsLong document and code processingModels for full codebase analysis

How it works

Query and key vectors in the attention mechanism are rotated by an angle proportional to the token position before computing the dot product. This makes attention between tokens depend on their relative distances rather than absolute positions.

Problem solved

Standard positional encoding (additive or sinusoidal) generalizes poorly to sequences longer than seen during training. RoPE encodes positions through matrix rotation, which naturally transfers to longer sequences.

Components

Frequency schedule (theta)Determines the rotation angle as a function of position and dimension index.

A set of frequencies θ_i = base^(-2i/d) (base default 10000) defining the rotation angle for each dimension pair. Low dimensions rotate quickly, high dimensions slowly, forming a multi-resolution position representation.

Official

Rotation matrixInjects absolute position information via rotation.

A block-diagonal rotation matrix that rotates pairs of adjacent vector dimensions by angle m·θ_i, where m is the token position. Implemented efficiently as element-wise operations on pairs (x, y) → (x·cos − y·sin, x·sin + y·cos).

INQuery/key vectors before rotation.
OUTRotated Q/K vectors of the same shape.
Application to Q and KEnsures attention depends on relative position.

RoPE is applied only to query and key vectors (not to value V) before computing the attention dot product. This makes the attention score depend on the relative position m−n of two tokens.

Implementation

Implementation pitfalls
Degradation when extrapolating beyond training contextHigh

RoPE trained on sequences up to N tokens degrades for sequences >N without extrapolation techniques (Position Interpolation, NTK-aware scaling, YaRN, LongRoPE). Naive context extension leads to chaotic attention distributions.

Fix:Apply Position Interpolation, NTK-aware base scaling, YaRN, or LongRoPE to extend context with brief fine-tuning.
float32 precision required for small anglesMedium

At large positions the rotation angles become very small — float16/bfloat16 computation can cause numerical errors and loss of positional information.

Fix:Compute cos/sin tables and the rotation itself in float32, then cast the result to bf16 after application.
Even head dimension requirementLow

RoPE rotates pairs of adjacent dimensions, so the attention head dimension must be even; an odd dimension requires special handling or partial RoPE.

Fix:Ensure an even head dimension or use a partial rotary_dim variant.

Evolution

Original paper · 2021 · arXiv:2104.09864 (cs.CL) · Jianlin Su
RoFormer: Enhanced Transformer with Rotary Position Embedding
Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, Yunfeng Liu
2021
RoPE introduced in the RoFormer paper
Inflection point

Su et al. propose encoding positions by rotating Q/K vectors, unifying absolute and relative position and enabling use with linear attention.

2021
Popularized by EleutherAI (GPT-J / GPT-NeoX)

The 'Rotary Embeddings: A Relative Revolution' post and the use of RoPE in GPT-J and GPT-NeoX spread the method across decoder-only LLMs.

2023
LLaMA makes RoPE the de facto standard
Inflection point

Adoption of RoPE in LLaMA, then Mistral, Qwen, Gemma, and DBRX establishes it as the default positional encoding in open LLMs.

2023
Position Interpolation extends RoPE context

Chen et al. show that scaling position indices (interpolation) extends the context window of RoPE-based models with minimal fine-tuning.

2023
YaRN — efficient RoPE context extension

YaRN combines NTK-aware scaling with attention temperature correction, extending RoPE model context more efficiently than plain interpolation.

2024
LongRoPE extends context to 2M+ tokens

LongRoPE uses evolutionary search over non-uniform RoPE scaling factors, extending the context window to over 2 million tokens.

Hyperparameters (configurable axes)

Rotary base (theta)Critical

Frequency base (default 10000). Increasing the base lengthens wavelengths and is the basis of NTK-aware scaling for context extension.

10000Default value from the RoFormer paper and most LLMs.
500000Increased base in long-context models (e.g. LLaMA 3).
Rotary dimension / percentageMedium

Fraction of the head dimension subjected to rotation. Full RoPE rotates 100% of dimensions; some models (GPT-NeoX) use partial RoPE (e.g. 25%).

1.0Full RoPE (LLaMA, Mistral).
0.25Partial RoPE in GPT-NeoX.
Head dimension (even)High

The attention head dimension must be even because RoPE rotates dimension pairs.

Computational complexity

Time complexity: O(n · d). Space complexity: O(n · d).

Execution paradigm

Primary mode
Dense

RoPE is a deterministic transformation with no routing or conditional activation.

Activation pattern
All paths active
Routing mechanism

Parallelism

Parallelism level
Fully parallel

The rotation is applied independently to each token and each dimension pair, so it is fully parallel.

Scope
TrainingInferenceAcross tokens

Hardware requirements

Primary

RoPE consists of element-wise trigonometric operations applied to each Q and K vector — fully GPU-accelerated, often fused with attention kernels (FlashAttention supports RoPE fusion).

Good fit

RoPE requires no specialized hardware — it is a cheap transformation that runs on any accelerator and on CPU.