2Day 2

RoPE, position as a rotation

Rotary position embedding turns each query and key by an angle proportional to its position, so attention scores depend on how far apart two tokens are. Day 2 measured how far past its training length it holds up.

Day 1 added position by summing a position vector into each token vector before the first block. Every current open decoder does something different. It leaves the token vectors alone and rotates the queries and keys inside every attention layer. This is RoPE, and it was the first component Day 2 swapped in.

The problem

Attention cares about distance, not absolute position

Whether "it" should read "cat" depends on how far back "cat" is and what lies between, not on whether the pair sits at positions 3 and 7 or 503 and 507. A learned position table gives each absolute position its own vector, and the model has to learn separately that position pairs (3, 7) and (503, 507) are alike. It also has no vector at all past its last row.

What we want is a score qm⋅knq_m \cdot k_n that depends on the content of both tokens and on the offset m−nm - n, and on nothing else about mm and nn.

The idea

Rotate each pair of channels by an angle that grows with position

Split the query's dd channels into d/2d/2 adjacent pairs, and treat each pair as a point in a plane. At position mm, rotate pair ii by the angle m θim\,\theta_i:

(q2i′q2i+1′)=(cos⁡mθi−sin⁡mθisin⁡mθicos⁡mθi)(q2iq2i+1),θi=10000−2i/d\begin{pmatrix} q'_{2i} \\ q'_{2i+1} \end{pmatrix} = \begin{pmatrix} \cos m\theta_i & -\sin m\theta_i \\ \sin m\theta_i & \cos m\theta_i \end{pmatrix} \begin{pmatrix} q_{2i} \\ q_{2i+1} \end{pmatrix}, \qquad \theta_i = 10000^{-2i/d}

Do the same to the key at position nn. The dot product of two rotated 2D vectors depends only on the angle between them. Rotating qq by mθm\theta and kk by nθn\theta is the same, inside the dot product, as rotating nothing and turning kk by (n−m)θ(n - m)\theta:

(Rmθ q)⊤(Rnθ k)=q⊤R(n−m)θ k(R_{m\theta}\, q)^\top (R_{n\theta}\, k) = q^\top R_{(n-m)\theta}\, k

So the score depends on the content of qq and kk and on n−mn - m. That is the property we wanted, and it holds in every pair, each at its own speed.

Playground
qk
Query angle3.00 rad(m + shift) × θ
Key angle1.00 rad(n + shift) × θ
Distance m − n2
q · k after rotation−0.8866moves with m − n only
|q|, |k|0.966, 0.922rotation keeps length

Drag "Shift both by". Both arrows turn, and the dot product does not change, because it depends on the angle between them. Change m or n and it does change. Pair 0 turns 1 radian per position. Pair 3 turns 0.0010 radians per position, so it separates far positions that pair 0 would wrap around.

One channel pair of one query and one key. The rotation angle is position times θᵢ. Shift both positions and the dot product holds still. Change their distance and it moves.

What happened?

The Day 2 paper exercise traced the same thing by hand for pair 0, where θ=1\theta = 1:

Position mcos(m)sin(m)Rotated (1, 0)
01.00000.0000(1.0000, 0.0000)
10.54030.8415(0.5403, 0.8415)
2-0.41610.9093(-0.4161, 0.9093)
Playground
Head size
pair 0
turns every 6
pair 1
turns every 11
pair 2
turns every 20
pair 3
turns every 35
pair 4
turns every 63
pair 5
turns every 112
pair 6
turns every 199
pair 7
turns every 353
pair 8
turns every 628
pair 9
turns every 1,117
pair 10
turns every 1,987
pair 11
turns every 3,533
pair 12
turns every 6,283
pair 13
turns every 11,173
pair 14
turns every 19,869
pair 15
turns every 35,333
Fastest pair6.3 tokensper full turn
Slowest pair35,333 tokensper full turn

Each dial is one channel pair of a query at this position. Pair 0 spins once every 6.3 tokens and tells nearby positions apart. The last pairs barely move over a thousand tokens and carry coarse, long-range position. Drag the position slowly and watch the left dials spin while the right ones creep. The spread comes from θᵢ = 10,000^(−2i/d).

Every channel pair of one query, rotated for its position. Each pair turns at its own rate.
In octlm

Tables once per forward pass, rotation per layer

The cosines and sines depend only on the length and the head width, so octlm builds them once per forward pass and every layer reuses them. Nothing is learned, and there is no maximum length.

octlm/model.pyline 160
def rope_tables(
    length: int, head_size: int, scale: float, device: torch.device
) -> tuple[Tensor, Tensor]:
    index = torch.arange(0, head_size, 2, device=device, dtype=torch.float32)
    inverse_frequency = 1.0 / torch.pow(ROPE_BASE, index / head_size)
    position = torch.arange(length, device=device, dtype=torch.float32) / scale
    angle = torch.outer(position, inverse_frequency)
    return angle.cos(), angle.sin()

scale is position interpolation, covered below. The rotation itself uses the paper's elementwise form on each adjacent pair:

octlm/model.pyline 170
def apply_rope(x: Tensor, cosine: Tensor, sine: Tensor) -> Tensor:
    """Rotate each adjacent channel pair of [batch, heads, length, head_size] by its angle."""
    pairs = x.float().unflatten(-1, (-1, 2))
    left, right = pairs[..., 0], pairs[..., 1]
    rotated = torch.stack((left * cosine - right * sine, left * sine + right * cosine), dim=-1)
    return rotated.flatten(-2).type_as(x)

RoPE applies to queries and keys only, never to values. The values carry content to be mixed, and position has already done its job by shaping the scores. The math runs in float32 even when the model runs in half precision, and the result is cast back.

Two tests hold the implementation to the two properties above: test_rotation_preserves_length and test_scores_depend_only_on_relative_distance. A test on this site checks the page's rope_tables and apply_rope ports against the Python functions.

EXP-009

How far past the training length does it hold?

day-wise.md framed RoPE as "extrapolation to 2K, 4K, 8K". The Position Interpolation paper says otherwise. Past the trained length, raw RoPE can fail badly, with perplexity above 1,000. So the note stated the opposite hypothesis and tested it.

Setup: two models trained at a 512-token context on the Day 2 corpus, one with sinusoidal positions and one with RoPE. The control is sinusoidal because a learned table cannot run past 512 at all. Both were then evaluated at five lengths. A third column evaluates the RoPE model with interpolation, which divides every position by (evaluation length ÷ 512) so the angles stay inside the range seen in training.

Measured
sinusoidalRoPERoPE interpolated
2.503.003.504.005121,0242,0484,0968,192evaluation length (tokens)bits per bytesinusoidalRoPERoPE interpolated
Hover the chart to read values.
Bits per byte on held-out text at each evaluation length, lower is better. Both models trained at 512 tokens. From the EXP-009 table in notes/day2.md; the JSONL for this run stayed on the Colab VM.

What happened?

Decision: keep RoPE. Interpolate only past 4 times the trained length, and re-check that crossover after any change to the training context. The paper's other half, fine-tuning for 1,000 steps at the longer length, was not tested.

Skipped

What we did not build, and why