RoPE, position as a rotation
Rotary position embedding turns each query and key by an angle proportional to its position, so attention scores depend on how far apart two tokens are. Day 2 measured how far past its training length it holds up.
Day 1 added position by summing a position vector into each token vector before the first block. Every current open decoder does something different. It leaves the token vectors alone and rotates the queries and keys inside every attention layer. This is RoPE, and it was the first component Day 2 swapped in.
Attention cares about distance, not absolute position
Whether "it" should read "cat" depends on how far back "cat" is and what lies between, not on whether the pair sits at positions 3 and 7 or 503 and 507. A learned position table gives each absolute position its own vector, and the model has to learn separately that position pairs (3, 7) and (503, 507) are alike. It also has no vector at all past its last row.
What we want is a score that depends on the content of both tokens and on the offset , and on nothing else about and .
Rotate each pair of channels by an angle that grows with position
Split the query's channels into adjacent pairs, and treat each pair as a point in a plane. At position , rotate pair by the angle :
Do the same to the key at position . The dot product of two rotated 2D vectors depends only on the angle between them. Rotating by and by is the same, inside the dot product, as rotating nothing and turning by :
So the score depends on the content of and and on . That is the property we wanted, and it holds in every pair, each at its own speed.
What happened?
- "Shift both by" turns both arrows by the same angle. The angle between them is fixed, so the dot product is fixed to four decimals.
- Changing one position changes the angle between them, and the dot product follows.
- The lengths of and never change. A rotation does not scale, so RoPE carries position without changing how large a score can get.
- Pair 0 turns a full radian per position and wraps around the circle about every 6 positions. It can tell neighbors apart and nothing more. Pair 3 turns so slowly that it separates positions dozens of tokens apart.
The Day 2 paper exercise traced the same thing by hand for pair 0, where :
| Position m | cos(m) | sin(m) | Rotated (1, 0) |
|---|---|---|---|
| 0 | 1.0000 | 0.0000 | (1.0000, 0.0000) |
| 1 | 0.5403 | 0.8415 | (0.5403, 0.8415) |
| 2 | -0.4161 | 0.9093 | (-0.4161, 0.9093) |
Tables once per forward pass, rotation per layer
The cosines and sines depend only on the length and the head width, so octlm builds them once per forward pass and every layer reuses them. Nothing is learned, and there is no maximum length.
def rope_tables(
length: int, head_size: int, scale: float, device: torch.device
) -> tuple[Tensor, Tensor]:
index = torch.arange(0, head_size, 2, device=device, dtype=torch.float32)
inverse_frequency = 1.0 / torch.pow(ROPE_BASE, index / head_size)
position = torch.arange(length, device=device, dtype=torch.float32) / scale
angle = torch.outer(position, inverse_frequency)
return angle.cos(), angle.sin()scale is position interpolation, covered below. The rotation itself uses the paper's elementwise form on each adjacent pair:
def apply_rope(x: Tensor, cosine: Tensor, sine: Tensor) -> Tensor:
"""Rotate each adjacent channel pair of [batch, heads, length, head_size] by its angle."""
pairs = x.float().unflatten(-1, (-1, 2))
left, right = pairs[..., 0], pairs[..., 1]
rotated = torch.stack((left * cosine - right * sine, left * sine + right * cosine), dim=-1)
return rotated.flatten(-2).type_as(x)RoPE applies to queries and keys only, never to values. The values carry content to be mixed, and position has already done its job by shaping the scores. The math runs in float32 even when the model runs in half precision, and the result is cast back.
Two tests hold the implementation to the two properties above: test_rotation_preserves_length and test_scores_depend_only_on_relative_distance. A test on this site checks the page's rope_tables and apply_rope ports against the Python functions.
How far past the training length does it hold?
day-wise.md framed RoPE as "extrapolation to 2K, 4K, 8K". The Position Interpolation paper says otherwise. Past the trained length, raw RoPE can fail badly, with perplexity above 1,000. So the note stated the opposite hypothesis and tested it.
Setup: two models trained at a 512-token context on the Day 2 corpus, one with sinusoidal positions and one with RoPE. The control is sinusoidal because a learned table cannot run past 512 at all. Both were then evaluated at five lengths. A third column evaluates the RoPE model with interpolation, which divides every position by (evaluation length ÷ 512) so the angles stay inside the range seen in training.
What happened?
- RoPE beats sinusoidal at every length, by 0.76 bits per byte at 2,048.
- Both models improve from 512 to 2,048. A longer window gives each prediction more context, and fewer positions sit at the start of a block with little context.
- Both turn back up after 4,096, which is 8 times the trained length.
- Nothing collapses. At 16 times the trained length RoPE is worse than at 512 but nowhere near the failure the paper reports. The note's explanation is that the paper's models trained at 2,048, so they met far larger absolute distances than a model trained at 512 meets at 8,192.
- Interpolation helps only at 8,192, by 0.032. At 1,024 to 4,096 it costs 0.02 to 0.04, because squeezing the angles below what training saw hurts while raw extrapolation still works.
Decision: keep RoPE. Interpolate only past 4 times the trained length, and re-check that crossover after any change to the training context. The paper's other half, fine-tuning for 1,000 steps at the longer length, was not tested.
What we did not build, and why
- Fine-tuning at the longer length. The Position Interpolation result depends on it. It belonged to a later phase, and that phase was dropped.
- NTK-aware or YaRN scaling. Both change the base frequency instead of dividing positions. Plain interpolation already answered the question at our lengths.
- A different base. octlm keeps 10,000. Qwen's config sets its own base, and Day 6 reads it from
config.json.