1Day 1

Embeddings and positions

A token ID becomes a vector by looking up one row of a learned table. Attention cannot see order on its own, so a second table, or a formula, adds the position.

The tokenizer turns text into integers. The first layer of the model turns each integer into a vector, and every later layer works on those vectors. This post covers that first layer and the second piece of information it has to add, which is where in the sequence each token sits.

Definition

An embedding is a row lookup

An embedding table is a learned matrix EE with one row per vocabulary entry and CC columns, where CC is the model width. Token ID ii becomes row EiE_i:

E∈RV×C,embed(i)=EiE \in \mathbb{R}^{V \times C}, \qquad \text{embed}(i) = E_i

For the Day 1 model, V=1,024V = 1{,}024 and C=128C = 128, so the table holds 131,072 numbers. A batch of token IDs shaped [B, T] becomes vectors shaped [B, T, C].

Mathematically the lookup is a one-hot vector times the matrix. In practice PyTorch's nn.Embedding indexes the row directly, which skips multiplying by 1,023 zeros. The gradient flows back only to the rows that were used, so a token that never appears in training keeps its initial random row.

The rows start random. Training moves them so that tokens used in similar contexts end up with similar vectors, because the layers above them are the same function for every token.

Playground
one-hot, 1 × V1 × 6
012345.001.00.00.00.00.00
table E, V × C6 × 5
012340−.98−.88.95.40.041−.19−.07−.52.11.462−.48−.69.53.04−.613−.26−.41.07−.41.964−.50.76−.62−.71−.695−.42.10.52−.85.83
one-hot × E, 1 × C1 × 5
01234−.19−.07−.52.11.46

Multiplying a one-hot vector by the table picks out one row, because every other row is multiplied by 0. That is why an embedding is written as a matrix product in papers and computed as an index in code: E[1] gives the same numbers with no multiplications. The gradient reaches only row 1.

Pick a token. The one-hot vector times the embedding table equals the highlighted row. Values are random, as they are at initialization.
The problem

Attention does not know the order

Attention compares every position's query with every earlier position's key. Nothing in that comparison says which key is one step back and which is fifty. The causal mask supplies one fact, that a key is not in the future. Beyond that, if you shuffled the earlier tokens, each position would receive the same weighted average.

That is fatal for language. "The dog bit the man" and "The man bit the dog" would look the same to the last position. So the model needs position added to the vector before attention sees it.

Learned positions

A second table, indexed by position

The simplest fix is a second table, P∈RTmax⁡×CP \in \mathbb{R}^{T_{\max} \times C}, with one row per position. The input to the first block is the sum:

ht(0)=Ext+Pth^{(0)}_t = E_{x_t} + P_t

This is what GPT-2 does and what octlm's Day 1 baseline does. The Day 1 table has 128 rows, one per position of its 128-token context.

Playground
Positions
token_embedding[ids]6 × 8
01234567thecatsatonthemat
position rows 0..56 × 8
01234567p0p1p2p3p4p5
sum = model input6 × 8
01234567thecatsatonthemat

Rows with the same word share one token row. Set pos 0 and pos 4 to "the": their token rows match, and the sum differs only by the position row. Without that row, attention could not tell the two apart.

Six token IDs, looked up in a small random token table, plus position rows. Pick tokens with the selects. Switch between learned rows (random here, as they are at initialization) and the fixed sinusoidal rows.

What happened?

A learned table has two limits:

  1. It stops at its last row. A model with a 128-row table has no vector for position 128. octlm raises an error rather than invent one. Day 2 recorded this as the reason the length experiment could not use learned positions as its control.
  2. Each row learns only from the positions training reaches. octlm packs training text into full-length blocks, so all 128 rows train. A model trained on shorter sequences would leave its later rows at their random starting values.
Sinusoidal positions

A formula instead of a table

The 2017 Transformer paper used fixed waves instead of a learned table. Column pair ii of position tt is a sine and a cosine at a frequency that falls with ii:

Pt, 2i=sin⁡ ⁣(t100002i/C),Pt, 2i+1=cos⁡ ⁣(t100002i/C)P_{t,\,2i} = \sin\!\left(\frac{t}{10000^{2i/C}}\right), \qquad P_{t,\,2i+1} = \cos\!\left(\frac{t}{10000^{2i/C}}\right)

The first pair turns once every 2π2\pi positions. The last pair turns about 10,000 times slower. Each position gets a unique pattern, and any length can be computed.

Playground
sinusoidal_positions(T, C)24 × 16
012345678910111213141501234567891011121314151617181920212223
cosine similarity between positions24 × 24
0123456789101112131415161718192021222301234567891011121314151617181920212223

Left: each row is one position. Columns come in sin and cos pairs, and the pair frequency falls from left to right, so the first columns flicker and the last barely move. Right: row i, column j compares position i with position j. The bright diagonal band says nearby positions look alike, and the band has the same width everywhere, because similarity depends on the distance i − j.

The sinusoidal table from octlm's formula, and the cosine similarity between every pair of positions. Change the width and the number of positions.

What happened?

That second point is why sinusoids were chosen. A fixed offset is a fixed linear map, so a model could learn to attend "3 back" once and use it anywhere. In practice the signal is added to the token vector and then mixed by the projections, which blurs it. Day 2 replaced both approaches with RoPE, which applies position inside attention itself.

In octlm

The code

The embedding step handles all three position types. Learned and sinusoidal positions are added here. RoPE adds nothing here and acts later inside attention.

octlm/model.pyline 401
def _embed(self, token_ids: Tensor) -> Tensor:
    x = self.token_embedding(token_ids)
    if self.position_embedding is not None:
        positions = torch.arange(token_ids.shape[1], device=token_ids.device)
        return x + self.position_embedding(positions)
    if self.config.position == "sinusoidal":
        table = sinusoidal_positions(token_ids.shape[1], self.config.d_model, token_ids.device)
        return x + table
    return x

The sinusoidal table is built from the sequence length on every forward pass, so it has no maximum length:

octlm/model.pyline 150
def sinusoidal_positions(length: int, width: int, device: torch.device) -> Tensor:
    position = torch.arange(length, device=device, dtype=torch.float32).unsqueeze(1)
    index = torch.arange(0, width, 2, device=device, dtype=torch.float32)
    angle = position / torch.pow(ROPE_BASE, index / width)
    table = torch.zeros(length, width, device=device)
    table[:, 0::2] = angle.sin()
    table[:, 1::2] = angle.cos()
    return table

A test on this site checks the TypeScript table on this page against this function to 1e-5.

EXP-005

What we measured, and what we chose not to build

Problem: we needed a way to inspect embeddings without a plotting dependency.

The benchmark prints the nearest token to token 0 and the nearest position to position 0 by cosine similarity. On the untrained benchmark model those neighbors are random, as they should be. The diagnostic only proves that the code path works. The Day 4 model is the first one trained long enough for its neighbors to mean something.

The Day 1 reading plan also covered pooling, which turns a sequence of vectors into one vector. Mean pooling averages [B, T, C] over TT into [B, C], and learned pooling adds weights to decide the mix. Both are for classifiers, which need one answer per sequence. A language model needs a prediction at every position, so neither belongs in its forward pass. The note keeps both as written exercises, and octlm has no pooling code.

Skipped

What we did not build, and why