Embeddings and positions
A token ID becomes a vector by looking up one row of a learned table. Attention cannot see order on its own, so a second table, or a formula, adds the position.
The tokenizer turns text into integers. The first layer of the model turns each integer into a vector, and every later layer works on those vectors. This post covers that first layer and the second piece of information it has to add, which is where in the sequence each token sits.
An embedding is a row lookup
An embedding table is a learned matrix with one row per vocabulary entry and columns, where is the model width. Token ID becomes row :
For the Day 1 model, and , so the table holds 131,072 numbers. A batch of token IDs shaped [B, T] becomes vectors shaped [B, T, C].
Mathematically the lookup is a one-hot vector times the matrix. In practice PyTorch's nn.Embedding indexes the row directly, which skips multiplying by 1,023 zeros. The gradient flows back only to the rows that were used, so a token that never appears in training keeps its initial random row.
The rows start random. Training moves them so that tokens used in similar contexts end up with similar vectors, because the layers above them are the same function for every token.
Attention does not know the order
Attention compares every position's query with every earlier position's key. Nothing in that comparison says which key is one step back and which is fifty. The causal mask supplies one fact, that a key is not in the future. Beyond that, if you shuffled the earlier tokens, each position would receive the same weighted average.
That is fatal for language. "The dog bit the man" and "The man bit the dog" would look the same to the last position. So the model needs position added to the vector before attention sees it.
A second table, indexed by position
The simplest fix is a second table, , with one row per position. The input to the first block is the sum:
This is what GPT-2 does and what octlm's Day 1 baseline does. The Day 1 table has 128 rows, one per position of its 128-token context.
What happened?
- The same word at two positions gets the same token row. Only the position row tells them apart in the sum.
- The sum keeps both signals in one vector of the same width. Nothing is concatenated, so the model width does not grow.
- Swapping to sinusoidal rows changes only the middle table. The token table is untouched.
A learned table has two limits:
- It stops at its last row. A model with a 128-row table has no vector for position 128. octlm raises an error rather than invent one. Day 2 recorded this as the reason the length experiment could not use learned positions as its control.
- Each row learns only from the positions training reaches. octlm packs training text into full-length blocks, so all 128 rows train. A model trained on shorter sequences would leave its later rows at their random starting values.
A formula instead of a table
The 2017 Transformer paper used fixed waves instead of a learned table. Column pair of position is a sine and a cosine at a frequency that falls with :
The first pair turns once every positions. The last pair turns about 10,000 times slower. Each position gets a unique pattern, and any length can be computed.
What happened?
- In the table, left columns change every row and right columns barely change. Fast pairs separate neighbors. Slow pairs separate distant positions.
- In the similarity map, the band along the diagonal has the same shape all the way down. The dot product of two sinusoidal position vectors is a sum of terms, so it depends only on the distance .
- Wider tables separate positions more sharply, because more frequencies take part.
That second point is why sinusoids were chosen. A fixed offset is a fixed linear map, so a model could learn to attend "3 back" once and use it anywhere. In practice the signal is added to the token vector and then mixed by the projections, which blurs it. Day 2 replaced both approaches with RoPE, which applies position inside attention itself.
The code
The embedding step handles all three position types. Learned and sinusoidal positions are added here. RoPE adds nothing here and acts later inside attention.
def _embed(self, token_ids: Tensor) -> Tensor:
x = self.token_embedding(token_ids)
if self.position_embedding is not None:
positions = torch.arange(token_ids.shape[1], device=token_ids.device)
return x + self.position_embedding(positions)
if self.config.position == "sinusoidal":
table = sinusoidal_positions(token_ids.shape[1], self.config.d_model, token_ids.device)
return x + table
return xThe sinusoidal table is built from the sequence length on every forward pass, so it has no maximum length:
def sinusoidal_positions(length: int, width: int, device: torch.device) -> Tensor:
position = torch.arange(length, device=device, dtype=torch.float32).unsqueeze(1)
index = torch.arange(0, width, 2, device=device, dtype=torch.float32)
angle = position / torch.pow(ROPE_BASE, index / width)
table = torch.zeros(length, width, device=device)
table[:, 0::2] = angle.sin()
table[:, 1::2] = angle.cos()
return tableA test on this site checks the TypeScript table on this page against this function to 1e-5.
What we measured, and what we chose not to build
Problem: we needed a way to inspect embeddings without a plotting dependency.
The benchmark prints the nearest token to token 0 and the nearest position to position 0 by cosine similarity. On the untrained benchmark model those neighbors are random, as they should be. The diagnostic only proves that the code path works. The Day 4 model is the first one trained long enough for its neighbors to mean something.
The Day 1 reading plan also covered pooling, which turns a sequence of vectors into one vector. Mean pooling averages [B, T, C] over into [B, C], and learned pooling adds weights to decide the mix. Both are for classifiers, which need one answer per sequence. A language model needs a prediction at every position, so neither belongs in its forward pass. The note keeps both as written exercises, and octlm has no pooling code.
What we did not build, and why
- Pooling layers. Explained above. Unused code is not written.
- Embedding plots. An untrained model's embedding space has no structure worth plotting.
- Separate input and output tables. octlm ties the output head to this same table. The baseline block covers why.