Building a small language modeland checking every step.
octlm is a decoder written from scratch in PyTorch. These posts explain each piece: what it does, how we built it, what it measured, and what we chose not to build. Every playground runs the same logic as the Python code, and tests check them against each other.
Day 0
Start here
2 posts
- What octlm is, and how it measuresThe goal, the rules every experiment follows, the laptop it started on, and the first two experiments, which built the lab before any model existed.
- What a language model predictsA language model assigns a probability to every possible next token. Cross entropy, perplexity and bits per byte are three ways of reading how much probability it gave the right one.
Day 1
Foundations and the Generation 0 Transformer
7 posts
- Characters as tokensThe simplest tokenizer gives every character its own ID. It is easy to build and easy to check, and it shows exactly what a better tokenizer has to fix.
- Byte-level BPE, merge by mergeStart from the 256 byte values, then repeatedly glue the most frequent adjacent pair into a new token. Nothing is ever unknown, and common fragments become single tokens.
- Measuring a tokenizerThe BPE model scored a perplexity eight times worse than the character model and was the better model. Bits per byte is the unit that makes the two comparable.
- Embeddings and positionsA token ID becomes a vector by looking up one row of a learned table. Attention cannot see order on its own, so a second table, or a formula, adds the position.
- Causal attention, by handHow each position in a sequence decides which earlier positions to read, traced step by step through octlm's attention layer with the weights PyTorch actually produced.
- The baseline blockA Transformer block is attention and a feed-forward layer, each wrapped in a norm and a residual connection. This post assembles the Day 1 decoder from those parts, including the initialization bug that stopped it learning.
- A training loop that resumes exactlyBatches, AdamW, gradient clipping, a warmup and cosine schedule, and checkpoints that restore every random generator, so a run split in two ends with the same weights as a run done in one go.
Day 2
Modern decoder core
8 posts
- A corpus big enough to measureDay 1 trained on 111 KB of planning notes. Day 2 needed held-out code and prose to compare architectures, so it built a 7 MB corpus from the Python standard library and six public-domain books.
- RoPE, position as a rotationRotary position embedding turns each query and key by an angle proportional to its position, so attention scores depend on how far apart two tokens are. Day 2 measured how far past its training length it holds up.
- RMSNorm against LayerNormRMSNorm drops LayerNorm's mean subtraction and its bias. On Day 2 it matched LayerNorm's quality inside the noise and ran slower, for a reason that had nothing to do with arithmetic.
- SwiGLU, a gated feed-forward layerSwiGLU replaces the feed-forward layer's single activation with a product of two projections, one of them passed through a smooth gate. It was the only single change on Day 2 that beat the noise.
- Grouped-query attention and the KV cacheGeneration stores every past key and value, and that cache grows with the number of key and value heads. Grouped-query attention shares them across query heads. Day 2 measured the saving exactly and the quality cost not at all, because the cost was below the noise.
- SDPA and FlashAttention's tilingPlain attention builds a T by T score matrix in memory. FlashAttention computes the same result a block at a time with a running maximum and sum, and never stores the square. Day 2 wrote the tiling by hand and measured PyTorch's kernels against the plain path.
- Where the norm goesThe 2017 Transformer normalized after each residual addition. Modern decoders normalize before each block instead. Day 2 swapped the order back and lost a full bit per byte.
- Reading a grid through seed noiseEight architectures, three seeds each, one budget. Before reading any gap between two rows, read how far one row moves when only the random seed changes. On Day 2 that one number decided almost everything.
Day 3
Multi-token prediction and the attention variants
5 posts
- Multi-token predictionTrain extra heads to predict the token two and three steps ahead from the same hidden state. The code is built and tested, the plan predicted no gain at 3.3M parameters, and the run moved to Day 5.
- Sparse attentionA sliding window lets each token read only its recent neighbors, and strided columns add a few long-range links. The masks are built and tested, and the run was dropped when the plan changed.
- Compressed attention and the copy probeKeep the recent keys exactly and replace older ones with block averages, for a fraction of the cache. Day 3 built it, fixed a subtle problem with RoPE, and designed a probe that a 3M model could actually pass.
- Multi-head latent attentionMLA caches one small shared vector per token instead of per-head keys and values. Against multi-head attention that is a large saving. Against the two-KV-head model octlm already had, the arithmetic said it was small, and the note wrote that down before building it.
- Why the plan changedOn 2026-09-23 the project stopped before its Day 3 runs and rewrote its plan. Two facts forced it. The Day 2 models had seen about 1 percent of the data they needed, and no model this lab can pretrain could drive a tool-use harness.
Day 4
Train the from-scratch model properly
4 posts
- TinyStories and a 6x faster encoderThe Day 4 model needs about 400 million tokens. The pure-Python BPE encoder was expected to be too slow for 2.2 GB of text. One dictionary from chunk to IDs made it six times faster, and the whole training split encoded in 16 minutes.
- Mixed precision on a T4Running the matrix math in 16-bit floats nearly doubled training speed with no measurable loss. The first attempt picked bfloat16, which the T4 only emulates, and ran slower than float32.
- The main run26 million parameters, the Day 2 modern stack, 393 million TinyStories tokens, 4.1 hours in float16 on a Kaggle T4. Validation loss fell from 1.25 to 0.756, 0.484 bits per byte, and the curve was still sloping gently when the budget ran out.
- SamplingA model outputs a distribution, and something has to pick a token from it. Greedy decoding repeats itself on stories. Temperature reshapes the distribution and top-k cuts its tail, and both are one line each. On the trained Day 4 model, both decoders wrote coherent stories, and sampling traded the loops for odder plot turns.