Reading a grid through seed noise
Eight architectures, three seeds each, one budget. Before reading any gap between two rows, read how far one row moves when only the random seed changes. On Day 2 that one number decided almost everything.
Day 2 ran every component comparison as one grid on a Colab T4. Each variant changes one thing about the Day 1 baseline, except modern, which combines the four keepers. Every variant trained with three different random seeds at the same budget. This post reads the whole grid, and the lesson it taught is about the reading, not about any one component.
What was held fixed
- Width 256, 4 layers, 8 query heads and a 256-token context, about 3.3 million parameters.
- 400 steps at batch 8, which is 819,200 training tokens per run.
- The same corpus, tokenizer and validation blocks for every run.
- Seeds 1337, 1338 and 1339, the same three for every variant.
- The grid's 24 runs took 249 seconds of training and evaluation on the T4, about 11 times faster per step than the laptop.
The grid reports bits per byte separately on held-out code and prose, the mean over three seeds, and the spread, which is the largest seed's result minus the smallest.
Read the spread before the gap
The baseline, with nothing changed except the seed, moved 0.0434 bits per byte on code across its three runs. The seed changes the initial weights and the order of the batches, and at 400 steps that is enough to move the result by 0.04. Any difference between two variants smaller than that could come from the seeds alone.
Read that way, the grid says:
| Variant | Change | Code gap vs baseline | Verdict |
|---|---|---|---|
| rmsnorm | RMSNorm | 0.0013 | Inside the noise |
| swiglu | SwiGLU | -0.0818 | Outside the noise, keep |
| post-norm | Post-norm | 0.9971 | Far outside, revert |
| gqa-4 | 4 KV heads | -0.0342 | Inside the noise |
| gqa-2 | 2 KV heads | 0.0110 | Inside the noise, keep for the cache |
| mqa-1 | 1 KV head | -0.0051 | Inside the noise |
| modern | RoPE, RMSNorm, SwiGLU, 2 KV heads | -0.1294 | Best mean, widest spread |
Negative is better. Two of six single changes produced a readable result, one good and one bad. The rest are indistinguishable from the baseline at this budget. "Indistinguishable" is not "equal": a real 0.01 improvement would look exactly like this.
modern against swiglu, unresolved
The combined stack has the best code and prose means in the grid, on a quarter of the baseline's cache. Its margin over swiglu alone is 0.048 on code. Its own seed spread is 0.1194, almost three times the baseline's. The grid ranks the two and cannot separate them.
The note's diagnosis is that the fix is steps, not seeds. More seeds would give a better estimate of a noisy number. More training would make the number less noisy, because an undertrained model's result depends heavily on where it started. At 3.27 million parameters and 2.96 million training tokens, each model saw about one token per parameter. Compute-optimal training wants about 20.
The grid is why the plan changed
Day 3 planned to fix this by raising the budget fivefold and remeasuring the noise first, as "Stage 0". Working out the numbers showed that even that budget would be about 18 times too small, and the corpus would repeat. Why the plan changed covers the arithmetic. Day 5 measures the seed spread again at a budget where the models are trained properly, and that spread becomes the smallest effect a comparison may claim.
The rule that survives is that every comparison states its noise floor first, and a gap inside it is recorded as unresolved with its numbers, not as a win.
Where these numbers live
The 24 runs happened on a Colab VM, and their JSONL never came back to the repository. runs/day2-variants.jsonl holds only two records from an earlier laptop run. The table in notes/day2.md is the record, and this page parses it at build time. The note warns that a local octlm.day2 report summarizes those two rows and must not be read as the Day 2 result.