2Day 2

Reading a grid through seed noise

Eight architectures, three seeds each, one budget. Before reading any gap between two rows, read how far one row moves when only the random seed changes. On Day 2 that one number decided almost everything.

Day 2 ran every component comparison as one grid on a Colab T4. Each variant changes one thing about the Day 1 baseline, except modern, which combines the four keepers. Every variant trained with three different random seeds at the same budget. This post reads the whole grid, and the lesson it taught is about the reading, not about any one component.

The setup

What was held fixed

The grid reports bits per byte separately on held-out code and prose, the mean over three seeds, and the spread, which is the largest seed's result minus the smallest.

Measured
Held-out split
baseline seed spread2.803.003.203.403.603.804.00code bits per byte, lower is betterbaselinermsnormswiglupost-normgqa-4gqa-2mqa-1modern
Hover a row to read its numbers.
All eight variants. The dot is the mean over three seeds. The bar's length is the spread across seeds, drawn centered on the mean because the note records the range, not where each seed fell. The gray band is the baseline's spread. Switch between code and prose. Hover a row for exact numbers.
The noise floor

Read the spread before the gap

The baseline, with nothing changed except the seed, moved 0.0434 bits per byte on code across its three runs. The seed changes the initial weights and the order of the batches, and at 400 steps that is enough to move the result by 0.04. Any difference between two variants smaller than that could come from the seeds alone.

Read that way, the grid says:

VariantChangeCode gap vs baselineVerdict
rmsnormRMSNorm0.0013Inside the noise
swigluSwiGLU-0.0818Outside the noise, keep
post-normPost-norm0.9971Far outside, revert
gqa-44 KV heads-0.0342Inside the noise
gqa-22 KV heads0.0110Inside the noise, keep for the cache
mqa-11 KV head-0.0051Inside the noise
modernRoPE, RMSNorm, SwiGLU, 2 KV heads-0.1294Best mean, widest spread

Negative is better. Two of six single changes produced a readable result, one good and one bad. The rest are indistinguishable from the baseline at this budget. "Indistinguishable" is not "equal": a real 0.01 improvement would look exactly like this.

Playground
Held-out text
baseline 2.9339swiglu 2.8521betterworse
Gap between means0.0818bits per byte
Half spreads added0.0260
The ranges do not overlap. The gap is bigger than the seed noise.
Pick any two variants from the grid. Each bar is that variant's seed range, the dot its mean. If the bars overlap, three seeds cannot separate them.
The open question

modern against swiglu, unresolved

The combined stack has the best code and prose means in the grid, on a quarter of the baseline's cache. Its margin over swiglu alone is 0.048 on code. Its own seed spread is 0.1194, almost three times the baseline's. The grid ranks the two and cannot separate them.

The note's diagnosis is that the fix is steps, not seeds. More seeds would give a better estimate of a noisy number. More training would make the number less noisy, because an undertrained model's result depends heavily on where it started. At 3.27 million parameters and 2.96 million training tokens, each model saw about one token per parameter. Compute-optimal training wants about 20.

What it led to

The grid is why the plan changed

Day 3 planned to fix this by raising the budget fivefold and remeasuring the noise first, as "Stage 0". Working out the numbers showed that even that budget would be about 18 times too small, and the corpus would repeat. Why the plan changed covers the arithmetic. Day 5 measures the seed spread again at a budget where the models are trained properly, and that spread becomes the smallest effect a comparison may claim.

The rule that survives is that every comparison states its noise floor first, and a gap inside it is recorded as unresolved with its numbers, not as a win.

A housekeeping note

Where these numbers live

The 24 runs happened on a Colab VM, and their JSONL never came back to the repository. runs/day2-variants.jsonl holds only two records from an earlier laptop run. The table in notes/day2.md is the record, and this page parses it at build time. The note warns that a local octlm.day2 report summarizes those two rows and must not be read as the Day 2 result.