The main run
26 million parameters, the Day 2 modern stack, 393 million TinyStories tokens, 4.1 hours in float16 on a Kaggle T4. Validation loss fell from 1.25 to 0.756, 0.484 bits per byte, and the curve was still sloping gently when the budget ran out.
Every model so far was too small or too undertrained to produce readable text or a readable architecture comparison. EXP-067 is the first run sized to do both. It ran once, on a Kaggle T4, for 4.1 hours, and this post walks through the setup, the curve it produced, and how to tell whether it was done.
The Day 2 modern stack at width 512
configs/day4.toml takes every Day 2 keeper: RoPE, RMSNorm, SwiGLU and two KV heads, through SDPA, with pre-norm residuals. Width 512, 8 layers, 8 query heads, a 512-token context and the 8,192-entry TinyStories tokenizer.
The embedding is 16 percent of the model, down from 24 percent at Day 1 size. SwiGLU's three matrices hold most of the rest.
15 tokens per parameter
24,000 steps of 32 blocks of 512 tokens is 393,216,000 tokens, about 15 per parameter. That is short of the compute-optimal 20 and far past Day 2's 0.25. The training split has 966 million tokens, so no token repeats.
The learning rate peaks at 0.0006 after a 500-step warmup and follows a cosine to 0.00006. Weight decay is 0.1 and gradients clip at 1, the same as every earlier day.
Kaggle, one notebook version, 12 hours
The run moved from Colab to Kaggle on 2026-09-25. Kaggle's documentation, checked that day, gives a 12-hour limit per GPU session, 20 GB of saved output in /kaggle/working, and a T4 x2 option. Save & Run All starts a clean session and runs the notebook top to bottom: prepare the data, train, write samples.
The trainer writes a checkpoint at every evaluation, every 1000 steps, so a failure loses at most that many. A checkpoint survives into a later session only after its version's output is saved and attached as input to the next one. The notebook made no promise to resume a version that failed before saving.
octlm used one GPU of the pair, cuda:0. The Colab timing, about 4.2 hours, was the starting estimate, and the first Kaggle interval had to confirm it.
A run that looked stuck
Version 3 of the notebook was cancelled at 787 seconds because its log had stopped moving. It was not stuck. The log ended after the encode benchmark, and the next step, encoding the 2.2 GB training split, printed nothing until it finished. That step took 934 seconds on Colab, so the run was about halfway through it.
The fix was one line: encode_split now prints the token count every time it flushes 16 million tokens to disk. Version 4 ran with that line and finished. A long step with no output looks exactly like a hung one, and the cost of telling them apart was a wasted session.
Hypothesis, measurement, stop conditions
- Problem. No model so far was trained enough to read an architecture effect or produce text.
- Hypothesis. About 26 million parameters on about 393 million tokens writes coherent short stories.
- Baseline. None at this scale. The exit check is absolute.
- Measurement. Validation loss and bits per byte on the first 256 non-overlapping 512-token blocks of the validation split, every 1,000 steps. The GPU name, PyTorch version and elapsed time from the first and last records.
- Stop. If validation loss has not improved for 3,000 steps, stop and keep the best checkpoint. If the loss diverges, halve the learning rate and restart. If Kaggle's first interval is much slower than Colab's, or the run cannot finish in 12 hours, stop and choose the length from the measured Kaggle speed.
Day 4's exit check asks for three things. Validation loss has flattened, ten fixed prompts produce stories a reader judges coherent, and bits per byte on held-out stories is recorded.
The curve
The run logged one record every 1,000 steps: the loss on that step's training batch, the validation loss and bits per byte on 131,072 held-out tokens, the learning rate and the gradient norm. All 24 records are below, read straight from the run's metrics.jsonl.
What happened?
- Validation loss fell from 1.2513 at step 1,000 to 0.7561 at step 24,000. Most of that drop happened in the first quarter of the run.
- Bits per byte fell from 0.8009 to 0.4839. The model spends under half a bit on each byte of a story it has never seen.
- Perplexity ended at 2.13. On average the model is as unsure as a fair choice between about two tokens.
- Train loss sits below and above validation loss by turns. It is one batch, so it jitters. The two never separate, which means no overfitting: every training token was new.
- The gradient norm stayed between 0.32 and 0.64 at every record, under the clip of 1. It drifted up a little as the learning rate fell. The loss never spiked, and the divergence stop never fired.
The note records nine of the checkpoints as a table, parsed here at build time:
| Step | Train loss | Validation loss | Bits per byte | Elapsed s |
|---|---|---|---|---|
| 1,000 | 1.2874 | 1.2513 | 0.8009 | 608 |
| 3,000 | 0.9606 | 1.0070 | 0.6445 | 1,836 |
| 6,000 | 0.8399 | 0.9180 | 0.5876 | 3,679 |
| 9,000 | 0.8298 | 0.8695 | 0.5565 | 5,522 |
| 12,000 | 0.8413 | 0.8375 | 0.5361 | 7,365 |
| 15,000 | 0.7696 | 0.8069 | 0.5165 | 9,208 |
| 18,000 | 0.7757 | 0.7816 | 0.5003 | 11,049 |
| 21,000 | 0.6988 | 0.7648 | 0.4895 | 12,890 |
| 24,000 | 0.7296 | 0.7561 | 0.4839 | 14,733 |
Flattened, not stopped
"Flat" needs a number. The stop condition says no improvement for 3,000 steps. That never happened. The last four 1,000-step drops in validation loss were 0.0055, 0.0045, 0.0025, 0.0017.
Each drop is smaller than the one before, and the last came while the learning rate sat at its floor of 0.00006. The model is still learning, slowly. A longer run would buy a little more, and it needs no new data: the model saw 393,216,000 of the 966 million encoded training tokens.
The exit check asked for a flattened curve, and this one qualifies. Decision: keep the step-24,000 checkpoint as the Day 4 model.
The same run on a different machine
Two numbers say Kaggle reproduced the Colab path. The step-1,000 validation loss was 1.2513, the same to four places as the Colab float16 run in EXP-066, 1.2513. Same seed and same data order gave the same result on a different machine.
The speed matched too. The whole run took 14,733 seconds for 24,000 steps, 0.614 seconds per step, against 0.625 on Colab. That is 26,690 tokens per second, or about 4.2 TFLOPs at 6 FLOPs per parameter per token. The 4.2-hour estimate held.
What we did not build, and why
- A longer run. The curve still slopes, but the exit check passed. Day 5's comparisons run at about 100 million tokens, a quarter of this budget, so a longer Day 4 run would not feed them.
- The second T4. Kaggle's T4 x2 option has two GPUs. Using both needs distributed data parallel code that no later day uses. octlm ran on
cuda:0alone. - The 4,096-entry vocabulary. Comparing it costs a second full run. It moved to Day 5, if GPU time allows.