4Day 4

The main run

26 million parameters, the Day 2 modern stack, 393 million TinyStories tokens, 4.1 hours in float16 on a Kaggle T4. Validation loss fell from 1.25 to 0.756, 0.484 bits per byte, and the curve was still sloping gently when the budget ran out.

Every model so far was too small or too undertrained to produce readable text or a readable architecture comparison. EXP-067 is the first run sized to do both. It ran once, on a Kaggle T4, for 4.1 hours, and this post walks through the setup, the curve it produced, and how to tell whether it was done.

The model

The Day 2 modern stack at width 512

configs/day4.toml takes every Day 2 keeper: RoPE, RMSNorm, SwiGLU and two KV heads, through SDPA, with pre-norm residuals. Width 512, 8 layers, 8 query heads, a 512-token context and the 8,192-entry TinyStories tokenizer.

Playground
Token embedding (shared with the output head)4,194,304
Learned position table0
Attention: query, key, value, output5,242,880
Feed-forward16,809,984
Norm gains and biases8,704
Parameters26,255,872
Feed-forward hidden width1,368⅔ × 4d, rounded to 8
float32 weights100.2 MiB
The Day 4 config, read from configs/day4.toml. The dry run reports 26,255,872 parameters, and so does this counter. Switch to Day 1 to compare.

The embedding is 16 percent of the model, down from 24 percent at Day 1 size. SwiGLU's three matrices hold most of the rest.

The budget

15 tokens per parameter

24,000 steps of 32 blocks of 512 tokens is 393,216,000 tokens, about 15 per parameter. That is short of the compute-optimal 20 and far past Day 2's 0.25. The training split has 966 million tokens, so no token repeats.

Playground
T4 throughput
Training compute6.2e+16 FLOPs6 × N × D
Tokens per parameter15.0compute-optimal is about 20
Share of the optimal budget75%
T4 time4.2 hat 4.1 TFLOPs, before overhead
Optimal tokens525M5.6 h on one T4
Compute and time for the main run. The measured 4.1 TFLOPs comes from EXP-066 on a Colab T4.

The learning rate peaks at 0.0006 after a 500-step warmup and follows a cosine to 0.00006. Weight decay is 0.1 and gradients clip at 1, the same as every earlier day.

Playground
0.00e+01.00e−42.00e−43.00e−44.00e−45.00e−46.00e−405,00010,00015,00020,000steplearning ratelearning rate
Hover the chart to read values.
Step 01.20e-6peak × 1 / warmup
Step 5006.00e-4peak
Step 122503.30e-4halfway down the cosine
Step 240006.00e-5floor
The Day 4 schedule. The warmup is about 2 percent of the run.
The host

Kaggle, one notebook version, 12 hours

The run moved from Colab to Kaggle on 2026-09-25. Kaggle's documentation, checked that day, gives a 12-hour limit per GPU session, 20 GB of saved output in /kaggle/working, and a T4 x2 option. Save & Run All starts a clean session and runs the notebook top to bottom: prepare the data, train, write samples.

The trainer writes a checkpoint at every evaluation, every 1000 steps, so a failure loses at most that many. A checkpoint survives into a later session only after its version's output is saved and attached as input to the next one. The notebook made no promise to resume a version that failed before saving.

octlm used one GPU of the pair, cuda:0. The Colab timing, about 4.2 hours, was the starting estimate, and the first Kaggle interval had to confirm it.

A run that looked stuck

Version 3 of the notebook was cancelled at 787 seconds because its log had stopped moving. It was not stuck. The log ended after the encode benchmark, and the next step, encoding the 2.2 GB training split, printed nothing until it finished. That step took 934 seconds on Colab, so the run was about halfway through it.

The fix was one line: encode_split now prints the token count every time it flushes 16 million tokens to disk. Version 4 ran with that line and finished. A long step with no output looks exactly like a hung one, and the cost of telling them apart was a wasted session.

The experiment

Hypothesis, measurement, stop conditions

Day 4's exit check asks for three things. Validation loss has flattened, ten fixed prompts produce stories a reader judges coherent, and bits per byte on held-out stories is recorded.

EXP-067

The curve

The run logged one record every 1,000 steps: the loss on that step's training batch, the validation loss and bits per byte on 131,072 held-out tokens, the learning rate and the gradient norm. All 24 records are below, read straight from the run's metrics.jsonl.

Playground
Metric
validationtrain batch
0.700.800.901.001.101.201.3005,00010,00015,00020,000stepnats per tokenvalidationtrain batch
Hover the chart to read values.
Validation loss, step 1,0001.2513
Validation loss, step 24,0000.7561
Average drop per 1,000 steps0.0215over steps 1,000 to 24,000

Train loss is one batch per record, so it jitters. Validation loss is the same 131,072 held-out tokens every time, so it is smooth. Slide the start to step 15,000 and the curve that looked flat from step 1,000 still slopes down. Gain per 1,000 uses a log scale, because the drops shrink by a factor of 100.

Every record of the Kaggle run. Switch the metric, or move the start of the x axis to zoom into the tail.

What happened?

The note records nine of the checkpoints as a table, parsed here at build time:

StepTrain lossValidation lossBits per byteElapsed s
1,0001.28741.25130.8009608
3,0000.96061.00700.64451,836
6,0000.83990.91800.58763,679
9,0000.82980.86950.55655,522
12,0000.84130.83750.53617,365
15,0000.76960.80690.51659,208
18,0000.77570.78160.500311,049
21,0000.69880.76480.489512,890
24,0000.72960.75610.483914,733
Is it done?

Flattened, not stopped

"Flat" needs a number. The stop condition says no improvement for 3,000 steps. That never happened. The last four 1,000-step drops in validation loss were 0.0055, 0.0045, 0.0025, 0.0017.

Each drop is smaller than the one before, and the last came while the learning rate sat at its floor of 0.00006. The model is still learning, slowly. A longer run would buy a little more, and it needs no new data: the model saw 393,216,000 of the 966 million encoded training tokens.

The exit check asked for a flattened curve, and this one qualifies. Decision: keep the step-24,000 checkpoint as the Day 4 model.

Kaggle against Colab

The same run on a different machine

Two numbers say Kaggle reproduced the Colab path. The step-1,000 validation loss was 1.2513, the same to four places as the Colab float16 run in EXP-066, 1.2513. Same seed and same data order gave the same result on a different machine.

The speed matched too. The whole run took 14,733 seconds for 24,000 steps, 0.614 seconds per step, against 0.625 on Colab. That is 26,690 tokens per second, or about 4.2 TFLOPs at 6 FLOPs per parameter per token. The 4.2-hour estimate held.

Skipped

What we did not build, and why