3Day 3

Why the plan changed

On 2026-09-23 the project stopped before its Day 3 runs and rewrote its plan. Two facts forced it. The Day 2 models had seen about 1 percent of the data they needed, and no model this lab can pretrain could drive a tool-use harness.

The first plan aimed to pretrain a 50 to 150 million parameter model and grow it, over 64 experiments, into a local coding and writing assistant. Day 3 had four experiments built and ready for a GPU. Before the first run, the question came up whether the project was on the right track. Working through our own numbers said it was not.

Fact 1

The Day 2 models were starved of data

Hoffmann et al. found that for a fixed compute budget, the best loss comes from training on about 20 tokens per parameter. Fewer tokens means a larger model than the budget can use.

The Day 2 grid trained models of about 3.3 million parameters for 400 steps of 8 blocks of 256 tokens, 819,200 tokens each. That is about 0.25 tokens per parameter, 1 percent of the compute-optimal budget. The Day 3 plan raised the steps fivefold, to 4.1 million tokens, which is still about 1.1 tokens per parameter. Its long-context config reached 8.2 million tokens, and the whole Day 2 corpus holds 2.96 million, so it would have repeated data.

That explains the Day 2 grid. A seed spread of 0.043 bits per byte, larger than most of the effects being measured, is what an undertrained model looks like. Where it ends up depends heavily on where it started. Stage 0 would have measured that spread again, at a budget still about 18 times too small.

Playground
T4 throughput
Training compute1.8e+13 FLOPs6 × N × D
Tokens per parameter0.22compute-optimal is about 20
Share of the optimal budget1%
T4 time0 minat 4.1 TFLOPs, before overhead
Optimal tokens75M7 min on one T4
Training compute is about 6 × parameters × tokens: 2 for the forward pass and 4 for the backward pass, per parameter per token. Time assumes one T4 at a sustained throughput. Presets are the runs this project did or considered.

What happened?

That last row is why pretraining stops near 20 million parameters.

Fact 2

The goal needs a model that follows instructions

The project's goal, in the user's words, is to learn machine learning "in depth" for placement season and to end with "a harness that can control my own model". A harness shows the model a task and some tools, parses each tool call it emits, runs the tool, and feeds back the result. That needs a model that follows instructions and produces well-formed tool calls. No model that can be pretrained on free Colab does either.

The decision

Keep the from-scratch model, then move the same code to Qwen

The new plan has two halves:

  1. Train the from-scratch model once, properly. About 20 million parameters, on TinyStories, with the Day 2 modern stack. TinyStories is a dataset on which models under about 30 million parameters write coherent English, and the run fits the measured compute. This is Day 4 and Day 5.
  2. Load Qwen into the same model code. octlm/model.py already has RoPE, RMSNorm, SwiGLU and grouped-query attention, the same parts Qwen uses. Match its logits, write LoRA, fine-tune it for tool calls, and build the harness and its eval around it. This is Day 6 to Day 9.

Two alternatives were rejected. Continuing the first plan spends the remaining time on variants the harness never uses and ends with a model too weak to drive it. Fine-tuning Qwen with an off-the-shelf trainer gives the harness a working model and leaves nothing to explain when an interviewer asks how attention, the KV cache or LoRA work.

What was dropped

Where each first-plan experiment went

Every first-plan experiment got a new number, a merge into another experiment, or a reason it was dropped. New numbers start at EXP-065, so no number ever means two things. The table is read from notes/day3.md when the site builds.

Map
Show
First planNow
Stage 0 noise floorReplaced by EXP-069 at the new scale
EXP-016 multi-token predictionCode kept. Runs as EXP-070
EXP-017 sparse attentionCode kept behind default-off flags. Not scheduled
EXP-018 compressed attention and copy probeCode kept behind default-off flags. Not scheduled
EXP-019 MLACode kept behind default-off flags. Not scheduled
EXP-020 to EXP-025, MoE, combined run, asymmetric compute, Muon, mHCDropped
EXP-026 mixed precisionEXP-066, float16 only
EXP-027 to EXP-029, batch size, schedule, checkpointingFolded into EXP-066 and EXP-067 where the run needs them
EXP-030 scaling checkDropped
EXP-031 to EXP-034, DDP, FSDP, tensor and pipeline parallelismDropped. One GPU
EXP-035 and EXP-036, naive generation and KV cacheEXP-071, then EXP-076 on Qwen
EXP-037 and EXP-038, prefill split and `torch.compile`Folded into EXP-071 and EXP-076 if time allows
EXP-039 quantizationEXP-072, int8 only
EXP-040 to EXP-042, batching, prefix cache, speculative decodingDropped. Cache reuse across turns lives in EXP-078
EXP-043 SFTEXP-082
EXP-044 LoRAEXP-080
EXP-045 and EXP-046, DPO and reward modelDropped
EXP-047 and EXP-048, GRPO and RLVRDay 9 option
EXP-049 to EXP-053, evals and regression gateReplaced by the EXP-079 harness eval and bits per byte
EXP-054 task classifierDropped
EXP-055 and EXP-056, hybrid retrievalDropped. The harness has a `grep` tool
EXP-057 tool loopEXP-078
EXP-058 tool-call fine-tuneEXP-081 and EXP-082
EXP-059 writing path, EXP-060 cache stackDropped
EXP-061 router and cost logDay 9 option
EXP-062 to EXP-064, long context, stability run, model cardDropped
Every first-plan experiment and where it went, read from notes/day3.md. Filter by outcome.

The dropped groups share reasons:

The rules

Conflicts the change created, and how they were recorded

AGENTS.md forbids starting a day before every exit check of the previous day passes. None of Day 3's passed. The user's request outranks the plan, so the Day 3 checks were recorded as withdrawn, not failed, and Day 4 starts from Day 2's passed checks.

AGENTS.md also says every tokenizer learns its merges from our own data. Qwen's tokenizer ships with its weights. The rule keeps applying to every tokenizer octlm trains. For Qwen, the checkpoint loader applies the hash check to tokenizer.json instead.