Roadmap
Nine days, one experiment at a time.
This page reads PLAN.md and the day notes at build time. A status changes when a note records a result. Days 1 to 3 followed the first plan. Days 4 to 9 follow the plan written on 2026-09-23.
Done · 16Built, not trained · 4Next · 0Planned · 15
- 1
Foundations and the Generation 0 Transformer
EXP-001Repository and measurementsDoneEXP-002Hardware limitDoneEXP-003Character tokenizerDoneEXP-004Byte-level BPEDoneEXP-005Embeddings and positionsDoneEXP-006Causal multi-head attentionDoneEXP-007Baseline blockDoneEXP-008Training and resumeDone
- 2
Modern decoder core
EXP-009The length sweepDoneEXP-010 to EXP-013 and EXP-015The variant gridDoneEXP-013KV cache bytes at the Day 2 shapeDoneEXP-014SDPA and the tiled sketchDone
- 3
Multi-token prediction and the attention variants
EXP-016Multi-token predictionBuilt, not trainedEXP-017Sparse attentionBuilt, not trainedEXP-018Compressed attention and the needle probeBuilt, not trainedEXP-019MLABuilt, not trained
- 4
Train the from-scratch model properly
EXP-065TinyStories corpus and tokenizerDoneEXP-066Mixed precisionDoneEXP-067The main runDoneEXP-068SamplingDone
- 5
Compare at the new scale and make generation fast
EXP-069Seed spreadPlannedEXP-070Multi-token prediction at depth 2 against the same controlPlannedEXP-071Naive generation against a KV cachePlannedEXP-072Int8 weight-only quantizationPlanned
- 6
Run Qwen through octlm
EXP-073Weight loadingPlannedEXP-074Logit parityPlannedEXP-075Tokenizer parityPlannedEXP-076KV-cached generation on QwenPlanned
- 7
Build the harness and its eval
EXP-077Tool-call formatPlannedEXP-078Tool loopPlannedEXP-079Eval set and baselinePlanned
- 8
LoRA and tool-use fine-tuning
EXP-080LoRA from scratchPlannedEXP-081Training tracesPlannedEXP-082SFT with LoRAPlannedEXP-083Merge and quantizePlanned
- 9
One optional extension
Pick at most one extension. Not scheduled yet.
Retired
Where the first plan's experiments went
The first plan had 64 experiments. When it changed, each one was kept, renumbered, folded in or dropped. This table comes from notes/day3.md.
| First plan | Now |
|---|---|
| Stage 0 noise floor | Replaced by EXP-069 at the new scale |
| EXP-016 multi-token prediction | Code kept. Runs as EXP-070 |
| EXP-017 sparse attention | Code kept behind default-off flags. Not scheduled |
| EXP-018 compressed attention and copy probe | Code kept behind default-off flags. Not scheduled |
| EXP-019 MLA | Code kept behind default-off flags. Not scheduled |
| EXP-020 to EXP-025, MoE, combined run, asymmetric compute, Muon, mHC | Dropped |
| EXP-026 mixed precision | EXP-066, float16 only |
| EXP-027 to EXP-029, batch size, schedule, checkpointing | Folded into EXP-066 and EXP-067 where the run needs them |
| EXP-030 scaling check | Dropped |
| EXP-031 to EXP-034, DDP, FSDP, tensor and pipeline parallelism | Dropped. One GPU |
| EXP-035 and EXP-036, naive generation and KV cache | EXP-071, then EXP-076 on Qwen |
| EXP-037 and EXP-038, prefill split and `torch.compile` | Folded into EXP-071 and EXP-076 if time allows |
| EXP-039 quantization | EXP-072, int8 only |
| EXP-040 to EXP-042, batching, prefix cache, speculative decoding | Dropped. Cache reuse across turns lives in EXP-078 |
| EXP-043 SFT | EXP-082 |
| EXP-044 LoRA | EXP-080 |
| EXP-045 and EXP-046, DPO and reward model | Dropped |
| EXP-047 and EXP-048, GRPO and RLVR | Day 9 option |
| EXP-049 to EXP-053, evals and regression gate | Replaced by the EXP-079 harness eval and bits per byte |
| EXP-054 task classifier | Dropped |
| EXP-055 and EXP-056, hybrid retrieval | Dropped. The harness has a `grep` tool |
| EXP-057 tool loop | EXP-078 |
| EXP-058 tool-call fine-tune | EXP-081 and EXP-082 |
| EXP-059 writing path, EXP-060 cache stack | Dropped |
| EXP-061 router and cost log | Day 9 option |
| EXP-062 to EXP-064, long context, stability run, model card | Dropped |