Roadmap

Nine days, one experiment at a time.

This page reads PLAN.md and the day notes at build time. A status changes when a note records a result. Days 1 to 3 followed the first plan. Days 4 to 9 follow the plan written on 2026-09-23.

Done · 16Built, not trained · 4Next · 0Planned · 15
  1. 1

    Foundations and the Generation 0 Transformer

  2. 2

    Modern decoder core

  3. 3

    Multi-token prediction and the attention variants

  4. 4

    Train the from-scratch model properly

  5. 5

    Compare at the new scale and make generation fast

    • EXP-069Seed spreadPlanned
    • EXP-070Multi-token prediction at depth 2 against the same controlPlanned
    • EXP-071Naive generation against a KV cachePlanned
    • EXP-072Int8 weight-only quantizationPlanned
  6. 6

    Run Qwen through octlm

    • EXP-073Weight loadingPlanned
    • EXP-074Logit parityPlanned
    • EXP-075Tokenizer parityPlanned
    • EXP-076KV-cached generation on QwenPlanned
  7. 7

    Build the harness and its eval

    • EXP-077Tool-call formatPlanned
    • EXP-078Tool loopPlanned
    • EXP-079Eval set and baselinePlanned
  8. 8

    LoRA and tool-use fine-tuning

    • EXP-080LoRA from scratchPlanned
    • EXP-081Training tracesPlanned
    • EXP-082SFT with LoRAPlanned
    • EXP-083Merge and quantizePlanned
  9. 9

    One optional extension

    Pick at most one extension. Not scheduled yet.

    Retired

    Where the first plan's experiments went

    The first plan had 64 experiments. When it changed, each one was kept, renumbered, folded in or dropped. This table comes from notes/day3.md.

    First planNow
    Stage 0 noise floorReplaced by EXP-069 at the new scale
    EXP-016 multi-token predictionCode kept. Runs as EXP-070
    EXP-017 sparse attentionCode kept behind default-off flags. Not scheduled
    EXP-018 compressed attention and copy probeCode kept behind default-off flags. Not scheduled
    EXP-019 MLACode kept behind default-off flags. Not scheduled
    EXP-020 to EXP-025, MoE, combined run, asymmetric compute, Muon, mHCDropped
    EXP-026 mixed precisionEXP-066, float16 only
    EXP-027 to EXP-029, batch size, schedule, checkpointingFolded into EXP-066 and EXP-067 where the run needs them
    EXP-030 scaling checkDropped
    EXP-031 to EXP-034, DDP, FSDP, tensor and pipeline parallelismDropped. One GPU
    EXP-035 and EXP-036, naive generation and KV cacheEXP-071, then EXP-076 on Qwen
    EXP-037 and EXP-038, prefill split and `torch.compile`Folded into EXP-071 and EXP-076 if time allows
    EXP-039 quantizationEXP-072, int8 only
    EXP-040 to EXP-042, batching, prefix cache, speculative decodingDropped. Cache reuse across turns lives in EXP-078
    EXP-043 SFTEXP-082
    EXP-044 LoRAEXP-080
    EXP-045 and EXP-046, DPO and reward modelDropped
    EXP-047 and EXP-048, GRPO and RLVRDay 9 option
    EXP-049 to EXP-053, evals and regression gateReplaced by the EXP-079 harness eval and bits per byte
    EXP-054 task classifierDropped
    EXP-055 and EXP-056, hybrid retrievalDropped. The harness has a `grep` tool
    EXP-057 tool loopEXP-078
    EXP-058 tool-call fine-tuneEXP-081 and EXP-082
    EXP-059 writing path, EXP-060 cache stackDropped
    EXP-061 router and cost logDay 9 option
    EXP-062 to EXP-064, long context, stability run, model cardDropped