2Day 2

A corpus big enough to measure

Day 1 trained on 111 KB of planning notes. Day 2 needed held-out code and prose to compare architectures, so it built a 7 MB corpus from the Python standard library and six public-domain books.

Day 2 compares six changes to the decoder, one at a time, against the Day 1 baseline. A comparison needs held-out text large enough that the difference between two models is bigger than the noise of evaluating them. The project's own planning documents, about 111 KB, were not. So before any Day 2 experiment ran, Day 2 built a corpus.

What it holds

Python code and nineteenth-century prose

The plan wanted a model that would eventually read code and edit prose, so the corpus has both, kept in separate categories so each can be evaluated alone.

Gutenberg IDTitleKind
84Frankenstein; or, The Modern PrometheusNarrative
1228On the Origin of SpeciesExpository
1342Pride and PrejudiceNarrative
1497The RepublicExpository
2701Moby Dick; or, The WhaleNarrative
2814DublinersNarrative
How it splits

By chunk, not by book

octlm/corpus.py cuts every source into chunks of about 64 KB on line boundaries, so no chunk ends in the middle of a line. A seeded shuffle then sends 10 percent of the chunks to validation.

octlm/corpus.pyline 60
def chunk(text: str, limit: int = CHUNK_BYTES) -> list[str]:
    """Split on line boundaries so no chunk ends mid-line."""
    chunks, current, size = [], [], 0
    for line in text.splitlines(keepends=True):
        current.append(line)
        size += len(line.encode())
        if size >= limit:
            chunks.append("".join(current))
            current, size = [], 0
    if current:
        chunks.append("".join(current))
    return chunks

Splitting by chunk instead of by whole book keeps every source present in both splits. Two neighboring chunks of one book share a style but share no text, so validation never contains a sentence the model trained on. Each chunk stays a separate document in the JSONL file, so the BPE trainer never merges across a chunk boundary.

Playground
Train6,918 KB
Validation860 KB
codeprose
Hold out 10 percent
stdlib
book 1
book 2
book 3
book 4
book 5
book 6
Sources in validation5 of 7
Validation chunks6of 63

Top: the real Colab build, from notes/day2.md. Bottom: an illustration with made-up chunk counts, orange squares held out. Holding out whole books would test one author the model never read. Holding out chunks keeps every source in both splits, and since chunks share no text, no validation sentence is ever trained on.

Top: bytes per split from the Colab build. Bottom: an illustration of the two ways to hold out 10 percent.
The numbers

The build every result uses

The corpus was first built on the laptop. When training moved to Colab, the corpus had to move too, because the code half is read from the running machine's standard library, and Colab runs Python 3.13 where the laptop runs 3.14. A different Python is a different corpus. The Colab build is the one every Day 2 training result uses:

SplitDocumentsBytesCodeProse
Train2056,918,2774,103,1832,815,094
Validation23860,062466,611393,451

The mix is 60 percent code and 40 percent prose by bytes. data/manifest.json records a SHA-256 fingerprint of each split, and the note tells every later reader to check them before comparing a new run, because Colab upgrades its Python image without warning.

The tokenizer was retrained on this corpus at 2,048 entries, from every tenth training document to keep the pair-recount cost down. It stopped at 1,787 merges and reached 0.431 tokens per byte. Cut into 256-token blocks, the training split is 11,508 blocks, about 2.96 million tokens, and validation is 1,528 blocks: 782 of code and 745 of prose.

Remember the 2.96 million. Why the plan changed comes back to it. The Day 2 models had 3.3 million parameters, and a model that size wants about 20 times more tokens than this.

Known limits

What the note says it does not cover

Day 2 needed a corpus large enough to separate architectures, not the final data mix. Whether it was large enough is the question the rest of Day 2 answers.