A corpus big enough to measure
Day 1 trained on 111 KB of planning notes. Day 2 needed held-out code and prose to compare architectures, so it built a 7 MB corpus from the Python standard library and six public-domain books.
Day 2 compares six changes to the decoder, one at a time, against the Day 1 baseline. A comparison needs held-out text large enough that the difference between two models is bigger than the noise of evaluating them. The project's own planning documents, about 111 KB, were not. So before any Day 2 experiment ran, Day 2 built a corpus.
Python code and nineteenth-century prose
The plan wanted a model that would eventually read code and edit prose, so the corpus has both, kept in separate categories so each can be evaluated alone.
- Code is every top-level module of the Python standard library installed on the machine, sorted, skipping test files and dunder files. It is Python-licensed.
- Prose is six Project Gutenberg books, all public domain in the United States. Four are narrative and two are expository, so the prose is not one author's voice.
| Gutenberg ID | Title | Kind |
|---|---|---|
| 84 | Frankenstein; or, The Modern Prometheus | Narrative |
| 1228 | On the Origin of Species | Expository |
| 1342 | Pride and Prejudice | Narrative |
| 1497 | The Republic | Expository |
| 2701 | Moby Dick; or, The Whale | Narrative |
| 2814 | Dubliners | Narrative |
By chunk, not by book
octlm/corpus.py cuts every source into chunks of about 64 KB on line boundaries, so no chunk ends in the middle of a line. A seeded shuffle then sends 10 percent of the chunks to validation.
def chunk(text: str, limit: int = CHUNK_BYTES) -> list[str]:
"""Split on line boundaries so no chunk ends mid-line."""
chunks, current, size = [], [], 0
for line in text.splitlines(keepends=True):
current.append(line)
size += len(line.encode())
if size >= limit:
chunks.append("".join(current))
current, size = [], 0
if current:
chunks.append("".join(current))
return chunksSplitting by chunk instead of by whole book keeps every source present in both splits. Two neighboring chunks of one book share a style but share no text, so validation never contains a sentence the model trained on. Each chunk stays a separate document in the JSONL file, so the BPE trainer never merges across a chunk boundary.
The build every result uses
The corpus was first built on the laptop. When training moved to Colab, the corpus had to move too, because the code half is read from the running machine's standard library, and Colab runs Python 3.13 where the laptop runs 3.14. A different Python is a different corpus. The Colab build is the one every Day 2 training result uses:
| Split | Documents | Bytes | Code | Prose |
|---|---|---|---|---|
| Train | 205 | 6,918,277 | 4,103,183 | 2,815,094 |
| Validation | 23 | 860,062 | 466,611 | 393,451 |
The mix is 60 percent code and 40 percent prose by bytes. data/manifest.json records a SHA-256 fingerprint of each split, and the note tells every later reader to check them before comparing a new run, because Colab upgrades its Python image without warning.
The tokenizer was retrained on this corpus at 2,048 entries, from every tenth training document to keep the pair-recount cost down. It stopped at 1,787 merges and reached 0.431 tokens per byte. Cut into 256-token blocks, the training split is 11,508 blocks, about 2.96 million tokens, and validation is 1,528 blocks: 782 of code and 745 of prose.
Remember the 2.96 million. Why the plan changed comes back to it. The Day 2 models had 3.3 million parameters, and a model that size wants about 20 times more tokens than this.
What the note says it does not cover
- The code is Python only. The TypeScript and JavaScript the first plan targeted are absent.
- The prose is nineteenth-century literature, not modern writing.
- The standard-library file set depends on the Python install. The fingerprints pin what was used. They do not promise that another machine rebuilds the same corpus.
- The plan asked for 50 percent code, 30 percent prose and 20 percent instruction data. There was no instruction data, and the code side ran out before its budget, which is the whole gap between 3 to 2 and 5 to 3.
Day 2 needed a corpus large enough to separate architectures, not the final data mix. Whether it was large enough is the question the rest of Day 2 answers.