TinyStories and a 6x faster encoder
The Day 4 model needs about 400 million tokens. The pure-Python BPE encoder was expected to be too slow for 2.2 GB of text. One dictionary from chunk to IDs made it six times faster, and the whole training split encoded in 16 minutes.
Day 2's corpus held 2.96 million training tokens. The Day 4 model, 26 million parameters, needs about 400 million to reach a readable result. EXP-065 got the data, trained the tokenizer, and encoded everything, and the first question was whether octlm's own Python encoder could do that in one session at all.
TinyStories
TinyStories is a synthetic dataset of short children's stories, generated by GPT-3.5 and GPT-4 with a deliberately small vocabulary. Its authors report that models under about 30 million parameters trained on it write coherent, grammatical stories, something that normally takes a far larger model on web text. That is the property the new plan needs from a 20M model.
octlm uses the V2 files, generated by GPT-4. The training file is 2,227,753,162 bytes and the validation file is 22.5 MB. The license is CDLA-Sharing-1.0.
A line holding only <|endoftext|> separates stories, and blank lines sometimes follow it. The validation file starts in the middle of a story. The note's decision was to split on the delimiter line, strip surrounding newlines, drop empty stories, and keep the truncated first story, because dropping it would be a special case for one story out of thousands.
def stories(path: Path) -> Iterator[str]:
lines: list[str] = []
with path.open(encoding="utf-8", newline="") as file:
for line in file:
if line.rstrip("\r\n") != DELIMITER:
lines.append(line)
continue
story = "".join(lines).strip("\n")
lines = []
if story:
yield story
story = "".join(lines).strip("\n")
if story:
yield storyEvery story is then encoded with <|bos|> and <|eos|> and appended to one flat file of 16-bit token IDs. Training reads random 513-token windows from that file through a memory map, so a window can span two stories with <|eos|> <|bos|> between them. The BPE merges still never cross a story boundary, because each story is encoded on its own.
8,192 entries from the first 10 million characters
The plan asked to choose between 4,096 and 8,192 entries by bits per byte. Bits per byte is a property of a trained model, so the choice would cost two full training runs. The TinyStories authors used a 10K vocabulary. The note took 8,192, records the compression it reaches, and moved the 4,096 comparison to Day 5 if GPU time allows.
BPE training recounts every pair after every merge, so its cost grows with the sample. The tokenizer trained on the first 10 million characters of the training split. TinyStories uses a small vocabulary, so a small sample covers it.
Every chunk always encodes the same way
Encoding a story means pre-tokenizing it into chunks, then running the merge loop on each chunk. The merge loop depends only on the chunk's bytes. It never looks at the neighbors, because merges never cross a chunk boundary. So the same chunk always produces the same IDs, and the answer can be stored after the first time.
What happened?
- The first sentence is almost all misses. The cache is empty.
- By the third sentence, most chunks are hits:
Lily,the,park, spaces and periods. - Paste a longer story and the hit rate climbs. TinyStories repeats a small vocabulary across 2.7 million stories, which is the case this cache is built for.
The change is a dictionary on the tokenizer and three lines in encode. The output does not change, only the time.
def encode(self, text: str, add_special_tokens: bool = False) -> list[int]:
cache = self.chunk_cache
output: list[int] = []
for chunk in pretokenize(text):
token_ids = cache.get(chunk)
if token_ids is None:
token_ids = cache[chunk] = self.encode_chunk(chunk)
output.extend(token_ids)
return [self.bos_id, *output, self.eos_id] if add_special_tokens else outputWhat we measured
Hypothesis, in two parts: the pure-Python encoder is too slow to encode 2.2 GB in one session, and the chunk cache fixes the merge cost while the per-character pre-tokenizer stays the bottleneck. Stop condition: if the training split takes more than two hours, stop and speed up the pre-tokenizer before training anything.
The benchmark encoded the first 2 million characters of the validation split twice, once calling the merge loop on every chunk, once through the cache starting empty. The Colab column is the note's EXP-065 table. The main run repeated the whole preparation on Kaggle two days later, and the Kaggle column is read from that run's prepare.jsonl and bpe.jsonl:
| Measure | Colab | Kaggle |
|---|---|---|
| Tokenizer sample | 10,000,355 characters, 12,328 stories | 10,000,355 characters, 12,328 stories |
| Tokenizer training | 303 s | 238 s |
| Encode without cache | 427 KB/s | 462 KB/s |
| Encode with cache | 2,559 KB/s, 6.0x faster | 2,909 KB/s, 6.3x faster |
| Train split | 2,717,495 stories, 966,389,646 tokens, 934 s | 2,717,495 stories, 966,389,646 tokens, 738 s |
| Valid split | 27,630 stories, 9,759,613 tokens, 10 s | 27,630 stories, 9,759,613 tokens, 7 s |
| Bytes per token | 2.262 on both splits | 2.262 on both splits |
What happened?
- The cache made encoding 6 times faster, from 427 KB/s to 2,559 KB/s.
- The training split, 966 million tokens, encoded in 934 seconds. That is 15.6 minutes against a two-hour limit. The first half of the hypothesis is false.
- The pre-tokenizer's cost was not measured on its own, so the second half stays open.
- Both splits compress to 2.262 bytes per token.
- 966 million tokens is 2.5 times what the main run consumes, so the model sees each token at most once.
- Kaggle produced the same story counts, token counts and bytes per token. Only the times differ: Kaggle's CPU encoded the training split in 738 seconds, and the cache's speedup there was 6.3x.
Decision: keep the cache.
What we did not build, and why
- A Rust or C tokenizer. The Python encoder met the stop condition with room to spare.
- Parallel encoding across processes. Same reason.
- The 4,096 vocabulary run. Moved to Day 5, if GPU time allows.
- Train-validation deduplication. TinyStories ships separate files, and the note did not check them for overlapping stories.