4Day 4

TinyStories and a 6x faster encoder

The Day 4 model needs about 400 million tokens. The pure-Python BPE encoder was expected to be too slow for 2.2 GB of text. One dictionary from chunk to IDs made it six times faster, and the whole training split encoded in 16 minutes.

Day 2's corpus held 2.96 million training tokens. The Day 4 model, 26 million parameters, needs about 400 million to reach a readable result. EXP-065 got the data, trained the tokenizer, and encoded everything, and the first question was whether octlm's own Python encoder could do that in one session at all.

The data

TinyStories

TinyStories is a synthetic dataset of short children's stories, generated by GPT-3.5 and GPT-4 with a deliberately small vocabulary. Its authors report that models under about 30 million parameters trained on it write coherent, grammatical stories, something that normally takes a far larger model on web text. That is the property the new plan needs from a 20M model.

octlm uses the V2 files, generated by GPT-4. The training file is 2,227,753,162 bytes and the validation file is 22.5 MB. The license is CDLA-Sharing-1.0.

A line holding only <|endoftext|> separates stories, and blank lines sometimes follow it. The validation file starts in the middle of a story. The note's decision was to split on the delimiter line, strip surrounding newlines, drop empty stories, and keep the truncated first story, because dropping it would be a special case for one story out of thousands.

octlm/day4.pyline 77
def stories(path: Path) -> Iterator[str]:
    lines: list[str] = []
    with path.open(encoding="utf-8", newline="") as file:
        for line in file:
            if line.rstrip("\r\n") != DELIMITER:
                lines.append(line)
                continue
            story = "".join(lines).strip("\n")
            lines = []
            if story:
                yield story
    story = "".join(lines).strip("\n")
    if story:
        yield story

Every story is then encoded with <|bos|> and <|eos|> and appended to one flat file of 16-bit token IDs. Training reads random 513-token windows from that file through a memory map, so a window can span two stories with <|eos|> <|bos|> between them. The BPE merges still never cross a story boundary, because each story is encoded on its own.

Diagram

The TinyStories V2 training file, GPT-4 stories separated by <|endoftext|> lines.

Stage 1 of 7
From the downloaded text file to the windows the trainer reads.
The tokenizer

8,192 entries from the first 10 million characters

The plan asked to choose between 4,096 and 8,192 entries by bits per byte. Bits per byte is a property of a trained model, so the choice would cost two full training runs. The TinyStories authors used a 10K vocabulary. The note took 8,192, records the compression it reaches, and moved the 4,096 comparison to Day 5 if GPU time allows.

BPE training recounts every pair after every merge, so its cost grows with the sample. The tokenizer trained on the first 10 million characters of the training split. TinyStories uses a small vocabulary, so a small sample covers it.

The encoder

Every chunk always encodes the same way

Encoding a story means pre-tokenizing it into chunks, then running the merge loop on each chunk. The merge loop depends only on the chunk's bytes. It never looks at the neighbors, because merges never cross a chunk boundary. So the same chunk always produces the same IDs, and the answer can be stored after the first time.

Playground
Once␠upon␠a␠time,␠there␠was␠a␠little␠girl␠named␠Lily.␠Lily␠liked␠to␠play␠in␠the␠park.␠One␠day,␠Lily␠saw␠a␠big␠dog␠in␠the␠park.␠The␠dog␠was␠happy␠and␠wanted␠to␠play.␠Lily␠and␠the␠dog␠played␠all␠day.␠They␠were␠very␠happy.
miss run the merge loop, store the IDshit copy the stored IDs
Chunks101
Distinct chunks33merge loops run
Hit rate67.3%

Pre-tokenization splits text into chunks and BPE never merges across a chunk edge, so a chunk always encodes to the same IDs. On one paragraph the hit rate is modest. Over 2.7 million stories drawn from a small vocabulary, almost every chunk is a word the encoder has already seen.

Each chunk of the text, colored by whether a cache that starts empty has seen it. Misses run the merge loop and store the result. Hits copy the stored IDs.

What happened?

The change is a dictionary on the tokenizer and three lines in encode. The output does not change, only the time.

octlm/tokenizer.pyline 218
def encode(self, text: str, add_special_tokens: bool = False) -> list[int]:
    cache = self.chunk_cache
    output: list[int] = []
    for chunk in pretokenize(text):
        token_ids = cache.get(chunk)
        if token_ids is None:
            token_ids = cache[chunk] = self.encode_chunk(chunk)
        output.extend(token_ids)
    return [self.bos_id, *output, self.eos_id] if add_special_tokens else output
EXP-065

What we measured

Hypothesis, in two parts: the pure-Python encoder is too slow to encode 2.2 GB in one session, and the chunk cache fixes the merge cost while the per-character pre-tokenizer stays the bottleneck. Stop condition: if the training split takes more than two hours, stop and speed up the pre-tokenizer before training anything.

The benchmark encoded the first 2 million characters of the validation split twice, once calling the merge loop on every chunk, once through the cache starting empty. The Colab column is the note's EXP-065 table. The main run repeated the whole preparation on Kaggle two days later, and the Kaggle column is read from that run's prepare.jsonl and bpe.jsonl:

MeasureColabKaggle
Tokenizer sample10,000,355 characters, 12,328 stories10,000,355 characters, 12,328 stories
Tokenizer training303 s238 s
Encode without cache427 KB/s462 KB/s
Encode with cache2,559 KB/s, 6.0x faster2,909 KB/s, 6.3x faster
Train split2,717,495 stories, 966,389,646 tokens, 934 s2,717,495 stories, 966,389,646 tokens, 738 s
Valid split27,630 stories, 9,759,613 tokens, 10 s27,630 stories, 9,759,613 tokens, 7 s
Bytes per token2.262 on both splits2.262 on both splits

What happened?

Decision: keep the cache.

Skipped

What we did not build, and why