Measuring a tokenizer
The BPE model scored a perplexity eight times worse than the character model and was the better model. Bits per byte is the unit that makes the two comparable.
Day 1 trained the same small model twice for the same 40 steps, once on character tokens and once on BPE tokens, and evaluated both on the same held-out file. Read by loss or perplexity, the character model won by a wide margin. Read by bits per byte, BPE won. Only one of those readings is a fair comparison.
Two models, three metrics, opposite verdicts
| Tokenizer | Validation loss | Perplexity | Bits per byte |
|---|---|---|---|
| Character | 3.4011 | 29.998 | 4.9068 |
| BPE, 1,024 entries | 5.4823 | 240.391 | 4.1497 |
All six numbers come from runs/day1-character.jsonl and runs/day1-bpe.jsonl at step 40.
A token is a different amount of text in each model
The character model predicts one byte per step. The BPE model predicts almost two bytes per step on this file, because a BPE token covers 1.90 bytes on average. Predicting two bytes at once is harder than predicting one, so the loss per prediction is naturally higher. Comparing per-token loss across tokenizers compares predictions of different sizes.
Bits per byte removes the size. Take the total loss across the whole file, which is the information the model needed for that file, and divide by the file's bytes:
Both models describe the same bytes, so this asks both the same question. How many bits does the model need per byte of this text?
Plug in the character model. One byte per token, so bits per byte is the loss divided by . For BPE, the per-token loss is larger, but each token pays for almost two bytes, and the division brings it under the character model.
What happened?
- At the measured losses, perplexity prefers the character model and bits per byte prefers BPE. The verdict line says which.
- Lower the BPE loss until its perplexity beats the character model. The BPE loss has to fall below the character model's loss, and by then its bits per byte is about half the character model's.
- The token count is where the difference lives. BPE used about half as many predictions for the same text.
Compare tokenizers only in bits per byte
After 40 steps, BPE improved bits per byte by 15.4 percent over the character model. That kept BPE for the Day 1 model and made a rule that every later comparison follows. Perplexity is fine inside one tokenizer, and across tokenizers only bits per byte counts. AGENTS.md carries the rule for the rest of the project.
The comparison has limits the note states. Forty steps on a tiny model and a tiny file says the direction, not the size, of the effect at scale. It does not choose a vocabulary size either. That choice needs two full training runs, which is why Day 4 took 8,192 entries from the TinyStories paper's 10K and recorded it as a choice rather than a result.
Bytes per token, the other number to track
Bytes per token measures compression, and it changes with the text, not only the tokenizer.
| Tokenizer | Text | Bytes per token |
|---|---|---|
| Character, Day 1 | day-wise.md | 1.00 |
| BPE 1,024, Day 1 | day-wise.md | 1.90 |
| BPE 2,048, Day 2 | code and prose corpus, Colab build | 2.32 |
| BPE 8,192, Day 4 | TinyStories | 2.26 |
The Day 2 row is the inverse of the 0.431 tokens per byte that notes/day2.md records. The Day 4 vocabulary is four times larger but compresses TinyStories less than Day 2's compressed its corpus. TinyStories has a small, simple vocabulary, and words are short. More merges cannot make a three-letter word longer than one token.
Better compression is good for speed, because a context window holds more text. It is not the same as a better model. Bits per byte is the model's number. Bytes per token is the tokenizer's.
What we did not measure, and why
- The final vocabulary size. Two planning documents suggested 16K to 32K. A choice like that needs trained models at each size, and Day 1 measured only what its small corpus could support.
- Per-token perplexity across tokenizers. Recorded in the runs and never used to compare them.