1Day 1

Measuring a tokenizer

The BPE model scored a perplexity eight times worse than the character model and was the better model. Bits per byte is the unit that makes the two comparable.

Day 1 trained the same small model twice for the same 40 steps, once on character tokens and once on BPE tokens, and evaluated both on the same held-out file. Read by loss or perplexity, the character model won by a wide margin. Read by bits per byte, BPE won. Only one of those readings is a fair comparison.

The result

Two models, three metrics, opposite verdicts

TokenizerValidation lossPerplexityBits per byte
Character3.401129.9984.9068
BPE, 1,024 entries5.4823240.3914.1497

All six numbers come from runs/day1-character.jsonl and runs/day1-bpe.jsonl at step 40.

Measured
Loss per token (nats)
characterBPE
3.504.004.505.005.506.002040stepnats per tokencharacterBPE
Hover the chart to read values.
Bits per byte
characterBPE
4.004.505.005.502040stepbits per bytecharacterBPE
Hover the chart to read values.
Validation at steps 20 and 40, the only two evaluations each run made. Left: loss per token. Right: bits per byte. The order of the two lines flips.
Why

A token is a different amount of text in each model

The character model predicts one byte per step. The BPE model predicts almost two bytes per step on this file, because a BPE token covers 1.90 bytes on average. Predicting two bytes at once is harder than predicting one, so the loss per prediction is naturally higher. Comparing per-token loss across tokenizers compares predictions of different sizes.

Bits per byte removes the size. Take the total loss across the whole file, which is the information the model needed for that file, and divide by the file's bytes:

bpb⁡=Ltoken⋅Ntokensln⁡2⋅Nbytes=Ltokenln⁡2⋅1bytes per token\operatorname{bpb} = \frac{\mathcal{L}_{\text{token}} \cdot N_{\text{tokens}}}{\ln 2 \cdot N_{\text{bytes}}} = \frac{\mathcal{L}_{\text{token}}}{\ln 2} \cdot \frac{1}{\text{bytes per token}}

Both models describe the same bytes, so this asks both the same question. How many bits does the model need per byte of this text?

Plug in the character model. One byte per token, so bits per byte is the loss divided by ln⁡2\ln 2. For BPE, the per-token loss is larger, but each token pays for almost two bytes, and the division brings it under the character model.

Playground
Character
Perplexity30.0
Tokens13,8241.00 bytes each
Bits per byte4.907
BPE 1,024
Perplexity240.4
Tokens7,1681.91 bytes each
Bits per byte4.150
By perplexity, Character looks better. By bits per byte, which both tokenizers share, BPE 1,024 is better by 0.757 bits.
The two runs, starting at their measured losses. Drag either loss and watch perplexity and bits per byte. The token and byte counts are fixed by each tokenizer's split of the validation text.

What happened?

Playground
2.0 bits per bytethis tokenizer
0.01.02.03.04.05.06.07.01.01.52.02.53.03.54.04.55.0bytes per tokenloss per token, nats2.0 bits per bytethis tokenizer
Hover the chart to read values.
Loss per token3.466 natsbpb × bytes per token × ln 2
Perplexity32.0
Same model, 1 byte per token1.386 natsperplexity 4.0

Every point on the line predicts text equally well. Only the tokenizer changes. A tokenizer that packs more bytes into each token must pay more loss per token to reach the same bits per byte, and perplexity grows exponentially with it. That is why the two metrics disagreed.

Hold the model's quality fixed in bits per byte and change only the tokenizer. The loss per token that means the same quality moves along the line.
Decision

Compare tokenizers only in bits per byte

After 40 steps, BPE improved bits per byte by 15.4 percent over the character model. That kept BPE for the Day 1 model and made a rule that every later comparison follows. Perplexity is fine inside one tokenizer, and across tokenizers only bits per byte counts. AGENTS.md carries the rule for the rest of the project.

The comparison has limits the note states. Forty steps on a tiny model and a tiny file says the direction, not the size, of the effect at scale. It does not choose a vocabulary size either. That choice needs two full training runs, which is why Day 4 took 8,192 entries from the TinyStories paper's 10K and recorded it as a choice rather than a result.

Compression

Bytes per token, the other number to track

Bytes per token measures compression, and it changes with the text, not only the tokenizer.

TokenizerTextBytes per token
Character, Day 1day-wise.md1.00
BPE 1,024, Day 1day-wise.md1.90
BPE 2,048, Day 2code and prose corpus, Colab build2.32
BPE 8,192, Day 4TinyStories2.26

The Day 2 row is the inverse of the 0.431 tokens per byte that notes/day2.md records. The Day 4 vocabulary is four times larger but compresses TinyStories less than Day 2's compressed its corpus. TinyStories has a small, simple vocabulary, and words are short. More merges cannot make a three-letter word longer than one token.

Better compression is good for speed, because a context window holds more text. It is not the same as a better model. Bits per byte is the model's number. Bytes per token is the tokenizer's.

Skipped

What we did not measure, and why