1Day 1

Characters as tokens

The simplest tokenizer gives every character its own ID. It is easy to build and easy to check, and it shows exactly what a better tokenizer has to fix.

A model reads integers, not text. A tokenizer is the pair of functions that turns text into a list of integer IDs and back. Its design decides how long every sequence is, what the model can represent at all, and what one prediction means.

Day 1 built two tokenizers. This post covers the first and simplest one, which is also the deliberate baseline for the second.

Definition

One character, one ID

The character tokenizer collects every distinct character in the training text, sorts them, and numbers them. Four special tokens come first:

IDTokenUsed for
0<|pad|>Filler at the end of a short block. Never a training target.
1<|bos|>Beginning of a document.
2<|eos|>End of a document.
3<|unk|>Any character the training text did not contain.

Encoding looks each character up. Decoding looks each ID up and joins the characters. That is the whole algorithm:

octlm/tokenizer.pyline 86
def encode(self, text: str, add_special_tokens: bool = False) -> list[int]:
    lookup = self.token_to_id
    tokens = [lookup.get(character, 3) for character in text]
    return [self.bos_id, *tokens, self.eos_id] if add_special_tokens else tokens

lookup.get(character, 3) is the important part. A character outside the vocabulary becomes ID 3, and there is no way back. Decoding ID 3 gives <|unk|>, not the original character.

Playground

A vocabulary learned from one file

EXP-003 trained the character tokenizer on the project's own PLAN.md. The playground below learns its vocabulary from today's PLAN.md at build time, the same way.

Playground
T42h56e53␠5m61o63d52e53l60␠5r66e53a49d52s67␠5t68e53x72t68␠5o63n62e53␠5c51h56a49r66a49c51t68e53r66␠5a49t68␠5a49␠5t68i57m61e53.10
Vocabulary764 specials + 72 characters
Tokens45one per character
UTF-8 bytes45
Unknown0.0%ID 3, <|unk|>

The vocabulary is every distinct character in the training text, sorted. Anything the training text never contained becomes ID 3, and the original character is gone. Try the Japanese or emoji preset.

Each box is one token with its ID underneath. Red boxes are <|unk|>: characters PLAN.md never used.

What happened?

EXP-003

What we measured

Problem: we needed a tokenizer to compare BPE against, trained on the same text.

The character tokenizer learned 86 entries from PLAN.md: 4 specials and 82 characters. On the held-out file, day-wise.md, it produced 13,939 tokens for 13,939 bytes. That file is plain ASCII, so one character is one byte and one token.

On an unseen Unicode sample, 77.8 percent of the characters mapped to <|unk|>.

Then a small model trained on each tokenizer's output for the same 40 steps. With the character tokenizer, validation loss reached 3.4011 nats per token, perplexity 29.998, and 4.9068 bits per byte. Those three numbers come from runs/day1-character.jsonl. Measuring a tokenizer sets them against BPE.

Decision: keep it as the baseline. It was never meant to be the model's tokenizer.

The problems

Why characters hurt, especially on code

Day 1's paper exercise asked for three reasons character tokens hurt a model that reads code. The note records these:

  1. Sequences get long. Every character is a position. A 4-space indent is 4 positions, and return is 6. Attention costs grow with the square of the length, and the context window fills with syntax.
  2. The model spends its capacity spelling. To predict values after sum(, a character model must predict v, then a, then l, each a separate decision. A tokenizer that has values as one token makes that one decision.
  3. Unknown characters destroy information. The vocabulary is fixed at training time. A new symbol, an accented name or an emoji in a string literal becomes <|unk|>, and the model cannot read or write it.

The small vocabulary has one real advantage. Every ID gets a lot of training data. The next tokenizer keeps that property for its base alphabet and fixes all three problems.

Playground
bytes6,241
characters5,776
chunks1,089
Sequence length in bytes79
Sequence length in characters76
Sequence length in chunks33

The bars are attention scores per head, the sequence length squared. A character tokenizer makes one position per character, so the same text costs several times the attention of a tokenizer that groups characters. Chunks are what pre-tokenization makes, a rough floor for what BPE can reach. Indentation is the worst case: four spaces are four positions.

One text measured three ways. Edit it, or paste some indented code, and watch the attention cost of each sequence length.
The rules

What every octlm tokenizer must preserve

AGENTS.md sets rules that both tokenizers follow, and tests enforce each one:

Skipped

What we did not build, and why