Characters as tokens
The simplest tokenizer gives every character its own ID. It is easy to build and easy to check, and it shows exactly what a better tokenizer has to fix.
A model reads integers, not text. A tokenizer is the pair of functions that turns text into a list of integer IDs and back. Its design decides how long every sequence is, what the model can represent at all, and what one prediction means.
Day 1 built two tokenizers. This post covers the first and simplest one, which is also the deliberate baseline for the second.
One character, one ID
The character tokenizer collects every distinct character in the training text, sorts them, and numbers them. Four special tokens come first:
| ID | Token | Used for |
|---|---|---|
| 0 | <|pad|> | Filler at the end of a short block. Never a training target. |
| 1 | <|bos|> | Beginning of a document. |
| 2 | <|eos|> | End of a document. |
| 3 | <|unk|> | Any character the training text did not contain. |
Encoding looks each character up. Decoding looks each ID up and joins the characters. That is the whole algorithm:
def encode(self, text: str, add_special_tokens: bool = False) -> list[int]:
lookup = self.token_to_id
tokens = [lookup.get(character, 3) for character in text]
return [self.bos_id, *tokens, self.eos_id] if add_special_tokens else tokenslookup.get(character, 3) is the important part. A character outside the vocabulary becomes ID 3, and there is no way back. Decoding ID 3 gives <|unk|>, not the original character.
A vocabulary learned from one file
EXP-003 trained the character tokenizer on the project's own PLAN.md. The playground below learns its vocabulary from today's PLAN.md at build time, the same way.
What happened?
- English and Python encode cleanly, because
PLAN.mdis English prose with code-like names in it. Every token is one character, so the token count equals the character count. - The Japanese preset is almost entirely unknown.
PLAN.mdcontains no Japanese, so those characters have no ID. - The emoji are unknown for the same reason. One emoji is one character here, but four UTF-8 bytes.
What we measured
Problem: we needed a tokenizer to compare BPE against, trained on the same text.
The character tokenizer learned 86 entries from PLAN.md: 4 specials and 82 characters. On the held-out file, day-wise.md, it produced 13,939 tokens for 13,939 bytes. That file is plain ASCII, so one character is one byte and one token.
On an unseen Unicode sample, 77.8 percent of the characters mapped to <|unk|>.
Then a small model trained on each tokenizer's output for the same 40 steps. With the character tokenizer, validation loss reached 3.4011 nats per token, perplexity 29.998, and 4.9068 bits per byte. Those three numbers come from runs/day1-character.jsonl. Measuring a tokenizer sets them against BPE.
Decision: keep it as the baseline. It was never meant to be the model's tokenizer.
Why characters hurt, especially on code
Day 1's paper exercise asked for three reasons character tokens hurt a model that reads code. The note records these:
- Sequences get long. Every character is a position. A 4-space indent is 4 positions, and
returnis 6. Attention costs grow with the square of the length, and the context window fills with syntax. - The model spends its capacity spelling. To predict
valuesaftersum(, a character model must predictv, thena, thenl, each a separate decision. A tokenizer that hasvaluesas one token makes that one decision. - Unknown characters destroy information. The vocabulary is fixed at training time. A new symbol, an accented name or an emoji in a string literal becomes
<|unk|>, and the model cannot read or write it.
The small vocabulary has one real advantage. Every ID gets a lot of training data. The next tokenizer keeps that property for its base alphabet and fixes all three problems.
What every octlm tokenizer must preserve
AGENTS.md sets rules that both tokenizers follow, and tests enforce each one:
- Exact round trip. Decoding the encoded text returns the same string. No lowercasing and no Unicode normalization, so
éwritten as one code point andéwritten aseplus a combining accent stay different. - Tabs, indentation and CRLF survive. Code depends on whitespace.
- Stable IDs. Special-token IDs never move between versions, and loading a tokenizer file whose specials differ fails.
- Ordinary text never becomes a special token. The literal characters
<|eos|>typed into a document are encoded as characters, never as ID 2. Otherwise a document could end itself early.
What we did not build, and why
- A larger character vocabulary. Training on more text would shrink the unknown rate, but no size of training text covers every character in Unicode. Bytes do, and BPE is built on bytes.
- Word-level tokens. A word vocabulary has the unknown-token problem far worse than characters, and code is full of names no dictionary contains.