What octlm is, and how it measures
The goal, the rules every experiment follows, the laptop it started on, and the first two experiments, which built the lab before any model existed.
octlm is a small decoder language model written from scratch in PyTorch. It exists to learn how a language model works, down to the arithmetic, and to be able to defend each piece with a number that came from running it.
This series follows the build in the order it happened. Day 0 is this post and the next one. They cover what the project is, how it measures, and what a language model is even trying to do. Day 1 builds the first working model. Day 2 replaces its parts with the ones current open models use. Day 3 builds four research ideas and then stops. Day 4 trains the model for real.
Build it, train it once, then move to a real model
The plan has three stages.
- Write a decoder from scratch, about 20 million parameters, and train it on children's stories until it writes coherent ones.
- Load Qwen, a real open model, into the same code and check that it produces the same logits as the reference implementation.
- Fine-tune Qwen with LoRA, written in octlm, to emit tool calls, and build a harness that runs it against a fixed set of tasks.
The first stage proves the internals. The second proves they are the same internals a production model uses. The third produces something measurable on top of them. The plan used to be bigger, and why the plan changed covers what was cut and why.
Every experiment is written down before it runs
Each experiment gets an ID, EXP-001 onward, and a section in the day's note with the same five parts:
- Problem. One sentence on what is wrong or unknown.
- Hypothesis. What we expect, stated so a result can prove it false.
- Baseline. The thing it is compared against, measured the same way.
- Measurement. The exact numbers that will decide it.
- Stop condition. When to give up or change course.
After the run, the section records the result, what failed, and a keep or revert decision. Failed ideas stay in the note with their numbers. Several posts in this series are about things that did not work, and those are the ones that taught the most.
Two more rules shape everything:
- Change one thing per comparison. If two parts change at once, a better number cannot be credited to either.
- No optimization claim without a baseline. "Faster" means a measured time against a measured time on the same machine.
A lab that can repeat its own results
Before any model, the repository needed to produce the same numbers twice. EXP-001 set that up:
- Python 3.14 and PyTorch 2.14 for CPU, pinned with
uv.lock, so a fresh checkout installs the same versions. - One TOML file per day holds every setting that varies. A SHA-256 of the settings becomes the config fingerprint, and a checkpoint refuses to resume under a different fingerprint.
- A dry run builds the model and prints its parameter count and output shape without training. For the Day 1 config it reports 541,952 parameters and logits shaped
[1, 128, 1024], meaning one sequence, 128 positions, a score for each of 1,024 vocabulary entries. - Every measurement is one JSON line in
runs/. The charts on this site read those files, and a test fails the site build if a file is malformed.
The parameter count is simple enough to do by hand, and the counter below does the same sums octlm does. A test checks it against parameter_count on the real PyTorch model for five configs.
What happened?
- The token embedding is the largest single part at Day 1 size, 1,024 entries of 128 numbers. The output head reuses the same matrix, so it adds nothing. The baseline block explains why.
- Each layer costs the same, so doubling the layers roughly doubles everything except the embedding.
- Switch positions to RoPE and the position table disappears. RoPE has no parameters.
How far can a laptop go?
The machine is an Intel i5-12450H with 12 logical CPUs, 15 GiB of RAM, and no usable GPU. EXP-002 asked at which context length a single forward pass of the Day 1 model shape stops being practical on this CPU.
octlm.bench builds the model at each length, runs one untrained forward pass on a sequence of zeros, and records the time and the process's peak memory.
def measure(context_length: int, device: torch.device) -> dict[str, object]:
torch.manual_seed(0)
config = DecoderConfig(
vocab_size=512,
context_length=context_length,
d_model=128,
n_heads=4,
n_layers=2,
)
model = Decoder(config).eval().to(device)
inputs = torch.zeros((1, context_length), dtype=torch.long, device=device)
with torch.inference_mode():
if device.type == "cuda":
model(inputs) # first call loads kernels, so time the second
torch.cuda.synchronize()
started = time.perf_counter()
model(inputs)
if device.type == "cuda":
torch.cuda.synchronize()
elapsed = time.perf_counter() - started
count = parameter_count(model)
token_vectors = F.normalize(model.token_embedding.weight.detach(), dim=1)
position_vectors = F.normalize(model.position_embedding.weight.detach(), dim=1)
token_neighbor = int((token_vectors[0] @ token_vectors[1:].T).argmax()) + 1
position_neighbor = (
int((position_vectors[0] @ position_vectors[1:].T).argmax()) + 1
if context_length > 1
else None
)
return {
"batch_size": 1,
"config_sha256": None,
"context_length": context_length,
"decode_tokens_per_second": None,
"device": str(device),
"dtype": "float32",
"elapsed_seconds": elapsed,
"experiment_id": "EXP-002",
"kv_cache_bytes": 0,
"loss": None,
"model_bytes": count * 4,
"nearest_position_to_zero": position_neighbor,
"nearest_token_to_zero": token_neighbor,
"parameter_count": count,
"peak_rss_bytes": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss * 1024,
"perplexity": None,
"python": platform.python_version(),
"run_id": f"{device.type}-context-{context_length}",
"schema": "octlm-bench-v1",
"time_to_first_token_ms": None,
"tokenizer_sha256": None,
"tokens_processed": context_length,
"torch": torch.__version__,
}What happened?
- From 128 to 1,024 tokens the time grows about as fast as the length. The feed-forward layers and projections cost the same per token, and at these lengths they dominate.
- From 1,024 to 2,048 the time quadruples, from 28.7 ms to 115.2 ms. Attention compares every position with every earlier one, so its cost grows with the square of the length, and at 2,048 it starts to show. Causal attention and SDPA come back to this square.
- Peak memory reached 504 MB by the end of the process.
These are smoke measurements from one run, and the note says so. They set 2,048 as the longest context Day 1 would claim to support, and nothing more.
Training moved off the laptop
Day 2 ran a grid of 24 training runs. Saturating every core for most of an hour overheated the laptop, so from Day 2 onward all training runs on a cloud T4 GPU, first on Colab and from Day 4 on Kaggle. The laptop keeps the short checks: the unit tests, a dry run, and measurements that do not need a GPU. Every result on this site names the machine it ran on.
What we did not build, and why
- A configuration framework. One TOML file per day and a frozen dataclass cover every setting. A framework would add code without adding a setting.
- Plots in the Python code. octlm writes JSON lines. This site draws them.
- A GPU from day one. Day 1 models trained in seconds on the CPU. The GPU arrived when a measured run made the laptop overheat.