0Day 0

What octlm is, and how it measures

The goal, the rules every experiment follows, the laptop it started on, and the first two experiments, which built the lab before any model existed.

octlm is a small decoder language model written from scratch in PyTorch. It exists to learn how a language model works, down to the arithmetic, and to be able to defend each piece with a number that came from running it.

This series follows the build in the order it happened. Day 0 is this post and the next one. They cover what the project is, how it measures, and what a language model is even trying to do. Day 1 builds the first working model. Day 2 replaces its parts with the ones current open models use. Day 3 builds four research ideas and then stops. Day 4 trains the model for real.

The goal

Build it, train it once, then move to a real model

The plan has three stages.

  1. Write a decoder from scratch, about 20 million parameters, and train it on children's stories until it writes coherent ones.
  2. Load Qwen, a real open model, into the same code and check that it produces the same logits as the reference implementation.
  3. Fine-tune Qwen with LoRA, written in octlm, to emit tool calls, and build a harness that runs it against a fixed set of tasks.

The first stage proves the internals. The second proves they are the same internals a production model uses. The third produces something measurable on top of them. The plan used to be bigger, and why the plan changed covers what was cut and why.

The method

Every experiment is written down before it runs

Each experiment gets an ID, EXP-001 onward, and a section in the day's note with the same five parts:

After the run, the section records the result, what failed, and a keep or revert decision. Failed ideas stay in the note with their numbers. Several posts in this series are about things that did not work, and those are the ones that taught the most.

Two more rules shape everything:

Diagram
↺

One sentence on what is wrong or unknown. EXP-002 asked how long a sequence the laptop can run before memory or time runs out.

Stage 1 of 8
The life of one experiment. Click a stage, or step through. The loop closes when a decision becomes the baseline of the next experiment.
EXP-001

A lab that can repeat its own results

Before any model, the repository needed to produce the same numbers twice. EXP-001 set that up:

The parameter count is simple enough to do by hand, and the counter below does the same sums octlm does. A test checks it against parameter_count on the real PyTorch model for five configs.

Playground
Token embedding (shared with the output head)131,072
Learned position table16,384
Attention: query, key, value, output131,072
Feed-forward262,144
Norm gains and biases1,280
Parameters541,952
Feed-forward hidden width5124d
float32 weights2.1 MiB
Parameters by part. The Day 1 preset matches the 541,952 the dry run reports. Every setting comes from configs/day1.toml.

What happened?

EXP-002

How far can a laptop go?

The machine is an Intel i5-12450H with 12 logical CPUs, 15 GiB of RAM, and no usable GPU. EXP-002 asked at which context length a single forward pass of the Day 1 model shape stops being practical on this CPU.

octlm.bench builds the model at each length, runs one untrained forward pass on a sequence of zeros, and records the time and the process's peak memory.

octlm/bench.pyline 16
def measure(context_length: int, device: torch.device) -> dict[str, object]:
    torch.manual_seed(0)
    config = DecoderConfig(
        vocab_size=512,
        context_length=context_length,
        d_model=128,
        n_heads=4,
        n_layers=2,
    )
    model = Decoder(config).eval().to(device)
    inputs = torch.zeros((1, context_length), dtype=torch.long, device=device)
    with torch.inference_mode():
        if device.type == "cuda":
            model(inputs)  # first call loads kernels, so time the second
            torch.cuda.synchronize()
        started = time.perf_counter()
        model(inputs)
        if device.type == "cuda":
            torch.cuda.synchronize()
        elapsed = time.perf_counter() - started
    count = parameter_count(model)
    token_vectors = F.normalize(model.token_embedding.weight.detach(), dim=1)
    position_vectors = F.normalize(model.position_embedding.weight.detach(), dim=1)
    token_neighbor = int((token_vectors[0] @ token_vectors[1:].T).argmax()) + 1
    position_neighbor = (
        int((position_vectors[0] @ position_vectors[1:].T).argmax()) + 1
        if context_length > 1
        else None
    )
    return {
        "batch_size": 1,
        "config_sha256": None,
        "context_length": context_length,
        "decode_tokens_per_second": None,
        "device": str(device),
        "dtype": "float32",
        "elapsed_seconds": elapsed,
        "experiment_id": "EXP-002",
        "kv_cache_bytes": 0,
        "loss": None,
        "model_bytes": count * 4,
        "nearest_position_to_zero": position_neighbor,
        "nearest_token_to_zero": token_neighbor,
        "parameter_count": count,
        "peak_rss_bytes": resource.getrusage(resource.RUSAGE_SELF).ru_maxrss * 1024,
        "perplexity": None,
        "python": platform.python_version(),
        "run_id": f"{device.type}-context-{context_length}",
        "schema": "octlm-bench-v1",
        "time_to_first_token_ms": None,
        "tokenizer_sha256": None,
        "tokens_processed": context_length,
        "torch": torch.__version__,
    }
Measured
0.020.040.060.080.0100.0120.01282565121,0242,048context length (tokens)millisecondsforward pass
Hover the chart to read values.
One forward pass at batch 1, float32, CPU. Five points from one run, recorded in notes/day1.md. There is no runs/ file for EXP-002, so these numbers come from the note.

What happened?

These are smoke measurements from one run, and the note says so. They set 2,048 as the longest context Day 1 would claim to support, and nothing more.

A later constraint

Training moved off the laptop

Day 2 ran a grid of 24 training runs. Saturating every core for most of an hour overheated the laptop, so from Day 2 onward all training runs on a cloud T4 GPU, first on Colab and from Day 4 on Kaggle. The laptop keeps the short checks: the unit tests, a dry run, and measurements that do not need a GPU. Every result on this site names the machine it ran on.

Skipped

What we did not build, and why