1Day 1

The baseline block

A Transformer block is attention and a feed-forward layer, each wrapped in a norm and a residual connection. This post assembles the Day 1 decoder from those parts, including the initialization bug that stopped it learning.

The previous posts built the pieces that turn text into vectors and let positions read each other. A decoder stacks the same block several times on top of those vectors, then turns the final vectors back into token scores. Day 1's baseline has two blocks, width 128 and four heads, and it follows the small GPT-2 layout.

The whole model

The forward pass, with shapes

For a batch of BB sequences of TT token IDs, with the Day 1 settings C=128C = 128 and V=1,024V = 1{,}024:

Diagram

Integers from the tokenizer, one per position.

Stage 1 of 6
The Day 1 forward pass. Click a stage for what it does and the shape it produces. Shapes use the Day 1 config.

Every block keeps the shape [B, T, C]. That is what lets blocks stack, because the output of one is a valid input to the next.

octlm/model.pyline 411
def forward(self, token_ids: Tensor, all_depths: bool = False) -> Tensor:
    if token_ids.ndim != 2:
        raise ValueError("token IDs must have shape [batch, sequence]")
    length = token_ids.shape[1]
    if self.position_embedding is not None and length > self.config.context_length:
        raise ValueError("sequence exceeds context_length")
    rope = None
    if self.config.position == "rope":
        rope = rope_tables(
            length, self.config.head_size, self.config.rope_scale, token_ids.device
        )
    x = self.dropout(self._embed(token_ids))
    for block in self.blocks:
        x = block(x, rope)
    hidden = self.final_norm(x)
    logits = self.lm_head(hidden)
    if not all_depths:
        return logits
    return torch.stack([logits, *(self.lm_head(a(hidden)) for a in self.mtp_adapters)])
The residual stream

Each block adds to the vector, it does not replace it

A block does not compute a new vector from scratch. It computes a correction and adds it:

x←x+Attention⁡(LN⁡(x)),x←x+FFN⁡(LN⁡(x))x \leftarrow x + \operatorname{Attention}(\operatorname{LN}(x)), \qquad x \leftarrow x + \operatorname{FFN}(\operatorname{LN}(x))

The running vector xx is called the residual stream. The addition has a concrete benefit for training. The gradient of x+f(x)x + f(x) with respect to xx is 1+f′(x)1 + f'(x), so every block passes at least the identity backward. A deep stack of additions gives the gradient a direct path from the loss to the embedding, and it does not shrink as blocks multiply.

octlm/model.pyline 353
class TransformerBlock(nn.Module):
    def __init__(self, config: DecoderConfig) -> None:
        super().__init__()
        self.pre_norm = config.residual == "pre"
        self.attention_norm = make_norm(config)
        self.attention = CausalSelfAttention(config)
        self.feed_forward_norm = make_norm(config)
        self.feed_forward = FeedForward(config)

    def forward(self, x: Tensor, rope: tuple[Tensor, Tensor] | None = None) -> Tensor:
        if self.pre_norm:
            x = x + self.attention(self.attention_norm(x), rope)
            return x + self.feed_forward(self.feed_forward_norm(x))
        x = self.attention_norm(x + self.attention(x, rope))
        return self.feed_forward_norm(x + self.feed_forward(x))

The norm sits inside the branch, before attention and before the feed-forward layer. That placement is called pre-norm. The 2017 paper put it after the addition, called post-norm, and Day 2 tested that choice in where the norm goes.

LayerNorm

Normalize each vector across its own features

LayerNorm rescales one position's vector to mean 0 and variance 1, then applies a learned gain γ\gamma and bias β\beta per feature:

LN⁡(x)=γ⊙x−μσ2+ϵ+β,μ=1C∑ixi,σ2=1C∑i(xi−μ)2\operatorname{LN}(x) = \gamma \odot \frac{x - \mu}{\sqrt{\sigma^2 + \epsilon}} + \beta, \qquad \mu = \frac{1}{C}\sum_{i} x_i, \quad \sigma^2 = \frac{1}{C}\sum_i (x_i - \mu)^2

The statistics come from the CC features of one vector. They never mix positions or sequences. The Day 1 reading plan made this point explicit, because BatchNorm, the older method, averages across the batch, and that would let one sequence's statistics depend on another's.

Playground
Show
input x
1.20
−0.40
2.10
0.30
−1.00
0.80
LayerNorm(x), gain 1, bias 0
0.69
−0.88
1.57
−0.20
−1.47
0.29
mean(x)0.50LayerNorm subtracts it
std(x)1.02LayerNorm divides by it
rms(x)1.14RMSNorm divides by it
mean of RMSNorm(x)0.44not forced to 0
A six-number vector and its LayerNorm, with gain 1 and bias 0. Use the buttons to shift or scale the input.

What happened?

Without a norm, activations can grow block by block, and softmax and the feed-forward layer behave differently at different scales. The norm gives every block an input on the same scale. Day 2 compares this with RMSNorm, which drops the mean.

Feed-forward

A per-position two-layer network

Attention moves information between positions. The feed-forward layer then processes each position alone. It widens the vector to 4C4C, applies a nonlinearity, and projects back:

FFN⁡(x)=Wdown GELU⁡(Wup x)\operatorname{FFN}(x) = W_{\text{down}}\, \operatorname{GELU}(W_{\text{up}}\, x)

For Day 1 that is 128 to 512 and back to 128. The two matrices hold 131,072 parameters per block, twice the attention projections. Most of a Transformer's parameters live here.

GELU is a smooth version of ReLU. It weights each input by the probability that a standard normal variable is below it, GELU⁡(x)=x⋅Φ(x)\operatorname{GELU}(x) = x \cdot \Phi(x). Large positive inputs pass through, large negative ones go to 0, and near 0 the curve is smooth and dips slightly below zero.

Curves
ReLUGELU
0.01.02.03.04.0−4.0−3.0−2.0−1.00.01.02.03.04.0inputoutputReLUGELU
Hover the chart to read values.
GELU against ReLU. The GELU line uses the tanh approximation, which stays within about 0.001 of the exact curve PyTorch computes.
octlm/model.pyline 337
class FeedForward(nn.Module):
    def __init__(self, config: DecoderConfig) -> None:
        super().__init__()
        hidden = config.hidden_size
        self.gated = config.feed_forward == "swiglu"
        self.up = nn.Linear(config.d_model, hidden, bias=False)
        self.gate = nn.Linear(config.d_model, hidden, bias=False) if self.gated else None
        self.down = nn.Linear(hidden, config.d_model, bias=False)
        self.dropout = nn.Dropout(config.dropout)

    def forward(self, x: Tensor) -> Tensor:
        if self.gate is None:
            return self.dropout(self.down(F.gelu(self.up(x))))
        return self.dropout(self.down(F.silu(self.gate(x)) * self.up(x)))

With feed_forward = "gelu" this is the Day 1 layer. The gate branch is SwiGLU, added on Day 2.

The output head

Tied weights, the embedding read backwards

The head turns each final vector into VV scores. octlm uses the embedding matrix for this, transposed:

z=h E⊤,z∈RVz = h\, E^\top, \qquad z \in \mathbb{R}^{V}

The score for token vv is the dot product between the final hidden vector and that token's embedding row. The model predicts token vv when its output vector points the same way as vv's input vector. One matrix serves both ends, which saves V×CV \times C parameters, 131,072 at Day 1 size and 4.2 million at Day 4 size, and gives rare tokens' rows a gradient from both directions. In code it is one line: self.lm_head.weight = self.token_embedding.weight.

EXP-007

The bug that stopped the model learning

Problem: the first test run of the baseline learned poorly. The test that overfits one block failed.

The cause was the initialization. PyTorch's nn.Embedding draws its weights from a standard normal, σ=1\sigma = 1. That is a reasonable input, because the first LayerNorm rescales it. With a tied head, the same rows are also the output weights. The final hidden vector has entries of about unit size after the last norm, and its dot product with a row of CC unit-variance numbers has a standard deviation of about C≈11\sqrt{C} \approx 11. Softmax over scores that far apart puts nearly all the probability on one arbitrary token.

Playground
Standard deviation of embedding and linear weights
Logit standard deviation11.31σ × √128
Top token probability0.780uniform is 0.0010
Expected loss at step 037.19 natsln 1024 = 6.93

The tied head scores each token by a dot product between the final hidden vector, about unit size per entry after the last norm, and that token's embedding row. With rows drawn at σ = 1, the scores spread by σ√128 ≈ 11, one random token gets most of the probability, and the loss starts far above the ln V a blank model should show. Average of 24 random draws, V = 1024, width 128, the Day 1 shape.

The expected loss of a fresh model with a tied head at each initial weight scale. A model that knows nothing should start at ln V.

What happened?

The fix draws every linear and embedding weight from N(0,0.022)\mathcal{N}(0, 0.02^2), the GPT-2 value, and sets biases to zero. The overfit test passed after the change. Decision: keep.

octlm/model.pyline 394
@staticmethod
def _initialize(module: nn.Module) -> None:
    if isinstance(module, (nn.Linear, nn.Embedding)):
        nn.init.normal_(module.weight, mean=0.0, std=0.02)
    if isinstance(module, nn.Linear) and module.bias is not None:
        nn.init.zeros_(module.bias)
Parameters

Where the 541,952 parameters live

Playground
Token embedding (shared with the output head)131,072
Learned position table16,384
Attention: query, key, value, output131,072
Feed-forward262,144
Norm gains and biases1,280
Parameters541,952
Feed-forward hidden width5124d
float32 weights2.1 MiB
The Day 1 config from configs/day1.toml. Change any field to see what moves.

Per block, attention holds 4C2=65,5364C^2 = 65{,}536 weights, the feed-forward layer 8C2=131,0728C^2 = 131{,}072, and the two norms 4C=5124C = 512. The linear layers have no biases. The token table adds VCVC, the position table Tmax⁡CT_{\max}C, and the final norm 2C2C.

Skipped

What we did not build, and why