The baseline block
A Transformer block is attention and a feed-forward layer, each wrapped in a norm and a residual connection. This post assembles the Day 1 decoder from those parts, including the initialization bug that stopped it learning.
The previous posts built the pieces that turn text into vectors and let positions read each other. A decoder stacks the same block several times on top of those vectors, then turns the final vectors back into token scores. Day 1's baseline has two blocks, width 128 and four heads, and it follows the small GPT-2 layout.
The forward pass, with shapes
For a batch of sequences of token IDs, with the Day 1 settings and :
Every block keeps the shape [B, T, C]. That is what lets blocks stack, because the output of one is a valid input to the next.
def forward(self, token_ids: Tensor, all_depths: bool = False) -> Tensor:
if token_ids.ndim != 2:
raise ValueError("token IDs must have shape [batch, sequence]")
length = token_ids.shape[1]
if self.position_embedding is not None and length > self.config.context_length:
raise ValueError("sequence exceeds context_length")
rope = None
if self.config.position == "rope":
rope = rope_tables(
length, self.config.head_size, self.config.rope_scale, token_ids.device
)
x = self.dropout(self._embed(token_ids))
for block in self.blocks:
x = block(x, rope)
hidden = self.final_norm(x)
logits = self.lm_head(hidden)
if not all_depths:
return logits
return torch.stack([logits, *(self.lm_head(a(hidden)) for a in self.mtp_adapters)])Each block adds to the vector, it does not replace it
A block does not compute a new vector from scratch. It computes a correction and adds it:
The running vector is called the residual stream. The addition has a concrete benefit for training. The gradient of with respect to is , so every block passes at least the identity backward. A deep stack of additions gives the gradient a direct path from the loss to the embedding, and it does not shrink as blocks multiply.
class TransformerBlock(nn.Module):
def __init__(self, config: DecoderConfig) -> None:
super().__init__()
self.pre_norm = config.residual == "pre"
self.attention_norm = make_norm(config)
self.attention = CausalSelfAttention(config)
self.feed_forward_norm = make_norm(config)
self.feed_forward = FeedForward(config)
def forward(self, x: Tensor, rope: tuple[Tensor, Tensor] | None = None) -> Tensor:
if self.pre_norm:
x = x + self.attention(self.attention_norm(x), rope)
return x + self.feed_forward(self.feed_forward_norm(x))
x = self.attention_norm(x + self.attention(x, rope))
return self.feed_forward_norm(x + self.feed_forward(x))The norm sits inside the branch, before attention and before the feed-forward layer. That placement is called pre-norm. The 2017 paper put it after the addition, called post-norm, and Day 2 tested that choice in where the norm goes.
Normalize each vector across its own features
LayerNorm rescales one position's vector to mean 0 and variance 1, then applies a learned gain and bias per feature:
The statistics come from the features of one vector. They never mix positions or sequences. The Day 1 reading plan made this point explicit, because BatchNorm, the older method, averages across the batch, and that would let one sequence's statistics depend on another's.
What happened?
- The output always has mean 0 and standard deviation 1, whatever the input.
- "Add 1 to every entry" leaves the output unchanged. LayerNorm removes any shift.
- "Double every entry" also leaves it unchanged. LayerNorm removes any scale.
- The order of the entries survives. Larger inputs stay larger.
Without a norm, activations can grow block by block, and softmax and the feed-forward layer behave differently at different scales. The norm gives every block an input on the same scale. Day 2 compares this with RMSNorm, which drops the mean.
A per-position two-layer network
Attention moves information between positions. The feed-forward layer then processes each position alone. It widens the vector to , applies a nonlinearity, and projects back:
For Day 1 that is 128 to 512 and back to 128. The two matrices hold 131,072 parameters per block, twice the attention projections. Most of a Transformer's parameters live here.
GELU is a smooth version of ReLU. It weights each input by the probability that a standard normal variable is below it, . Large positive inputs pass through, large negative ones go to 0, and near 0 the curve is smooth and dips slightly below zero.
class FeedForward(nn.Module):
def __init__(self, config: DecoderConfig) -> None:
super().__init__()
hidden = config.hidden_size
self.gated = config.feed_forward == "swiglu"
self.up = nn.Linear(config.d_model, hidden, bias=False)
self.gate = nn.Linear(config.d_model, hidden, bias=False) if self.gated else None
self.down = nn.Linear(hidden, config.d_model, bias=False)
self.dropout = nn.Dropout(config.dropout)
def forward(self, x: Tensor) -> Tensor:
if self.gate is None:
return self.dropout(self.down(F.gelu(self.up(x))))
return self.dropout(self.down(F.silu(self.gate(x)) * self.up(x)))With feed_forward = "gelu" this is the Day 1 layer. The gate branch is SwiGLU, added on Day 2.
Tied weights, the embedding read backwards
The head turns each final vector into scores. octlm uses the embedding matrix for this, transposed:
The score for token is the dot product between the final hidden vector and that token's embedding row. The model predicts token when its output vector points the same way as 's input vector. One matrix serves both ends, which saves parameters, 131,072 at Day 1 size and 4.2 million at Day 4 size, and gives rare tokens' rows a gradient from both directions. In code it is one line: self.lm_head.weight = self.token_embedding.weight.
The bug that stopped the model learning
Problem: the first test run of the baseline learned poorly. The test that overfits one block failed.
The cause was the initialization. PyTorch's nn.Embedding draws its weights from a standard normal, . That is a reasonable input, because the first LayerNorm rescales it. With a tied head, the same rows are also the output weights. The final hidden vector has entries of about unit size after the last norm, and its dot product with a row of unit-variance numbers has a standard deviation of about . Softmax over scores that far apart puts nearly all the probability on one arbitrary token.
What happened?
- At the top token takes most of the probability and the starting loss is far above . The model begins confident and wrong, and its first job is to undo that.
- At the scores spread by about 0.23, the distribution is nearly uniform, and the loss starts at .
The fix draws every linear and embedding weight from , the GPT-2 value, and sets biases to zero. The overfit test passed after the change. Decision: keep.
@staticmethod
def _initialize(module: nn.Module) -> None:
if isinstance(module, (nn.Linear, nn.Embedding)):
nn.init.normal_(module.weight, mean=0.0, std=0.02)
if isinstance(module, nn.Linear) and module.bias is not None:
nn.init.zeros_(module.bias)Where the 541,952 parameters live
Per block, attention holds weights, the feed-forward layer , and the two norms . The linear layers have no biases. The token table adds , the position table , and the final norm .
What we did not build, and why
- Biases on the linear layers. octlm follows current decoders and uses none. The norms keep their bias on Day 1 because PyTorch's
LayerNormhas one. - Dropout. The layers accept it, and every config sets it to 0. These models are undertrained, not overfit.
- Scaled initialization for residual projections. GPT-2 also shrinks the output projections by . At two layers the difference is small, and the change would need its own comparison.