2Day 2

Where the norm goes

The 2017 Transformer normalized after each residual addition. Modern decoders normalize before each block instead. Day 2 swapped the order back and lost a full bit per byte.

A Transformer block has two parts, attention and a feed-forward layer, and each is wrapped in a residual addition and a norm. The order of those three operations has a name. The 2017 paper added first and normalized after, called post-norm. GPT-2 and almost everything since normalize first, inside the branch, called pre-norm. EXP-015 tested whether that choice matters at our scale.

The two layouts

Norm inside the branch, or on the main path

Diagram
Placement
x inx outNormAttention+NormFeed-forward+
x = x + Attention(Norm(x)); x = x + FFN(Norm(x))

The straight line on the left never passes through a norm. Every block adds to it, and the gradient reaches the embedding through additions only. octlm uses this placement.

One block in each layout. The blue path is the residual connection. In pre-norm it carries x untouched from input to output. In post-norm it runs into a norm after every addition.

What changes

Both are one line in octlm:

octlm/model.pyline 353
class TransformerBlock(nn.Module):
    def __init__(self, config: DecoderConfig) -> None:
        super().__init__()
        self.pre_norm = config.residual == "pre"
        self.attention_norm = make_norm(config)
        self.attention = CausalSelfAttention(config)
        self.feed_forward_norm = make_norm(config)
        self.feed_forward = FeedForward(config)

    def forward(self, x: Tensor, rope: tuple[Tensor, Tensor] | None = None) -> Tensor:
        if self.pre_norm:
            x = x + self.attention(self.attention_norm(x), rope)
            return x + self.feed_forward(self.feed_forward_norm(x))
        x = self.attention_norm(x + self.attention(x, rope))
        return self.feed_forward_norm(x + self.feed_forward(x))
The theory

Why post-norm needs warmup

Xiong et al. analyzed both layouts at initialization. In post-norm, the gradients of the parameters near the output are large, whatever the depth. A large learning rate at step 0 applied to large gradients makes the first updates unstable, which is why the original Transformer needed a learning-rate warmup to train at all. In pre-norm, the gradient scales with 1/L1/\sqrt{L} and stays well behaved, and the paper shows pre-norm can train without warmup.

The note drew a methodological rule from this. Removing warmup is a separate change, so it never travels with a residual change. EXP-015 kept Day 1's warmup for both layouts. Only the placement moved.

Playground
Placement
x₀outf1(N(x))+f2(N(x))+f3(N(x))+f4(N(x))+
out = x₀ + f1 + f2 + f3 + f4
Norms on the main path0
x₀ reaches the outputunchanged

In pre-norm the output is a plain sum. The gradient of the output with respect to x₀ always includes an identity term, however deep the stack.

The stack unrolled. Add sublayers and compare what reaches the output in each placement.
EXP-015

One bit per byte worse

Hypothesis: at four layers with warmup, post-norm should still train, because that is the regime where the paper says it works. The question was how much it costs.

Measured
Held-out split
baseline seed spread2.803.003.203.403.603.804.00code bits per byte, lower is betterbaselinermsnormswiglupost-normgqa-4gqa-2mqa-1modern
Hover a row to read its numbers.
Mean bits per byte over three seeds with the seed spread as a bar, baseline spread shaded. The post-norm row sits far to the right of everything else. From notes/day2.md.

What happened?

The prediction held in the weak sense. It trained, badly. Four hundred steps with a 40-step warmup was not enough for post-norm to catch up, and nothing in the grid suggests more steps would close a full bit.

Decision: revert. The Day 1 pre-norm placement stays in every model.

Skipped

What we did not build, and why