Where the norm goes
The 2017 Transformer normalized after each residual addition. Modern decoders normalize before each block instead. Day 2 swapped the order back and lost a full bit per byte.
A Transformer block has two parts, attention and a feed-forward layer, and each is wrapped in a residual addition and a norm. The order of those three operations has a name. The 2017 paper added first and normalized after, called post-norm. GPT-2 and almost everything since normalize first, inside the branch, called pre-norm. EXP-015 tested whether that choice matters at our scale.
Norm inside the branch, or on the main path
What changes
- In pre-norm, the residual stream is the embedding plus every block's correction. Nothing on the main path rescales it. The gradient of the loss reaches the first layer through additions alone.
- In post-norm, every sum is normalized before it moves on. The main path goes through norms, and the gradient has to pass back through each of them.
Both are one line in octlm:
class TransformerBlock(nn.Module):
def __init__(self, config: DecoderConfig) -> None:
super().__init__()
self.pre_norm = config.residual == "pre"
self.attention_norm = make_norm(config)
self.attention = CausalSelfAttention(config)
self.feed_forward_norm = make_norm(config)
self.feed_forward = FeedForward(config)
def forward(self, x: Tensor, rope: tuple[Tensor, Tensor] | None = None) -> Tensor:
if self.pre_norm:
x = x + self.attention(self.attention_norm(x), rope)
return x + self.feed_forward(self.feed_forward_norm(x))
x = self.attention_norm(x + self.attention(x, rope))
return self.feed_forward_norm(x + self.feed_forward(x))Why post-norm needs warmup
Xiong et al. analyzed both layouts at initialization. In post-norm, the gradients of the parameters near the output are large, whatever the depth. A large learning rate at step 0 applied to large gradients makes the first updates unstable, which is why the original Transformer needed a learning-rate warmup to train at all. In pre-norm, the gradient scales with and stays well behaved, and the paper shows pre-norm can train without warmup.
The note drew a methodological rule from this. Removing warmup is a separate change, so it never travels with a residual change. EXP-015 kept Day 1's warmup for both layouts. Only the placement moved.
One bit per byte worse
Hypothesis: at four layers with warmup, post-norm should still train, because that is the regime where the paper says it works. The question was how much it costs.
What happened?
- Post-norm is 1.00 bits per byte worse on code and 0.70 worse on prose.
- Its seed spread is small, 0.0047 on code. It is not unstable across seeds. It reliably lands in a worse place.
- Parameter count is identical, 3,740,160, because the same norms exist in both layouts. Step time is about the same.
The prediction held in the weak sense. It trained, badly. Four hundred steps with a 40-step warmup was not enough for post-norm to catch up, and nothing in the grid suggests more steps would close a full bit.
Decision: revert. The Day 1 pre-norm placement stays in every model.
What we did not build, and why
- Removing warmup from pre-norm. The theory says pre-norm can drop it. The note deferred the question, because changing warmup would invalidate every Day 2 comparison. It never came back as its own experiment, and the Day 4 model keeps a 500-step warmup.
- Newer residual variants.
day-wise.mdlisted "newer residual variants" such as DeepNorm and hyper-connections. mHC was a Day 3 experiment and was dropped with the rest of Day 3b.