Multi-token prediction
Train extra heads to predict the token two and three steps ahead from the same hidden state. The code is built and tested, the plan predicted no gain at 3.3M parameters, and the run moved to Day 5.
Next-token prediction gives the model one target per position. Multi-token prediction (MTP) gives it several targets, the next token, the one after, and so on, each from its own head on the same trunk. The extra targets are extra training signal, and at inference time an extra head can draft tokens ahead for speculative decoding. Day 3 built it, wrote down a negative hypothesis, and stopped before the run.
One trunk, several future targets
The trunk is the usual stack of blocks. It produces one final hidden vector per position. Head maps to a distribution over the token steps ahead, . The training loss averages the heads:
Depth is ordinary next-token training. Each extra depth loses positions at the end of the block, because their targets fall past the last token.
What happened?
- At depth 2, position 2 must predict both the next token and the one after it. The second target requires planning past the immediate next token.
- The table's right edge loses one cell per extra depth. Those positions have no target that far ahead.
- Each extra depth adds one 256 × 256 adapter, 65,536 parameters, about 2 percent of the model.
Parallel heads or sequential modules
The two papers Day 3 read build MTP differently:
- Gloeckle et al. (2024) put independent heads on one shared trunk, each a Transformer layer feeding a shared unembedding. The heads run in parallel.
- DeepSeek-V3 uses sequential modules. Each is a full Transformer block plus a projection that mixes the trunk's hidden state with the embedding of the next true token, so every depth keeps a complete causal chain. It reports that the first extra depth's prediction is accepted over 80 percent of the time in speculative decoding, worth about 1.8 times faster decoding.
The note chose Gloeckle's parallel form. A full Transformer block per depth would add almost a model's worth of parameters at 3.3 million, and the comparison would measure capacity instead of the objective. It then made one more deviation and recorded it: each depth gets one linear adapter instead of a Transformer layer, the cheapest thing that keeps the heads distinct.
Why the heads cannot all be tied
The Day 3 plan first proposed tying every head to the embedding matrix, as the main head is, so MTP would add no parameters. The note caught the error before any code. Heads that share the unembedding and read the same hidden state compute identical logits. Depth 2 would predict token again, not . Each depth needs its own transformation before the shared unembedding, and octlm's model says so where it builds them:
def __init__(self, config: DecoderConfig) -> None:
super().__init__()
config.validate()
self.config = config
self.token_embedding = nn.Embedding(config.vocab_size, config.d_model)
self.position_embedding = (
nn.Embedding(config.context_length, config.d_model)
if config.position == "learned"
else None
)
self.dropout = nn.Dropout(config.dropout)
self.blocks = nn.ModuleList(TransformerBlock(config) for _ in range(config.n_layers))
self.final_norm = make_norm(config)
self.lm_head = nn.Linear(config.d_model, config.vocab_size, bias=False)
self.lm_head.weight = self.token_embedding.weight
# One adapter per extra depth. Heads sharing both the trunk state and the tied
# unembedding would compute identical logits, so each depth needs its own map.
self.mtp_adapters = nn.ModuleList(
nn.Linear(config.d_model, config.d_model, bias=False)
for _ in range(config.mtp_depth - 1)
)
self.apply(self._initialize)The loss shifts the targets by for depth and drops the positions that fall off the end:
def mtp_loss(model: Decoder, inputs: Tensor, targets: Tensor, pad_id: int) -> Tensor:
"""Depth k predicts token t+k, so depth k reads targets shifted by k and loses k positions."""
depth = model.config.mtp_depth
if depth == 1:
logits = model(inputs)
return F.cross_entropy(logits.flatten(0, 1), targets.flatten(), ignore_index=pad_id)
stack = model(inputs, all_depths=True)
losses = []
for k, logits in enumerate(stack):
span = targets.shape[1] - k
shifted = targets[:, k:]
losses.append(
F.cross_entropy(logits[:, :span].flatten(0, 1), shifted.flatten(), ignore_index=pad_id)
)
return torch.stack(losses).mean()With mtp_depth = 1 this is exactly the ordinary cross entropy, and a test checks that the logits match the pre-Day-3 model. That identity is the guarantee that adding the feature changed nothing for every model that does not use it.
Written negative, before the run
Gloeckle et al. report a scale coupling. MTP's gains grow with model size, and small models see muted or harmful effects on some benchmarks. At 7B on 200B tokens, their 2-future model matched the baseline and the 4-future model regressed. The gains that hold are on code and at 3B parameters and above.
So the note wrote the hypothesis in the negative. MTP does not improve bits per byte at 3.3M parameters. What the run would actually measure is the cost, and the depth-2 agreement rate, which is how often the depth-2 head's top guess is the true token two steps ahead, measured by mtp_agreement in octlm/day3.py. That rate is the number that says whether a free draft head for speculative decoding is worth building.
Stop condition: depths 1, 2 and 3, three seeds each. Keep only if depth 2 beats the noise floor on code, or if its agreement rate justifies a draft head. Otherwise revert and record the scale finding as the reason.
Not run. Moved to Day 5 as EXP-070
The plan changed on 2026-09-23, before Stage 0. The code stayed. Why the plan changed explains the decision. EXP-070 runs depth 2 against the same control on the 20M model, after EXP-069 measures the seed spread at that scale, and the hypothesis stays negative.
What we did not build, and why
- DeepSeek's sequential modules. Explained above: they would turn an objective comparison into a capacity comparison.
- Speculative decoding. A draft head is only useful if the agreement rate is high, and nothing measured it. Speculative decoding was dropped from the plan with the rest of the serving work.
- The small-scale MTP paper. The note listed "Babies Learn to Look Ahead" to read before the run. It was not read, and the note says so.