SwiGLU, a gated feed-forward layer
SwiGLU replaces the feed-forward layer's single activation with a product of two projections, one of them passed through a smooth gate. It was the only single change on Day 2 that beat the noise.
The feed-forward layer holds about two thirds of a Transformer block's parameters. Day 1's version widens the vector, applies GELU to each hidden unit, and projects back. SwiGLU, used by LLaMA, Qwen and most current decoders, changes what happens in the middle.
Two projections, one of them a gate
SwiGLU computes two projections of the input instead of one. One goes through the SiLU function, also called Swish, and the result multiplies the other elementwise:
Read each hidden unit as a valve. The gate projection decides how open the valve is, and the up projection decides what flows through. With GELU, the same number decides both whether a unit fires and how strongly. With a gate, the model can learn those two things separately, and the product lets a unit respond to a combination of input features rather than a threshold on one.
What happened?
- SiLU looks like GELU: near 0 for large negative inputs, near the identity for large positive ones, with a small dip below 0 in between.
- A gate value near −4 closes the unit whatever the up value is. A large gate passes the up value almost unchanged, sign included. GELU can only pass positive values through.
- At every width, SwiGLU's hidden size is about two thirds of GELU's, and the two parameter counts differ by well under 1 percent.
Three matrices, so two thirds of the width
SwiGLU has three weight matrices where GELU has two. At the same hidden width it would have 50 percent more feed-forward parameters, and any improvement could come from the extra weights. The GLU-variants paper therefore cuts the hidden width to . octlm rounds that to a multiple of 8, which GPU matrix kernels prefer:
@property
def hidden_size(self) -> int:
"""SwiGLU spends three matrices instead of two, so 2/3 holds the parameter count."""
full = self.d_model * self.ff_multiplier
if self.feed_forward == "gelu":
return full
return max(8, round(full * 2 / 3 / 8) * 8)At that gives 680 instead of 1,024, and at it gives 1,368 instead of 2,048. EXP-011 compared at matched parameter count for exactly this reason, and the note says why. Matching the width instead would only measure that we added weights.
The layer itself is the same class as Day 1's, with the gate branch switched on:
class FeedForward(nn.Module):
def __init__(self, config: DecoderConfig) -> None:
super().__init__()
hidden = config.hidden_size
self.gated = config.feed_forward == "swiglu"
self.up = nn.Linear(config.d_model, hidden, bias=False)
self.gate = nn.Linear(config.d_model, hidden, bias=False) if self.gated else None
self.down = nn.Linear(hidden, config.d_model, bias=False)
self.dropout = nn.Dropout(config.dropout)
def forward(self, x: Tensor) -> Tensor:
if self.gate is None:
return self.dropout(self.down(F.gelu(self.up(x))))
return self.dropout(self.down(F.silu(self.gate(x)) * self.up(x)))The one swap that cleared the noise
Hypothesis: SwiGLU improves bits per byte at equal parameters. The paper reports a perplexity of 1.944 against GELU's 1.983 after 65,536 steps.
What happened?
- On code, SwiGLU improves on the baseline by 0.0818 bits per byte. The baseline's own seed spread is 0.0434, so the gain is about twice the noise.
- On prose it improves by 0.0196, a smaller gain against a much tighter prose spread.
- Its own seeds agree closely, with a code spread of 0.0085, a fifth of the baseline's.
- It costs 11 percent more time per step. Three smaller matrix multiplies are slower than two larger ones at this size, and there is one more elementwise product.
Decision: keep. Of the six single-component changes on Day 2, this is the only one whose effect sat outside the seed noise. SwiGLU is part of the modern stack and of the Day 4 model.
What is known, and what is not
The GLU-variants paper tests several gates and reports that the gated forms beat the ungated ones. On why, its conclusion famously offers no explanation and attributes the success to "divine benevolence". A common reading is that the multiplicative interaction adds expressiveness per parameter: a unit can compute something like "feature A, but only when feature B", which a single threshold needs more units to approximate. That reading is an intuition, not a result this project measured.
What we did not build, and why
- Other GLU variants. GEGLU and ReGLU use GELU or ReLU as the gate. The paper finds them close to SwiGLU, and the models Day 6 must load use SwiGLU.
- A width sweep. Only the parameter-matched width was run.