2Day 2

SwiGLU, a gated feed-forward layer

SwiGLU replaces the feed-forward layer's single activation with a product of two projections, one of them passed through a smooth gate. It was the only single change on Day 2 that beat the noise.

The feed-forward layer holds about two thirds of a Transformer block's parameters. Day 1's version widens the vector, applies GELU to each hidden unit, and projects back. SwiGLU, used by LLaMA, Qwen and most current decoders, changes what happens in the middle.

Definition

Two projections, one of them a gate

SwiGLU computes two projections of the input instead of one. One goes through the SiLU function, also called Swish, and the result multiplies the other elementwise:

FFN⁡SwiGLU(x)=Wdown(SiLU⁡(Wgate x)⊙Wup x),SiLU⁡(z)=z⋅σ(z)\operatorname{FFN}_{\text{SwiGLU}}(x) = W_{\text{down}}\left(\operatorname{SiLU}(W_{\text{gate}}\, x) \odot W_{\text{up}}\, x\right), \qquad \operatorname{SiLU}(z) = z \cdot \sigma(z)

Read each hidden unit as a valve. The gate projection decides how open the valve is, and the up projection decides what flows through. With GELU, the same number decides both whether a unit fires and how strongly. With a gate, the model can learn those two things separately, and the product lets a unit respond to a combination of input features rather than a threshold on one.

Playground
ReLUGELUSiLU (Swish)
0.01.02.03.04.0−4.0−3.0−2.0−1.00.01.02.03.04.0inputoutputReLUGELUSiLU (Swish)
Hover the chart to read values.
SiLU(gate)1.226how open the gate is
SiLU(gate) × up−0.981what reaches the down projection
GELU(up), no gate−0.170the Day 1 unit
Model width d
GELU hidden1,0242 × 256 × 1024 = 524,288
SwiGLU hidden6803 × 256 × 680 = 522,240
Difference−0.39%rounding to a multiple of 8
Top: the three activation curves. Middle: one hidden unit, with its gate and up values set by hand. Bottom: the hidden width that keeps SwiGLU's parameter count equal to GELU's.

What happened?

Playground
SiLU(gate) × up9 × 9
−3−2.25−1.5−0.7500.751.52.2534−12−8.84−5.89−2.95.002.955.898.84123−8.57−6.43−4.29−2.14.002.144.296.438.572−5.28−3.96−2.64−1.32.001.322.643.965.281−2.19−1.64−1.10−.55.00.551.101.642.190.00.00.00.00.00.00.00.00.00−1.81.61.40.20.00−.20−.40−.61−.81−2.72.54.36.18.00−.18−.36−.54−.72−3.43.32.21.11.00−.11−.21−.32−.43−4.22.16.11.05.00−.05−.11−.16−.22
GELU(up), no gate9 × 9
−3−2.25−1.5−0.7500.751.52.2534−.00−.03−.10−.17.00.581.402.223.003−.00−.03−.10−.17.00.581.402.223.002−.00−.03−.10−.17.00.581.402.223.001−.00−.03−.10−.17.00.581.402.223.000−.00−.03−.10−.17.00.581.402.223.00−1−.00−.03−.10−.17.00.581.402.223.00−2−.00−.03−.10−.17.00.581.402.223.00−3−.00−.03−.10−.17.00.581.402.223.00−4−.00−.03−.10−.17.00.581.402.223.00

Rows are the gate value, columns the up value. Hover a cell to read it. On the left, a negative gate zeroes the whole row, and a positive gate lets both signs of up through, scaled. On the right the gate does not exist, so every row is the same: the unit's output depends on one number and is almost never negative. The gated unit responds to a pair of inputs.

One hidden unit's output over a grid of inputs. Left: the gated SwiGLU unit. Right: the ungated GELU unit, which ignores the gate axis.
Matching parameters

Three matrices, so two thirds of the width

SwiGLU has three weight matrices where GELU has two. At the same hidden width it would have 50 percent more feed-forward parameters, and any improvement could come from the extra weights. The GLU-variants paper therefore cuts the hidden width to 23⋅4C\tfrac{2}{3} \cdot 4C. octlm rounds that to a multiple of 8, which GPU matrix kernels prefer:

octlm/model.pyline 49
@property
def hidden_size(self) -> int:
    """SwiGLU spends three matrices instead of two, so 2/3 holds the parameter count."""
    full = self.d_model * self.ff_multiplier
    if self.feed_forward == "gelu":
        return full
    return max(8, round(full * 2 / 3 / 8) * 8)

At C=256C = 256 that gives 680 instead of 1,024, and at C=512C = 512 it gives 1,368 instead of 2,048. EXP-011 compared at matched parameter count for exactly this reason, and the note says why. Matching the width instead would only measure that we added weights.

The layer itself is the same class as Day 1's, with the gate branch switched on:

octlm/model.pyline 337
class FeedForward(nn.Module):
    def __init__(self, config: DecoderConfig) -> None:
        super().__init__()
        hidden = config.hidden_size
        self.gated = config.feed_forward == "swiglu"
        self.up = nn.Linear(config.d_model, hidden, bias=False)
        self.gate = nn.Linear(config.d_model, hidden, bias=False) if self.gated else None
        self.down = nn.Linear(hidden, config.d_model, bias=False)
        self.dropout = nn.Dropout(config.dropout)

    def forward(self, x: Tensor) -> Tensor:
        if self.gate is None:
            return self.dropout(self.down(F.gelu(self.up(x))))
        return self.dropout(self.down(F.silu(self.gate(x)) * self.up(x)))
EXP-011

The one swap that cleared the noise

Hypothesis: SwiGLU improves bits per byte at equal parameters. The paper reports a perplexity of 1.944 against GELU's 1.983 after 65,536 steps.

Measured
Held-out split
baseline seed spread2.803.003.203.403.603.804.00code bits per byte, lower is betterbaselinermsnormswiglupost-normgqa-4gqa-2mqa-1modern
Hover a row to read its numbers.
Mean bits per byte over three seeds, with the spread across seeds drawn as a bar centered on the mean. The gray band is the baseline's seed spread. From the variant-grid table in notes/day2.md.

What happened?

Decision: keep. Of the six single-component changes on Day 2, this is the only one whose effect sat outside the seed noise. SwiGLU is part of the modern stack and of the Day 4 model.

Why it works

What is known, and what is not

The GLU-variants paper tests several gates and reports that the gated forms beat the ungated ones. On why, its conclusion famously offers no explanation and attributes the success to "divine benevolence". A common reading is that the multiplicative interaction adds expressiveness per parameter: a unit can compute something like "feature A, but only when feature B", which a single threshold needs more units to approximate. That reading is an intuition, not a result this project measured.

Skipped

What we did not build, and why