4Day 4

Sampling

A model outputs a distribution, and something has to pick a token from it. Greedy decoding repeats itself on stories. Temperature reshapes the distribution and top-k cuts its tail, and both are one line each. On the trained Day 4 model, both decoders wrote coherent stories, and sampling traded the loops for odder plot turns.

A language model predicts a probability for every token in the vocabulary. Generation turns that into text one token at a time. It picks a token, appends it, and runs the model again. Up to Day 3, octlm always picked the most likely token. Day 4 added sampling, because the exit check judges the model's stories, and greedy decoding would make it judge the decoding instead.

The problem

Greedy decoding loops

Greedy decoding takes the argmax at every step. It is deterministic, and on story data it repeats itself. Once a phrase like "and they were very happy" becomes the most likely continuation of itself, the model says it again, and again. Sampling breaks the loop by sometimes picking a less likely token, and two settings control how often and how far.

Temperature

Divide the scores before the softmax

Temperature τ\tau divides every logit before the softmax:

pi=ezi/τ∑jezj/τp_i = \frac{e^{z_i / \tau}}{\sum_j e^{z_j / \tau}}

At τ=1\tau = 1 the distribution is the model's own. Below 1 the gaps between scores grow, and the distribution sharpens toward the top token. Above 1 the gaps shrink, and it flattens toward uniform. As τ\tau approaches 0 it becomes greedy, which is why octlm treats temperature 0 as argmax rather than dividing by zero.

Top-k

Keep only the k best tokens

Top-k sets every logit outside the kk largest to −∞-\infty before the softmax, so those tokens get exactly zero probability. Temperature alone never removes a token. With thousands of unlikely tokens, their combined probability can be large enough that the model regularly picks nonsense. Top-k cuts that tail.

Playground

Context: The dog was very. These are made-up scores for eight candidate tokens.

TokenLogitProbabilitypDrawn in 200
" happy"3.10.44995
" sad"1.20.04211
" big"2.20.14628
" little"2.60.24052
" red"0.40.0152
" very"1.90.10012
" the"−0.50.0050
" ran"−1.20.0020
Entropy2.09 bitsuniform over 8 is 3.00
Eight candidate next tokens with made-up scores. The probabilities use the same function as octlm's sampling_distribution, and a test checks them against PyTorch at four temperatures and four values of k. The draw count uses a seeded generator.

What happened?

Chart
" happy"" little"" big"" sad"
0.00.20.40.60.81.00.51.01.52.0temperatureprobability" happy"" little"" big"" sad"
Hover the chart to read values.

The same eight scores as the playground above, four of them drawn. At low temperature the top token takes almost everything, and sampling turns into greedy decoding. As the temperature rises the curves pull together toward 1/8, a uniform guess. Hover to read the probabilities at one temperature.

Probability of four candidate tokens as the temperature rises, top-k off.
In octlm

Two functions and a generator

Decoder.generate gained temperature, top_k and a torch.Generator. The default temperature is 0, so every earlier day's greedy output is unchanged.

octlm/model.pyline 452
def sample_token(
    logits: Tensor, temperature: float, top_k: int, generator: torch.Generator | None
) -> Tensor:
    if temperature == 0:
        return logits.argmax(dim=-1, keepdim=True)
    probabilities = sampling_distribution(logits, temperature, top_k)
    return torch.multinomial(probabilities.cpu(), 1, generator=generator).to(logits.device)
octlm/model.pyline 461
def sampling_distribution(logits: Tensor, temperature: float, top_k: int) -> Tensor:
    if top_k:
        cutoff = logits.topk(min(top_k, logits.shape[-1]), dim=-1).values[:, -1:]
        logits = logits.masked_fill(logits < cutoff, float("-inf"))
    return torch.softmax(logits / temperature, dim=-1)

sampling_distribution is the deterministic half. It filters, scales and applies softmax. The draw itself uses torch.multinomial with the seeded generator, so a given seed reproduces the same story. This page can check the distribution against PyTorch and cannot check the draws, because PyTorch's random number generator cannot be matched in the browser.

octlm/model.pyline 431
@torch.inference_mode()
def generate(
    self,
    token_ids: Tensor,
    max_new_tokens: int,
    temperature: float = 0.0,
    top_k: int = 0,
    generator: torch.Generator | None = None,
) -> Tensor:
    if token_ids.ndim != 2 or token_ids.shape[0] != 1:
        raise ValueError("generation accepts one sequence")
    if temperature < 0 or top_k < 0:
        raise ValueError("temperature and top_k must not be negative")
    for _ in range(max_new_tokens):
        context = token_ids[:, -self.config.context_length :]
        logits = self(context)[:, -1].float()
        next_token = sample_token(logits, temperature, top_k, generator)
        token_ids = torch.cat((token_ids, next_token), dim=1)
    return token_ids

Generation still reruns the whole context for every new token. The KV cache that avoids this is EXP-071 on Day 5.

EXP-068

Twenty stories from the trained model

octlm.day4 samples loaded the step-24,000 checkpoint from the main run and wrote ten fixed prompts twice: once greedy, once at temperature 0.8 with top-k 40, seed 0. Each continuation is 200 new tokens, so most stories stop mid-sentence. All twenty are below, read from the run's samples.jsonl.

Playground
Greedy

Once upon a time, there was a little girl named Lily. She loved to play with her toys and eat yummy food. One day, Lily found a big box in her room. She was very excited to see what was inside. Lily opened the box and saw a lot of yummy food. There were apples, bananas, and carrots. She wanted to eat all of them. But her mom said, "No, Lily, you must share with your brother." Lily did not want to share with her brother. She wanted all the food for herself. Lily was sad and started to cry. Her mom …

Repeated 4-grams0 of 880.0% of the generated text
Temperature 0.8, top-k 40, seed 0

Once upon a time, there was a little girl named Lily. She loved to play with her mom's makeup. One day, Lily found a big box of makeup in the makeup. She was very happy. Lily put on the makeup and went outside. She saw her friend, Tom, playing with a ball. Tom saw Lily and said, "You look funny, Lily!" Lily smiled and said, "You look silly, Tom!" Lily and Tom played all day with the makeup. They had a lot of fun. When they were tired, they went back to their homes. When they got home, Lily …

Repeated 4-grams1 of 851.2% of the generated text
Prompt given to the modelmarked words belong to a 4-word run that already appeared earlier in the same story
Pick a prompt to read both continuations side by side. Marked words belong to a four-word run that already appeared earlier in the same story, counted on this page from the generated text only.

What happened?

The note's reading, confirmed by the user on 2026-09-28: both sets pass as coherent for a 26M TinyStories model. That closes the last Day 4 exit check.

Decision: keep temperature 0.8 and top-k 40 as the default for generation.

Skipped

What we did not build, and why