Lesson 3 of 5
Generating text
A language model writes by running its prediction in a loop. It computes the distribution over the next character, chooses one, appends it and reads it back in. How it chooses (greedy, sampling, a temperature) decides what the text looks like.
Run the generation loop by hand, from logits through a softmax with a temperature to the chosen character.
Draw a sample from a distribution with one uniform random number and the running totals of the probabilities.
Explain what greedy decoding and sampling each do, and why sampled text looks more like the training text.
Apply a temperature to logits or to probabilities, and say what values below and above 1 do.

Suppose a trained character model has read "th". It hands back a probability for every character: e 0.59, a 0.21, t 0.07, and small amounts for the rest. That is a prediction, not text. To write, something has to choose one character, add it to the text, and ask the model again about the longer text, over and over.
The previous lesson built the probabilities and the loss that trains them. This lesson builds the loop that writes with them, and looks closely at the one decision inside it: which character to take. Always the favorite? A random draw? Something in between? Every chat model writes its replies with this same loop, one token at a time, and the temperature setting of a language-model API is one of the knobs you will meet here.
The generation loop
Why a loop? Because the model only ever predicts one character ahead. To get a second character, the first one has to become part of the text the model reads. So generation runs the prediction again and again:
- Feed a starting text, the prompt (older papers call it the prime), through the model, updating the state.
- Compute , the distribution over the next character.
- Choose a character from and append it to the output.
- Feed that character in as the next input and go back to step 2.
Step 4 is where generation differs from training. With teacher forcing, the input at every step was the true character from the text. Here there is no true text, so the model reads its own choices, and a poor choice stays in the text it reads from then on.
The diagram below runs this loop three times. It uses a tiny model with 6 characters, space, a, e, h, n and t, in that order, starting from the prompt "th". Its logits are made up for the example (a trained RNN would compute them from its state), so every number can be checked by hand. The top row holds the text so far. The chart below it shows the round in progress: first the logits, then the probabilities, then their running totals with a random number drawn across them.
Greedy or sampled
The loop is fixed; the choice in step 3 is ours. There are two basic ways to make it.
Greedy decoding always takes the most likely character. It is deterministic: the same prompt always gives the same text. And it tends to fall into loops. Greedy decoding with the bigram table from the previous lesson, starting from "T", produces "The the the the the the the": after "e" the most likely character is a space, after a space it is "t", after "t" it is "h", after "h" it is "e", and the cycle never breaks.
Greedy also does not find the most likely text, even though it takes the most likely character at every step. A two-character example shows why. Say the first character is A with probability 0.6 or B with 0.4. After A the model is unsure: its best next character has probability 0.4. After B it is confident: its best next character has probability 0.9. Greedy takes A, then the best continuation, for a two-character text of probability . Starting with B instead gives . The locally best first choice led to a less likely text. Finding the most likely text in general means searching over whole sequences, which is not what a single greedy pass does.
Sampling draws the character at random with the model's probabilities, so a character with probability 0.3 is chosen 30% of the time. Sampled text is more varied, and on average it looks more like the training text. To see why, we need to know what a trained model's probabilities mean.
What the probabilities mean. A trained model's is its estimate of how often character follows contexts like the one it is in. The reason is the loss. Cross-entropy rewards honest estimates: on average over many occurrences of a context, the loss is lowest when the model's probabilities match the true frequencies.
Here are numbers first. Suppose that after some context the next character is "e" 30% of the time and something else 70% of the time. A model that says for "e" pays in the 30% of cases where "e" comes and in the other 70%, so on average nats. Saying costs , and saying costs . The honest 0.3 is cheapest. The box below shows this holds for any true frequency.
Go slower: Why cross-entropy rewards the true frequencies
Take the simplest case, two possible next characters. Suppose that in some context the first one truly follows a fraction of the time, and the model says . A fraction of the time the loss is , and the rest of the time it is , so the expected loss is Differentiate term by term. The derivative of is . The derivative of is , by the chain rule on . So Set it to zero and solve for , one move per line:
The second derivative is , a sum of two positive terms, so curves upward everywhere and is its minimum. The same holds with any number of outcomes: the expected cross-entropy is smallest when the predicted distribution equals the true one.
So training pushes the probabilities toward the real frequencies, at least on text like the training text. Sampling from them reproduces those frequencies, to the extent the model has learned them: if "e" follows a context 30% of the time in Shakespeare, a well-trained model samples "e" after it about 30% of the time. Greedy decoding throws that information away and keeps only the favorite.
Drawing a sample
Sampling needs a way to pick a character with exactly its probability, using the one kind of randomness a computer offers directly: a uniform random number between 0 and 1. The diagram did it with running totals. Here is why that works.
Lay the probabilities end to end along , each as a stretch as long as its probability. For the stretches are , and . Their right ends are the running totals , and . Draw uniformly from and return the character whose stretch contains it, which is the first character whose running total is above . For example is not below 0.5 but is below 0.8, so it picks the second character.
Why is that fair? A uniform lands in any stretch with probability equal to the stretch's length, and each length is exactly that character's probability. The first character is picked whenever , which happens half the time; the second whenever , which happens of the time; the third the remaining .
Try one draw yourself. The trap is to compare with the probabilities instead of their running totals.
A model's next-character probabilities are for a, for e and for o, in that order. The sampler draws uniformly from and uses the running-totals method of the lesson. Which character does it return?
Choose one answer, then check.
Temperature
Greedy decoding is dull and repetitive; plain sampling sometimes picks an unlikely character that derails the text. Temperature moves between the two. Divide the logits by a number before the softmax:
Here is the ordinary softmax probability (), and is the probability after the temperature. With nothing changes. With the large probabilities grow at the expense of the small ones, and as all the probability moves to the most likely character: greedy decoding. With the distribution flattens toward uniform. (We write because is the sequence length; the instrument below calls it .)
The second form of the formula, with powers of the probabilities, is not obvious from the first. The box shows why they agree.
Go slower: Why dividing the logits is the same as a power of the probabilities
Softmax gives with . Take logs: , so . The logits are the log-probabilities plus one constant , the same for every .
Divide by and exponentiate, one move at a time: The second step splits , and the third uses .
The softmax divides each of these by their sum. The factor is in every term, so it cancels: So you can apply a temperature either to logits or to probabilities. In code, work with and a stable softmax: raising probabilities to the power directly underflows to 0 for small .
A worked example with .
- At the power is , so each probability is squared: . These sum to , and dividing each by gives .
- At the power is , so each is square-rooted: . These sum to , and dividing gives .
The order never changes; only how decisive the distribution is. The first character went from 0.5 to 0.658 when sharpened and to 0.415 when flattened.
The instrument below lets you sample from a distribution and change its temperature. Its three logits, for (e, a, i), give probabilities of about 0.5, 0.3 and 0.2, the worked example's. The blue bars at the top are the logits and the white bars are the probabilities. Sample draws characters and keeps a running count, drawn as hollow bars beside the white ones, and the strip under the buttons shows the characters drawn.
Check what a temperature below 1 does to a distribution you have not seen yet.
A model's next-character probabilities are . What does sampling at temperature do?
Choose one answer, then check.
Now put the lesson into code: a temperature that cannot underflow, a sampler that uses exactly one random number per draw, and a generation loop over the bigram table from the previous lesson. You will see greedy decoding fall into "the the the" for yourself.
Generate text from a bigram model.
apply_temperature(probs, temperature): the reshaped distribution, computed in log space so that a tiny temperature cannot underflow. Temperature 0 is greedy.sample_index(probs, rng): one draw with one uniform random number, by the running-sum method of the lesson.generate(P, stoi, itos, prime, n, temperature, rng): continue a prime forncharacters, each drawn from the row of the bigram table for the previous character.
The code at the bottom builds the Shakespeare bigram table and prints samples at temperatures 1, 0.5 and 0.
raise NotImplementedError, then press Run tests. Each check says what it expects.To recap: generation is a loop of predict, choose, append; greedy takes the favorite and repeats itself; sampling draws with the model's probabilities, which training makes into honest frequencies; and a temperature sharpens or flattens those probabilities before the draw. The bigram table writes poor English because one character of context is very little. The next lesson trains the RNN, whose state can carry much more.