Lesson 4 of 5
Preferences, rewards and adapters
Supervised fine-tuning can only imitate. Preference methods learn from comparisons, reinforcement learning learns from answers a program can check, and LoRA adapts a model by training a small change to each weight matrix. This lesson writes out each objective and its gradient, trains an adapter, and ends with what a post-training company does all day.
Write the reward-model loss and the RLHF objective, and derive the DPO loss and its gradient.
Compute group-relative advantages for reinforcement learning with verifiable rewards, and show why subtracting a baseline keeps the direction of the gradient.
Explain how safety tuning and distillation reuse these objectives, and derive the gradient of the distillation loss.
Derive the gradients of a LoRA adapter, train one, and explain why a low-rank update is often enough.
Explain to a client what a post-training company does all day.

In the previous lesson you fine-tuned a small Shakespeare model on chat replies. It learned how a reply starts, and it lost some of its Shakespeare along the way. The method you used, supervised fine-tuning, has two limits. It can only imitate: someone has to write every ideal answer, the model never sees what a worse answer looks like, and it can only become as good as the people who wrote the examples. And the way you ran it trains every weight, which for a large model takes several times the memory of the weights themselves.
The methods in this lesson get around those limits. First, learning from comparisons, because picking the better of two answers is quicker than writing an ideal one: a reward model, RLHF and DPO. Second, learning from answers a program can check, which is central to training reasoning models. Third, two jobs that reuse the same objectives: safety tuning and distillation. Fourth, LoRA, which fine-tunes by training a small, low-rank change to each weight matrix. The lesson ends with what a post-training company does all day, including how it measures whether a change helped.
The pipeline widget from the big picture opens here at stage 6, Preferences, the first stage on the rail whose loss is not the cross-entropy of a given text. As before, the panel under the rail says what goes in and comes out, shows the loss in a pink box, and gives a small example.
The sections follow the rail: preferences, then checkable rewards, then two more uses of the same objectives, then LoRA, which works with any of them.
Learning from preferences
Comparing two answers is quicker than writing an ideal one, and it teaches something imitation cannot: which of two plausible answers is better. Show a person (or a judging model) two answers to the same prompt and ask which is better. Call the preferred one and the other (for "win" and "lose").
The data diagram from the previous lesson opens below at step 3, one such comparison. The track at the top lights the current stage, lime marks the chosen answer, and coral the rejected one.
A reward model turns comparisons into a score. It is a network, often the SFT model with its output layer replaced by a single number, that maps a prompt and a response to a reward , where are its parameters. The Bradley-Terry model says the preferred response wins with probability , where is the sigmoid. Training maximizes the likelihood of the observed choices:
Step 3 of the diagram is one such pair. The rewards are for the chosen answer and for the rejected one. Their difference is , so the reward model gives the observed choice the probability , and the loss is . Training pushes the difference up, which pushes the loss toward 0. Only the difference matters: adding the same constant to every reward changes nothing, so rewards have no fixed zero point.
RLHF (reinforcement learning from human feedback) then trains the language model, now called the policy , to produce high-reward responses without drifting far from a frozen reference model , usually the SFT model:
Here is the average over responses sampled from the policy, is the Kullback-Leibler divergence, which is zero when the two distributions agree and positive otherwise, and sets how much drift costs. The penalty matters. A reward model is only an imperfect stand-in for human judgment, and a policy pushed hard against it finds responses that score well for the wrong reasons, an effect called reward hacking. Sampling is not differentiable, so the gradient is estimated from sampled responses with the policy gradient, derived in the section on verifiable rewards below. The classic algorithm for this step is PPO (Schulman and colleagues, 2017), used in the SFT, reward model, PPO pipeline that Ouyang and colleagues (2022) described for instruction-following models. It is demanding: training needs the policy, the reference, the reward model and a value model (an estimate of expected reward) at once.
The RLHF objective as a tether. The reward pulls the policy up the hill, and the KL penalty pulls it back toward the reference model, harder the further it drifts. The best policy sits in between, and closer to the post when β is larger.
DPO (direct preference optimization, Rafailov and colleagues, 2023) targets the same optimum without a reward model and without sampling during training. Write for the log-probability of a whole response, the sum of its tokens' log-probabilities (the chain rule from the big picture). The loss for one pair is
Each log-ratio measures how much the model has raised a response's probability compared with where it started. The loss is small when the preferred response has been raised more than the rejected one. It is a classification loss on fixed pairs of text, which is why DPO is so much simpler to run than RLHF.
Go slower: Where the DPO loss comes from
Fix one prompt and drop it from the notation. The RLHF objective for a distribution over responses is
Step 1: the best policy. Define , where makes it sum to 1. In words: take the reference model and multiply each response's probability by , so better responses gain; a small trusts the reward a lot, a large stays close to the reference. We now show that is the maximizer. Take logs of the definition, , and solve for the reference: . Write the penalty's logarithm as a difference and substitute: Split the bracket into three sums. The first two combine into one logarithm, and the last uses : Multiply by and subtract from the reward term: The two reward sums cancel, and the remaining sum is a KL divergence: The second term does not depend on . A KL divergence is never negative and is zero only when the two distributions are equal, so the first term is at most 0 and reaches 0 exactly at . So the maximizer is .
Step 2: solve for the reward. From the definition of , .
Step 3: substitute into Bradley-Terry. The probability that beats is . Both responses answer the same prompt, so the terms cancel:
Step 4: fit the policy directly. Replace the unknown with the model and minimize the negative log-likelihood of the observed preferences. That is . The reward model has disappeared: the policy's own log-ratios play its part.
The gradient explains how DPO behaves. Write for the quantity inside the sigmoid, so the loss for one pair is . Its derivative with respect to is , derived line by line below. Inside , only the policy's log-probabilities depend on , since the reference model is frozen, so the chain rule gives
Go slower: The gradient of the DPO loss
The sigmoid's derivative, from the derivatives lesson: .
Through the logarithm. The chain rule with :
Substitute and cancel :
Rewrite . Use and put everything over one denominator: Multiply the top and the bottom by : So .
The gradient of . Expand the log-ratios: The two reference terms are constants, so
Multiply the two pieces, and , to get the formula above.
A gradient-descent step moves against this, so it raises and lowers , and each of those is a sum over the response's tokens, trained through the same per-token gradients as cross-entropy. The pair's weight is : above one half when the model currently favors the rejected response relative to the reference (), approaching 1 the more wrongly it ranks the pair, and close to 0 once it strongly favors the preferred one.
To recap: a reward model learns scores whose difference predicts which answer people prefer, and RLHF optimizes the policy against it while the KL penalty holds it near the reference. DPO applies the same sigmoid-of-a-difference loss directly to the policy's log-ratios. The next exercise computes one DPO loss.
A DPO run uses . For one prompt, the preferred response has log-probability under the policy and under the frozen reference. The rejected response has under the policy and under the reference. Compute the DPO loss for this pair, to 3 decimal places.
Type a number: 0.25, -2, 3/4 and sqrt(2) all work. Enter checks.
Rewards you can check
Preferences need a judge, a person or a model. Some tasks need only a checker: a math problem with a known answer, code with unit tests, a puzzle with a verifiable solution. Reinforcement learning with verifiable rewards uses the checker as the reward, as in step 4 of the data diagram above. For each prompt, sample attempts from the current model, score each with if it passes and if not, and update the model to make passing attempts more likely. Reasoning models are trained this way at large scale: the model writes out intermediate steps before its answer, only the answer is checked, and the steps that led to correct answers become more likely.
The update comes from the policy gradient. In words: the gradient of the average reward equals the average, over attempts sampled from the model, of each attempt's reward times the gradient of that attempt's log-probability. (The box below derives it.)
In practice the reward is replaced by an advantage: how much better the attempt did than a baseline. With the group's average as the baseline, where , and the step is .
Step 4 of the diagram is an example. Four attempts at "What is 7 × 8?" get rewards . Their mean is , and subtracting it from each reward gives the advantages : the two successes are made more likely and the two failures less likely. If all four pass, or all four fail, every advantage is 0 and the prompt teaches nothing, so useful training prompts are the ones the model solves only some of the time. And because , each term is the cross-entropy gradient on the model's own attempt, scaled by its advantage: imitate your own successes, move away from your own failures.
Group-relative advantages. Left: attempts that pass sit above the group's average and become more likely, and attempts that fail sit below it and become less likely. Right: when every attempt passes, each one equals the average, so nothing changes.
Go slower: The policy gradient and the baseline
The log-derivative trick. Since , we have . So The reward itself is never differentiated, which is why a pass or fail checker is enough. Averaging over sampled attempts estimates this expectation.
A baseline changes nothing on average. For any number that does not depend on , The second step is the log-derivative trick backwards, , and the third swaps the sum and the gradient. So subtracting from every reward keeps the expected gradient and can greatly reduce its noise: without it, a prompt where every attempt scores 1 would push all of them up for no reason.
The group mean as a baseline. includes attempt itself, so it is not quite independent of . Write and split the mean into attempt and the others: The attempts are sampled independently, so for the average of is , and by the baseline result with . So the true gradient scaled by a positive constant. The direction is unchanged.
Production algorithms add more. PPO-style methods clip each update so that a single batch cannot move the policy far. GRPO (Shao and colleagues, 2024) uses the group-relative advantage above, divided by the group's standard deviation, and many setups keep a KL penalty toward a reference model.
To recap: with a checker, the model learns from its own attempts, pushing each attempt's log-probability up or down in proportion to how much better it did than the group's average. The next question asks you to recognize this stage from a description.
A team has 5,000 math word problems with known final answers. For each problem the model writes eight step-by-step solutions. A script checks each final number, and training makes the solutions that got it right more likely and the others less likely. Which stage is this?
Choose one answer, then check.
Two more tools: safety tuning and distillation
Two more jobs reuse the same objectives: teaching a model where the boundaries are, and moving what a large model knows into a small one.
Safety tuning uses all three kinds of training signal you have now seen: demonstrations of declining harmful requests and helping with harmless ones (a few are in chat-sft.txt, the file you fine-tuned on in the previous lesson), preference pairs that penalize both harmful compliance and needless refusal, and rule-based or model-graded rewards. The hard part is the boundary, which is why evaluations measure over-refusal as well as harm.
Distillation trains a smaller student to imitate a larger teacher: either SFT on the teacher's generated responses, or matching the teacher's whole next-token distribution at every position with , an approach popularized by Hinton, Vinyals and Dean (2015). The gradient with respect to the student's logits is : the of module 6 with the one-hot target replaced by the teacher's probabilities. The derivative of a log-softmax is , where is 1 when and 0 otherwise. So
using in the last step. A soft target carries more information per example than a single correct token, since it also says which wrong answers were nearly right.
To recap: safety tuning aims demonstrations, preferences and rewards at one boundary, and is measured on both sides of it. Distillation is cross-entropy with the teacher's distribution in place of the one-hot target, so its gradient is .
LoRA: fine-tuning a few directions
Every method so far can train every weight of the model, which is called full fine-tuning. Adam keeps two running averages for every weight it trains (you implemented them in module 6), so the memory for gradients and optimizer state is several times the memory of the weights themselves. A common mixed-precision setup keeps 16-bit weights and gradients plus 32-bit copies of the weights and both averages: about 16 bytes per trained parameter, against 2 for the weights alone. For large models that is the binding constraint. LoRA (low-rank adaptation, Hu and colleagues, 2021), which you met in module 3, freezes each pretrained matrix and trains only a low-rank change. In the column convention of that lesson, with of shape starting at zero and of shape starting random, where the rank is small, often 8 to 64.
The diagram from that lesson draws a small case: a pretrained matrix and a rank-2 update. Each square is one entry, coral for positive and blue for negative, and under each matrix are its shape and how many numbers it holds or trains.
In code, with rows as tokens, every matrix is transposed and the layer is
where is for a batch of tokens, is now , (for "down", shape ) squeezes each input to numbers, (for "up", shape ) starts at zero, and is a fixed scale ( in the paper's notation).
The gradients follow from the linear-layer rule of backpropagation: for with upstream gradient , the weight gradient is and the input gradient is . Let and . The adapter path is a linear layer followed by a linear layer , so the rule applied twice, from the output back, gives
At the start , so : only moves on the first step, and starts moving once is nonzero. If both started at zero, both gradients would be zero forever. That is why one factor starts random.
The savings. A rank- adapter on a matrix trains numbers. Suppose a model has 32 layers, each with four attention matrices, and gets rank-16 adapters on all of them. That is trainable numbers, against in those matrices: . After training, the change can be merged into the weights, , so the adapted model runs exactly as fast as the original. Or the adapters can be kept separate, so one base model serves many customers, each with a small adapter.
What rank can and cannot do. Whatever the input, the adapter's contribution to an output row, , is a combination of the rows of , so it lies in a space of at most dimensions. If the change a task needs has rank at most , the adapter can learn it exactly. If not, the best it can do, when the inputs are spread evenly in every direction, is the best rank- approximation of the needed change, which module 4 showed is the truncated SVD. The code exercise below lets you watch gradient descent find exactly that. Whether real adaptations are low rank is an empirical question. For many tasks rank 8 to 64 comes close to full fine-tuning; published comparisons have also found that LoRA tends to learn less of a genuinely new domain than full fine-tuning, and to forget less of the old one.
To recap: LoRA freezes and trains a rank- change , which costs numbers instead of , can be merged afterwards, and can represent any change of rank at most . First count the savings for a real configuration.
A model has 40 layers. LoRA adapters of rank 8 are added to two matrices in each layer, the query and value projections, each . How many numbers does the run train in total? Enter the whole number.
Type a number: 0.25, -2, 3/4 and sqrt(2) all work. Enter checks.
Then train an adapter and see the rank limit for yourself.
Train a LoRA adapter on a frozen layer and see what its rank can and cannot learn. In code the layer is X @ W0 + scale * (X @ down) @ up, with down of shape (d_in, r) and up of shape (r, d_out).
lora_loss_and_grads(X, Y, W0, down, up, scale): the mean squared error of the adapted layer and its gradients fordownandup, from the linear-layer rule.W0gets no gradient.train_lora(X, Y, W0, r, steps, lr, scale, rng):downstarts random andupat zero, then plain gradient descent. Returndown,upand the losses.
The experiment builds a layer whose task needs a change of rank 2 with singular values 3 and 1, and inputs spread evenly in every direction. It trains adapters of rank 1, 2 and 3. Predict the final losses before you run it.
raise NotImplementedError, then press Run tests. Each check says what it expects.What a post-training company does
A company that does only post-training, like the one in the big picture, works on five things, and only one of them is running training code.
- Data. Collecting demonstrations from domain experts, generating candidates with strong models and filtering them, collecting preference pairs, and writing the guidelines that say what a good answer is. Most of the quality comes from here.
- Evaluation. Held-out test sets for the customer's tasks, automatic checkers where answers can be verified, rubrics and model-based judges where they cannot (with checks that the judge agrees with people), and regression tests on general abilities to catch forgetting. The three losses of your fine-tuning run in the previous lesson (the trained replies, the held-out replies and Shakespeare) were a tiny version of this.
- Training runs. SFT and LoRA or full fine-tuning on an open-weight base model, with sweeps over the learning rate, the number of steps and the data mixture. Then preference optimization or reinforcement learning where it pays off.
- Reward and verification infrastructure. Reward models, unit-test sandboxes for code, answer checkers, and the rules for safety behavior.
- Serving. Merging or hot-swapping adapters, quantizing, and meeting latency and cost targets, because a model that is too slow or expensive is not a product.
The last exercise asks you to say all of this to someone who has to decide whether to pay for it.
A client asks why they should pay a post-training company rather than train their own model, and what that company would actually do for them. Write the explanation you would give in a meeting: what pretraining and post-training are, what the company's work consists of, and what post-training can and cannot deliver.
Saved in this browser as you type.