ml.lab
Python sleeps until you run code
09 From LSTMs to today's models

Lesson 4 of 5

Preferences, rewards and adapters

Supervised fine-tuning can only imitate. Preference methods learn from comparisons, reinforcement learning learns from answers a program can check, and LoRA adapts a model by training a small change to each weight matrix. This lesson writes out each objective and its gradient, trains an adapter, and ends with what a post-training company does all day.

About 60 minutes
By the end you can
  • Write the reward-model loss and the RLHF objective, and derive the DPO loss and its gradient.

  • Compute group-relative advantages for reinforcement learning with verifiable rewards, and show why subtracting a baseline keeps the direction of the gradient.

  • Explain how safety tuning and distillation reuse these objectives, and derive the gradient of the distillation loss.

  • Derive the gradients of a LoRA adapter, train one, and explain why a low-rank update is often enough.

  • Explain to a client what a post-training company does all day.

In the previous lesson you fine-tuned a small Shakespeare model on chat replies. It learned how a reply starts, and it lost some of its Shakespeare along the way. The method you used, supervised fine-tuning, has two limits. It can only imitate: someone has to write every ideal answer, the model never sees what a worse answer looks like, and it can only become as good as the people who wrote the examples. And the way you ran it trains every weight, which for a large model takes several times the memory of the weights themselves.

The methods in this lesson get around those limits. First, learning from comparisons, because picking the better of two answers is quicker than writing an ideal one: a reward model, RLHF and DPO. Second, learning from answers a program can check, which is central to training reasoning models. Third, two jobs that reuse the same objectives: safety tuning and distillation. Fourth, LoRA, which fine-tunes by training a small, low-rank change to each weight matrix. The lesson ends with what a post-training company does all day, including how it measures whether a change helped.

The pipeline widget from the big picture opens here at stage 6, Preferences, the first stage on the rail whose loss is not the cross-entropy of a given text. As before, the panel under the rail says what goes in and comes out, shows the loss in a pink box, and gives a small example.

The sections follow the rail: preferences, then checkable rewards, then two more uses of the same objectives, then LoRA, which works with any of them.

Learning from preferences

Comparing two answers is quicker than writing an ideal one, and it teaches something imitation cannot: which of two plausible answers is better. Show a person (or a judging model) two answers to the same prompt xx and ask which is better. Call the preferred one ywy_w and the other yly_l (for "win" and "lose").

The data diagram from the previous lesson opens below at step 3, one such comparison. The track at the top lights the current stage, lime marks the chosen answer, and coral the rejected one.

A reward model turns comparisons into a score. It is a network, often the SFT model with its output layer replaced by a single number, that maps a prompt and a response to a reward rϕ(x,y)r_\phi(x,\allowbreak y), where ϕ\phi are its parameters. The Bradley-Terry model says the preferred response wins with probability σ(rϕ(x,yw)−rϕ(x,yl))\sigma(r_\phi(x,\allowbreak y_w) - r_\phi(x,\allowbreak y_l)), where σ\sigma is the sigmoid. Training maximizes the likelihood of the observed choices:

LRM(ϕ)=−ln⁡σ(rϕ(x,yw)−rϕ(x,yl)).L_{\text{RM}}(\phi) = -\ln \sigma\big(r_\phi(x, y_w) - r_\phi(x, y_l)\big).

Step 3 of the diagram is one such pair. The rewards are 1.21.2 for the chosen answer and 0.40.4 for the rejected one. Their difference is 0.80.8, so the reward model gives the observed choice the probability σ(0.8)=1/(1+e−0.8)≈1/1.449≈0.690\sigma(0.8) = 1/(1 + e^{-0.8}) \approx 1/1.449 \approx 0.690, and the loss is −ln⁡0.690≈0.371-\ln 0.690 \approx 0.371. Training pushes the difference up, which pushes the loss toward 0. Only the difference matters: adding the same constant to every reward changes nothing, so rewards have no fixed zero point.

RLHF (reinforcement learning from human feedback) then trains the language model, now called the policy πθ\pi_\theta, to produce high-reward responses without drifting far from a frozen reference model πref\pi_{\text{ref}}, usually the SFT model:

max⁡θ  Ey∼πθ(⋅∣x)[rϕ(x,y)]−β KL⁡(πθ(⋅∣x) ∥ πref(⋅∣x)).\max_\theta\; \mathbb{E}_{y \sim \pi_\theta(\cdot \mid x)}\big[r_\phi(x, y)\big] - \beta\,\operatorname{KL}\big(\pi_\theta(\cdot \mid x)\,\big\|\,\pi_{\text{ref}}(\cdot \mid x)\big).

Here Ey∼πθ\mathbb{E}_{y \sim \pi_\theta} is the average over responses sampled from the policy, KL⁡(π ∥ πref)=∑yπ(y)ln⁡π(y)πref(y)\operatorname{KL}(\pi \,\|\, \pi_{\text{ref}}) = \sum_y \pi(y)\ln\frac{\pi(y)}{\pi_{\text{ref}}(y)} is the Kullback-Leibler divergence, which is zero when the two distributions agree and positive otherwise, and β>0\beta > 0 sets how much drift costs. The penalty matters. A reward model is only an imperfect stand-in for human judgment, and a policy pushed hard against it finds responses that score well for the wrong reasons, an effect called reward hacking. Sampling is not differentiable, so the gradient is estimated from sampled responses with the policy gradient, derived in the section on verifiable rewards below. The classic algorithm for this step is PPO (Schulman and colleagues, 2017), used in the SFT, reward model, PPO pipeline that Ouyang and colleagues (2022) described for instruction-following models. It is demanding: training needs the policy, the reference, the reward model and a value model (an estimate of expected reward) at once.

A glowing ball rests partway up a slope of lime contour rings that rise to a peak. An amber elastic band, visibly stretched, ties the ball back to a white post at the foot of the slopeThe RLHF objective as a tether. The reward pulls the policy up the hill, and the KL penalty pulls it back toward the reference model, harder the further it drifts. The best policy sits in between, and closer to the post when β is larger.

DPO (direct preference optimization, Rafailov and colleagues, 2023) targets the same optimum without a reward model and without sampling during training. Write ln⁡π(y∣x)\ln\pi(y \mid x) for the log-probability of a whole response, the sum of its tokens' log-probabilities (the chain rule from the big picture). The loss for one pair is

LDPO(θ)=−ln⁡σ ⁣(β[ln⁡πθ(yw∣x)πref(yw∣x)−ln⁡πθ(yl∣x)πref(yl∣x)]).L_{\text{DPO}}(\theta) = -\ln\sigma\!\left(\beta\left[\ln\frac{\pi_\theta(y_w \mid x)}{\pi_{\text{ref}}(y_w \mid x)} - \ln\frac{\pi_\theta(y_l \mid x)}{\pi_{\text{ref}}(y_l \mid x)}\right]\right).

Each log-ratio measures how much the model has raised a response's probability compared with where it started. The loss is small when the preferred response has been raised more than the rejected one. It is a classification loss on fixed pairs of text, which is why DPO is so much simpler to run than RLHF.

Go slower: Where the DPO loss comes from

Fix one prompt xx and drop it from the notation. The RLHF objective for a distribution π\pi over responses is J(π)=∑yπ(y) r(y)−β∑yπ(y)ln⁡π(y)πref(y).J(\pi) = \sum_y \pi(y)\,r(y) - \beta\sum_y \pi(y)\ln\frac{\pi(y)}{\pi_{\text{ref}}(y)} .

Step 1: the best policy. Define π∗(y)=πref(y) er(y)/β/Z\pi^*(y) = \pi_{\text{ref}}(y)\,e^{r(y)/\beta}/Z, where Z=∑yπref(y) er(y)/βZ = \sum_y \pi_{\text{ref}}(y)\,e^{r(y)/\beta} makes it sum to 1. In words: take the reference model and multiply each response's probability by er(y)/βe^{r(y)/\beta}, so better responses gain; a small β\beta trusts the reward a lot, a large β\beta stays close to the reference. We now show that π∗\pi^* is the maximizer. Take logs of the definition, ln⁡π∗(y)=ln⁡πref(y)+r(y)/β−ln⁡Z\ln\pi^*(y) = \ln\pi_{\text{ref}}(y) + r(y)/\beta - \ln Z, and solve for the reference: ln⁡πref(y)=ln⁡π∗(y)−r(y)/β+ln⁡Z\ln\pi_{\text{ref}}(y) = \ln\pi^*(y) - r(y)/\beta + \ln Z. Write the penalty's logarithm as a difference and substitute: ∑yπ(y)ln⁡π(y)πref(y)=∑yπ(y)[ln⁡π(y)−ln⁡π∗(y)+r(y)β−ln⁡Z].\sum_y \pi(y)\ln\frac{\pi(y)}{\pi_{\text{ref}}(y)} = \sum_y \pi(y)\Big[\ln\pi(y) - \ln\pi^*(y) + \frac{r(y)}{\beta} - \ln Z\Big]. Split the bracket into three sums. The first two combine into one logarithm, and the last uses ∑yπ(y)=1\sum_y \pi(y) = 1: =∑yπ(y)ln⁡π(y)π∗(y)+1β∑yπ(y) r(y)−ln⁡Z.= \sum_y \pi(y)\ln\frac{\pi(y)}{\pi^*(y)} + \frac{1}{\beta}\sum_y \pi(y)\,r(y) - \ln Z . Multiply by β\beta and subtract from the reward term: J(π)=∑yπ(y) r(y)−β∑yπ(y)ln⁡π(y)π∗(y)−∑yπ(y) r(y)+βln⁡Z.J(\pi) = \sum_y \pi(y)\,r(y) - \beta\sum_y \pi(y)\ln\frac{\pi(y)}{\pi^*(y)} - \sum_y \pi(y)\,r(y) + \beta\ln Z . The two reward sums cancel, and the remaining sum is a KL divergence: J(π)=−β KL⁡(π ∥ π∗)+βln⁡Z.J(\pi) = -\beta\,\operatorname{KL}(\pi \,\|\, \pi^*) + \beta\ln Z . The second term does not depend on π\pi. A KL divergence is never negative and is zero only when the two distributions are equal, so the first term is at most 0 and reaches 0 exactly at π=π∗\pi = \pi^*. So the maximizer is π∗\pi^*.

Step 2: solve for the reward. From the definition of π∗\pi^*, r(y)=βln⁡π∗(y)πref(y)+βln⁡Zr(y) = \beta\ln\frac{\pi^*(y)}{\pi_{\text{ref}}(y)} + \beta\ln Z.

Step 3: substitute into Bradley-Terry. The probability that ywy_w beats yly_l is σ(r(yw)−r(yl))\sigma(r(y_w) - r(y_l)). Both responses answer the same prompt, so the βln⁡Z\beta\ln Z terms cancel: σ ⁣(βln⁡π∗(yw)πref(yw)−βln⁡π∗(yl)πref(yl)).\sigma\!\left(\beta\ln\frac{\pi^*(y_w)}{\pi_{\text{ref}}(y_w)} - \beta\ln\frac{\pi^*(y_l)}{\pi_{\text{ref}}(y_l)}\right).

Step 4: fit the policy directly. Replace the unknown π∗\pi^* with the model πθ\pi_\theta and minimize the negative log-likelihood of the observed preferences. That is LDPOL_{\text{DPO}}. The reward model has disappeared: the policy's own log-ratios play its part.

The gradient explains how DPO behaves. Write mm for the quantity inside the sigmoid, so the loss for one pair is −ln⁡σ(m)-\ln\sigma(m). Its derivative with respect to mm is −σ(−m)-\sigma(-m), derived line by line below. Inside mm, only the policy's log-probabilities depend on θ\theta, since the reference model is frozen, so the chain rule gives

∇θLDPO=−β σ(−m) [∇θln⁡πθ(yw∣x)−∇θln⁡πθ(yl∣x)].\nabla_\theta L_{\text{DPO}} = -\beta\,\sigma(-m)\,\big[\nabla_\theta\ln\pi_\theta(y_w \mid x) - \nabla_\theta\ln\pi_\theta(y_l \mid x)\big].
Go slower: The gradient of the DPO loss

The sigmoid's derivative, from the derivatives lesson: σ′(m)=σ(m) (1−σ(m))\sigma'(m) = \sigma(m)\,(1 - \sigma(m)).

Through the logarithm. The chain rule with dduln⁡u=1/u\frac{d}{du}\ln u = 1/u: ddm[−ln⁡σ(m)]=−σ′(m)σ(m).\frac{d}{dm}\big[-\ln\sigma(m)\big] = -\frac{\sigma'(m)}{\sigma(m)} .

Substitute σ′\sigma' and cancel σ(m)\sigma(m): −σ(m) (1−σ(m))σ(m)=−(1−σ(m)).-\frac{\sigma(m)\,(1 - \sigma(m))}{\sigma(m)} = -(1 - \sigma(m)) .

Rewrite 1−σ(m)1 - \sigma(m). Use σ(m)=1/(1+e−m)\sigma(m) = 1/(1 + e^{-m}) and put everything over one denominator: 1−σ(m)=1+e−m−11+e−m=e−m1+e−m.1 - \sigma(m) = \frac{1 + e^{-m} - 1}{1 + e^{-m}} = \frac{e^{-m}}{1 + e^{-m}} . Multiply the top and the bottom by eme^{m}: e−m1+e−m=1em+1=σ(−m).\frac{e^{-m}}{1 + e^{-m}} = \frac{1}{e^{m} + 1} = \sigma(-m) . So ddm[−ln⁡σ(m)]=−σ(−m)\frac{d}{dm}\big[-\ln\sigma(m)\big] = -\sigma(-m).

The gradient of mm. Expand the log-ratios: m=β[ln⁡πθ(yw∣x)−ln⁡πref(yw∣x)−ln⁡πθ(yl∣x)+ln⁡πref(yl∣x)].m = \beta\big[\ln\pi_\theta(y_w \mid x) - \ln\pi_{\text{ref}}(y_w \mid x) - \ln\pi_\theta(y_l \mid x) + \ln\pi_{\text{ref}}(y_l \mid x)\big]. The two reference terms are constants, so ∇θm=β[∇θln⁡πθ(yw∣x)−∇θln⁡πθ(yl∣x)].\nabla_\theta m = \beta\big[\nabla_\theta\ln\pi_\theta(y_w \mid x) - \nabla_\theta\ln\pi_\theta(y_l \mid x)\big].

Multiply the two pieces, −σ(−m)-\sigma(-m) and ∇θm\nabla_\theta m, to get the formula above.

A gradient-descent step moves against this, so it raises ln⁡πθ(yw∣x)\ln\pi_\theta(y_w \mid x) and lowers ln⁡πθ(yl∣x)\ln\pi_\theta(y_l \mid x), and each of those is a sum over the response's tokens, trained through the same per-token gradients as cross-entropy. The pair's weight is σ(−m)\sigma(-m): above one half when the model currently favors the rejected response relative to the reference (m<0m < 0), approaching 1 the more wrongly it ranks the pair, and close to 0 once it strongly favors the preferred one.

To recap: a reward model learns scores whose difference predicts which answer people prefer, and RLHF optimizes the policy against it while the KL penalty holds it near the reference. DPO applies the same sigmoid-of-a-difference loss directly to the policy's log-ratios. The next exercise computes one DPO loss.

Work it outOne DPO loss

A DPO run uses β=0.1\beta = 0.1. For one prompt, the preferred response has log-probability −40-40 under the policy and −42-42 under the frozen reference. The rejected response has −38-38 under the policy and −37-37 under the reference. Compute the DPO loss for this pair, to 3 decimal places.

loss

Type a number: 0.25, -2, 3/4 and sqrt(2) all work. Enter checks.

Rewards you can check

Preferences need a judge, a person or a model. Some tasks need only a checker: a math problem with a known answer, code with unit tests, a puzzle with a verifiable solution. Reinforcement learning with verifiable rewards uses the checker as the reward, as in step 4 of the data diagram above. For each prompt, sample GG attempts from the current model, score each with Ri=1R_i = 1 if it passes and 00 if not, and update the model to make passing attempts more likely. Reasoning models are trained this way at large scale: the model writes out intermediate steps before its answer, only the answer is checked, and the steps that led to correct answers become more likely.

The update comes from the policy gradient. In words: the gradient of the average reward equals the average, over attempts sampled from the model, of each attempt's reward times the gradient of that attempt's log-probability. (The box below derives it.)

∇θ Ey∼πθ[R(y)]=Ey∼πθ[R(y) ∇θln⁡πθ(y)].\nabla_\theta\, \mathbb{E}_{y \sim \pi_\theta}\big[R(y)\big] = \mathbb{E}_{y \sim \pi_\theta}\big[R(y)\,\nabla_\theta \ln\pi_\theta(y)\big].

In practice the reward is replaced by an advantage: how much better the attempt did than a baseline. With the group's average as the baseline, Ai=Ri−RˉA_i = R_i - \bar{R} where Rˉ=1G∑jRj\bar{R} = \frac{1}{G}\sum_j R_j, and the step is θ←θ+η 1G∑iAi ∇θln⁡πθ(yi)\theta \leftarrow \theta + \eta\,\frac{1}{G}\sum_i A_i\,\nabla_\theta\ln\pi_\theta(y_i).

Step 4 of the diagram is an example. Four attempts at "What is 7 × 8?" get rewards (1,0,0,1)(1,\allowbreak 0,\allowbreak 0,\allowbreak 1). Their mean is Rˉ=(1+0+0+1)/4=0.5\bar{R} = (1 + 0 + 0 + 1)/4 = 0.5, and subtracting it from each reward gives the advantages (+0.5,−0.5,−0.5,+0.5)(+0.5,\allowbreak -0.5,\allowbreak -0.5,\allowbreak +0.5): the two successes are made more likely and the two failures less likely. If all four pass, or all four fail, every advantage is 0 and the prompt teaches nothing, so useful training prompts are the ones the model solves only some of the time. And because ln⁡πθ(y)=∑tln⁡πθ(yt∣x,y<t)\ln\pi_\theta(y) = \sum_t \ln\pi_\theta(y_t \mid x,\allowbreak y_{<t}), each term is the cross-entropy gradient on the model's own attempt, scaled by its advantage: imitate your own successes, move away from your own failures.

Two panels. Left: eight thin paths fan out from one point; three end on a high lime line and glow brighter, five end on a low coral line and fade, and a faint amber level line runs between the two lines, closer to the coral one. Right: all eight paths end on the lime line, the amber level line lies on top of it, and every path keeps the same brightnessGroup-relative advantages. Left: attempts that pass sit above the group's average and become more likely, and attempts that fail sit below it and become less likely. Right: when every attempt passes, each one equals the average, so nothing changes.

Go slower: The policy gradient and the baseline

The log-derivative trick. Since ∇ln⁡π=∇π/π\nabla\ln\pi = \nabla\pi/\pi, we have ∇π=π ∇ln⁡π\nabla\pi = \pi\,\nabla\ln\pi. So ∇θ∑yπθ(y)R(y)=∑yR(y) ∇θπθ(y)=∑yπθ(y) R(y) ∇θln⁡πθ(y)=Ey∼πθ[R(y) ∇θln⁡πθ(y)].\nabla_\theta\sum_y \pi_\theta(y)R(y) = \sum_y R(y)\,\nabla_\theta\pi_\theta(y) = \sum_y \pi_\theta(y)\,R(y)\,\nabla_\theta\ln\pi_\theta(y) = \mathbb{E}_{y\sim\pi_\theta}\big[R(y)\,\nabla_\theta\ln\pi_\theta(y)\big]. The reward itself is never differentiated, which is why a pass or fail checker is enough. Averaging R(yi)∇ln⁡πθ(yi)R(y_i)\nabla\ln\pi_\theta(y_i) over sampled attempts estimates this expectation.

A baseline changes nothing on average. For any number bb that does not depend on yy, Ey∼πθ[b ∇θln⁡πθ(y)]=b∑yπθ(y) ∇θln⁡πθ(y)=b∑y∇θπθ(y)=b ∇θ∑yπθ(y)=b ∇θ1=0.\mathbb{E}_{y\sim\pi_\theta}\big[b\,\nabla_\theta\ln\pi_\theta(y)\big] = b\sum_y \pi_\theta(y)\,\nabla_\theta\ln\pi_\theta(y) = b\sum_y \nabla_\theta\pi_\theta(y) = b\,\nabla_\theta\sum_y\pi_\theta(y) = b\,\nabla_\theta 1 = 0 . The second step is the log-derivative trick backwards, π ∇ln⁡π=∇π\pi\,\nabla\ln\pi = \nabla\pi, and the third swaps the sum and the gradient. So subtracting bb from every reward keeps the expected gradient and can greatly reduce its noise: without it, a prompt where every attempt scores 1 would push all of them up for no reason.

The group mean as a baseline. Rˉ\bar{R} includes attempt ii itself, so it is not quite independent of yiy_i. Write gi=∇θln⁡πθ(yi)\mathbf{g}_i = \nabla_\theta\ln\pi_\theta(y_i) and split the mean into attempt ii and the others: (Ri−Rˉ) gi=(1−1G)Ri gi−1G∑j≠iRj gi.(R_i - \bar{R})\,\mathbf{g}_i = \Big(1 - \frac{1}{G}\Big)R_i\,\mathbf{g}_i - \frac{1}{G}\sum_{j \ne i} R_j\,\mathbf{g}_i . The attempts are sampled independently, so for j≠ij \ne i the average of Rj giR_j\,\mathbf{g}_i is E[Rj] E[gi]\mathbb{E}[R_j]\,\mathbb{E}[\mathbf{g}_i], and E[gi]=0\mathbb{E}[\mathbf{g}_i] = \mathbf{0} by the baseline result with b=1b = 1. So E[(Ri−Rˉ) gi]=G−1G E[Ri gi]=G−1G ∇θ Ey∼πθ[R(y)],\mathbb{E}\big[(R_i - \bar{R})\,\mathbf{g}_i\big] = \frac{G - 1}{G}\,\mathbb{E}\big[R_i\,\mathbf{g}_i\big] = \frac{G - 1}{G}\,\nabla_\theta\,\mathbb{E}_{y\sim\pi_\theta}\big[R(y)\big], the true gradient scaled by a positive constant. The direction is unchanged.

Production algorithms add more. PPO-style methods clip each update so that a single batch cannot move the policy far. GRPO (Shao and colleagues, 2024) uses the group-relative advantage above, divided by the group's standard deviation, and many setups keep a KL penalty toward a reference model.

To recap: with a checker, the model learns from its own attempts, pushing each attempt's log-probability up or down in proportion to how much better it did than the group's average. The next question asks you to recognize this stage from a description.

Quick checkName the stage

A team has 5,000 math word problems with known final answers. For each problem the model writes eight step-by-step solutions. A script checks each final number, and training makes the solutions that got it right more likely and the others less likely. Which stage is this?

Choose one answer, then check.

Two more tools: safety tuning and distillation

Two more jobs reuse the same objectives: teaching a model where the boundaries are, and moving what a large model knows into a small one.

Safety tuning uses all three kinds of training signal you have now seen: demonstrations of declining harmful requests and helping with harmless ones (a few are in chat-sft.txt, the file you fine-tuned on in the previous lesson), preference pairs that penalize both harmful compliance and needless refusal, and rule-based or model-graded rewards. The hard part is the boundary, which is why evaluations measure over-refusal as well as harm.

Distillation trains a smaller student to imitate a larger teacher: either SFT on the teacher's generated responses, or matching the teacher's whole next-token distribution q\mathbf{q} at every position with L=−∑vqvln⁡pvL = -\sum_v q_v\ln p_v, an approach popularized by Hinton, Vinyals and Dean (2015). The gradient with respect to the student's logits z\mathbf{z} is p−q\mathbf{p} - \mathbf{q}: the p−y\mathbf{p} - \mathbf{y} of module 6 with the one-hot target replaced by the teacher's probabilities. The derivative of a log-softmax is ∂ln⁡pv/∂zj=δvj−pj\partial \ln p_v/\partial z_j = \delta_{vj} - p_j, where δvj\delta_{vj} is 1 when v=jv = j and 0 otherwise. So

∂L∂zj=−∑vqv(δvj−pj)=−qj+pj∑vqv=pj−qj,\frac{\partial L}{\partial z_j} = -\sum_v q_v(\delta_{vj} - p_j) = -q_j + p_j\sum_v q_v = p_j - q_j ,

using ∑vqv=1\sum_v q_v = 1 in the last step. A soft target carries more information per example than a single correct token, since it also says which wrong answers were nearly right.

To recap: safety tuning aims demonstrations, preferences and rewards at one boundary, and is measured on both sides of it. Distillation is cross-entropy with the teacher's distribution in place of the one-hot target, so its gradient is p−q\mathbf{p} - \mathbf{q}.

LoRA: fine-tuning a few directions

Every method so far can train every weight of the model, which is called full fine-tuning. Adam keeps two running averages for every weight it trains (you implemented them in module 6), so the memory for gradients and optimizer state is several times the memory of the weights themselves. A common mixed-precision setup keeps 16-bit weights and gradients plus 32-bit copies of the weights and both averages: about 16 bytes per trained parameter, against 2 for the weights alone. For large models that is the binding constraint. LoRA (low-rank adaptation, Hu and colleagues, 2021), which you met in module 3, freezes each pretrained matrix W0\mathbf{W}_0 and trains only a low-rank change. In the column convention of that lesson, ΔW=BA\Delta\mathbf{W} = \mathbf{B}\mathbf{A} with B\mathbf{B} of shape dout×rd_{\text{out}} \times r starting at zero and A\mathbf{A} of shape r×dinr \times d_{\text{in}} starting random, where the rank rr is small, often 8 to 64.

The diagram from that lesson draws a small case: a 6×66 \times 6 pretrained matrix and a rank-2 update. Each square is one entry, coral for positive and blue for negative, and under each matrix are its shape and how many numbers it holds or trains.

In code, with rows as tokens, every matrix is transposed and the layer is

P=XW0+s (XD) U,\mathbf{P} = \mathbf{X}\mathbf{W}_0 + s\,(\mathbf{X}\mathbf{D})\,\mathbf{U},

where X\mathbf{X} is (B,din)(B,\allowbreak d_{\text{in}}) for a batch of BB tokens, W0\mathbf{W}_0 is now din×doutd_{\text{in}} \times d_{\text{out}}, D=A⊤\mathbf{D} = \mathbf{A}^\top (for "down", shape din×rd_{\text{in}} \times r) squeezes each input to rr numbers, U=B⊤\mathbf{U} = \mathbf{B}^\top (for "up", shape r×doutr \times d_{\text{out}}) starts at zero, and ss is a fixed scale (α/r\alpha/r in the paper's notation).

The gradients follow from the linear-layer rule of backpropagation: for Y=XW\mathbf{Y} = \mathbf{X}\mathbf{W} with upstream gradient ∂L/∂Y\partial L/\partial\mathbf{Y}, the weight gradient is X⊤(∂L/∂Y)\mathbf{X}^\top(\partial L/\partial\mathbf{Y}) and the input gradient is (∂L/∂Y) W⊤(\partial L/\partial\mathbf{Y})\,\mathbf{W}^\top. Let G=∂L/∂P\mathbf{G} = \partial L/\partial\mathbf{P} and H=XD\mathbf{H} = \mathbf{X}\mathbf{D}. The adapter path is a linear layer H=XD\mathbf{H} = \mathbf{X}\mathbf{D} followed by a linear layer s HUs\,\mathbf{H}\mathbf{U}, so the rule applied twice, from the output back, gives

∂L∂U=s H⊤G,∂L∂H=s GU⊤,∂L∂D=X⊤∂L∂H=s X⊤GU⊤.\frac{\partial L}{\partial\mathbf{U}} = s\,\mathbf{H}^\top\mathbf{G}, \qquad \frac{\partial L}{\partial\mathbf{H}} = s\,\mathbf{G}\mathbf{U}^\top, \qquad \frac{\partial L}{\partial\mathbf{D}} = \mathbf{X}^\top\frac{\partial L}{\partial\mathbf{H}} = s\,\mathbf{X}^\top\mathbf{G}\mathbf{U}^\top .

At the start U=0\mathbf{U} = \mathbf{0}, so ∂L/∂D=0\partial L/\partial\mathbf{D} = \mathbf{0}: only U\mathbf{U} moves on the first step, and D\mathbf{D} starts moving once U\mathbf{U} is nonzero. If both started at zero, both gradients would be zero forever. That is why one factor starts random.

The savings. A rank-rr adapter on a din×doutd_{\text{in}} \times d_{\text{out}} matrix trains r(din+dout)r(d_{\text{in}} + d_{\text{out}}) numbers. Suppose a model has 32 layers, each with four 4096×40964096 \times 4096 attention matrices, and gets rank-16 adapters on all of them. That is 32×4×16×8192=16,777,21632 \times 4 \times 16 \times 8192 = 16{,}777{,}216 trainable numbers, against 32×4×40962=2,147,483,64832 \times 4 \times 4096^2 = 2{,}147{,}483{,}648 in those matrices: 0.78%0.78\%. After training, the change can be merged into the weights, W0+s DU\mathbf{W}_0 + s\,\mathbf{D}\mathbf{U}, so the adapted model runs exactly as fast as the original. Or the adapters can be kept separate, so one base model serves many customers, each with a small adapter.

What rank rr can and cannot do. Whatever the input, the adapter's contribution to an output row, s (xD) Us\,(\mathbf{x}\mathbf{D})\,\mathbf{U}, is a combination of the rr rows of U\mathbf{U}, so it lies in a space of at most rr dimensions. If the change a task needs has rank at most rr, the adapter can learn it exactly. If not, the best it can do, when the inputs are spread evenly in every direction, is the best rank-rr approximation of the needed change, which module 4 showed is the truncated SVD. The code exercise below lets you watch gradient descent find exactly that. Whether real adaptations are low rank is an empirical question. For many tasks rank 8 to 64 comes close to full fine-tuning; published comparisons have also found that LoRA tends to learn less of a genuinely new domain than full fine-tuning, and to forget less of the old one.

To recap: LoRA freezes W0\mathbf{W}_0 and trains a rank-rr change BA\mathbf{B}\mathbf{A}, which costs r(din+dout)r(d_{\text{in}} + d_{\text{out}}) numbers instead of dindoutd_{\text{in}}d_{\text{out}}, can be merged afterwards, and can represent any change of rank at most rr. First count the savings for a real configuration.

Work it outTrainable parameters in a LoRA run

A model has 40 layers. LoRA adapters of rank 8 are added to two matrices in each layer, the query and value projections, each 5120×51205120 \times 5120. How many numbers does the run train in total? Enter the whole number.

trainable parameters

Type a number: 0.25, -2, 3/4 and sqrt(2) all work. Enter checks.

Then train an adapter and see the rank limit for yourself.

Code itTrain a LoRA adapter

Train a LoRA adapter on a frozen layer and see what its rank can and cannot learn. In code the layer is X @ W0 + scale * (X @ down) @ up, with down of shape (d_in, r) and up of shape (r, d_out).

  1. lora_loss_and_grads(X, Y, W0, down, up, scale): the mean squared error of the adapted layer and its gradients for down and up, from the linear-layer rule. W0 gets no gradient.
  2. train_lora(X, Y, W0, r, steps, lr, scale, rng): down starts random and up at zero, then plain gradient descent. Return down, up and the losses.

The experiment builds a layer whose task needs a change of rank 2 with singular values 3 and 1, and inputs spread evenly in every direction. It trains adapters of rank 1, 2 and 3. Predict the final losses before you run it.

⌘+Enter runsPython sleeps until you run code
Write your code where the starter says raise NotImplementedError, then press Run tests. Each check says what it expects.

What a post-training company does

A company that does only post-training, like the one in the big picture, works on five things, and only one of them is running training code.

  • Data. Collecting demonstrations from domain experts, generating candidates with strong models and filtering them, collecting preference pairs, and writing the guidelines that say what a good answer is. Most of the quality comes from here.
  • Evaluation. Held-out test sets for the customer's tasks, automatic checkers where answers can be verified, rubrics and model-based judges where they cannot (with checks that the judge agrees with people), and regression tests on general abilities to catch forgetting. The three losses of your fine-tuning run in the previous lesson (the trained replies, the held-out replies and Shakespeare) were a tiny version of this.
  • Training runs. SFT and LoRA or full fine-tuning on an open-weight base model, with sweeps over the learning rate, the number of steps and the data mixture. Then preference optimization or reinforcement learning where it pays off.
  • Reward and verification infrastructure. Reward models, unit-test sandboxes for code, answer checkers, and the rules for safety behavior.
  • Serving. Merging or hot-swapping adapters, quantizing, and meeting latency and cost targets, because a model that is too slow or expensive is not a product.

The last exercise asks you to say all of this to someone who has to decide whether to pay for it.

In your own wordsPost-training, for a client

A client asks why they should pay a post-training company rather than train their own model, and what that company would actually do for them. Write the explanation you would give in a meeting: what pretraining and post-training are, what the company's work consists of, and what post-training can and cannot deliver.

A few sentences first: 0 of 60 characters.

Saved in this browser as you type.

Next: Vision and what comes next