Lesson 4 of 4
Rank and low-rank updates
Rank counts how many dimensions survive a transformation. It decides when equations can be solved, and it explains why low-rank updates such as LoRA can fine-tune a huge model while training a tiny share of its numbers.
Find the column space, null space and rank of a small matrix, and use rank-nullity.
Say whether has no solution, one, or infinitely many, from the column space and the null space.
Build rank-one and rank- matrices from outer products, and count the numbers needed to store them.
Follow a LoRA update entry by entry, and count its trainable parameters.

Back to the two microphones from the last lesson. In the bad arrangement, microphone 2 hears exactly half as much of each voice as microphone 1:
(The second column is half the first, so copies of it are copies of the first.) Two voices go in, but every pair of recordings is a multiple of : one number's worth of information comes out. And some pairs of voices vanish completely. Voices give recordings .
The determinant said "zero, information is lost". This lesson says how much: one dimension survives, and one is crushed. That count is the rank. Then comes the twist. A matrix that keeps only a few dimensions is also a matrix that takes very few numbers to write down, and that is exactly the idea behind LoRA, one of the most widely used ways to fine-tune large language models cheaply.
Column space, null space and rank
Some matrices cannot be undone. To say how much they lose, we need names for what they keep and what they crush.
One word first: the dimension of a line, plane or larger flat space through the origin is the number of vectors in a basis for it (module 1): 1 for a line, 2 for a plane, for all of , the space of lists of numbers. Every basis of the same space has the same number of vectors, so the dimension is well defined. Now three definitions, each shown on the bad microphone matrix.
- The column space of is the set of all outputs . Since every output is a combination of the columns, it is the span of the columns (module 1). For the microphones: every multiple of , a line.
- The null space of is the set of all inputs that sends to . It is never empty, since , and it is a flat space through the origin: if sends and to , it sends every combination of them to too. For the microphones: the inputs with , which are the multiples of , another line.
- The rank of is the dimension of its column space: the largest number of columns you can pick that are independent of each other. Rank is the number of dimensions that survive the transformation. For the microphones: 1.
In 2D there are three cases. Rank 2: the plane maps onto the plane, the determinant is nonzero, and only goes to . Rank 1: the plane collapses onto a line, and a whole line of inputs goes to . Rank 0: the zero matrix sends everything to the origin.
The widget below shows the bad microphone matrix at work, so you can see both spaces. The blue and coral lines are where the grid lands: all on one line, the column space. The blue and coral arrows are the two columns, and . The white dashed arrow is an input you can drag, starting at the vanishing voices ; its image is drawn as a solid white arrow and given in the readout. The det A readout is there too.
To recap: the column space is where outputs can land, the null space is what gets crushed to zero, and the rank is the dimension of the column space. Next, a 3D example shows how those two dimensions are tied together.
Rank-nullity: what survives plus what is crushed
In 2D the numbers were small enough to see at a glance. In 3D the same bookkeeping reveals a rule that holds for every matrix.
Take the matrix whose columns are the three vectors from module 1's section on independence, , and :
What survives. The third column is the sum of the first two, , so it adds no new direction. Any output is
Call the two coefficients and . Then every output is : its third entry is always the sum of the first two. So the column space is the plane , and the rank is 2. All of 3D space is squashed onto that plane.
What is crushed. The relation says . Check it row by row: , and . Nothing else is crushed except multiples of it: if then , and since and are independent both coefficients are 0. So and , which makes . The null space is the line of multiples of .
A shadow is a good picture of this kind of map. Sunlight casts every point of a 3D object onto a flat wall: 3D goes in, a 2D plane comes out, and every point along one ray of light lands on the same spot of the shadow.
A shadow squashes 3D onto a plane. A whole line of points, one ray of light, lands on each point of the wall: two dimensions survive and one is crushed.
Now count: 2 dimensions survive (the plane) and 1 is crushed (the line), and , the number of input dimensions. That is always true.
Go slower: Why rank-nullity holds
Let the null space have dimension , with basis (if , start with no vectors).
Build a basis of all inputs. Keep adding vectors of that are not yet in the span of the ones you have. Each new vector is independent of the earlier ones, and cannot hold more than independent vectors, so this stops, and it stops exactly when the span is all of . The result is a basis of , so it has vectors: . We show that is a basis of the column space, so the rank is .
They reach every output. Any output is for some . Write in the basis. By linearity, Every is , because is in the null space, so the first sum vanishes:
They are independent. Suppose . By linearity this says , so is in the null space. Then it is a combination of the null-space basis: it equals for some numbers . Move everything to one side: This is a combination of basis vectors equal to zero, and basis vectors are independent, so every coefficient is zero. In particular every .
So the column space has a basis of vectors: .
In short: every input dimension is either kept (it counts toward the rank) or crushed (it counts toward the null space), never both and never neither.
Solving with a singular matrix
In the last lesson, an invertible meant exactly one solution to . What happens when cannot be inverted? The column space and the null space answer it. With the 3x3 matrix above:
- lies in the plane , since . And works: . So does for every number , because adding a null-space vector adds to the output. For example gives , and . Infinitely many solutions: one particular solution plus the whole null space.
- has but . It is outside the column space, and no input lands there. No solution.
One more fact, stated without proof: the largest number of independent rows always equals the largest number of independent columns. That is surprising, since the rows and the columns live in different spaces, but it means rank is the same whether you look at rows or columns. An matrix therefore has rank at most the smaller of and : a matrix has only 3 rows, so its rank is at most 3, even though it has 1,000 columns. A matrix whose rank reaches that limit is called full rank.
For a square matrix, all of the following say the same thing, and each one implies the others:
- ;
- has an inverse;
- the columns are independent;
- the rank is ;
- the null space contains only ;
- has exactly one solution for every .
Now read a rank off the columns of a wide matrix: walk through them and ask whether each one adds a new direction.
The matrix has columns . What is its rank?
Choose one answer, then check.
Next, go the other way: find a vector that a singular matrix crushes. Look for a relation between its columns, as we did for .
The matrix is singular. Find a nonzero vector with , scaled so that its third entry is 1. Enter as a column of three entries, top to bottom.
One entry per box, top to bottom. 0.25, -2, 3/4 and sqrt(2) all work. Enter moves to the next empty box and checks once all are filled.
How a computer finds the rank
Spotting relations by eye works for small integer matrices. A computer needs a procedure. Walk through the columns and keep each one that adds a new direction. A column adds a new direction when it has a nonzero part perpendicular to all the columns kept so far, and projection from the dot product lesson finds that part: keep a set of perpendicular unit vectors that span the kept columns, remove the new column's shadow on each of them, and see whether anything is left.
On the 3x3 example:
- is the first column, so keep it. Its unit vector is .
- . Its dot product with is , so its shadow on is . What is left is , which is not zero, so keep . Its unit vector is divided by its length, .
- . Its dot product with is , so its shadow on is , which leaves . That leftover is exactly the vector left over in step 2, so it points along : its shadow on is all of it, and nothing is left. Skip .
Two columns kept: rank 2. np.linalg.matrix_rank uses a sturdier method based on singular values, which arrive in module 4. Now write the procedure yourself.
Compute rank the way the lesson defines it: count the columns that add a new direction.
independent_columns(A, tol)scans the columns from left to right and returns the indices of those that are not combinations of the columns kept before them.rank(A, tol)is how many there are.nullity(A, tol)is the dimension of the null space, from rank-nullity.
To decide whether a column adds a new direction, use projection from the dot product lesson: keep a list of perpendicular unit vectors spanning the kept columns, remove the new column's shadow on each of them, and see whether anything is left. Because of rounding, "nothing left" means shorter than tol times the column's own length.
raise NotImplementedError, then press Run tests. Each check says what it expects.Rank one and rank k
So far low rank has meant something lost. Now it becomes something gained: a matrix of low rank takes very few numbers to write down.
The simplest nonzero matrices have rank 1. Take a column vector with entries and a column vector with entries. Their outer product is the matrix whose entry in row and column is . (The transpose is laid on its side as a row, so is a column times a row, as in module 2.) For and , row 1 is and row 2 is :
Every column is a multiple of (3, 1 and 2 times it) and every row is a multiple of (1 and 2 times it). So the column space is the line through and the rank is 1. Storing it takes the numbers of and instead of all entries. For a rank-one matrix that is 2,000 numbers instead of 1,000,000.
Sums of outer products. A sum of outer products,
where is with columns and is with columns , has rank at most . The reason: column of is (with meaning entry of ), so every column is a combination of the same vectors. The equality on the right is the outer-product view of matrix multiplication from module 2: a product is the sum of (column of the left factor) times (row of the right factor). Storing this way takes numbers.
The converse is also true: every matrix of rank is a sum of outer products. Pick independent columns of . They span the column space, because independent vectors inside a space of dimension reach all of it. So every column of is a combination of them: column is for some numbers . Collect those numbers in a matrix ; then with , which is outer products, one for each column of with the matching row of .
One more fact carries the rest of this lesson. If is and is , then has rank at most , however large and are. The reason is the column view: each column of is times a column of , which is a combination of the columns of . All outputs live in a space of dimension at most .
The run below builds a matrix from two outer products, draws each layer as a heatmap (one colored cell per entry), and checks the ranks and the storage count.
Each rank-one layer is a pattern of stripes: one fixed row shape, repeated down the rows with different strengths. The sum has rank 2 and costs numbers instead of 48.
A picture that is not built from a few layers can still be approximated by a few. The widget below does that for a grid of brightness values, a smiley face, which is a matrix with 256 entries between 0 (dark) and 1 (bright). is a sum of rank-one layers. The widget picks the best possible layers, using the singular value decomposition from module 4; for now take that choice as given and watch what rank buys.
On screen: on the left, in the middle, and on the right, which is what is still missing (coral where is brighter, blue where overshoots, dark where they agree). Under them are the slider for , the Build it up button and a sentence that says what changed. Below that, the last layer kept is drawn as a column times a row. The readouts give the relative error (the size of as a share of the size of ), the numbers stored, , and the rank of A.
The face is exact at rank 8, not 16, because it is symmetric from left to right: column equals its mirror column, so at most 8 columns can be independent. Notice the price at that point. Eight layers of a picture store numbers, exactly as many as the picture itself. Factoring saves storage only while : here , so . Low rank pays off when is much smaller than both dimensions. For a small picture that is barely the case. For a weight matrix in a language model the break-even rank is , and the ranks people use for LoRA are more like 8 or 16.
LoRA: a low-rank change to the weights
Everything in this lesson comes together in one of the most widely used ways to fine-tune large language models. Fine-tuning means taking a pretrained model and adjusting its weights for a new task.
A pretrained layer has a weight matrix of shape : in the column convention, it maps an input with numbers to the output with numbers. Full fine-tuning adjusts all entries. LoRA (low-rank adaptation, Hu and colleagues, 2021) freezes and learns only a change of the form
with a small rank . (, the Greek capital delta, means "change in". and are the names the LoRA paper uses for the two factors; this is not the matrix of the earlier sections.) The adapted layer computes . Two facts from this lesson apply at once: has rank at most , and the two factors hold only numbers.
The diagram follows one tiny LoRA update through five steps: a weight matrix and an update of rank 2. Each square is one entry, colored by its value (coral for positive, blue for negative, deeper for larger), with the number written on it. The shape and the count of numbers are written under each matrix ('s are to its right). Move through the steps with Next and Back, the step dots, or the arrow keys.
Three details follow from this lesson.
- Training starts from the pretrained model. starts at zero (and at small random values), so at the start and the adapted layer gives exactly .
- The update is restricted, not small. It has rank at most , so whatever the input, it adds a vector from one fixed -dimensional subspace, the column space of . But the entries of and can be large, so along those few directions the change can be big.
- It costs nothing at run time. After training, can be added up into one matrix (step 4 of the diagram), so the adapted model runs as fast as the original. Or and can be kept apart, so one base model can serve many small adapters.
Count the trainable numbers of an update yourself, this time for a rectangular matrix.
A weight matrix has shape . A LoRA update uses rank , with of shape and of shape . How many numbers does the update train? Type the whole number without commas.
Type a number: 0.25, -2, 3/4 and sqrt(2) all work. Enter checks.
Now build a LoRA-style adapter in code, in the row convention numpy uses, and measure why the factored form is also cheaper to run.
Build a LoRA-style adapter for one layer. In code a layer computes X @ W, with X of shape (batch, d_in) and W of shape (d_in, d_out) (the row convention from module 2). So the two factors come in the transposed order compared with the lesson's math: down has shape (d_in, r) and squeezes each input row to numbers, then up has shape (r, d_out). In the lesson's notation down is and up is , because . The change to the weights is down @ up, a (d_in, d_out) matrix of rank at most .
lora_params(d_in, d_out, r): how many numbers the two factors hold.lora_init(d_in, d_out, r, rng):downrandom with standard deviation 0.01,upall zeros, so the adapted layer starts out identical to the original.lora_forward(X, W, down, up, scale): the adapted layer's output,X @ (W + scale * down @ up), computed without ever buildingdown @ up.merge(W, down, up, scale): the single matrix to ship after training.update_madds(batch, d_in, d_out, r): the multiply-adds of the update path when grouped as(X @ down) @ up, and whendown @ upis formed first and then multiplied byX. Return them as a pair.
raise NotImplementedError, then press Run tests. Each check says what it expects.To recap the lesson: rank counts the dimensions that survive; rank plus nullity is the number of columns; a rank- matrix is outer products and costs numbers; and LoRA trains exactly such a matrix as the change to a frozen layer. Finish by explaining what that restriction does and does not mean.
A colleague says: "LoRA with rank 8 on a 4096 by 4096 layer trains less than half a percent of the parameters, so it can only make tiny changes to the model." Explain what rank measures, what a rank-8 update can and cannot do to a layer, and where your colleague is right and where they are wrong.
Saved in this browser as you type.