ml.lab
Python sleeps until you run code
02 Matrices are machines

Lesson 4 of 4

Layers and parameter counts

A dense layer is a matrix product plus a bias, and a network is a stack of them with nonlinearities in between. This lesson shows why the nonlinearities are needed, then counts parameters and multiply-adds from a small network up to a 7B model.

About 45 minutes
By the end you can
  • Write a dense layer and a two-layer network in the row-batch convention, and follow the shapes through them.

  • Show, by hand and in algebra, that two layers without a nonlinearity collapse into one.

  • Count the parameters and multiply-adds of a stack of dense layers.

  • Count the parameters of a published 7B model and say what the count implies for memory and compute.

  • Implement dense layers, a two-layer forward pass, a parameter counter and the collapse in Python.

The last lesson ended with one layer, Y = X @ W + b, run on a whole batch at once. A network stacks several of them. Two questions follow, and this lesson answers both.

First, what does a stack compute, and why does every network put a nonlinearity between its layers? Without one, a hundred layers turn out to be no better than one, and you will see exactly why. Second, how big is a network: how many numbers must it store, and how much arithmetic does one example cost? The same counting that sizes a small digit classifier also explains a phrase you hear every day, "a 7B model".

A layer is a matrix product plus a bias

Every network in the rest of the course is built from one piece, so it is worth writing it down carefully.

A dense layer (also called fully connected) in the code convention is

dense(X)=XW+b,\text{dense}(\mathbf{X}) = \mathbf{X}\mathbf{W} + \mathbf{b},

with X\mathbf{X} of shape (batch,din)(\text{batch},\allowbreak d_{\text{in}}), W\mathbf{W} of shape (din,dout)(d_{\text{in}},\allowbreak d_{\text{out}}), and the bias b\mathbf{b} added to every row by broadcasting. Every output feature reads every input feature, which is where the name "dense" comes from. A two-layer network puts a nonlinearity ϕ\phi between two of them, applied to every entry separately:

H=ϕ(XW1+b1),Y=HW2+b2.\mathbf{H} = \phi(\mathbf{X}\mathbf{W}_1 + \mathbf{b}_1), \qquad \mathbf{Y} = \mathbf{H}\mathbf{W}_2 + \mathbf{b}_2 .

A common choice is the ReLU, ϕ(z)=max⁡(0,z)\phi(z) = \max(0,\allowbreak z): it keeps positive numbers and turns negative ones into 0. Follow the shapes: X\mathbf{X} is (batch,din)(\text{batch},\allowbreak d_{\text{in}}) and W1\mathbf{W}_1 is (din,h)(d_{\text{in}},\allowbreak h), so the hidden layer H\mathbf{H} is (batch,h)(\text{batch},\allowbreak h); with W2\mathbf{W}_2 of shape (h,dout)(h,\allowbreak d_{\text{out}}) the output is (batch,dout)(\text{batch},\allowbreak d_{\text{out}}). The hidden width hh is a free choice; it is the inner size that has to match on both sides.

A worked example with one input, so every number is visible. Take a network with 2 inputs, 2 hidden units and 1 output:

W1=[12−11],b1=(0,1),W2=[21],b2=−1,\mathbf{W}_1 = \begin{bmatrix} 1 & 2 \\ -1 & 1 \end{bmatrix}, \quad \mathbf{b}_1 = (0, 1), \quad \mathbf{W}_2 = \begin{bmatrix} 2 \\ 1 \end{bmatrix}, \quad \mathbf{b}_2 = -1,

and the input row x=(0,2)\mathbf{x} = (0,\allowbreak 2).

  1. Layer 1 product. Entry jj of xW1\mathbf{x}\mathbf{W}_1 is x\mathbf{x} dotted with column jj of W1\mathbf{W}_1: ((0)(1)+(2)(−1), (0)(2)+(2)(1))=(−2,2)\big((0)(1) + (2)(-1),\allowbreak \ (0)(2) + (2)(1)\big) = (-2,\allowbreak 2).
  2. Add the bias. (−2+0, 2+1)=(−2,3)(-2 + 0,\allowbreak \ 2 + 1) = (-2,\allowbreak 3).
  3. ReLU. (max⁡(0,−2), max⁡(0,3))=(0,3)(\max(0,\allowbreak -2),\allowbreak \ \max(0,\allowbreak 3)) = (0,\allowbreak 3). The negative value became 0.
  4. Layer 2. (0)(2)+(3)(1)=3(0)(2) + (3)(1) = 3, plus b2=−1b_2 = -1, gives the output 22.

Why the nonlinearity? Leave out step 3 and see what happens. Layer 2 then acts on (−2,3)(-2,\allowbreak 3) directly: (−2)(2)+(3)(1)−1=−4+3−1=−2(-2)(2) + (3)(1) - 1 = -4 + 3 - 1 = -2. Now try a single dense layer with weight W1W2\mathbf{W}_1\mathbf{W}_2 and bias b1W2+b2\mathbf{b}_1\mathbf{W}_2 + \mathbf{b}_2 (the algebra below says where these come from). The weight is (2×2)(2×1)=2×1(2 \times 2)(2 \times 1) = 2 \times 1, with entries (1)(2)+(2)(1)=4(1)(2) + (2)(1) = 4 and (−1)(2)+(1)(1)=−1(-1)(2) + (1)(1) = -1. The bias is (0)(2)+(1)(1)−1=0(0)(2) + (1)(1) - 1 = 0. That one layer gives xW+b=(0)(4)+(2)(−1)+0=−2\mathbf{x}\mathbf{W} + b = (0)(4) + (2)(-1) + 0 = -2: exactly what the two layers gave without the ReLU. With the ReLU the answer was 2 instead, because the ReLU zeroed the negative hidden value.

Here is the same collapse in general, for any input:

(XW1+b1)W2+b2=X(W1W2)+(b1W2+b2).(\mathbf{X}\mathbf{W}_1 + \mathbf{b}_1)\mathbf{W}_2 + \mathbf{b}_2 = \mathbf{X}(\mathbf{W}_1\mathbf{W}_2) + (\mathbf{b}_1\mathbf{W}_2 + \mathbf{b}_2).

That is a single dense layer, with weight W1W2\mathbf{W}_1\mathbf{W}_2 and bias b1W2+b2\mathbf{b}_1\mathbf{W}_2 + \mathbf{b}_2. Stacking a hundred layers without nonlinearities still gives one layer, so depth buys nothing. Neurons and layers picks up exactly here: the nonlinearity is what lets each layer bend space so that the next one can do something new.

Go slower: Collapsing two layers, step by step

In the row convention, treat each bias as a row vector: b1\mathbf{b}_1 is 1×h1 \times h and b2\mathbf{b}_2 is 1×dout1 \times d_{\text{out}}, added to every row. Look at a single row of X\mathbf{X} and call it x\mathbf{x}, a 1×din1 \times d_{\text{in}} row, as in the worked example.

Step 1. The first layer gives the row xW1+b1\mathbf{x}\mathbf{W}_1 + \mathbf{b}_1, of shape 1×h1 \times h.

Step 2. Multiply by W2\mathbf{W}_2. Matrix multiplication distributes over addition, (P+Q)R=PR+QR(\mathbf{P} + \mathbf{Q})\mathbf{R} = \mathbf{P}\mathbf{R} + \mathbf{Q}\mathbf{R}, because every entry of the product is a sum that splits in two: ∑k(Pik+Qik)Rkj=∑kPikRkj+∑kQikRkj\sum_k (P_{ik} + Q_{ik})R_{kj} = \sum_k P_{ik}R_{kj} + \sum_k Q_{ik}R_{kj}. So (xW1+b1)W2=(xW1)W2+b1W2.(\mathbf{x}\mathbf{W}_1 + \mathbf{b}_1)\mathbf{W}_2 = (\mathbf{x}\mathbf{W}_1)\mathbf{W}_2 + \mathbf{b}_1\mathbf{W}_2 .

Step 3. Regroup with associativity, from Matrix multiplication: (xW1)W2=x(W1W2)(\mathbf{x}\mathbf{W}_1)\mathbf{W}_2 = \mathbf{x}(\mathbf{W}_1\mathbf{W}_2).

Step 4. Add the second bias. The row is x(W1W2)+(b1W2+b2)\mathbf{x}(\mathbf{W}_1\mathbf{W}_2) + (\mathbf{b}_1\mathbf{W}_2 + \mathbf{b}_2), which is one dense layer with W=W1W2\mathbf{W} = \mathbf{W}_1\mathbf{W}_2 and b=b1W2+b2\mathbf{b} = \mathbf{b}_1\mathbf{W}_2 + \mathbf{b}_2.

Step 5, shapes. W1W2\mathbf{W}_1\mathbf{W}_2 is (din×h)(h×dout)=din×dout(d_{\text{in}} \times h)(h \times d_{\text{out}}) = d_{\text{in}} \times d_{\text{out}}, and b1W2\mathbf{b}_1\mathbf{W}_2 is (1×h)(h×dout)=1×dout(1 \times h)(h \times d_{\text{out}}) = 1 \times d_{\text{out}}, the right shape for a bias. Every row gets the same treatment, so the whole batch collapses the same way.

Now run a small network both ways yourself, with the ReLU and without it, and collapse it.

On paperTwo layers, with and without the ReLU

A two-layer network in the code convention has W1=[2−111]\mathbf{W}_1 = \begin{bmatrix} 2 & -1 \\ 1 & 1 \end{bmatrix}, b1=(−1,0)\mathbf{b}_1 = (-1,\allowbreak 0), W2=[31]\mathbf{W}_2 = \begin{bmatrix} 3 \\ 1 \end{bmatrix} and b2=1b_2 = 1. The input is the row x=(1,−2)\mathbf{x} = (1,\allowbreak -2).

  1. Compute xW1+b1\mathbf{x}\mathbf{W}_1 + \mathbf{b}_1, apply the ReLU to each entry, then multiply by W2\mathbf{W}_2 and add b2b_2.
  2. Compute the output again with the ReLU left out.
  3. Collapse the network without the ReLU into one layer, W=W1W2\mathbf{W} = \mathbf{W}_1\mathbf{W}_2 and b=b1W2+b2b = \mathbf{b}_1\mathbf{W}_2 + b_2, and check that xW+b\mathbf{x}\mathbf{W} + b matches part 2.

To check, enter the output of part 1, then the output of part 2.

Work it on real paper: writing each step is the point. Then check your final answer here and compare your working with the walk-through.

Output with the ReLU (part 1), then without it (part 2)

One entry per box, top to bottom. 0.25, -2, 3/4 and sqrt(2) all work. Enter moves to the next empty box and checks once all are filled.

Counting parameters and work

The numbers a network stores decide how much memory it needs, and, as you will see, nearly the same count decides how much arithmetic it does. So it pays to count them.

A dense layer from dind_{\text{in}} to doutd_{\text{out}} has din×doutd_{\text{in}} \times d_{\text{out}} weights and doutd_{\text{out}} biases. Together these are its parameters, the values training adjusts. A forward pass over a batch costs batch×din×dout\text{batch} \times d_{\text{in}} \times d_{\text{out}} multiply-adds for the matrix product, by the mnpmnp rule from Matrix multiplication with m=batchm = \text{batch}, n=dinn = d_{\text{in}} and p=doutp = d_{\text{out}}; the bias adds only batch×dout\text{batch} \times d_{\text{out}} additions on top.

The shape widget from the last lesson counts both for you, per layer and in total. Here it holds a tiny network, 4 features in, a hidden layer of 3 and 2 outputs, run on a batch of 2 examples. The table under the formula writes each layer's count out, such as 4⋅3+3=154 \cdot 3 + 3 = 15 for the parameters of layer 1.

A worked example at a realistic size: a network for small digit images with 28×28=78428 \times 28 = 784 pixels, one hidden layer of 128 units and 10 outputs, one per digit.

  • Layer 1: 784×128=100,352784 \times 128 = 100{,}352 weights plus 128 biases, 100,480100{,}480 parameters.
  • Layer 2: 128×10=1,280128 \times 10 = 1{,}280 weights plus 10 biases, 1,2901{,}290 parameters.
  • Total: 100,480+1,290=101,770100{,}480 + 1{,}290 = 101{,}770 parameters.

The forward pass for one example costs 100,352+1,280=101,632100{,}352 + 1{,}280 = 101{,}632 multiply-adds, one per weight, and a batch of 64 costs 64×101,632=6,504,44864 \times 101{,}632 = 6{,}504{,}448.

Now count a slightly deeper network yourself.

Work it outCount the parameters

A network takes 32 input features, has two hidden layers of 100 units each, and outputs 10 scores. Every layer is dense with a bias. How many parameters does it have? Enter a whole number.

parameters

Type a number: 0.25, -2, 3/4 and sqrt(2) all work. Enter checks.

The dense layer in the next exercise checks its shapes before it computes anything, a habit that pays off for the rest of the course. In Backpropagation you will derive the gradient of a layer's weights entry by entry, and shapes will guide and check the result: the gradient of W must have the same shape as W, and only one arrangement of the available matrices produces it. A recurrent network (Sequences and state) adds a time axis to every array, and an LSTM (The forward and backward pass) stacks the weights of four gates into one matrix and splits the result apart again. Each of those steps is a shape to get right.

Now turn both sections of this lesson into code: a dense layer that checks its shapes, a two-layer forward pass, a parameter counter and the collapse.

Code itDense layers and a two-layer network

Build the pieces of a small network in the row-batch convention that code uses.

  • dense(X, W, b) returns X @ W + b for X of shape (batch, d_in), W of shape (d_in, d_out) and b of shape (d_out,). It raises ValueError if W does not fit X, or if b does not have exactly d_out entries (a one-entry bias would otherwise broadcast silently).
  • mlp_forward(X, params) applies dense, then relu, then dense again.
  • count_params(sizes) counts the weights and biases for layer widths such as [784, 128, 10].
  • collapse(W1, b1, W2, b2) returns the single layer (W, b) that equals two dense layers with no nonlinearity between them.
⌘+Enter runsPython sleeps until you run code
Write your code where the starter says raise NotImplementedError, then press Run tests. Each check says what it expects.

A 7B model, counted

The same arithmetic scales all the way up. Here is a published design, the first LLaMA model from Meta (2023) at its 7B size: width 4,096, 32 layers and a vocabulary of 32,000 tokens. Each layer has four 4096×40964096 \times 4096 attention matrices (Attention explains what they do) and an MLP made of three matrices with 4096×110084096 \times 11008 entries each, a gated variant of the two-matrix MLP above. There are no biases. Each layer also has two small vectors of normalization weights, 4,096 numbers each, and there is one more such vector after the last layer.

Count it one piece at a time. A 4096×40964096 \times 4096 matrix holds 40962=16,777,2164096^2 = 16{,}777{,}216 numbers and a 4096×110084096 \times 11008 matrix holds 45,088,76845{,}088{,}768, so:

PieceShapesParameters
Attention, per layer4×(4096×4096)4 \times (4096 \times 4096)67,108,86467{,}108{,}864
MLP, per layer3×(4096×11008)3 \times (4096 \times 11008)135,266,304135{,}266{,}304
All 32 layers32×202,375,16832 \times 202{,}375{,}1686,476,005,3766{,}476{,}005{,}376
Token embeddings and output layer2×(32000×4096)2 \times (32000 \times 4096)262,144,000262{,}144{,}000
Normalization weights65×409665 \times 4096266,240266{,}240
Total6,738,415,6166{,}738{,}415{,}616

The third row adds the first two, 67,108,864+135,266,304=202,375,16867{,}108{,}864 + 135{,}266{,}304 = 202{,}375{,}168 per layer, and multiplies by 32. The normalization row counts 2×32+1=652 \times 32 + 1 = 65 vectors. The token embedding is a table with one row of 4,096 numbers for each of the 32,000 tokens, and the output layer is a matrix that turns the final 4,096-number vector into one score per token. The total is about 6.7 billion parameters, called "7B". Almost all of them sit in weight matrices, and the MLPs alone hold 32×135,266,304≈4.332 \times 135{,}266{,}304 \approx 4.3 billion, nearly two thirds of the total.

The code below redoes the count, converts it to memory at 16 bits (2 bytes) per parameter, and counts the multiply-adds for one token. Expect a total of 6,738,415,616, about 13.5 GB, and about 6.61 billion multiply-adds.

⌘+Enter runs · edit freelyPython sleeps until you run code

To recap: a dense layer is a matrix product plus a bias; without a nonlinearity between them, layers collapse into one; a layer stores dindout+doutd_{\text{in}}d_{\text{out}} + d_{\text{out}} numbers whatever the batch; and a forward pass costs about one multiply-add per weight per example, which is why a parameter count tells you both memory and compute. Finish with mixed reps from this lesson and the last one: shapes, transposes, broadcasting and counts, with new numbers each time.

Practice setShape reps
3 correct in a row completes the set. A miss starts the count again.

In numpy, A has shape (10, 12) and b has shape (10,). What is the shape of A + b?

Next: the module checkpoint