Lesson 4 of 4
Layers and parameter counts
A dense layer is a matrix product plus a bias, and a network is a stack of them with nonlinearities in between. This lesson shows why the nonlinearities are needed, then counts parameters and multiply-adds from a small network up to a 7B model.
Write a dense layer and a two-layer network in the row-batch convention, and follow the shapes through them.
Show, by hand and in algebra, that two layers without a nonlinearity collapse into one.
Count the parameters and multiply-adds of a stack of dense layers.
Count the parameters of a published 7B model and say what the count implies for memory and compute.
Implement dense layers, a two-layer forward pass, a parameter counter and the collapse in Python.

The last lesson ended with one layer, Y = X @ W + b, run on a whole batch at once. A network stacks several of them. Two questions follow, and this lesson answers both.
First, what does a stack compute, and why does every network put a nonlinearity between its layers? Without one, a hundred layers turn out to be no better than one, and you will see exactly why. Second, how big is a network: how many numbers must it store, and how much arithmetic does one example cost? The same counting that sizes a small digit classifier also explains a phrase you hear every day, "a 7B model".
A layer is a matrix product plus a bias
Every network in the rest of the course is built from one piece, so it is worth writing it down carefully.
A dense layer (also called fully connected) in the code convention is
with of shape , of shape , and the bias added to every row by broadcasting. Every output feature reads every input feature, which is where the name "dense" comes from. A two-layer network puts a nonlinearity between two of them, applied to every entry separately:
A common choice is the ReLU, : it keeps positive numbers and turns negative ones into 0. Follow the shapes: is and is , so the hidden layer is ; with of shape the output is . The hidden width is a free choice; it is the inner size that has to match on both sides.
A worked example with one input, so every number is visible. Take a network with 2 inputs, 2 hidden units and 1 output:
and the input row .
- Layer 1 product. Entry of is dotted with column of : .
- Add the bias. .
- ReLU. . The negative value became 0.
- Layer 2. , plus , gives the output .
Why the nonlinearity? Leave out step 3 and see what happens. Layer 2 then acts on directly: . Now try a single dense layer with weight and bias (the algebra below says where these come from). The weight is , with entries and . The bias is . That one layer gives : exactly what the two layers gave without the ReLU. With the ReLU the answer was 2 instead, because the ReLU zeroed the negative hidden value.
Here is the same collapse in general, for any input:
That is a single dense layer, with weight and bias . Stacking a hundred layers without nonlinearities still gives one layer, so depth buys nothing. Neurons and layers picks up exactly here: the nonlinearity is what lets each layer bend space so that the next one can do something new.
Go slower: Collapsing two layers, step by step
In the row convention, treat each bias as a row vector: is and is , added to every row. Look at a single row of and call it , a row, as in the worked example.
Step 1. The first layer gives the row , of shape .
Step 2. Multiply by . Matrix multiplication distributes over addition, , because every entry of the product is a sum that splits in two: . So
Step 3. Regroup with associativity, from Matrix multiplication: .
Step 4. Add the second bias. The row is , which is one dense layer with and .
Step 5, shapes. is , and is , the right shape for a bias. Every row gets the same treatment, so the whole batch collapses the same way.
Now run a small network both ways yourself, with the ReLU and without it, and collapse it.
A two-layer network in the code convention has , , and . The input is the row .
- Compute , apply the ReLU to each entry, then multiply by and add .
- Compute the output again with the ReLU left out.
- Collapse the network without the ReLU into one layer, and , and check that matches part 2.
To check, enter the output of part 1, then the output of part 2.
Work it on real paper: writing each step is the point. Then check your final answer here and compare your working with the walk-through.
One entry per box, top to bottom. 0.25, -2, 3/4 and sqrt(2) all work. Enter moves to the next empty box and checks once all are filled.
Counting parameters and work
The numbers a network stores decide how much memory it needs, and, as you will see, nearly the same count decides how much arithmetic it does. So it pays to count them.
A dense layer from to has weights and biases. Together these are its parameters, the values training adjusts. A forward pass over a batch costs multiply-adds for the matrix product, by the rule from Matrix multiplication with , and ; the bias adds only additions on top.
The shape widget from the last lesson counts both for you, per layer and in total. Here it holds a tiny network, 4 features in, a hidden layer of 3 and 2 outputs, run on a batch of 2 examples. The table under the formula writes each layer's count out, such as for the parameters of layer 1.
A worked example at a realistic size: a network for small digit images with pixels, one hidden layer of 128 units and 10 outputs, one per digit.
- Layer 1: weights plus 128 biases, parameters.
- Layer 2: weights plus 10 biases, parameters.
- Total: parameters.
The forward pass for one example costs multiply-adds, one per weight, and a batch of 64 costs .
Now count a slightly deeper network yourself.
A network takes 32 input features, has two hidden layers of 100 units each, and outputs 10 scores. Every layer is dense with a bias. How many parameters does it have? Enter a whole number.
Type a number: 0.25, -2, 3/4 and sqrt(2) all work. Enter checks.
The dense layer in the next exercise checks its shapes before it computes anything, a habit that pays off for the rest of the course. In Backpropagation you will derive the gradient of a layer's weights entry by entry, and shapes will guide and check the result: the gradient of W must have the same shape as W, and only one arrangement of the available matrices produces it. A recurrent network (Sequences and state) adds a time axis to every array, and an LSTM (The forward and backward pass) stacks the weights of four gates into one matrix and splits the result apart again. Each of those steps is a shape to get right.
Now turn both sections of this lesson into code: a dense layer that checks its shapes, a two-layer forward pass, a parameter counter and the collapse.
Build the pieces of a small network in the row-batch convention that code uses.
dense(X, W, b)returnsX @ W + bforXof shape(batch, d_in),Wof shape(d_in, d_out)andbof shape(d_out,). It raisesValueErrorifWdoes not fitX, or ifbdoes not have exactlyd_outentries (a one-entry bias would otherwise broadcast silently).mlp_forward(X, params)appliesdense, thenrelu, thendenseagain.count_params(sizes)counts the weights and biases for layer widths such as[784, 128, 10].collapse(W1, b1, W2, b2)returns the single layer(W, b)that equals two dense layers with no nonlinearity between them.
raise NotImplementedError, then press Run tests. Each check says what it expects.A 7B model, counted
The same arithmetic scales all the way up. Here is a published design, the first LLaMA model from Meta (2023) at its 7B size: width 4,096, 32 layers and a vocabulary of 32,000 tokens. Each layer has four attention matrices (Attention explains what they do) and an MLP made of three matrices with entries each, a gated variant of the two-matrix MLP above. There are no biases. Each layer also has two small vectors of normalization weights, 4,096 numbers each, and there is one more such vector after the last layer.
Count it one piece at a time. A matrix holds numbers and a matrix holds , so:
| Piece | Shapes | Parameters |
|---|---|---|
| Attention, per layer | ||
| MLP, per layer | ||
| All 32 layers | ||
| Token embeddings and output layer | ||
| Normalization weights | ||
| Total |
The third row adds the first two, per layer, and multiplies by 32. The normalization row counts vectors. The token embedding is a table with one row of 4,096 numbers for each of the 32,000 tokens, and the output layer is a matrix that turns the final 4,096-number vector into one score per token. The total is about 6.7 billion parameters, called "7B". Almost all of them sit in weight matrices, and the MLPs alone hold billion, nearly two thirds of the total.
The code below redoes the count, converts it to memory at 16 bits (2 bytes) per parameter, and counts the multiply-adds for one token. Expect a total of 6,738,415,616, about 13.5 GB, and about 6.61 billion multiply-adds.
To recap: a dense layer is a matrix product plus a bias; without a nonlinearity between them, layers collapse into one; a layer stores numbers whatever the batch; and a forward pass costs about one multiply-add per weight per example, which is why a parameter count tells you both memory and compute. Finish with mixed reps from this lesson and the last one: shapes, transposes, broadcasting and counts, with new numbers each time.
In numpy, A has shape (10, 12) and b has shape (10,). What is the shape of A + b?