Which points can a few vectors reach, and in how many ways? Span, linear independence and bases answer those two questions, and they show that the coordinates of a vector depend on the basis you describe it in.
About 40 minutes
By the end you can
Solve for the coefficients that reach any target from two vectors in the plane, and recognize when no single answer exists.
Decide whether a set of vectors in the plane spans the whole plane, only a line, or only the origin.
Decide whether a set of vectors is linearly independent, with either of the two equivalent tests.
Find the coordinates of a vector in a basis, and explain why they depend on the basis.
In Arrows and lists you reached the target (5,5) with the moves a=(2,1) and b=(1,3): two copies of a and one of b, so (5,5)=2a+b. The target (0,5) needed −a+2b. Two questions come next.
Can these two moves reach every point of the plane? What about other pairs of moves?
When a point can be reached, is the recipe the only one, or could two different recipes land on the same point?
This lesson answers both. The answers give four words the rest of the course uses constantly: span, linear independence, basis and coordinates. Along the way a number appears, a1b2−a2b1, that comes back in module 3 as the determinant.
Solving two equations for each new target repeats the same work every time. If we do the algebra once, with letters in place of the numbers, we answer the first question for every target at once.
Write the two moves with letters, a=(a1,a2) and b=(b1,b2), and the target as t=(t1,t2). We want c1a+c2b=t. Scaling gives c1a=(a1c1,a2c1) and c2b=(b1c2,b2c2), and adding them entry by entry gives the left side, (a1c1+b1c2,a2c1+b2c2). So there is one equation per entry, exactly as before:
a1c1+b1c2=t1(1)a2c1+b2c2=t2(2)
Eliminating one unknown at a time (the box below does it line by line) gives
Both match what you found by hand. Call the denominator D=a1b2−a2b1. Whenever D=0, the formulas give an answer for every target, and the box shows it is the only answer. When D=0 the formulas would divide by zero, and that happens exactly when one vector is a multiple of the other.
Go slower: Solving for the coefficients in general
Start from equations (1) and (2).
Eliminate c2. Multiply both sides of (1) by b2, and both sides of (2) by b1. Now both equations contain the same c2 term, b1b2c2:
a1b2c1+b1b2c2=t1b2,a2b1c1+b1b2c2=t2b1.
Subtract the second equation from the first, left side from left side and right side from right side:
a1b2c1−a2b1c1+b1b2c2−b1b2c2=t1b2−t2b1.
The two c2 terms cancel:
a1b2c1−a2b1c1=t1b2−t2b1.
Factor c1 out of the left side:
(a1b2−a2b1)c1=t1b2−t2b1.
Eliminate c1. The same moves with the roles swapped. Multiply (1) by a2 and (2) by a1, so both contain the c1 term a1a2c1:
a1a2c1+a2b1c2=a2t1,a1a2c1+a1b2c2=a1t2.
Subtract the first equation from the second:
a1a2c1−a1a2c1+a1b2c2−a2b1c2=a1t2−a2t1.
The two c1 terms cancel:
a1b2c2−a2b1c2=a1t2−a2t1.
Factor c2 out of the left side:
(a1b2−a2b1)c2=a1t2−a2t1.
Divide. With D=a1b2−a2b1=0, divide both results by D to get the two formulas. So any solution must have these values: there is at most one.
Check that they solve (1). Put both formulas into the left side of (1), over the common denominator D:
a1c1+b1c2=Da1(t1b2−t2b1)+b1(a1t2−a2t1).
Multiply out the brackets:
=Da1b2t1−a1b1t2+a1b1t2−a2b1t1.
The two middle terms cancel. Factor t1 out of the two that are left:
=Dt1(a1b2−a2b1)=Dt1D=t1.
Check that they solve (2). The same steps. Put both formulas into the left side of (2):
a2c1+b2c2=Da2(t1b2−t2b1)+b2(a1t2−a2t1).
Multiply out the brackets:
=Da2b2t1−a2b1t2+a1b2t2−a2b2t1.
This time the first and last terms cancel. Factor t2 out of the two that are left:
=Dt2(a1b2−a2b1)=Dt2D=t2.
So every target t is reachable, in exactly one way.
When D=0. First, a multiple always gives D=0: if b=ka for some number k, then D=a1(ka2)−a2(ka1)=ka1a2−ka1a2=0.
The converse holds too. Suppose D=0 and a1=0. Set k=b1/a1, so that b1=ka1. D=0 says a1b2−a2b1=0; add a2b1 to both sides to get a1b2=a2b1. Replace b1 by ka1: a1b2=ka1a2. Divide both sides by a1: b2=ka2. So b=(ka1,ka2)=ka. If a1=0 but a2=0, then D=(0)b2−a2b1=−a2b1, and since a2=0, D=0 forces b1=0. Set k=b2/a2, so that b2=ka2. Then b=(0,ka2)=k(0,a2)=ka. If a=0, then a=0b is a multiple of b.
So D=0 exactly when one vector is a multiple of the other, which puts both on one line through the origin. Every combination then stays on that line: with b=ka, c1a+c2b=c1a+c2ka=(c1+c2k)a (and the same with the roles of a and b swapped). A target off the line cannot be reached at all. A target on the line, say sa, is reached by every pair with c1+c2k=s, so there are infinitely many answers. Either way there is no single answer for a formula to give, which is why it would divide by zero.
The number D=a1b2−a2b1 comes back in module 3 as the determinant, where it measures the signed area of the parallelogram built on a and b. Zero area means the parallelogram has collapsed onto a line: the same fact, seen from another side.
Now try the whole method on new vectors: first by elimination, then with the formula, and then on a pair whose D is zero.
On paperFind the coefficients
Let a=(1,2) and b=(3,−1).
Find c1 and c2 with c1a+c2b=(−1,5). Write one equation per entry and solve by elimination.
Check your answer with the lesson's formulas c1=a1b2−a2b1t1b2−t2b1 and c2=a1b2−a2b1a1t2−a2t1, and by building the combination.
Now try to write (3,1) as a combination of a=(1,2) and d=(−2,−4). What goes wrong, and what does the picture look like?
Enter c1 and c2 from part 1 in the check box, c1 on top.
Work it on real paper: writing each step is the point. Then check your final answer here and compare your working with the walk-through.
c1 and c2 from part 1
One entry per box, top to bottom. 0.25, -2, 3/4 and sqrt(2) all work. Enter moves to the next empty box and checks once all are filled.
The first question asked which targets a set of moves can reach. The collection of all of them gets a name, because we will ask about it constantly.
The span of some vectors is the set of all their linear combinations. Setting every coefficient to zero lands on the origin, so a span always contains the origin. That is why spans are lines and planes through the origin, never shifted ones.
The formula settles what the span of two vectors looks like in the plane:
If they do not lie on a common line through the origin (D=0), they span the whole plane.
If they lie on a common line (one is a multiple of the other) and are not both zero, they span only that line.
If both are the zero vector, they span only the origin.
The figure shows the span directly. It draws a and b, shades everything their combinations reach, and names it in the corner: "span: plane", "span: line" or "span: origin". While the span is the plane, faint grid lines mark whole steps of a and b. The white point p is a target you can drag. When it is reachable, dashed arrows show a recipe (scaled copies of the vectors, walked tip to tail from the origin to p). When it is not, a dashed segment shows the gap.
The idea is the same in more dimensions, with more room. In three dimensions, two vectors that do not lie on a common line span a plane through the origin, not all of space. You need a third vector that points out of that plane to reach everything.
In three dimensions, two arrows that do not share a line span a flat sheet through the origin. A third arrow that leaves the sheet adds the missing direction.
Spanning is all or nothing, but some pairs need very large coefficients to reach ordinary points. The next question applies the test D=0 to four pairs, including one that is nearly parallel.
Quick checkWhich pair fills the plane
Which pair of vectors spans the whole plane?
Choose one answer, then check.
To recap: the span of some vectors is every point their combinations reach, and it always contains the origin. Two vectors in the plane span all of it exactly when D=a1b2−a2b1=0, that is, when they do not share a line through the origin.
The second question asked whether a point can have two different recipes. That happens exactly when one of the moves is redundant, so this section is about redundancy.
Take (1,0,1), (0,1,1) and (1,1,2). None of them is a multiple of another, yet the third is the sum of the first two:
(1,0,1)+(0,1,1)=(1+0,0+1,1+1)=(1,1,2).
The third vector adds no new direction: anything you can build with all three, you can build with the first two alone. In space, the first two span a plane through the origin, and the third lies inside that plane instead of leaving it the way the third arrow in the picture above does. And points now have more than one recipe. The point (1,1,2) itself is 0(1,0,1)+0(0,1,1)+1(1,1,2), and it is also 1(1,0,1)+1(0,1,1)+0(1,1,2).
Vectors are linearly independent when this cannot happen: none of them is a linear combination of the others. In plain words, each one points somewhere the others cannot reach. Vectors that are not independent are dependent.
An equivalent test is often easier to check: the only way to combine them into the zero vector is with every coefficient equal to zero. For the three vectors above,
a combination that gives zero with coefficients that are not all zero, so the three are dependent.
Why do the two tests agree? Take three vectors (the same two moves work for any number).
Suppose some combination c1v1+c2v2+c3v3=0 has a coefficient that is not zero, say c3=0. Move the last term to the other side: c1v1+c2v2=−c3v3. Divide both sides by −c3: v3=−c3c1v1−c3c2v2. So v3 is a combination of the others.
In the other direction, suppose v3=xv1+yv2. Subtract v3 from both sides: xv1+yv2−v3=0, a combination that gives zero with the coefficient −1, which is not zero.
Two consequences are worth keeping:
In the plane, two vectors are independent exactly when they do not lie on a common line: the D=0 condition from before.
There can be at most n independent vectors in Rn. For example, three vectors in the plane are always dependent. If two of them are independent, those two span the plane, so the third is a combination of them. If no two are independent, every pair lies on a common line, so all three lie on one line and one of them is a multiple of another. Module 3 returns to this count as the rank of a matrix.
Independence is not a pairwise test, and the next question is built around that trap. Look for a combination before you look for parallel pairs.
Quick checkIndependent or not
Are the vectors (1,0,2), (0,1,−1) and (2,3,1) linearly independent?
Put the two answers together. Enough moves to reach everything (they span), with none wasted (they are independent), gives every point exactly one recipe. That recipe works like an address, and this section is about those addresses.
A basis is a set of vectors that is independent and spans the whole space. The unit steps from the last lesson,
e1=(1,0),e2=(0,1),
are the standard basis of the plane. They span it, since any (x,y)=xe1+ye2, and they are independent, since D=(1)(1)−(0)(0)=1=0. With them, (3,2)=3e1+2e2: the entries of a vector are its coefficients in the standard basis. That is so familiar that it is easy to forget it is a choice. (In Rn the standard basis has n vectors, each with a single 1 and zeros elsewhere.)
Any two non-parallel vectors in the plane are also a basis, since D=0 means they span and are independent. Each basis lays its own grid over the plane. The figure compares two descriptions of one point. It is the combination figure from the last lesson without a target: c1 copies of a, then c2 copies of b, with the result in white. A switch draws the grid built from a and b: lines through every whole step of a and of b, the way the ordinary grid has lines through every whole step of e1 and e2.
The coordinates of a vector in a basis are the coefficients that build it from that basis. In the standard basis, the arrow from the origin to (5,5) has coordinates (5,5). In the basis a=(2,1), b=(1,3), the same arrow is 2a+1b, so its coordinates there are (2,1). Same arrow, different numbers. Coordinates describe a vector relative to a basis. Change the basis and the numbers change while the arrow stays put.
A basis is useful because coordinates in it are never ambiguous: each vector has exactly one set.
Go slower: Why coordinates in a basis are unique
Suppose a vector v could be built two ways from the basis b1,…,bn:
v=c1b1+⋯+cnbnandv=d1b1+⋯+dnbn.
Subtract the second equation from the first. The left side becomes v−v=0. On the right, the rules of arithmetic from the last lesson let us pair up the terms for each basis vector, cibi−dibi=(ci−di)bi:
0=(c1−d1)b1+⋯+(cn−dn)bn.
The basis is independent, so the only combination that gives 0 has every coefficient equal to zero: c1−d1=0, and so on up to cn−dn=0. So ci=di for every i, and the two ways were the same way.
Spanning guarantees that coordinates exist. Independence guarantees there is only one set of them.
The next exercise goes the other way round: you are given the coordinates in a basis and asked for the ordinary entries.
Work it outFrom basis coordinates to ordinary ones
A vector has coordinates (3,2) in the basis p=(1,1), q=(−1,2). Find its ordinary entries, that is, its coordinates in the standard basis e1,e2, and enter them top to bottom.
standard coordinates
One entry per box, top to bottom. 0.25, -2, 3/4 and sqrt(2) all work. Enter moves to the next empty box and checks once all are filled.
This idea pays off in module 4. There the key move is to pick a basis in which a matrix only stretches each basis vector. In those coordinates, applying the matrix a hundred times is as easy as raising a few numbers to the hundredth power. The same calculation explains why gradients can vanish or explode as they flow back through a recurrent network in module 7.