ml.lab
Python sleeps until you run code
01 Vectors and the dot product

Lesson 2 of 4

Span, basis and coordinates

Which points can a few vectors reach, and in how many ways? Span, linear independence and bases answer those two questions, and they show that the coordinates of a vector depend on the basis you describe it in.

About 40 minutes
By the end you can
  • Solve for the coefficients that reach any target from two vectors in the plane, and recognize when no single answer exists.

  • Decide whether a set of vectors in the plane spans the whole plane, only a line, or only the origin.

  • Decide whether a set of vectors is linearly independent, with either of the two equivalent tests.

  • Find the coordinates of a vector in a basis, and explain why they depend on the basis.

In Arrows and lists you reached the target (5,5)(5,\allowbreak 5) with the moves a=(2,1)\mathbf{a} = (2,\allowbreak 1) and b=(1,3)\mathbf{b} = (1,\allowbreak 3): two copies of a\mathbf{a} and one of b\mathbf{b}, so (5,5)=2a+b(5,\allowbreak 5) = 2\mathbf{a} + \mathbf{b}. The target (0,5)(0,\allowbreak 5) needed −a+2b-\mathbf{a} + 2\mathbf{b}. Two questions come next.

  • Can these two moves reach every point of the plane? What about other pairs of moves?
  • When a point can be reached, is the recipe the only one, or could two different recipes land on the same point?

This lesson answers both. The answers give four words the rest of the course uses constantly: span, linear independence, basis and coordinates. Along the way a number appears, a1b2−a2b1a_1 b_2 - a_2 b_1, that comes back in module 3 as the determinant.

One formula for every target

Solving two equations for each new target repeats the same work every time. If we do the algebra once, with letters in place of the numbers, we answer the first question for every target at once.

Write the two moves with letters, a=(a1,a2)\mathbf{a} = (a_1,\allowbreak a_2) and b=(b1,b2)\mathbf{b} = (b_1,\allowbreak b_2), and the target as t=(t1,t2)\mathbf{t} = (t_1,\allowbreak t_2). We want c1a+c2b=tc_1\mathbf{a} + c_2\mathbf{b} = \mathbf{t}. Scaling gives c1a=(a1c1, a2c1)c_1\mathbf{a} = (a_1 c_1,\allowbreak \ a_2 c_1) and c2b=(b1c2, b2c2)c_2\mathbf{b} = (b_1 c_2,\allowbreak \ b_2 c_2), and adding them entry by entry gives the left side, (a1c1+b1c2, a2c1+b2c2)(a_1 c_1 + b_1 c_2,\allowbreak \ a_2 c_1 + b_2 c_2). So there is one equation per entry, exactly as before:

a1c1+b1c2=t1(1)a2c1+b2c2=t2(2)a_1 c_1 + b_1 c_2 = t_1 \quad (1) \qquad\qquad a_2 c_1 + b_2 c_2 = t_2 \quad (2)

Eliminating one unknown at a time (the box below does it line by line) gives

c1=t1b2−t2b1a1b2−a2b1,c2=a1t2−a2t1a1b2−a2b1.c_1 = \frac{t_1 b_2 - t_2 b_1}{a_1 b_2 - a_2 b_1}, \qquad c_2 = \frac{a_1 t_2 - a_2 t_1}{a_1 b_2 - a_2 b_1}.

Try it on the last lesson's numbers, a=(2,1)\mathbf{a} = (2,\allowbreak 1) and b=(1,3)\mathbf{b} = (1,\allowbreak 3). The shared denominator is (2)(3)−(1)(1)=6−1=5(2)(3) - (1)(1) = 6 - 1 = 5. For the target (5,5)(5,\allowbreak 5):

c1=(5)(3)−(5)(1)5=15−55=2,c2=(2)(5)−(1)(5)5=10−55=1.c_1 = \frac{(5)(3) - (5)(1)}{5} = \frac{15 - 5}{5} = 2, \qquad c_2 = \frac{(2)(5) - (1)(5)}{5} = \frac{10 - 5}{5} = 1.

For the target (0,5)(0,\allowbreak 5):

c1=(0)(3)−(5)(1)5=−55=−1,c2=(2)(5)−(1)(0)5=105=2.c_1 = \frac{(0)(3) - (5)(1)}{5} = \frac{-5}{5} = -1, \qquad c_2 = \frac{(2)(5) - (1)(0)}{5} = \frac{10}{5} = 2.

Both match what you found by hand. Call the denominator D=a1b2−a2b1D = a_1 b_2 - a_2 b_1. Whenever D≠0D \ne 0, the formulas give an answer for every target, and the box shows it is the only answer. When D=0D = 0 the formulas would divide by zero, and that happens exactly when one vector is a multiple of the other.

Go slower: Solving for the coefficients in general

Start from equations (1) and (2).

Eliminate c2c_2. Multiply both sides of (1) by b2b_2, and both sides of (2) by b1b_1. Now both equations contain the same c2c_2 term, b1b2 c2b_1 b_2\,c_2: a1b2 c1+b1b2 c2=t1b2,a2b1 c1+b1b2 c2=t2b1.a_1 b_2\,c_1 + b_1 b_2\,c_2 = t_1 b_2,\allowbreak \qquad a_2 b_1\,c_1 + b_1 b_2\,c_2 = t_2 b_1. Subtract the second equation from the first, left side from left side and right side from right side: a1b2 c1−a2b1 c1+b1b2 c2−b1b2 c2=t1b2−t2b1.a_1 b_2\,c_1 - a_2 b_1\,c_1 + b_1 b_2\,c_2 - b_1 b_2\,c_2 = t_1 b_2 - t_2 b_1. The two c2c_2 terms cancel: a1b2 c1−a2b1 c1=t1b2−t2b1.a_1 b_2\,c_1 - a_2 b_1\,c_1 = t_1 b_2 - t_2 b_1. Factor c1c_1 out of the left side: (a1b2−a2b1) c1=t1b2−t2b1.(a_1 b_2 - a_2 b_1)\,c_1 = t_1 b_2 - t_2 b_1.

Eliminate c1c_1. The same moves with the roles swapped. Multiply (1) by a2a_2 and (2) by a1a_1, so both contain the c1c_1 term a1a2 c1a_1 a_2\,c_1: a1a2 c1+a2b1 c2=a2t1,a1a2 c1+a1b2 c2=a1t2.a_1 a_2\,c_1 + a_2 b_1\,c_2 = a_2 t_1,\allowbreak \qquad a_1 a_2\,c_1 + a_1 b_2\,c_2 = a_1 t_2. Subtract the first equation from the second: a1a2 c1−a1a2 c1+a1b2 c2−a2b1 c2=a1t2−a2t1.a_1 a_2\,c_1 - a_1 a_2\,c_1 + a_1 b_2\,c_2 - a_2 b_1\,c_2 = a_1 t_2 - a_2 t_1. The two c1c_1 terms cancel: a1b2 c2−a2b1 c2=a1t2−a2t1.a_1 b_2\,c_2 - a_2 b_1\,c_2 = a_1 t_2 - a_2 t_1. Factor c2c_2 out of the left side: (a1b2−a2b1) c2=a1t2−a2t1.(a_1 b_2 - a_2 b_1)\,c_2 = a_1 t_2 - a_2 t_1.

Divide. With D=a1b2−a2b1≠0D = a_1 b_2 - a_2 b_1 \ne 0, divide both results by DD to get the two formulas. So any solution must have these values: there is at most one.

Check that they solve (1). Put both formulas into the left side of (1), over the common denominator DD: a1c1+b1c2=a1 (t1b2−t2b1)+b1 (a1t2−a2t1)D.a_1 c_1 + b_1 c_2 = \frac{a_1\,(t_1 b_2 - t_2 b_1) + b_1\,(a_1 t_2 - a_2 t_1)}{D}. Multiply out the brackets: =a1b2t1−a1b1t2+a1b1t2−a2b1t1D.= \frac{a_1 b_2 t_1 - a_1 b_1 t_2 + a_1 b_1 t_2 - a_2 b_1 t_1}{D}. The two middle terms cancel. Factor t1t_1 out of the two that are left: =t1 (a1b2−a2b1)D=t1 DD=t1.= \frac{t_1\,(a_1 b_2 - a_2 b_1)}{D} = \frac{t_1\,D}{D} = t_1.

Check that they solve (2). The same steps. Put both formulas into the left side of (2): a2c1+b2c2=a2 (t1b2−t2b1)+b2 (a1t2−a2t1)D.a_2 c_1 + b_2 c_2 = \frac{a_2\,(t_1 b_2 - t_2 b_1) + b_2\,(a_1 t_2 - a_2 t_1)}{D}. Multiply out the brackets: =a2b2t1−a2b1t2+a1b2t2−a2b2t1D.= \frac{a_2 b_2 t_1 - a_2 b_1 t_2 + a_1 b_2 t_2 - a_2 b_2 t_1}{D}. This time the first and last terms cancel. Factor t2t_2 out of the two that are left: =t2 (a1b2−a2b1)D=t2 DD=t2.= \frac{t_2\,(a_1 b_2 - a_2 b_1)}{D} = \frac{t_2\,D}{D} = t_2. So every target t\mathbf{t} is reachable, in exactly one way.

When D=0D = 0. First, a multiple always gives D=0D = 0: if b=k a\mathbf{b} = k\,\mathbf{a} for some number kk, then D=a1(ka2)−a2(ka1)=ka1a2−ka1a2=0D = a_1(k a_2) - a_2(k a_1) = k a_1 a_2 - k a_1 a_2 = 0.

The converse holds too. Suppose D=0D = 0 and a1≠0a_1 \ne 0. Set k=b1/a1k = b_1 / a_1, so that b1=ka1b_1 = k a_1. D=0D = 0 says a1b2−a2b1=0a_1 b_2 - a_2 b_1 = 0; add a2b1a_2 b_1 to both sides to get a1b2=a2b1a_1 b_2 = a_2 b_1. Replace b1b_1 by ka1k a_1: a1b2=ka1a2a_1 b_2 = k a_1 a_2. Divide both sides by a1a_1: b2=ka2b_2 = k a_2. So b=(ka1,ka2)=k a\mathbf{b} = (k a_1,\allowbreak k a_2) = k\,\mathbf{a}. If a1=0a_1 = 0 but a2≠0a_2 \ne 0, then D=(0) b2−a2b1=−a2b1D = (0)\,b_2 - a_2 b_1 = -a_2 b_1, and since a2≠0a_2 \ne 0, D=0D = 0 forces b1=0b_1 = 0. Set k=b2/a2k = b_2 / a_2, so that b2=ka2b_2 = k a_2. Then b=(0,ka2)=k (0,a2)=k a\mathbf{b} = (0,\allowbreak k a_2) = k\,(0,\allowbreak a_2) = k\,\mathbf{a}. If a=0\mathbf{a} = \mathbf{0}, then a=0 b\mathbf{a} = 0\,\mathbf{b} is a multiple of b\mathbf{b}.

So D=0D = 0 exactly when one vector is a multiple of the other, which puts both on one line through the origin. Every combination then stays on that line: with b=k a\mathbf{b} = k\,\mathbf{a}, c1a+c2b=c1a+c2k a=(c1+c2k) ac_1\mathbf{a} + c_2\mathbf{b} = c_1\mathbf{a} + c_2 k\,\mathbf{a} = (c_1 + c_2 k)\,\mathbf{a} (and the same with the roles of a\mathbf{a} and b\mathbf{b} swapped). A target off the line cannot be reached at all. A target on the line, say s as\,\mathbf{a}, is reached by every pair with c1+c2k=sc_1 + c_2 k = s, so there are infinitely many answers. Either way there is no single answer for a formula to give, which is why it would divide by zero.

The number D=a1b2−a2b1D = a_1 b_2 - a_2 b_1 comes back in module 3 as the determinant, where it measures the signed area of the parallelogram built on a\mathbf{a} and b\mathbf{b}. Zero area means the parallelogram has collapsed onto a line: the same fact, seen from another side.

Now try the whole method on new vectors: first by elimination, then with the formula, and then on a pair whose DD is zero.

On paperFind the coefficients

Let a=(1,2)\mathbf{a} = (1,\allowbreak 2) and b=(3,−1)\mathbf{b} = (3,\allowbreak -1).

  1. Find c1c_1 and c2c_2 with c1a+c2b=(−1,5)c_1\mathbf{a} + c_2\mathbf{b} = (-1,\allowbreak 5). Write one equation per entry and solve by elimination.
  2. Check your answer with the lesson's formulas c1=t1b2−t2b1a1b2−a2b1c_1 = \frac{t_1 b_2 - t_2 b_1}{a_1 b_2 - a_2 b_1} and c2=a1t2−a2t1a1b2−a2b1c_2 = \frac{a_1 t_2 - a_2 t_1}{a_1 b_2 - a_2 b_1}, and by building the combination.
  3. Now try to write (3,1)(3,\allowbreak 1) as a combination of a=(1,2)\mathbf{a} = (1,\allowbreak 2) and d=(−2,−4)\mathbf{d} = (-2,\allowbreak -4). What goes wrong, and what does the picture look like?

Enter c1c_1 and c2c_2 from part 1 in the check box, c1c_1 on top.

Work it on real paper: writing each step is the point. Then check your final answer here and compare your working with the walk-through.

c1 and c2 from part 1

One entry per box, top to bottom. 0.25, -2, 3/4 and sqrt(2) all work. Enter moves to the next empty box and checks once all are filled.

Span: everything you can reach

The first question asked which targets a set of moves can reach. The collection of all of them gets a name, because we will ask about it constantly.

The span of some vectors is the set of all their linear combinations. Setting every coefficient to zero lands on the origin, so a span always contains the origin. That is why spans are lines and planes through the origin, never shifted ones.

The formula settles what the span of two vectors looks like in the plane:

  • If they do not lie on a common line through the origin (D≠0D \ne 0), they span the whole plane.
  • If they lie on a common line (one is a multiple of the other) and are not both zero, they span only that line.
  • If both are the zero vector, they span only the origin.

The figure shows the span directly. It draws a\mathbf{a} and b\mathbf{b}, shades everything their combinations reach, and names it in the corner: "span: plane", "span: line" or "span: origin". While the span is the plane, faint grid lines mark whole steps of a\mathbf{a} and b\mathbf{b}. The white point p\mathbf{p} is a target you can drag. When it is reachable, dashed arrows show a recipe (scaled copies of the vectors, walked tip to tail from the origin to p\mathbf{p}). When it is not, a dashed segment shows the gap.

The idea is the same in more dimensions, with more room. In three dimensions, two vectors that do not lie on a common line span a plane through the origin, not all of space. You need a third vector that points out of that plane to reach everything.

Two arrows from one origin lying in a flat translucent sheet through that origin, and a third arrow leaving the sheetIn three dimensions, two arrows that do not share a line span a flat sheet through the origin. A third arrow that leaves the sheet adds the missing direction.

Spanning is all or nothing, but some pairs need very large coefficients to reach ordinary points. The next question applies the test D≠0D \ne 0 to four pairs, including one that is nearly parallel.

Quick checkWhich pair fills the plane

Which pair of vectors spans the whole plane?

Choose one answer, then check.

To recap: the span of some vectors is every point their combinations reach, and it always contains the origin. Two vectors in the plane span all of it exactly when D=a1b2−a2b1≠0D = a_1 b_2 - a_2 b_1 \ne 0, that is, when they do not share a line through the origin.

Linear independence: no wasted moves

The second question asked whether a point can have two different recipes. That happens exactly when one of the moves is redundant, so this section is about redundancy.

Take (1,0,1)(1,\allowbreak 0,\allowbreak 1), (0,1,1)(0,\allowbreak 1,\allowbreak 1) and (1,1,2)(1,\allowbreak 1,\allowbreak 2). None of them is a multiple of another, yet the third is the sum of the first two:

(1,0,1)+(0,1,1)=(1+0, 0+1, 1+1)=(1,1,2).(1, 0, 1) + (0, 1, 1) = (1 + 0,\ 0 + 1,\ 1 + 1) = (1, 1, 2).

The third vector adds no new direction: anything you can build with all three, you can build with the first two alone. In space, the first two span a plane through the origin, and the third lies inside that plane instead of leaving it the way the third arrow in the picture above does. And points now have more than one recipe. The point (1,1,2)(1,\allowbreak 1,\allowbreak 2) itself is 0 (1,0,1)+0 (0,1,1)+1 (1,1,2)0\,(1,\allowbreak 0,\allowbreak 1) + 0\,(0,\allowbreak 1,\allowbreak 1) + 1\,(1,\allowbreak 1,\allowbreak 2), and it is also 1 (1,0,1)+1 (0,1,1)+0 (1,1,2)1\,(1,\allowbreak 0,\allowbreak 1) + 1\,(0,\allowbreak 1,\allowbreak 1) + 0\,(1,\allowbreak 1,\allowbreak 2).

Vectors are linearly independent when this cannot happen: none of them is a linear combination of the others. In plain words, each one points somewhere the others cannot reach. Vectors that are not independent are dependent.

An equivalent test is often easier to check: the only way to combine them into the zero vector is with every coefficient equal to zero. For the three vectors above,

1⋅(1,0,1)+1⋅(0,1,1)−1⋅(1,1,2)=(1+0−1, 0+1−1, 1+1−2)=(0,0,0),1\cdot(1, 0, 1) + 1\cdot(0, 1, 1) - 1\cdot(1, 1, 2) = (1 + 0 - 1,\ 0 + 1 - 1,\ 1 + 1 - 2) = (0, 0, 0),

a combination that gives zero with coefficients that are not all zero, so the three are dependent.

Why do the two tests agree? Take three vectors (the same two moves work for any number).

  • Suppose some combination c1v1+c2v2+c3v3=0c_1\mathbf{v}_1 + c_2\mathbf{v}_2 + c_3\mathbf{v}_3 = \mathbf{0} has a coefficient that is not zero, say c3≠0c_3 \ne 0. Move the last term to the other side: c1v1+c2v2=−c3v3c_1\mathbf{v}_1 + c_2\mathbf{v}_2 = -c_3\mathbf{v}_3. Divide both sides by −c3-c_3: v3=−c1c3v1−c2c3v2\mathbf{v}_3 = -\frac{c_1}{c_3}\mathbf{v}_1 - \frac{c_2}{c_3}\mathbf{v}_2. So v3\mathbf{v}_3 is a combination of the others.
  • In the other direction, suppose v3=x v1+y v2\mathbf{v}_3 = x\,\mathbf{v}_1 + y\,\mathbf{v}_2. Subtract v3\mathbf{v}_3 from both sides: x v1+y v2−v3=0x\,\mathbf{v}_1 + y\,\mathbf{v}_2 - \mathbf{v}_3 = \mathbf{0}, a combination that gives zero with the coefficient −1-1, which is not zero.

Two consequences are worth keeping:

  • In the plane, two vectors are independent exactly when they do not lie on a common line: the D≠0D \ne 0 condition from before.
  • There can be at most nn independent vectors in Rn\mathbb{R}^n. For example, three vectors in the plane are always dependent. If two of them are independent, those two span the plane, so the third is a combination of them. If no two are independent, every pair lies on a common line, so all three lie on one line and one of them is a multiple of another. Module 3 returns to this count as the rank of a matrix.

Independence is not a pairwise test, and the next question is built around that trap. Look for a combination before you look for parallel pairs.

Quick checkIndependent or not

Are the vectors (1,0,2)(1,\allowbreak 0,\allowbreak 2), (0,1,−1)(0,\allowbreak 1,\allowbreak -1) and (2,3,1)(2,\allowbreak 3,\allowbreak 1) linearly independent?

Choose one answer, then check.

Basis and coordinates

Put the two answers together. Enough moves to reach everything (they span), with none wasted (they are independent), gives every point exactly one recipe. That recipe works like an address, and this section is about those addresses.

A basis is a set of vectors that is independent and spans the whole space. The unit steps from the last lesson,

e1=(1,0),e2=(0,1),\mathbf{e}_1 = (1, 0), \qquad \mathbf{e}_2 = (0, 1),

are the standard basis of the plane. They span it, since any (x,y)=x e1+y e2(x,\allowbreak y) = x\,\mathbf{e}_1 + y\,\mathbf{e}_2, and they are independent, since D=(1)(1)−(0)(0)=1≠0D = (1)(1) - (0)(0) = 1 \ne 0. With them, (3,2)=3 e1+2 e2(3,\allowbreak 2) = 3\,\mathbf{e}_1 + 2\,\mathbf{e}_2: the entries of a vector are its coefficients in the standard basis. That is so familiar that it is easy to forget it is a choice. (In Rn\mathbb{R}^n the standard basis has nn vectors, each with a single 11 and zeros elsewhere.)

Any two non-parallel vectors in the plane are also a basis, since D≠0D \ne 0 means they span and are independent. Each basis lays its own grid over the plane. The figure compares two descriptions of one point. It is the combination figure from the last lesson without a target: c1c_1 copies of a\mathbf{a}, then c2c_2 copies of b\mathbf{b}, with the result in white. A switch draws the grid built from a\mathbf{a} and b\mathbf{b}: lines through every whole step of a\mathbf{a} and of b\mathbf{b}, the way the ordinary grid has lines through every whole step of e1\mathbf{e}_1 and e2\mathbf{e}_2.

The coordinates of a vector in a basis are the coefficients that build it from that basis. In the standard basis, the arrow from the origin to (5,5)(5,\allowbreak 5) has coordinates (5,5)(5,\allowbreak 5). In the basis a=(2,1)\mathbf{a} = (2,\allowbreak 1), b=(1,3)\mathbf{b} = (1,\allowbreak 3), the same arrow is 2a+1b2\mathbf{a} + 1\mathbf{b}, so its coordinates there are (2,1)(2,\allowbreak 1). Same arrow, different numbers. Coordinates describe a vector relative to a basis. Change the basis and the numbers change while the arrow stays put.

A basis is useful because coordinates in it are never ambiguous: each vector has exactly one set.

Go slower: Why coordinates in a basis are unique

Suppose a vector v\mathbf{v} could be built two ways from the basis b1,…,bn\mathbf{b}_1,\allowbreak \ldots,\allowbreak \mathbf{b}_n: v=c1b1+⋯+cnbnandv=d1b1+⋯+dnbn.\mathbf{v} = c_1\mathbf{b}_1 + \cdots + c_n\mathbf{b}_n \qquad\text{and}\qquad \mathbf{v} = d_1\mathbf{b}_1 + \cdots + d_n\mathbf{b}_n. Subtract the second equation from the first. The left side becomes v−v=0\mathbf{v} - \mathbf{v} = \mathbf{0}. On the right, the rules of arithmetic from the last lesson let us pair up the terms for each basis vector, cibi−dibi=(ci−di) bic_i\mathbf{b}_i - d_i\mathbf{b}_i = (c_i - d_i)\,\mathbf{b}_i: 0=(c1−d1) b1+⋯+(cn−dn) bn.\mathbf{0} = (c_1 - d_1)\,\mathbf{b}_1 + \cdots + (c_n - d_n)\,\mathbf{b}_n. The basis is independent, so the only combination that gives 0\mathbf{0} has every coefficient equal to zero: c1−d1=0c_1 - d_1 = 0, and so on up to cn−dn=0c_n - d_n = 0. So ci=dic_i = d_i for every ii, and the two ways were the same way.

Spanning guarantees that coordinates exist. Independence guarantees there is only one set of them.

The next exercise goes the other way round: you are given the coordinates in a basis and asked for the ordinary entries.

Work it outFrom basis coordinates to ordinary ones

A vector has coordinates (3,2)(3,\allowbreak 2) in the basis p=(1,1)\mathbf{p} = (1,\allowbreak 1), q=(−1,2)\mathbf{q} = (-1,\allowbreak 2). Find its ordinary entries, that is, its coordinates in the standard basis e1,e2\mathbf{e}_1,\allowbreak \mathbf{e}_2, and enter them top to bottom.

standard coordinates

One entry per box, top to bottom. 0.25, -2, 3/4 and sqrt(2) all work. Enter moves to the next empty box and checks once all are filled.

This idea pays off in module 4. There the key move is to pick a basis in which a matrix only stretches each basis vector. In those coordinates, applying the matrix a hundred times is as easy as raising a few numbers to the hundredth power. The same calculation explains why gradients can vanish or explode as they flow back through a recurrent network in module 7.

Next: The dot product