Back to the on-screen lesson ·
Extremes of a function whose point is not free but tied to a curve or surface: tangency read as parallel gradients, the system that follows, the comparison that picks the winner, and the multiplier as the price of the constraint.
Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.
By the end of this lesson you will be able to set up and solve a constrained extremum problem with the conditions $\nabla f = \lambda \nabla g$ and $g = k$, explain from the tangency picture why the two gradients must be parallel at an answer, compare the candidates to decide which is the maximum and which the minimum, and read the multiplier as the rate at which the best value changes when the constraint is loosened.
The gradient from lesson 18, and in particular the fact that it is perpendicular to a level curve; and the free extrema of lesson 19, where a candidate was a point with $\nabla f = \mathbf{0}$. Here the point is not free to go anywhere: it has to stay on a curve, and the condition for a best point changes accordingly — from the gradient vanishes to the gradient has nowhere left to push.
A constrained extremum is a largest or smallest value of $f$ among the points satisfying $g = k$. The curve or surface $g = k$ is the constraint, and $k$ is its level. The number $\lambda$ in $\nabla f = \lambda \nabla g$ is the Lagrange multiplier. Two curves are tangent at a point when they have the same tangent line there — equivalently, when their normals are parallel. A condition is necessary when every answer satisfies it and sufficient when everything satisfying it is an answer.
The picture first. Draw the level curves of $f$ and, over them, the constraint curve $g = k$. Walk along the constraint. As you cross level curve after level curve, $f$ is changing. You are at a best point exactly when you have stopped crossing them — when the constraint curve is tangent to a level curve of $f$.
Tangency of two curves means their normals are parallel. The gradient is normal to a level curve. So at a constrained extremum,
$$\nabla f = \lambda\,\nabla g \quad\text{for some number }\lambda, \qquad g = k.$$
That is the whole method, and the picture is the whole reason for it.
The same thing said with directions. The directions that keep you on the constraint are exactly those perpendicular to $\nabla g$. If $\nabla f$ had any component along such a direction, stepping that way would change $f$ while staying feasible, and the point was not optimal. So at an optimum $\nabla f$ is perpendicular to everything perpendicular to $\nabla g$, which makes it a multiple of $\nabla g$.
Solving. In two variables the conditions are three equations,
$$f_x = \lambda g_x, \qquad f_y = \lambda g_y, \qquad g(x, y) = k,$$
in three unknowns $x$, $y$, $\lambda$. A common first move is to divide the first by the second to eliminate $\lambda$ — legitimate only when the denominators are non-zero, which is worth saying out loud.
More variables, more constraints. For $f(x, y, z)$ on $g = k$ the conditions are four equations in four unknowns. With two constraints $g = k$ and $h = c$, the gradient of the objective must lie in the plane spanned by the two constraint gradients:
$$\nabla f = \lambda \nabla g + \mu \nabla h.$$
What the multiplier is worth. Let $v(k)$ be the best value as a function of the level. Then $\lambda = v'(k)$: the multiplier is the rate at which the answer improves per unit of loosened constraint. It is a price, in units of objective per unit of constraint, and it is often the most useful number the calculation produces.
Another way: picture
A contour map with a footpath drawn across it. Walking the path you cross contour after contour, gaining height; where the path runs along a contour instead of across it, you have stopped gaining, and that is the highest point of the walk. Tangency of path and contour is the condition, and parallel normals is that condition written in vectors.
Another way: steps
To extremise $f$ subject to $g = k$:
A constrained problem can sometimes be solved without multipliers: solve the constraint for one variable and substitute, turning it into a free one-variable problem. For $x + y = 10$ maximising $xy$, substitution gives $x(10 - x)$ in one line, and no multiplier is needed.
Three things make the multiplier method the better habit. First, most constraints cannot be solved for a variable in closed form — $x^3 + y^3 + xy = 7$ resists it — and the multiplier method never asks. Second, substitution destroys symmetry: a problem symmetric in $x$ and $y$ becomes an asymmetric one-variable problem, and the symmetric answer that was obvious from the picture has to be rediscovered by algebra. Third, and most practically, substitution throws away $\lambda$, and $\lambda$ is the answer to what is one more unit of the constraint worth, which is the question a real user of the calculation asks next.
There is a fourth, quieter reason. Solving the constraint for $y$ tacitly assumes it can be done on the whole curve, and at a point where the curve turns back on itself it cannot. The multiplier conditions are local and need no such assumption.
Reading the conditions as sufficient. They are necessary. They produce candidates, and comparing $f$ at the candidates is what decides. A point where the gradients are parallel can easily be neither a maximum nor a minimum.
Forgetting to add the constraint. Two equations in three unknowns have a whole family of solutions; the constraint is the third equation, not something already used up in forming $g$.
Dividing by a gradient component that might vanish. Eliminating $\lambda$ by division loses exactly the solutions where the denominator is zero, and those are often the interesting ones.
Assuming an answer exists. On an unbounded constraint — a straight line, a hyperbola — a maximum may simply not exist, and the conditions will still hand over a point. Checking that the constraint set is closed and bounded is what turns a candidate into an answer.
Treating $\lambda$ as a quantity of something. It is a rate: objective per unit of constraint, with the units to match.
The condition is $\nabla f = \lambda \nabla g$, and $\lambda$ is doing real work. It absorbs the difference in scale between the two gradients, which have no reason to be the same length and in general are not even measured in the same units: if $f$ is a profit in pounds and $g$ is a weight in kilograms, $\nabla f$ is in pounds per metre and $\nabla g$ in kilograms per metre.
That is also why $\lambda$ has a meaning rather than being scaffolding. Its units are the objective's over the constraint's — pounds per kilogram, in the example — and it is exactly the exchange rate between them. Writing the units next to the multiplier, every time, is what keeps it honest: a multiplier with no units is a number that is about the right size and will be compared against costs it has no business being compared against.
Maximise $A = xy$ subject to $2x + 2y = 40$, so $g = 2x + 2y$ and $k = 40$.
The constraint is already in the right form.
$\nabla A = \langle y, x \rangle$ and $\nabla g = \langle 2, 2 \rangle$, so $y = 2\lambda$ and $x = 2\lambda$, giving $x = y$.
Parallel gradients force the two sides equal.
The constraint then gives $x = y = 10$ and $A = 100$: the square. And $\lambda = 5$, so one extra metre of fencing would buy about five more square metres of area — the answer to the question a builder actually asks.
The optimum and its price come out of the same system.
Minimise $f = x^2 + y^2 + z^2$ subject to $x + 2y + 2z = 9$. Minimising the square of the distance avoids a square root and has the same minimiser.
Squaring is a legitimate simplification because it is increasing.
$\langle 2x, 2y, 2z \rangle = \lambda \langle 1, 2, 2 \rangle$ gives $x = \lambda/2$, $y = \lambda$, $z = \lambda$.
Three equations, one for each variable.
Substituting into the constraint: $\lambda/2 + 2\lambda + 2\lambda = 9$, so $\lambda = 2$ and the point is $(1, 2, 2)$, at distance $3$. Notice that the answer came out parallel to the plane's normal vector $\langle 1, 2, 2 \rangle$, which is the geometry of lesson 6 reappearing as the solution of an optimisation.
The shortest route to a plane is along its normal.
Extremise $f = x + y$ subject to $x^2 + 4y^2 = 8$. Then $\langle 1, 1 \rangle = \lambda \langle 2x, 8y \rangle$.
One equation per variable, plus the constraint to come.
So $2\lambda x = 1$ and $8\lambda y = 1$, giving $x = 4y$. The constraint becomes $16y^2 + 4y^2 = 8$, so $y^2 = 2/5$.
That gives two candidates, $y = \pm\sqrt{2/5}$ with $x = 4y$, and $f = 5y$ at each — so the maximum is $5\sqrt{2/5}$ and the minimum its negative. Two candidates, not one: the conditions could not distinguish them, and evaluating $f$ did it in a line.
Put the steps of the method of Lagrange multipliers in order, for extremising $f$ subject to a constraint with level $3$.
Number the steps in order (write the number in the box):
Match each piece of a constrained problem to what it is, for $f$ extremised subject to $g(x, y) = 9$.
| The direction the objective increases fastest in | A vector perpendicular to the constraint curve | The curve every candidate must lie on | The rate at which the best value changes as the level moves | |
|---|---|---|---|---|
| $\nabla f$ | ||||
| $\nabla g$ | ||||
| $g(x, y) = 9$ | ||||
| $\lambda$ |
Maximise $f(x, y) = 6x + 8y$ on the circle $x^2 + y^2 = 4$.
Answer:
Maximise $f(x, y) = xy$ subject to $x + y = 18$, with $x$ and $y$ positive.
Answer:
Put the steps of the method of Lagrange multipliers in order, for extremising $f$ subject to a constraint with level $7$.
Number the steps in order (write the number in the box):
On a closed constraint curve the conditions $\nabla f = \lambda \nabla g$ hold at two points: at $P$, $f = 5$; at $Q$, $f = 8$. What follows?
Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.
Match each piece of a constrained problem to what it is, for $f$ extremised subject to $g(x, y) = 2$.
| The direction the objective increases fastest in | A vector perpendicular to the constraint curve | The curve every candidate must lie on | The rate at which the best value changes as the level moves | |
|---|---|---|---|---|
| $\nabla f$ | ||||
| $\nabla g$ | ||||
| $g(x, y) = 2$ | ||||
| $\lambda$ |
You can turn a constrained problem into a system of equations, solve it for candidates, and compare their values to find the answer — and you can say what the multiplier prices. Next: adding up a function over a region, starting with the simplest region there is.
10. Your turn: extremes of a linear function on an ellipse, step 3