Back to the on-screen lesson ·
The OLS slope as co-movement over spread, the intercept through the point of means, and the same slope from a correlation.
Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.
By the end of this lesson you will be able to compute the OLS slope and intercept by hand from data points, from summary sums or from a correlation, and check the fit.
From the last lesson you know the population regression line is the best linear approximation to the conditional expectation function. In statistics you drew a least-squares line through a scatterplot, probably with a calculator.
This lesson opens the calculator. Ordinary least squares, or OLS, is the recipe that turns a sample of points into estimates of the intercept and slope, and every quantity in it can be computed by hand from sums you already know how to take.
| Term | What it means |
|---|---|
| OLS | Ordinary least squares: the intercept and slope that minimize the sum of squared residuals. |
| Residual | $\hat u_i = y_i - \hat y_i$: the vertical gap between a point and the fitted line. |
| Fitted value | $\hat y_i = \hat\beta_0 + \hat\beta_1 x_i$: the line's prediction at $x_i$. |
| Deviation | A value minus its sample mean, such as $x_i - \bar x$. |
| $S_{xx}$ | $\sum (x_i - \bar x)^2$: the total squared spread of $x$ about its mean. |
| $S_{xy}$ | $\sum (x_i - \bar x)(y_i - \bar y)$: how $x$ and $y$ move together. |
| Estimator | A rule that turns a sample into a guess about a population quantity. |
OLS chooses the line $\hat y = \hat\beta_0 + \hat\beta_1 x$ that makes the sum of squared residuals $\sum (y_i - \hat\beta_0 - \hat\beta_1 x_i)^2$ as small as possible. Setting the derivatives with respect to both coefficients to zero gives two equations, the normal equations, and their solution is
$$\hat\beta_1 = \frac{\sum (x_i - \bar x)(y_i - \bar y)}{\sum (x_i - \bar x)^2} = \frac{S_{xy}}{S_{xx}}, \qquad \hat\beta_0 = \bar y - \hat\beta_1 \bar x.$$
The numerator measures how $x$ and $y$ move together: a point above average in both adds a positive product, a point above in one and below in the other adds a negative one. The denominator measures how much $x$ varies. Their ratio is how much $y$ moves per unit of $x$, on average, in the sample.
The intercept formula says the fitted line always passes through the point of means $(\bar x, \bar y)$. Dividing top and bottom of the slope by $n - 1$ turns them into the sample covariance and variance, which gives the equivalent form $\hat\beta_1 = r\,s_y / s_x$.
Another way: action
Plot five points and the point of means. Draw any line through the point of means and tilt it until the squared vertical gaps are smallest. That tilt is the OLS slope.
Another way: steps
Write the sum of squared residuals as a function of the two coefficients, $SSR(b_0, b_1) = \sum (y_i - b_0 - b_1 x_i)^2$. It is a bowl-shaped function, so its minimum is where both partial derivatives are zero.
The derivative with respect to $b_0$ is $-2\sum (y_i - b_0 - b_1 x_i)$. Setting it to zero says the residuals sum to zero, and dividing by $n$ gives $\bar y = b_0 + b_1 \bar x$: the line passes through the point of means.
The derivative with respect to $b_1$ is $-2\sum x_i (y_i - b_0 - b_1 x_i)$. Setting it to zero says the residuals are uncorrelated with $x$ in the sample. Substituting the first equation into the second and rearranging gives $b_1 = S_{xy}/S_{xx}$. These two properties — residuals sum to zero, residuals uncorrelated with $x$ — are the sample versions of the CEF error facts from the last lesson.
Squaring residuals has two advantages: positive and negative misses cannot cancel, and the minimization has a clean closed-form answer. It also ties OLS to the conditional mean, since the mean is the number that minimizes squared error.
The cost is sensitivity to outliers. A point far from the line contributes the square of its distance, so a single unusual observation, such as one firm with enormous advertising, can tilt the whole line toward itself. A careful analyst plots the data and asks whether any single point is driving the slope. When one is, the question is whether it is an error to fix or a real observation that tells you something important.
The slope is measured in units of $y$ per unit of $x$. If sales are in thousands of units and advertising in thousands of dollars, a slope of $2.5$ means that weeks with one thousand dollars more advertising had, on average, two thousand five hundred more units of sales.
The intercept is the fitted value at $x = 0$. Sometimes that is meaningful — sales with no advertising — and sometimes it is far outside the data, such as the predicted wage of someone with zero years of schooling. When no observation has $x$ near zero, the intercept is just the number that makes the line pass through the point of means, and it should not be interpreted on its own.
Most importantly, the slope describes the sample. Whether more advertising causes more sales depends on why advertising varied, which is the question of the first lesson.
The last lesson described the population regression line, with coefficients written $\beta_0$ and $\beta_1$ and no hats. Nobody observes the population. We observe a sample, and OLS applied to it gives estimates, written with hats, that change from one sample to the next.
That distinction is the reason the rest of the course needs statistics. Draw five different weeks and you will get a different slope. The question is how far a particular sample's slope is likely to be from the population slope, and under what conditions its average across many samples equals the population slope at all. The first of those is about standard errors, and the second is about bias.
OLS has a strong property here: if the population CEF is linear and the sample is random, the OLS slope is centered on the population slope. What it cannot do is turn a descriptive population slope into a causal one. Unbiased for the regression line and unbiased for the causal effect are different claims, and later lessons keep them carefully apart, because a regression can be a perfectly good estimate of the wrong thing.
Take five weeks with advertising $x = 0, 1, 2, 3, 4$ and sales $y = 4, 8, 7, 11, 15$.
Check it: at $x = 2$ the line gives $9$, the mean of $y$, as it must.
Run three checks after every fit.
The first check is the fastest and catches the most common intercept error: forgetting to multiply the slope by $\bar x$.
A classic dataset in introductory econometrics covers $420$ California school districts. Regressing average fifth-grade test scores on the student-teacher ratio gives a slope of about $-2.3$: districts with one more student per teacher score about $2.3$ points lower on average. With an average ratio near $20$ and average scores near $654$, the fitted line is roughly $\hat y = 699 - 2.3x$.
Read it carefully. Reducing class size by two students is associated with about $4.6$ points higher scores, which is roughly a fifth of a standard deviation of district scores. Whether a school board that hires more teachers would actually gain those points is a causal question. Districts with small classes also tend to be richer, with more parental support and fewer students still learning English, and those differences raise scores by themselves.
The simple slope is the starting point of a famous debate. Later lessons add the share of English learners and family income to the regression and watch the slope shrink, and a randomized experiment in Tennessee, Project STAR, measured the causal effect of smaller classes directly. Its answer, a few percentile points for students placed in classes of about fifteen instead of twenty-two, is smaller than the raw slope suggests but real.
Real-estate websites estimate a home's price from its features. The simplest version regresses sale price on floor area. Suppose a sample of recent sales in one Ohio suburb gives an average size of $2{,}000$ square feet, an average price of $300{,}000$ dollars, $S_{xx}$ of $40$ million and $S_{xy}$ of $6$ billion (square feet times dollars).
The slope is $6{,}000{,}000{,}000 / 40{,}000{,}000 = 150$ dollars per square foot, and the intercept is $300{,}000 - 150 \times 2{,}000 = 0$. A $2{,}400$ square-foot house is predicted to sell for $360{,}000$ dollars.
Here the question is prediction, not causation, so the regression is doing exactly what it is good at. An appraiser does not need to know whether adding a room would raise the price by the same amount; they need the best guess about what similar houses sell for, and the conditional mean supplies it.
A common mistake is to report the correlation as the slope. They share a sign but not a scale: the correlation has no units and lies between $-1$ and $1$, while the slope is in units of $y$ per unit of $x$ and can be any size.
A second mistake is dividing by $S_{yy}$ instead of $S_{xx}$. Regressing $y$ on $x$ divides by the spread of the regressor, $x$; swapping the roles gives the slope of a different regression, $x$ on $y$, which is not its reciprocal.
A third is interpreting the intercept when $x = 0$ is far outside the data. Then the intercept is an extrapolation with no direct meaning.
Write the slope formula.
$\hat\beta_1 = S_{xy}/S_{xx}$
From the normal equations.
Read the sums.
$S_{xx} = 40, \ S_{xy} = 60$
From a sample of weeks.
Divide the two sums.
$60/40 = 1.5$
Units of $y$ per unit of $x$.
Read the means.
$\bar x = 6, \ \bar y = 24$
For the intercept.
Compute the intercept.
$24 - 1.5 \times 6 = 15$
Through the point of means.
Read the correlation.
$r = 0.6$
Unit-free co-movement.
Read the standard deviations.
$s_y = 10, \ s_x = 4$
The data's own units.
Write the rescaling.
$\hat\beta_1 = r\,s_y/s_x$
Covariance over variance.
Form the ratio of spreads.
$10/4 = 2.5$
Units of $y$ per unit of $x$.
Multiply by the correlation.
$0.6 \times 2.5 = 1.5$
The slope.
Check the sign.
$r > 0 \Rightarrow \hat\beta_1 > 0$
Always the same sign.
List the data.
$x: 1,2,3,4,5; \ y: 10,12,12,15,16$
Five observations.
Compute the means.
$\bar x = 3, \ \bar y = 13$
Sum and divide by five.
Find the deviations.
$x: -2,-1,0,1,2; \ y: -3,-1,-1,2,3$
Subtract the means.
Sum the squared x deviations.
$S_{xx} = 4+1+0+1+4 = 10$
Spread in $x$.
Sum the cross-products.
$S_{xy} = 6+1+0+2+6 = 15$
Co-movement.
Divide for the slope.
$15/10 = 1.5$
Per unit of $x$.
Compute the intercept.
$13 - 1.5 \times 3 = 8.5$
So $\hat y = 8.5 + 1.5x$.
Divide the sums.
$50/20 = 2.5$
The slope.
Multiply by the mean of x.
$2.5 \times 4 = 10$
The slope's share of $\bar y$.
Subtract from the mean of y.
In a sample, $S_{xx} = \sum(x_i - \bar x)^2 = 50$ and $S_{xy} = \sum(x_i - \bar x)(y_i - \bar y) = 150$. What is the OLS slope?
Answer: slope
Complete the worked solution: a sample has $S_{xx} = 40$, $S_{xy} = 120$, $\bar x = 5$ and $\bar y = 28$. Find the slope, the slope's contribution at the mean of $x$, and the intercept.
Divide the two sums.
$\hat\beta_1 = 120 / 40 =$ a
Cross-products over squared deviations.
Multiply the slope by the mean of x.
$\hat\beta_1 \bar x =$ b
How much of $\bar y$ the slope accounts for.
Subtract from the mean of y.
$\hat\beta_0 = 28 - \hat\beta_1 \bar x =$ c
The line passes through the point of means.
An OLS regression of $y$ on $x$ has slope $4$. The sample means are $\bar x = 4$ and $\bar y = 35$. What is the OLS intercept?
Answer: intercept
Two variables have correlation $r = 0.4$. The standard deviation of $y$ is $20$ and of $x$ is $5$. What is the OLS slope of $y$ on $x$?
Answer: slope
For the five points $x = 0, 1, 2, 3, 4$ and $y = 10, 10, 14, 17, 19$, fill in $\bar x$, $\bar y$, $S_{xx}$, $S_{xy}$, the slope and the intercept.
| value | |
|---|---|
| mean of x | |
| mean of y | |
| S_xx | |
| S_xy | |
| slope | |
| intercept |
A firm records weekly advertising $x$ (thousands of dollars) and sales $y$ (thousands of units) for five weeks: $x = 0, 1, 2, 3, 4$ and $y = 10, 10, 14, 17, 19$, in that order. What is the OLS slope of sales on advertising?
Answer: slope
Across a sample of school districts, class size $x$ (students per teacher) and average test score $y$ give $S_{xx} = 50$ and $S_{xy} = -150$. What is the OLS slope of test scores on class size, in points per student?
Answer: points per student
Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.
For the five points $x = 2, 4, 6, 8, 10$ and $y = 18, 20, 25, 28, 29$, fill in $\bar x$, $\bar y$, $S_{xx}$, $S_{xy}$, the slope and the intercept.
| value | |
|---|---|
| mean of x | |
| mean of y | |
| S_xx | |
| S_xy | |
| slope | |
| intercept |
You can fit a regression line by hand. Explain to someone why the fitted line always passes through the point of means.
17. Your turn: S_xx = 20, S_xy = 50, mean of x = 4, mean of y = 17., step 3
$17 - 10 = 7$
The intercept.