Back to the on-screen lesson ·

Ordinary least squares with one regressor

The OLS slope as co-movement over spread, the intercept through the point of means, and the same slope from a correlation.

Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.

1. What you will learn

By the end of this lesson you will be able to compute the OLS slope and intercept by hand from data points, from summary sums or from a correlation, and check the fit.

2. What you already have

From the last lesson you know the population regression line is the best linear approximation to the conditional expectation function. In statistics you drew a least-squares line through a scatterplot, probably with a calculator.

This lesson opens the calculator. Ordinary least squares, or OLS, is the recipe that turns a sample of points into estimates of the intercept and slope, and every quantity in it can be computed by hand from sums you already know how to take.

3. Terms to use precisely

TermWhat it means
OLSOrdinary least squares: the intercept and slope that minimize the sum of squared residuals.
Residual$\hat u_i = y_i - \hat y_i$: the vertical gap between a point and the fitted line.
Fitted value$\hat y_i = \hat\beta_0 + \hat\beta_1 x_i$: the line's prediction at $x_i$.
DeviationA value minus its sample mean, such as $x_i - \bar x$.
$S_{xx}$$\sum (x_i - \bar x)^2$: the total squared spread of $x$ about its mean.
$S_{xy}$$\sum (x_i - \bar x)(y_i - \bar y)$: how $x$ and $y$ move together.
EstimatorA rule that turns a sample into a guess about a population quantity.

4. The slope is co-movement divided by spread

OLS chooses the line $\hat y = \hat\beta_0 + \hat\beta_1 x$ that makes the sum of squared residuals $\sum (y_i - \hat\beta_0 - \hat\beta_1 x_i)^2$ as small as possible. Setting the derivatives with respect to both coefficients to zero gives two equations, the normal equations, and their solution is

$$\hat\beta_1 = \frac{\sum (x_i - \bar x)(y_i - \bar y)}{\sum (x_i - \bar x)^2} = \frac{S_{xy}}{S_{xx}}, \qquad \hat\beta_0 = \bar y - \hat\beta_1 \bar x.$$

The numerator measures how $x$ and $y$ move together: a point above average in both adds a positive product, a point above in one and below in the other adds a negative one. The denominator measures how much $x$ varies. Their ratio is how much $y$ moves per unit of $x$, on average, in the sample.

The intercept formula says the fitted line always passes through the point of means $(\bar x, \bar y)$. Dividing top and bottom of the slope by $n - 1$ turns them into the sample covariance and variance, which gives the equivalent form $\hat\beta_1 = r\,s_y / s_x$.

Another way: action

Plot five points and the point of means. Draw any line through the point of means and tilt it until the squared vertical gaps are smallest. That tilt is the OLS slope.

Another way: steps

  1. Compute the means $\bar x$ and $\bar y$.
  2. Subtract the means to get each point's deviations.
  3. Sum the squared $x$ deviations ($S_{xx}$) and the cross-products ($S_{xy}$).
  4. Divide $S_{xy}$ by $S_{xx}$ for the slope.
  5. Use the means for the intercept: $\hat\beta_0 = \bar y - \hat\beta_1 \bar x$.

5. Why these formulas minimize the squared residuals

Write the sum of squared residuals as a function of the two coefficients, $SSR(b_0, b_1) = \sum (y_i - b_0 - b_1 x_i)^2$. It is a bowl-shaped function, so its minimum is where both partial derivatives are zero.

The derivative with respect to $b_0$ is $-2\sum (y_i - b_0 - b_1 x_i)$. Setting it to zero says the residuals sum to zero, and dividing by $n$ gives $\bar y = b_0 + b_1 \bar x$: the line passes through the point of means.

The derivative with respect to $b_1$ is $-2\sum x_i (y_i - b_0 - b_1 x_i)$. Setting it to zero says the residuals are uncorrelated with $x$ in the sample. Substituting the first equation into the second and rearranging gives $b_1 = S_{xy}/S_{xx}$. These two properties — residuals sum to zero, residuals uncorrelated with $x$ — are the sample versions of the CEF error facts from the last lesson.

6. Why squares, and what they cost

Squaring residuals has two advantages: positive and negative misses cannot cancel, and the minimization has a clean closed-form answer. It also ties OLS to the conditional mean, since the mean is the number that minimizes squared error.

The cost is sensitivity to outliers. A point far from the line contributes the square of its distance, so a single unusual observation, such as one firm with enormous advertising, can tilt the whole line toward itself. A careful analyst plots the data and asks whether any single point is driving the slope. When one is, the question is whether it is an error to fix or a real observation that tells you something important.

7. Units and interpretation

The slope is measured in units of $y$ per unit of $x$. If sales are in thousands of units and advertising in thousands of dollars, a slope of $2.5$ means that weeks with one thousand dollars more advertising had, on average, two thousand five hundred more units of sales.

The intercept is the fitted value at $x = 0$. Sometimes that is meaningful — sales with no advertising — and sometimes it is far outside the data, such as the predicted wage of someone with zero years of schooling. When no observation has $x$ near zero, the intercept is just the number that makes the line pass through the point of means, and it should not be interpreted on its own.

Most importantly, the slope describes the sample. Whether more advertising causes more sales depends on why advertising varied, which is the question of the first lesson.

8. Sample estimates and population targets

The last lesson described the population regression line, with coefficients written $\beta_0$ and $\beta_1$ and no hats. Nobody observes the population. We observe a sample, and OLS applied to it gives estimates, written with hats, that change from one sample to the next.

That distinction is the reason the rest of the course needs statistics. Draw five different weeks and you will get a different slope. The question is how far a particular sample's slope is likely to be from the population slope, and under what conditions its average across many samples equals the population slope at all. The first of those is about standard errors, and the second is about bias.

OLS has a strong property here: if the population CEF is linear and the sample is random, the OLS slope is centered on the population slope. What it cannot do is turn a descriptive population slope into a causal one. Unbiased for the regression line and unbiased for the causal effect are different claims, and later lessons keep them carefully apart, because a regression can be a perfectly good estimate of the wrong thing.

9. Working an OLS fit, step by step

Take five weeks with advertising $x = 0, 1, 2, 3, 4$ and sales $y = 4, 8, 7, 11, 15$.

  1. Means. $\bar x = 10/5 = 2$ and $\bar y = 45/5 = 9$.
  2. Deviations. In $x$: $-2, -1, 0, 1, 2$. In $y$: $-5, -1, -2, 2, 6$.
  3. Sums. $S_{xx} = 4 + 1 + 0 + 1 + 4 = 10$. $S_{xy} = 10 + 1 + 0 + 2 + 12 = 25$.
  4. Slope. $25/10 = 2.5$.
  5. Intercept. $9 - 2.5 \times 2 = 4$, so $\hat y = 4 + 2.5x$.

Check it: at $x = 2$ the line gives $9$, the mean of $y$, as it must.

10. How to check an OLS answer

Run three checks after every fit.

  1. Does the line pass through the means? Plug $\bar x$ into your fitted line; you must get $\bar y$ exactly.
  2. Does the slope's sign match the picture? If $y$ rises with $x$ in the data, $S_{xy}$ and the slope must be positive. The slope always has the same sign as the correlation.
  3. Do the residuals sum to zero? Compute each $y_i - \hat y_i$ and add them. Anything but zero, up to rounding, means an arithmetic slip.

The first check is the fastest and catches the most common intercept error: forgetting to multiply the slope by $\bar x$.

11. In the world: class size and test scores

A classic dataset in introductory econometrics covers $420$ California school districts. Regressing average fifth-grade test scores on the student-teacher ratio gives a slope of about $-2.3$: districts with one more student per teacher score about $2.3$ points lower on average. With an average ratio near $20$ and average scores near $654$, the fitted line is roughly $\hat y = 699 - 2.3x$.

Read it carefully. Reducing class size by two students is associated with about $4.6$ points higher scores, which is roughly a fifth of a standard deviation of district scores. Whether a school board that hires more teachers would actually gain those points is a causal question. Districts with small classes also tend to be richer, with more parental support and fewer students still learning English, and those differences raise scores by themselves.

The simple slope is the starting point of a famous debate. Later lessons add the share of English learners and family income to the regression and watch the slope shrink, and a randomized experiment in Tennessee, Project STAR, measured the causal effect of smaller classes directly. Its answer, a few percentile points for students placed in classes of about fifteen instead of twenty-two, is smaller than the raw slope suggests but real.

12. In the world: pricing a house

Real-estate websites estimate a home's price from its features. The simplest version regresses sale price on floor area. Suppose a sample of recent sales in one Ohio suburb gives an average size of $2{,}000$ square feet, an average price of $300{,}000$ dollars, $S_{xx}$ of $40$ million and $S_{xy}$ of $6$ billion (square feet times dollars).

The slope is $6{,}000{,}000{,}000 / 40{,}000{,}000 = 150$ dollars per square foot, and the intercept is $300{,}000 - 150 \times 2{,}000 = 0$. A $2{,}400$ square-foot house is predicted to sell for $360{,}000$ dollars.

Here the question is prediction, not causation, so the regression is doing exactly what it is good at. An appraiser does not need to know whether adding a room would raise the price by the same amount; they need the best guess about what similar houses sell for, and the conditional mean supplies it.

13. The slope is not the correlation

A common mistake is to report the correlation as the slope. They share a sign but not a scale: the correlation has no units and lies between $-1$ and $1$, while the slope is in units of $y$ per unit of $x$ and can be any size.

A second mistake is dividing by $S_{yy}$ instead of $S_{xx}$. Regressing $y$ on $x$ divides by the spread of the regressor, $x$; swapping the roles gives the slope of a different regression, $x$ on $y$, which is not its reciprocal.

A third is interpreting the intercept when $x = 0$ is far outside the data. Then the intercept is an extrapolation with no direct meaning.

14. Slope from summary sums

  1. Write the slope formula.

    $\hat\beta_1 = S_{xy}/S_{xx}$

    From the normal equations.

  2. Read the sums.

    $S_{xx} = 40, \ S_{xy} = 60$

    From a sample of weeks.

  3. Divide the two sums.

    $60/40 = 1.5$

    Units of $y$ per unit of $x$.

  4. Read the means.

    $\bar x = 6, \ \bar y = 24$

    For the intercept.

  5. Compute the intercept.

    $24 - 1.5 \times 6 = 15$

    Through the point of means.

15. Slope from a correlation

  1. Read the correlation.

    $r = 0.6$

    Unit-free co-movement.

  2. Read the standard deviations.

    $s_y = 10, \ s_x = 4$

    The data's own units.

  3. Write the rescaling.

    $\hat\beta_1 = r\,s_y/s_x$

    Covariance over variance.

  4. Form the ratio of spreads.

    $10/4 = 2.5$

    Units of $y$ per unit of $x$.

  5. Multiply by the correlation.

    $0.6 \times 2.5 = 1.5$

    The slope.

  6. Check the sign.

    $r > 0 \Rightarrow \hat\beta_1 > 0$

    Always the same sign.

16. A full fit from five points

  1. List the data.

    $x: 1,2,3,4,5; \ y: 10,12,12,15,16$

    Five observations.

  2. Compute the means.

    $\bar x = 3, \ \bar y = 13$

    Sum and divide by five.

  3. Find the deviations.

    $x: -2,-1,0,1,2; \ y: -3,-1,-1,2,3$

    Subtract the means.

  4. Sum the squared x deviations.

    $S_{xx} = 4+1+0+1+4 = 10$

    Spread in $x$.

  5. Sum the cross-products.

    $S_{xy} = 6+1+0+2+6 = 15$

    Co-movement.

  6. Divide for the slope.

    $15/10 = 1.5$

    Per unit of $x$.

  7. Compute the intercept.

    $13 - 1.5 \times 3 = 8.5$

    So $\hat y = 8.5 + 1.5x$.

17. Your turn: S_xx = 20, S_xy = 50, mean of x = 4, mean of y = 17.

  1. Divide the sums.

    $50/20 = 2.5$

    The slope.

  2. Multiply by the mean of x.

    $2.5 \times 4 = 10$

    The slope's share of $\bar y$.

  3. Your turn: work this step out. Its working is at the end of the packet.

    Subtract from the mean of y.

18. Guided practice

In a sample, $S_{xx} = \sum(x_i - \bar x)^2 = 50$ and $S_{xy} = \sum(x_i - \bar x)(y_i - \bar y) = 150$. What is the OLS slope?

Answer: slope

19. Guided practice

Complete the worked solution: a sample has $S_{xx} = 40$, $S_{xy} = 120$, $\bar x = 5$ and $\bar y = 28$. Find the slope, the slope's contribution at the mean of $x$, and the intercept.

  1. Divide the two sums.

    $\hat\beta_1 = 120 / 40 =$ a

    Cross-products over squared deviations.

  2. Multiply the slope by the mean of x.

    $\hat\beta_1 \bar x =$ b

    How much of $\bar y$ the slope accounts for.

  3. Subtract from the mean of y.

    $\hat\beta_0 = 28 - \hat\beta_1 \bar x =$ c

    The line passes through the point of means.

20. Guided practice

An OLS regression of $y$ on $x$ has slope $4$. The sample means are $\bar x = 4$ and $\bar y = 35$. What is the OLS intercept?

Answer: intercept

21. Practice

Two variables have correlation $r = 0.4$. The standard deviation of $y$ is $20$ and of $x$ is $5$. What is the OLS slope of $y$ on $x$?

Answer: slope

22. Practice

For the five points $x = 0, 1, 2, 3, 4$ and $y = 10, 10, 14, 17, 19$, fill in $\bar x$, $\bar y$, $S_{xx}$, $S_{xy}$, the slope and the intercept.

value
mean of x
mean of y
S_xx
S_xy
slope
intercept

23. Practice

A firm records weekly advertising $x$ (thousands of dollars) and sales $y$ (thousands of units) for five weeks: $x = 0, 1, 2, 3, 4$ and $y = 10, 10, 14, 17, 19$, in that order. What is the OLS slope of sales on advertising?

Answer: slope

24. Somewhere new

Across a sample of school districts, class size $x$ (students per teacher) and average test score $y$ give $S_{xx} = 50$ and $S_{xy} = -150$. What is the OLS slope of test scores on class size, in points per student?

Answer: points per student

25. Lesson test

Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.

26. Test question

For the five points $x = 2, 4, 6, 8, 10$ and $y = 18, 20, 25, 28, 29$, fill in $\bar x$, $\bar y$, $S_{xx}$, $S_{xy}$, the slope and the intercept.

value
mean of x
mean of y
S_xx
S_xy
slope
intercept

27. What you can do now

You can fit a regression line by hand. Explain to someone why the fitted line always passes through the point of means.

Working for the steps left to you

17. Your turn: S_xx = 20, S_xy = 50, mean of x = 4, mean of y = 17., step 3

$17 - 10 = 7$

The intercept.