Back to the on-screen lesson ·

Statistics: shape, spread and lines of fit

Describe the shape, centre and spread of one variable, and fit, read and judge a line through a scatter plot using its slope and residuals.

Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.

1. What you will learn

In this lesson you learn to say three things about a set of numbers — its shape, its centre and its spread — and why a long tail makes the mean and the median disagree. You then move to two variables: how a line of best fit summarises a scatter plot, what its slope means, what a residual measures, and how a residual plot tells you whether a line was the right model in the first place.

2. What you bring to this

You already find a mean by adding and dividing, and a median by putting the values in order and taking the middle one. You have seen histograms and box plots before. What is new is treating the shape of a distribution as information in its own right, and knowing which summary to trust when the shape is lopsided.

3. Words you will need

Distribution: the pattern of how values are spread across their range.

Skew: a long tail on one side. Named for the side the tail is on: a right-skewed distribution has its tail on the right.

Standard deviation: roughly, the typical distance of a value from the mean. Bigger means more spread out.

Interquartile range (IQR): the width of the middle half of the data — the width of the box in a box plot.

Outlier: a value far from the rest.

4. Centre, spread and shape

A distribution needs at least three things said about it. Shape: is it roughly symmetric with a single peak, skewed with a long tail on one side, bimodal with two peaks — usually two groups mixed together — or uniform? Centre: the mean adds every value and divides, so every value's size counts; the median only counts positions, so extreme values move it hardly at all. That is why a right tail pulls the mean above the median while the median stays put, and why the median is the honest summary of a skewed data set. Spread: the standard deviation measures the typical distance from the mean, and the interquartile range measures the width of the middle half. Two data sets can share a centre exactly and differ completely in spread — which is what a pair of box plots with the same median and very different box widths is showing you.

Another way: diagram

Three histograms in a row: a symmetric bell with mean and median marked together; a right-skewed one with the median left of the mean and an arrow showing the tail dragging the mean; and a bimodal one with two peaks.

Another way: story

Imagine the histogram as a plank of wood carrying the data as weights. The mean is where it balances, so one heavy weight far out tips it a long way; the median is the point with equal numbers either side, and does not care how far out that weight is.

5. Three things that trip people up

"Same mean, same data." Two classes can both average $70$ with one scoring $68$ to $72$ and the other $40$ to $100$. The centre and the spread are independent questions.

"Right-skewed means the peak is on the right." It means the tail is on the right, so the peak is on the left. House prices are the standard example.

"The mean is always the best summary." With a long tail or an outlier, the mean is dragged away from where most of the data actually is, and the median describes a typical value far better.

6. House prices have a long right tail: mean or median larger?

  1. A few very expensive houses sit far to the right of everything else.

    That is what a right tail is.

  2. The median counts positions, so those few houses move it by only a place or two.

  3. The mean adds their full size, so it is dragged up: the mean is larger.

7. Two classes average $70$; class A has standard deviation $5$, class B $15$

  1. Equal means say the two classes sit in the same place.

    Centre agrees.

  2. Standard deviation is typical distance from the mean, so class B's scores lie about three times as far out.

    Spread differs.

  3. Class B is the more spread out — a mix of very high and very low scores that happens to average the same.

8. Your turn: adding one very large value to a data set — what moves more?

  1. The median shifts by at most one position in the ordered list.

  2. Your turn: work this step out. Its working is at the end of the packet.

    The mean adds the whole of the new value before dividing, so the mean moves much more.

9. Guided practice

A histogram of $99$ house prices has a long tail stretching to the right. Which is larger, the mean or the median?

10. Guided practice

Two classes both have a mean score of $73$. Class A has a standard deviation of $8$ and class B one of $14$. Whose scores are more spread out?

11. Practice

Find the mean of $49$, $61$, $44$, $62$ and $49$.

answer

12. Practice

Two box plots both have their median at $61$, but one box is $12$ units wide and the other is $36$ units wide. What differs between the two data sets?

13. What you bring to this

You already write and read $y = mx + b$, and you already interpret a slope as a rate. Fitting a line to data uses exactly that, with one change of meaning: the line no longer passes through the points, it passes among them, and the distances left over are the interesting part.

14. Words you will need

Scatter plot: one dot per case, with two measurements as its coordinates.

Line of best fit: a line drawn to summarise the trend in a scatter plot.

Prediction: the $y$ the line gives for a chosen $x$.

Residual: actual $-$ predicted. Positive above the line, negative below.

Residual plot: the residuals plotted against $x$. A shapeless band means a line was a reasonable model; a pattern means it was not.

Extrapolation: using the line outside the range of $x$ it was fitted over.

15. A line among the points, and what it misses

A line of best fit summarises the trend of a scatter plot, and it is read like any other line: the slope is the change in $y$ associated with one more unit of $x$ — 'each extra hour of study goes with about $1.8$ more points' — and the intercept is the value the line gives at $x = 0$, which sometimes makes sense and sometimes does not. Substituting an $x$ gives a prediction. What the line misses at each point is the residual, actual minus predicted: positive above the line, negative below. Plotting the residuals against $x$ is the honest test of the model. If a line really was the right shape, what is left over should look like a shapeless band; a clear curve or a fan in the residual plot says the data has a structure the line cannot follow. And a fitted line speaks only for the range it was fitted over — pushing it far outside that range produces arithmetic, not information.

Another way: diagram

A scatter plot with a fitted line through the middle of the cloud and a short vertical segment drawn from each point to the line, labelled 'residuals'; one point above the line marked '$+4$' and one below marked '$-3$'.

Another way: story

The line is the story the data mostly tells; the residuals are what each case did that the story did not account for. Looking at the residuals is how you find out whether the story was the right one to tell.

16. Three things that trip people up

"The line goes through the points." It goes among them. A fitted line passes through almost none of its data exactly, and that is the point: the misses are the residuals.

"Residual is predicted minus actual." It is actual minus predicted. Get it backwards and every sign is wrong, so points above the line look as though they were below it.

"If the arithmetic works, the prediction is good." Substituting $x = 60$ into a line fitted over $x = 0$ to $10$ gives a number, and the number is meaningless. A model describes the range it was built from.

17. $y = 1.8x + 12$: predict the score for $10$ hours of study

  1. Substitute $x = 10$: $y = 1.8 \times 10 + 12$.

    The hours go in place of $x$.

  2. $18 + 12 = 30$, so the line predicts a score of $30$.

18. The model predicted $30$ and the student scored $34$

  1. Residual is actual minus predicted: $34 - 30$.

    That order, always.

  2. The residual is $+4$: the student did better than the line said, so the point sits above it.

    Positive means above.

19. Your turn: a point lies below the line of best fit. What sign is its residual?

  1. Below the line means the actual value is less than the prediction.

  2. Your turn: work this step out. Its working is at the end of the packet.

    Actual minus predicted is therefore negative.

20. Guided practice

A line of best fit is $y = 3x + 18$, where $x$ is hours of study and $y$ is a test score. What score does it predict for $20$ hours?

Answer:

21. Guided practice

A model predicts $70$ for a student, and the student actually scores $69$. What is the residual?

Answer:

22. Practice

In the fitted line $y = 3.2x + 32$, with $x$ hours of study and $y$ a test score, what does $3.2$ mean?

23. Practice

A residual plot for a fitted line over $77$ data points shows a clear curve rather than a shapeless band. What does that suggest?

24. Somewhere new

A line of best fit $y = 3.8x + 32$ was worked out from students who studied between $0$ and $10$ hours, where $y$ is a test score out of $100$. A newspaper uses it for someone who studied $60$ hours and reports a predicted score of $260$. What is wrong?

25. Lesson test

Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.

26. Test question

Four histograms, each of $352$ measurements, are described below. Match each description to the word for its shape.

bell-shapedright-skewedbimodaluniform
symmetric, with one central peak
a long thin tail stretching to the right
two clear peaks with a dip between them
every bar the same height

27. Test question

A fitted line predicts $69$ for a point whose actual value is $67$. Is the point above or below the line?

28. What you can do now

You can describe a distribution's shape, centre and spread and use a fitted line with its residuals. Explain why a right-skewed data set has a mean above its median, and what a curved residual plot is telling you.

Working for the steps left to you

8. Your turn: adding one very large value to a data set — what moves more?, step 2

19. Your turn: a point lies below the line of best fit. What sign is its residual?, step 2