Back to the on-screen lesson ·
Shape, centre and spread; the mean and the median, the standard deviation and the interquartile range, the five-number summary and the outlier fences.
Paper packet. Every task here also exists on screen, where it is checked automatically; answers written on paper are not assessed by Nydus. When you are back at a device, enter your answers there.
By the end of this lesson you can take a list of numbers and say what it looks like: its shape, where it sits, how far it reaches, and whether anything in it stands apart. You can compute a mean, a median, a standard deviation, the quartiles and the interquartile range, and — the part that matters more than any of the arithmetic — you can say which of them to quote, because a single extreme value moves two of these measures a long way and the other two hardly at all.
You can find a mean, read a bar chart and put numbers in order. What is new is that a list of numbers has a shape, and that describing it well means saying three things rather than one — and choosing measures that survive an extreme value rather than being dragged by it.
Distribution: the pattern of values a variable takes — which values, and how often.
Centre: where the data sits, measured by the mean $\bar{x}$ or the median.
Spread: how far it reaches, measured by the standard deviation $s$ or the interquartile range $IQR = Q_3 - Q_1$.
Quartiles: $Q_1$ and $Q_3$, the values a quarter and three quarters of the way through the sorted data.
Resistant: unmoved by an extreme value. The median and the IQR are resistant; the mean and the standard deviation are not.
Outlier: a value far enough from the rest to be worth looking at — by convention, beyond $Q_1 - 1.5\,IQR$ or $Q_3 + 1.5\,IQR$.
A distribution is the pattern of values a variable takes, and describing one means saying three things, in context. Shape: symmetric, or skewed left or right, with one peak or several, and any values standing well apart. Centre: the mean $\bar{x}$, which is the total divided by the count, or the median, which is the middle of the sorted values. Spread: the standard deviation $s$, which is the square root of the average squared distance from the mean, or the interquartile range $IQR = Q_3 - Q_1$, the width of the middle half.
The choice between the two pairs is the whole of the practical skill. The mean and the standard deviation use the size of every value, so one extreme reading drags them; the median and the IQR use only position, so one extreme reading moves them barely at all. Use the mean and $s$ when the shape is roughly symmetric; use the median and the IQR when it is skewed or has outliers. And a value is flagged as an outlier when it falls below $Q_1 - 1.5\,IQR$ or above $Q_3 + 1.5\,IQR$ — the fences that a boxplot's whiskers stop at.
Another way: picture
Picture a histogram with a long tail stretching to the right: most of the bars crowd the left, a few short ones trail away. The median stays with the crowd; the mean slides out toward the tail, as though the tail were hanging off the end of a balanced ruler.
Another way: steps
The standard deviation looks like an odd recipe until you ask what else could work. Averaging the plain distances from the mean fails: they cancel to zero every time, by the definition of the mean. Averaging the absolute distances works, and is a real measure, but it has no algebra — you cannot combine two of them. Averaging the squares cancels nothing, and squares do combine: the variances of independent quantities add, which is the fact the rest of this course is built on. The price is that the average square is in squared units, so the final step takes the square root and brings the answer back to pounds, seconds or degrees.
| Data | Mean | Median | $s$ | $IQR$ |
|---|---|---|---|---|
| 7, 9, 10, 11, 13 | 10 | 10 | 2 | 2 |
| 7, 9, 10, 11, 63 | 20 | 10 | ~21 | 2 |
One value changed, and the two resistant measures did not notice.
"The median is the middle of the list." It is the middle of the sorted list. Reading the fourth number of an unsorted seven is the single most common error in this topic, and it looks right every time.
"An outlier is a mistake, so remove it." The rule flags a value for attention. Sometimes it is a broken sensor; sometimes it is the most important reading you have. Deleting data because it is inconvenient is how analyses go wrong.
"Variance and standard deviation are the same thing." The variance is in squared units — squared pounds, squared seconds — which mean nothing. The standard deviation is its square root, back in the units of the data. Stopping before the root gives a number roughly the square of the right one.
Already sorted, and there are seven values, so the median is the fourth: $78$.
The median is the middle of the sorted data.
$Q_1$ is the middle of $62, 70, 75$, so $Q_1 = 70$; $Q_3$ is the middle of $80, 84, 95$, so $Q_3 = 84$.
Each quartile is the median of a half.
$IQR = 84 - 70 = 14$, so the fences are $70 - 21 = 49$ and $84 + 21 = 105$: no outliers.
$1.5 \times 14 = 21$.
Roughly symmetric, centred at $78$, with a middle half $14$ marks wide and no outliers.
Shape, centre, spread, outliers — in context.
Incomes in thousands: $30, 32, 35, 38, 40, 200$. The total is $375$, so the mean is $375 \div 6 = 62.5$.
Add and divide.
Six values, so the median is the average of $35$ and $38$: $36.5$.
Even count: average the middle pair.
The mean is nearly double the median, and nobody in the list earns anything near $62.5$. Report the median.
The mean is not wrong; it is answering a different question.
Six values, so the median is the average of $9$ and $12$, which is $10.5$.
Even count.
$Q_1 = 8$ and $Q_3 = 15$, so $IQR = 7$ and the upper fence is $15 + 10.5$.
That fence is $25.5$, and $40$ is above it: an outlier, and the shape is skewed right, so the median and IQR are the measures to quote.
Find the mean of $23, 27, 11, 24, 15$.
answer
The values $11, 19, 2, 13, 5$ have mean $10$. Find their population standard deviation.
answer
For the data $21, 12, 29, 17, 33, 15, 23$, give $Q_1$, the median and $Q_3$.
Q1 = q1, median = me, Q3 = q3
A data set has $Q_1 = 10$ and $Q_3 = 22$. What is the interquartile range?
answer
With $Q_1 = 6$ and $Q_3 = 18$, above what value is a reading flagged as an outlier by the $1.5 \times IQR$ rule?
answer
A town's incomes are mostly near $18$ thousand, but a few reach $61$ thousand. Which is larger, the mean or the median?
Complete the five-number summary of $16, 7, 22, 12, 26, 10, 18$.
| Measure | Value |
|---|---|
| Minimum | |
| Q1 | |
| Median | |
| Q3 | |
| Maximum |
Five sensors read $20, 29, 11, 23, 17$ (mean $20$, median $20$). A sixth sensor fails and reports $80$. By how much does the **mean** move more than the **median** does?
answer
Lesson test: one question per skill, one attempt each, no hints. Your answers are checked when you submit.
Match each measure to the question it answers.
| Where is the centre? | How spread out is it? | Is one tail longer? | |
|---|---|---|---|
| Median | |||
| Interquartile range | |||
| Skewness |
You can describe a distribution by shape, centre and spread, and choose resistant measures when it is skewed. Without looking: which two of the four measures survive one wild value, and why does the standard deviation end with a square root?
9. Your turn: 5, 8, 9, 12, 15, 40, step 2
9. Your turn: 5, 8, 9, 12, 15, 40, step 3