Mean, median and mode
Three answers to one question: which single number best describes a whole data set. The mean shares the total out evenly, the median stands exactly in the middle, and the mode points at the most frequent value — and which one you pick can change the conclusion.
Before you start
This topic builds on earlier ideas. Before you start, it's worth working through the lessons below — they'll make everything click:
- DivisionDivision is the inverse of multiplication — divide the dividend by the divisor to get the quotient. Learn the names, the link to multiplication, division with a remainder and why you must never divide by zero.
- DecimalsA decimal writes part of a whole with a decimal point instead of a fraction bar. Learn place value, converting a fraction to a decimal, adding, multiplying and rounding.
All formulas
Arithmetic mean
the sum of all the data divided by how many there are
Total from the mean
this is how a missing value is recovered when the mean is known
Median — an odd sample
once sorted, it is exactly the middle number
Median — an even sample
the mean of the two middle values
Mode
the number that occurs most often — there may be none at all
Weighted mean
every value counts as many times as its weight says
A data set is usually a long list of numbers, and the question is always the same: in one number — what does it look like? Descriptive statistics gives three answers, and each measures something different. All three are worth knowing, because the gap between them is often the whole point of a problem.
The arithmetic mean
The mean is the sum of all the data divided by how many there are:
The symbol is read "x bar". The idea is one of sharing out: if everyone had the same, everyone would have exactly the mean.
A second, equally useful formula follows straight from the first — multiply both sides by :
Knowing the mean means knowing the total. That is all it takes to recover a missing number.
A frequency diagram
Instead of writing every observation out, you can show how many times each value occurs. Here is the spread of test grades in a class of 24:
The mean of such a diagram is found by multiplying every grade by its frequency:
That is already a weighted mean, with the frequencies as the weights:
The ordinary arithmetic mean is its special case — the one where every weight is the same.
Notice one thing: is not a real grade at all. A mean need not belong to the data set, and usually does not.
The median
The median is the middle value of the data once it is sorted. The sorting is part of the recipe, not decoration — without it the "middle" number means nothing.
For an odd sample size there is a single middle:
For an even one there are two, so we average them:
The mode
The mode is the value that occurs most often. On a frequency diagram you can see it immediately — it is the tallest bar. In the class drawn above the mode is the grade : eight students got it.
The mode is the only one of the three measures that need not exist:
- when every value occurs once, there is no mode;
- when two values share the highest frequency, there are two (a bimodal set).
It has one advantage the others lack: it works where the data are not numbers at all. The commonest shoe size or the commonest car colour is a mode, and nobody can average colours.
When the mean misleads
Take a company of seven and their salaries, in thousands:
Both measures:
"People here earn 7 thousand on average" is true and misleading at the same time: six out of seven earn less. The reason is simple — the mean uses the size of every number, so one extreme value drags it across. The median sees only order: if the boss earned 220 thousand instead of 22, the median would still be 5.
That is the whole reason salaries, house prices and waiting times are reported as medians rather than means.
Three measures on one drawing
A dot plot shows it best, with one dot per observation. Take the set :
For this set:
The gap between the mean and the median is information in itself: it says the data is skewed and that an outlier is sitting somewhere in it.
Which measure to use
| measure | what it measures | when it is good |
|---|---|---|
| mean | the "shared out evenly" level | symmetric data, no extremes |
| median | the middle of the sorted data | data with outliers |
| mode | the most frequent value | categories, non-numeric data |
The honest thing is to quote two of them. If the mean and the median are close, the data is reasonably symmetric; if they are far apart, neither one alone describes anything.
Exercises
The five kinds of question match the five things this lesson teaches. A prompt starting with and a colon asks for the mean of the data that follows. A prompt of the form with an x inside the data asks for the missing value — use the total . The Median and Mode prompts give the data in random order, so the first step is always to sort it. Every answer is a number; with an even sample the median may end in a half — type it, for instance 6.5.
Practice
Work through a set of exercises — they get harder as you go. At the end you'll see your score and the mistakes worth reviewing.
Common mistakes
- A median without sorting — the middle number of an unsorted list is not the median, just an arbitrary value.
- Dropping the half — with an even sample the median often ends in .5, and rounding it to something "nicer" is wrong.
- Dividing by the wrong count — divide by the number of observations, not by the number of distinct values.
- Taking the mode to be the largest value — the mode is the most frequent one, not the biggest.
- Confusing the measures — "average" in everyday speech often means the median; a problem asks for whatever it names.
- Trusting the mean alone — in a set with an outlier the mean describes the data worse than the median does.
Formula card
Topic: Mean, median and mode
Arithmetic mean
the sum of all the data divided by how many there are
Total from the mean
this is how a missing value is recovered when the mean is known
Median — an odd sample
once sorted, it is exactly the middle number
Median — an even sample
the mean of the two middle values
Mode
the number that occurs most often — there may be none at all
Weighted mean
every value counts as many times as its weight says
