Correlation and linear regression
Two quantities measured on the same objects: price and floor area, speed and fuel consumption. The correlation coefficient measures how strongly they move together, the regression line describes how, and the coefficient of determination says how much of it was really explained. Plus the most famous sentence in statistics: correlation is not causation.
Before you start
This topic builds on earlier ideas. Before you start, it's worth working through the lessons below — they'll make everything click:
- Hypothesis testingA new page variant converts at 4.8 per cent instead of 4.0 — a real improvement or chance? A statistical test turns that question into arithmetic: a null hypothesis, a test statistic, a rejection region and a p-value. Together with what a p-value does NOT say, because it is the most abused number in science.
- Linear functionA linear function y = ax + b draws a straight line. Meet the meaning of the slope and the intercept, the zero, the formula of a line through two points, the general form, and the conditions for parallel and perpendicular lines.
Where this is used
Real situations where you count exactly the way this lesson teaches:
- Flat price against floor areaSixty transactions in one district give the line ŷ = 9.2x + 45, where x is the floor area in square metres and ŷ the price in thousands. The model values a 50-square-metre flat at 9.2 · 50 + 45 = 505 thousand. The coefficient of determination is 0.78, so floor area explains 78 per cent of the variation in price — the other 22 per cent is the floor, the condition, the aspect, and how much of a hurry the seller was in.
- Fuel consumption against speedMeasurements on a route at speeds from 70 to 130 km/h give ŷ = 0.08v + 1.6 litres per 100 km. At 110 km/h the model predicts 10.4 litres and that is credible. Substituting 200 km/h gives 17.6 litres and is worthless: air resistance grows with the square of speed, so outside the measured range the line stops being a model of anything.
- A laboratory calibration curveA spectrophotometer is calibrated with five standards at concentrations from 0.2 to 2.0 mg/l. The fitted line has r² = 0.999, so absorbance predicts concentration to within a thousandth of the variation. A sample whose absorbance falls outside the range of the standards must be diluted and measured again — the standard forbids reading it off by extrapolation, because beyond that range the detector stops being linear.
- Forecasting demandA café plots daily ice-cream sales against temperature and gets a = 42 servings per degree with r = 0.86. For a day forecast at 28 degrees, 6 degrees above the summer average, the model predicts 252 servings more than usual. The coefficient of determination of 0.74 says, however, that a quarter of the variation is unexplained — the day of the week and the rain are separate factors that temperature does not stand in for.
All formulas
Sum of products of deviations
positive when the deviations lean the same way
Pearson correlation coefficient
always between −1 and 1; no units
Least-squares line
the line always passes through the point (x̄, ȳ)
Fitted value
the model’s prediction at a given x
Residuals
it is their squares the method minimises
Coefficient of determination
the fraction of the variation in y the model explains
Significance test for a correlation
t distribution with n − 2 degrees of freedom
Everything else in this branch spoke about one quantity: its mean, its spread, its distribution. This lesson is about two, measured on the same objects — floor area and price, speed and fuel consumption, temperature and ice-cream sales.
There are two questions. Whether the quantities move together — that is what the correlation coefficient answers. How — that is what the regression line answers.
The scatter plot
As usual in this branch, it starts with a drawing. Each object contributes one point :
The scatter plot is the first step, not an ornament. Every number below can be computed without looking at the picture, and every one of them can then lie — which is what the section on abuses is about.
Pearson's correlation coefficient
The starting point is the sum of products of deviations from the two means:
A product is positive when both coordinates deviate the same way (both above their mean or both below) and negative when they deviate oppositely. The sum therefore says whether agreement outweighs disagreement — but it depends on the units, so it has to be normalised:
Built this way, the coefficient has three properties worth remembering: it always lies in , it has no units, and it does not change when the data is rescaled (converting a price from pounds to euros does not move at all).
| reading | |
|---|---|
| weak or none | |
| moderate | |
| marked | |
| strong | |
| very strong |
The regression line
Correlation says "strongly" but does not say "by how much". That is what the regression line answers — one chosen line , picked so that the sum of the squared vertical distances of the points from it is smallest. Hence the name: the method of least squares.
The second formula says more than "how to compute ": rearranged it reads , i.e. the regression line always passes through the point . That makes a convenient check on the arithmetic.
A warning about reading the slope. "A one-unit rise in corresponds to a rise in of on average" is a correct sentence. "A one-unit rise in causes a rise in of " is a statement about causation, and regression does not license it — as we shall see.
Residuals
The points do not lie on the line and are not meant to. The difference between an observation and a prediction is called a residual:
| total |
The residuals always sum to zero — that follows directly from the formula for — so the sum is not the point. The pattern is: residuals scattered without structure say the linear model fits; residuals bending into an arc say the relationship is curved; a fan widening to the right says the spread grows with . Those are faults that no single number will show.
The coefficient of determination
The square of the correlation coefficient has its own, very practical reading:
For our data , so : floor area, speed or whatever plays the part of accounts for 81 per cent of the variation in , and 19 per cent stays in the residuals.
It is worth seeing how fast that falls:
| explained | ||
|---|---|---|
A correlation of sounds solid and explains less than half the variation. A correlation of , which the table above calls "moderate", explains nine per cent — practically nothing.
Prediction and the limits of extrapolation
Any can be substituted into a regression line to get a number. Not every such number means anything.
- Interpolation — predicting for an inside the range of the data. There the model has support in observations and is credible.
- Extrapolation — predicting outside it. The line continues there only because lines always continue, not because the relationship stays linear.
Is the correlation significant
A coefficient computed from a sample is — like every statistic in this branch — a random number. The question "is there any correlation in the population at all" is a hypothesis test from the previous lesson, with :
The statistic follows a Student t distribution with degrees of freedom. In practice it is handier to read off the critical at :
| significant from | |
|---|---|
The table says two things at once and both matter. A high correlation from a handful of points means almost nothing — our at barely clears the boundary of . And conversely: a low correlation from a thousand points can be significant and entirely unimportant — significance says "this is not noise", while says the model explains four per cent.
Where this mathematics gets abused
This is the lesson after which most people draw the wrong conclusion, so there are five abuses.
1. Correlation read as causation
The most famous sentence in statistics and still the most often broken. The same correlation between and is explained by four different things:
- affects — what we usually assume;
- affects — reverse causation ("hospitals are dangerous, more people die in them");
- both depend on a third variable;
- the coincidence is chance, especially when many pairs were compared at once.
The correction: the arithmetic cannot tell those four apart, because all it sees is a pair of numbers per object. What can settle it is a randomised experiment or domain knowledge — never the coefficient itself, even at .
2. The lurking variable
A special and very common case of the above. Ice-cream sales correlate with drownings because both rise in hot weather. The number of firefighters at a fire correlates with the damage because both depend on the size of the fire. A lurking variable manufactures a correlation that no direct link accounts for.
The correction: before believing in a relationship, list the candidates for a third variable. The formal tool is a partial correlation or a model that controls for it; the informal one is the question "what else changes along with both".
3. Extrapolating beyond the data
A model fitted between and knows nothing about . The classic examples are dramatic: linearly extrapolating 100-metre sprint records predicts a negative time, and a child's growth predicts a two-metre five-year-old.
The correction: the range of the data is part of the model and is quoted together with the equation of the line. A high does not license extrapolation — it speaks only to the quality of the fit where the points were.
4. The coefficient read without the plot
measures linear dependence only. Points lying exactly on a parabola can give , and a single outlier can lift from zero to or drop it from to zero.
The correction: the scatter plot is compulsory, not optional. Four data sets with identical , , and regression line but entirely different shapes are the classic Anscombe quartet — built precisely to make this visible.
5. Significance mistaken for strength
On a large sample any non-zero becomes significant: at the boundary is , at it is . A headline reading "researchers found a significant association" therefore says nothing about how strong that association is.
The correction: read , not the p-value. A significant correlation of explains one four-hundredth of the variation and predicts nothing.
Exercises
The first set works on a five-point table, exactly like the examples above. The r prompt gives five pairs and asks for the correlation coefficient — compute it through the deviations from the means, since always lands on the middle point. The "regression line" prompt gives the same data and asks for both coefficients at once: write them as a = …, b = …; the parts are named, so the order does not matter.
Practice
Work through a set of exercises — they get harder as you go. At the end you'll see your score and the mistakes worth reviewing.
The second set is authored, because it asks about quantities the generator deliberately does not compute: is squared, and a prediction with and known is a substitution into a linear function — both already have generators in other branches. They stay here as questions about interpretation, alongside residuals and the arithmetic of extrapolation.
Practice
Work through a set of exercises — they get harder as you go. At the end you'll see your score and the mistakes worth reviewing.
The last question is a trap and is meant to be one: the arithmetic gives and is correct, while the result is worthless, because lies outside the range of the fit. Judging "may this be extrapolated" has no number for an answer and stays content — it is the third point in the section on abuses.
Common mistakes
- Inferring a cause from a correlation alone — the same number is explained by four different causal arrangements, and the arithmetic does not tell them apart.
- Forgetting the lurking variable — ice-cream sales and drownings correlate through the heat; look for a third quantity before believing in a link.
- Reading without a scatter plot — the coefficient measures only linear dependence, and a single outlier can invert it entirely.
- Confusing with — a correlation of explains per cent of the variation, not ; that is the difference between "a strong link" and "less than half".
- Extrapolating beyond the data — lines always continue, data ends where it ends, and a high does not change that.
- Measuring distances perpendicular to the line rather than vertically — the method minimises deviations in , because is what is being predicted; the perpendicular gives a different line and answers a different question.
- Treating significance as strength — at an of is significant, and it explains four ten-thousandths of the variation.
- Swapping the roles of and — the regression of on and of on are two different lines; they coincide only when .
Formula card
Topic: Correlation and regression
Sum of products of deviations
positive when the deviations lean the same way
Pearson correlation coefficient
always between −1 and 1; no units
Least-squares line
the line always passes through the point (x̄, ȳ)
Fitted value
the model’s prediction at a given x
Residuals
it is their squares the method minimises
Coefficient of determination
the fraction of the variation in y the model explains
Significance test for a correlation
t distribution with n − 2 degrees of freedom
