Statistics and probability

The average person does not exist

Jul 30, 2026·16 min read·3050 words
Luminous histogram bars settling into the soft bell of a distribution, one outlying point brighter than the rest, against a cosmic nebula in violet and magenta

The sentence "the average person earns X" is convenient, quotable and — with a little bad luck — describes no real human being at all. Not because somebody botched the addition. All of descriptive statistics rests on a single trade: you give up information about the structure of a set and receive, in exchange, one number that fits in a headline. The trade is often an excellent deal, but it is never free. This piece is the narrative layer above our statistics and probability branch — what exactly gets lost in that trade, who first noticed, and where intuition stops working.

Seven numbers, three answers

The shortest possible proof that "the typical value" is not one thing fits into seven numbers. Take the set X = {2, 2, 2, 5, 7, 10, 77}.

MeasureValueWhat it actually measures
Mode (Mo)2the value that occurs most often (three times)
Median (Me)5the middle element once sorted (the fourth)
Mean ()15105 / 7 — the balance point of the set

Three correctly computed measures, three completely different numbers, one set. More to the point: six of the seven observations lie below the mean. The number 15 is neither typical, nor central, nor anywhere near anything the set actually contains. It is, however, the point at which the set would balance if you laid it on a plank.

This is not a pathology of a made-up example. The distribution of wages, house prices or waiting times is right-skewed: a minority reaches values many times higher than everyone else, and it is that minority which drags the mean along. This is why statistical offices publish both figures side by side — the median tends to sit some ten to twenty percent below the mean, and roughly two-thirds of employees earn less than the mean. The median has one property the mean never has: it splits the population exactly in half. The mechanics of both, together with the mode and the question of which to reach for, are laid out in the lesson on mean, median and mode.

Before anyone had computed anything

Statistics did not begin with formulas. It began with lists of the dead. In 1662 the London tradesman John Graunt published Natural and Political Observations Made upon the Bills of Mortality — an analysis of the weekly burial returns kept in London as an early warning system for plague. Graunt did something that looks obvious now and was new then: he treated a chaotic heap of individual fates as an object with its own stable structure. He noticed that the ratio of boys to girls born held steady year after year (he put it at roughly 14 to 13), and he built a table of how many out of a hundred born survive each successive decade.

It is worth being careful with the superlatives here, because popular retellings like to multiply them. Graunt's table rested largely on assumptions rather than on actual data about the age of the dead — nobody was recording that. The first mortality table built on real age records is usually taken to be Edmond Halley's Breslau table of 1693. The share of the work owed to William Petty, Graunt's friend and the man who coined "political arithmetick", is also disputed. What stands undisputed is the actual breakthrough: the idea that you can say something certain about a population while knowing nothing certain about any single person in it.

The average man and his shadow

Two centuries later the Belgian astronomer Adolphe Quetelet took the next step — and along the way handed statistics its most durable interpretive error. In an 1835 work he introduced l'homme moyen, the "average man": an entity whose height, chest circumference and even propensity to crime are the population averages. He also proposed the ratio of body mass to the square of height — the Quetelet index, renamed BMI by Ancel Keys in 1972.

The charge laid against l'homme moyen is reification: granting an abstraction the status of a thing. The average man does not exist, because nobody has every trait sitting exactly at the population average — and the more traits you include, the more certain it becomes that the set of such people is empty. The second charge is graver: in places Quetelet treated averageness as an ideal, deviation from which is a defect. The same mistake is made today by anyone reading BMI as a diagnosis for an individual patient — the index was designed to describe populations, and nineteenth-century European ones at that. The ratio is itself a mixed quantity in which mass and height must first be brought into one system of units; kilograms and pounds are handled by the mass converter.

The bell curve comes from the same era, but not from demography. The normal distribution grew out of measurement error in astronomy and geodesy: when a result is shaped by many small independent disturbances, their sum settles into that characteristic bell. The name stuck to Gauss (1809), although the curve was first described by Abraham de Moivre in 1733 as an approximation to the binomial distribution, and developed independently by Laplace — a textbook case of Stigler's law, under which a discovery rarely carries the name of its discoverer. The struggle with inconsistent meridian measurements, which we cover in the history of the metre, was one of the proving grounds where that theory matured.

What a standard deviation really tells you

Two sets can share an identical mean and have nothing else in common. Describing that difference is the job of dispersion, and its standard measure is built in three steps: take the deviations from the mean, square them, average them — then take the square root.

The square is not decoration. Without it the whole calculation collapses, because by the very definition of the mean the deviations sum to zero — positives and negatives cancel to the last decimal. Squaring settles the signs, but it does something else too: it punishes large deviations disproportionately. An observation 10 away weighs a hundred times more than one 1 away.

And here comes the sentence that popular explanations repeat most often and that happens to be false: that the standard deviation is "the average distance of an observation from the mean". It is not. Average distance is an entirely different measure — the mean absolute deviation. For our seven numbers it comes to 17.71, whereas the population standard deviation is 25.47. The gap is neither coincidence nor rounding: the root of the mean of squares is always at least as large as the mean of the absolute values, and the more ragged the set, the further the two drift apart.

If absolute values would settle the signs just as well — why did the square win? For reasons that are mathematical without being necessary: the square is differentiable everywhere (the absolute value is not, at zero), and the variances of independent quantities simply add, which cannot be said of mean absolute deviations. That is a choice of a tool with good properties, not a law of nature.

The famous n − 1 divisor deserves its own dose of honesty. For our set, dividing the sum of squares 4540 by 7 gives a population variance of 648.57, and by 6 a sample variance of 756.67. The reason is not cosmetic: the points of a sample are by construction closer to their own mean than to the unknown population mean, precisely because the sample mean is what minimises the sum of squared deviations. Dividing by n would therefore give a systematically deflated estimator; computing the mean spends one degree of freedom and n − 1 remain. But — and this is the part textbooks drop — Bessel's correction yields an unbiased variance, not an unbiased standard deviation. The square root is a concave function, so s still underestimates σ. Besides, unbiasedness is one criterion among several: if minimising mean squared error were the goal, a divisor of n + 1 can do better. Why squares, where the root comes from and what exactly the n − 1 does are unpacked in the lesson on variance and standard deviation.

The year 1654, and what came before it

The birth of probability theory is conventionally dated to the summer 1654 correspondence between Blaise Pascal and Pierre de Fermat. The starting point was the problem of points: how do you fairly divide the pot of a game interrupted before anyone has won? The answer they worked out — divide according to each player's chance of finishing the game, not according to points already scored — is in essence the notion of expected value.

The date, though, is a tidying convention rather than a breakthrough from nothing. The problem of points had been circulating through Italian treatises since at least Luca Pacioli (1494); Tartaglia and Cardano both wrestled with it. Cardano wrote Liber de ludo aleae around 1564 — but it was not printed until 1663, which is to say after Pascal and Fermat, so it could not have influenced their work. The closure came three centuries later still: in 1933 Andrey Kolmogorov, in Grundbegriffe der Wahrscheinlichkeitsrechnung, grounded probability in measure theory, replacing intuitive definitions with three axioms. Only from that point on is "probability lies in [0, 1]" a theorem rather than an observation.

Counting, before you divide

Laplace's classical definition is deceptively simple: P(A) = |A| / |Ω|, favourable outcomes over all outcomes. The whole weight rests on two things that are easy to miss. First, on the assumption of equal likelihood — every elementary outcome must be equally probable. A fair die and a shuffled deck satisfy it by definition; weather, prices and diseases almost never do. Second, on being able to count both quantities, and that is combinatorics.

StructureOrderRepetitionFormulaExample
Permutationmattersnon!lining up 8 runners in lanes
Arrangement, no repetitionmattersnon! / (n − k)!a chair and a deputy from 10 people
Arrangement with repetitionmattersyesnᵏa four-digit PIN from 0–9
Combinationirrelevantnon! / (k!(n − k)!)six numbers from 49 in a lottery

The whole table reduces to two questions asked in order: does swapping two elements give a different result, and may an element be taken twice. A lottery draw is a combination, because the order the balls come out in does not matter: C(49, 6) = 13,983,816. A single ticket therefore carries a probability of 1 / 13,983,816 ≈ 7.15 × 10⁻⁸ — roughly one in fourteen million. All four structures, with worked examples, are covered in combinatorics, and Laplace's definition together with its limits in the lesson on classical probability.

Three times intuition gets it wrong

The birthday problem. How many people do you need before the chance that two of them share a birthday passes one half? The answer is 23, and it sounds absurd until you restate the question. Compute the complement — everyone has a different date — and for n = 23 it comes to 0.4927, so the probability we want is 0.5073. The key is that instinctively we ask "does anyone share my birthday", when the real question is about any pair at all. In a group of 23 there are C(23, 2) = 253 pairs. At 50 people the probability reaches 97%.

The Monty Hall problem. Three doors, a prize behind one. You pick a door, the host opens one of the others to reveal an empty one, then offers you the switch. Switching raises your chance from 1/3 to 2/3, a result that infuriated people with doctorates for years. The whole thing turns on the fact that the host is not choosing at random — he knows where the prize is and is obliged to open an empty door. If your first pick was wrong (and it is wrong two times in three), then after the elimination exactly the prize remains. Switch, and you win precisely whenever you started out wrong.

Simpson's paradox. This is the most dangerous of the three, because it is not about games but about real decisions. It says that a relationship visible within subgroups can reverse once they are pooled. The canonical example was described in 1975 in Science by Peter Bickel, Eugene Hammel and J. William O'Connell, who examined graduate admissions at Berkeley in 1973. In aggregate, about 44% of men and 35% of women were admitted out of more than twelve thousand applicants — a figure that looks like hard proof of discrimination.

DeptApplicants (M)Admitted (M)Applicants (W)Admitted (W)
A82562.1%10882.4%
B56063.0%2568.0%
C32536.9%59334.1%
D41733.1%37534.9%
E19127.7%39323.9%
F3735.9%3417.0%
Total269144.5%183530.4%

Across the six largest departments women have the higher admission rate in four of them, and yet in the summary row they lose by fourteen percentage points. There is no contradiction and no arithmetic error — there is a confounding variable. Women applied more often to departments with low admission rates for everyone (see C, E, F), men to the easier ones (A, B). The pooled average mixes "a candidate's chance" with "a department's difficulty" and returns a number that describes neither.

Two caveats, because this story tends to be told too smoothly. The departments in the original paper are anonymous — the popular "women chose humanities, men chose the sciences" is a later gloss, not a finding of the authors. And the authors did not write that there was no bias: they wrote that once department is properly accounted for, a small deviation appears in favour of women, and the real question moves one level up — why the fields women apply to are chronically underfunded. The effect carries Edward Simpson's name (1951), though Pearson (1899) and Yule (1903) had described it earlier.

The aircraft nobody ever saw

During the Second World War analysts studied the distribution of bullet holes on bombers returning from missions, so as to add armour where the holes were thickest. Abraham Wald of the Statistical Research Group pointed out something that was not in the data: the sample consisted exclusively of machines that had come back. Areas with no holes on the survivors were not areas that went unhit — they were areas that, once hit, kept the aircraft from returning. The armour belonged where the data fell silent.

The story is true, though it circulates in a souped-up version. In his 1943 memoranda Wald did not write "armour the engines", nor did he draw the famous dotted bomber silhouette — he built a method for estimating vulnerability from the damage seen on survivors. The drawing and the punchline about engines are the work of later popularisers. The mechanism, though, was aptly named: survivorship bias is inference from a sample that selection has stripped of its most informative cases. The same bias underwrites every list of billionaires' habits and every "this method works, just ask the people it worked for".

A catalogue of traps

TrapWhat it isCounterexample
Correlation as causationco-occurrence taken for agencyice cream sales and drownings — the culprit is temperature
Gambler's fallacybelieving a streak "has to" reverseafter five heads, tails is still 0.5
Law of small numberscategorical conclusions from a handful of casesa small sample carries enormous random variation
Prosecutor's fallacyconfusing "probability of A given B" with its reversesee below
Mean as the typical casea balance point taken for a portraitx̄ = 15 in a set where six of seven numbers are smaller

Two of these deserve unpacking, because they cost the most. The gambler's fallacy even has a date: on 18 August 1913, in the casino at Monte Carlo, the roulette ball is said to have landed on black 26 times in a row, and players lost fortunes betting red "because it's due". The probability of such a run on a single-zero wheel is (18/37)²⁶, roughly one case in 137 million — and even so, the chance of red on the twenty-seventh spin remains exactly what it was on the first. The ball has no memory; only the player does.

The prosecutor's fallacy is the confusion of a conditional probability with its reverse. Suppose a DNA profile occurs in one person in a million, and a match is found. The probability of such a match in an innocent person is 0.000001, which sounds like a verdict. But the court's question runs the other way: what is the probability that a person with a matching profile is guilty? If the pool of possible suspects numbers a million people, then besides the perpetrator roughly one other person carries the profile — and the trace alone gives something like 50%, not 99.9999%. Same number, two different questions, two incomparable conclusions.

It is worth adding that another warning against reading summary tables alone is Anscombe's quartet: four datasets with nearly identical means, variances and correlation, and completely different shapes on a plot. We take it apart in the piece on why a function's graph can never double back.

What is strict and what is convention

Most of the confusion around statistics comes from conflating two layers: theorems that follow from the axioms, and conventions adopted for convenience.

StrictConventional
Kolmogorov's axioms: P(Ω) = 1, additivity for disjoint events, P(A) ∈ [0, 1]which measure of location goes in the headline (mean or median)
The counting formulas and P(A′) = 1 − P(A)n − 1 as the default divisor — unbiasedness is one criterion among several
50.73% for 23 people, 2/3 in the Monty Hall problemthe significance threshold α = 0.05 — a disciplinary convention, not a constant of nature
Deviations from the mean summing to zerothe square rather than the absolute value as the penalty for deviating

The first column is not up for negotiation. The second is — which is exactly why it has to be stated. A report that says "the mean" without saying why not the median is hiding a decision, not reporting a fact.

One number is the start of a question

Reducing a crowd to a single number is one of the finest inventions in the history of science: without it you cannot compare two countries, two treatments or two quarters. But every such number is a summary, and a summary always has an author and always leaves something out. Graunt extracted a regularity from a list of burials, Quetelet turned it into a man who does not exist, Wald counted the aircraft nobody ever saw, and Bickel showed that the same admissions round gives two opposite answers depending on whether you look at departments or at the total.

The moral is not "don't trust statistics". It is this: for every single number, three questions are worth asking — how widely the data are spread, what dropped out of the sample before anyone counted it, and whether the subgroups say the same thing as the whole. Only then does one number stop being a conclusion and become the start of a question.

Further reading

  • John Graunt, Natural and Political Observations Made upon the Bills of Mortality (1662) — the source text, in the public domain.
  • P. J. Bickel, E. A. Hammel, J. W. O'Connell, Sex Bias in Graduate Admissions: Data from Berkeley, Science 187 (1975), pp. 398–404 — the original behind the Simpson's paradox example.
  • A. N. Kolmogorov, Grundbegriffe der Wahrscheinlichkeitsrechnung (1933) — thirty pages that closed three centuries of argument.
  • Abraham Wald, A Method of Estimating Plane Vulnerability Based on Damage of Survivors (memoranda from 1943, published 1980) — survivorship bias in the author's own version, without the later embellishments.
  • Amos Tversky, Daniel Kahneman, Belief in the Law of Small Numbers, Psychological Bulletin 76 (1971) — where conclusions drawn from three observations come from.
  • Darrell Huff, How to Lie with Statistics (1954) — old, mean-spirited and still accurate.
  • Stephen M. Stigler, The History of Statistics: The Measurement of Uncertainty before 1900 (1986) — the standard history of the discipline, Stigler's law included.
Learn

The Statistics and probability branch

Open the branch