Method 02
Reading statistics without being misled
How to interrogate a published figure: what it measures, who it excludes, how certain it is, and what it cannot establish however large it looks.
01A figure is the end of a process
Every published number is the output of decisions: what to count, who to ask, how to define a category, what to do with cases that do not fit, when to stop collecting. Reading a statistic critically means recovering enough of those decisions to know what the number can support.
This is not scepticism for its own sake. Official statistics are generally produced carefully and documented well, and the documentation is where the answers are. The failure is usually not in the production of the figure but in its reuse by someone who never read the note attached to it.
02What exactly is being measured
Start with the definition, because words that look self-explanatory rarely are. Unemployment can mean people claiming a payment or people meeting a labour-force survey definition of seeking and available for work, and the two move differently. Waiting time can be measured from referral or from a decision to treat. A homelessness count may exclude several categories of people without shelter.
The practical test is to ask who is inside the measure and who is outside it, then to ask whether the excluded group is relevant to the claim being made. Most misleading uses of official figures are technically accurate statements about a narrower measure than the reader imagines.
03Counts, rates and the denominator problem
A count answers how many and never whether that is many. Converting to a rate requires a denominator that matches the numerator in population, period and geography: the people who could have been affected, the cases actually handled, the premises actually registered.
Denominator errors are hard to see because the result looks like analysis. A rate per head of population where the numerator counts only adults; a 2016 census denominator against 2024 incidents; a numerator on old boundaries and a denominator on new ones. Each produces a number that is confidently wrong, and none of them announces itself in the output.
04Percentages, points and small bases
Percentage change divides the difference by the original value, which has two consequences worth stating plainly. A fall and a rise of equal percentage do not cancel: 100 to 80 is a fall of a fifth, and recovering to 100 is a rise of a quarter. And a large percentage from a small base is nearly meaningless — three cases becoming nine is a 200 per cent rise and six cases.
Percentage points are a separate unit. A share moving from 40 to 44 per cent has risen four percentage points, which is a ten per cent increase in the share. Both are true; using them interchangeably is how a modest change becomes a headline.
The rule that removes most of this risk is to publish the underlying counts alongside any percentage. A reader who can see three and nine will not be misled by 200 per cent.
05Uncertainty is a property of the figure
Any figure derived from a sample is an estimate with an interval around it, determined by the sample size, the observed proportion and the confidence level. Quoting the point estimate alone implies a precision the method cannot deliver, and it is the reason so many reported "changes" are noise.
The calculator below applies the standard expression for a proportion. Its most useful lesson concerns subgroups: an estimate that is solid for a national sample of a thousand becomes very wide when the claim is about a subgroup of eighty within it, and a difference between two such subgroups is usually not establishable at all.
Note also what the interval does not cover. It describes sampling variation only. Non-response, question wording, ordering effects, coverage of the sampling frame and the mode of interview are additional sources of error that no formula quantifies, which is why two well-run surveys can differ by more than their stated margins.
Margin of error for a proportion
margin of error = z × √( p × (1 − p) / n )
range = p − margin … p + margin
z = 1.645 (90%) · 1.96 (95%) · 2.576 (99%)
Arithmetic illustration only. It covers sampling variation and nothing else: non-response, question wording, ordering effects and coverage of the sampling frame add further error that no formula quantifies. It assesses no real survey and produces no advice.
06Seasonality, revisions and vintages
Many series move with the calendar, so a month-on-month comparison in unadjusted data mixes the seasonal pattern with any real change. Publishers therefore offer seasonally adjusted series and year-on-year comparisons; the requirement on the reporter is to choose one basis, state it and stay with it.
Revisions are a feature of good statistical practice, not a failure of it. Early estimates are refined as returns arrive, which means a difference between two publications may be the revision rather than the world. Any chart should say which vintage of the data it uses, and provisional values should be marked rather than drawn identically to final ones.
07Comparing places and periods
Cross-country comparison is where definitions diverge most. Harmonised European series exist precisely because national measures are not comparable, and using a national figure for one country beside a harmonised figure for another produces a difference that is an artefact of the sources.
Within a country, comparability breaks at boundary changes, at reporting-system changes and at the introduction of new categories. A time series that crosses such a change is two series unless the publisher has restated the earlier values, and restatement should be verified in the documentation rather than assumed.
08Association, cause and the honest sentence
Two series moving together supports a question, not a conclusion. The standard alternatives must be considered and, where they cannot be excluded, stated: a common third factor, a shared trend, reverse causation, selection in how the data was gathered, or coincidence in a short series.
Sentence construction is where this discipline is applied or lost. "Rose alongside" is defensible; "linked to" is vague enough to be read as causal and should be avoided; "caused" requires a design that administrative data almost never provides. The verb should not claim more than the method.
09Risk in terms a reader can use
Relative change without absolute figures is unusable. A doubled risk means something different when it runs from one in a million to two in a million than when it runs from one in fifty to one in twenty-five. Give both, on the same denominator, and let the reader judge the magnitude.
The same applies to differences described as significant. Statistical significance concerns the probability of observing a result if there were no effect; it says nothing about whether the effect is large enough to matter. A significant but tiny difference is not news, and a large difference that fails a significance test is not established.
10Rankings, league tables and the noise in them
A ranking of places, institutions or authorities by a rate is one of the most reliably misleading forms a figure can take, and it is also one of the most requested. The mechanism is straightforward: where the units differ in size, the smallest ones produce the most extreme rates in both directions, because a single additional case moves a rate calculated on a small denominator much further than it moves one calculated on a large denominator. The top and bottom of such a table are therefore populated by the smallest units, whatever the underlying performance.
The test that settles it is to plot each unit's rate against its denominator and look at the shape. If the spread narrows as the denominator grows, and the extremes belong to the smallest units, the ranking is largely describing sample size. Statistical publishers deal with this by grouping small units, by suppressing rates below a minimum denominator, or by publishing an interval around each rate, and any of the three is preferable to an ordered list.
Where a ranking has to be reported, two sentences repair most of the damage: give the denominator alongside each rate, and state which differences are within the range that could be produced by ordinary variation. A table in which the first and eighth positions cannot be distinguished should not be described as a table of the best and the worst.
11An interrogation sequence for any figure
Most of these questions are answered by the publisher's own documentation in a few minutes. The remainder are answered by asking the publisher directly, which is a normal and generally welcomed enquiry rather than an imposition.
- 1What exactly is measured, and who is excluded?
- 2Full count or sample? If a sample, how large, and what is the interval?
- 3What is the denominator, from which year, on which boundaries?
- 4Adjusted or unadjusted; which vintage of the series?
- 5Has a definition, category or boundary changed within the period?
- 6Are counts published alongside percentages?
- 7What alternative explanations exist for the pattern?
- 8Does the language of the draft match the strength of the evidence?
12Presenting figures with the right precision
Precision should not exceed the method. A survey estimate reported to one decimal place asserts a resolution that the sampling error contradicts, and rounding to whole numbers is more honest as well as more readable. Every figure needs its unit, period and source, and where the number is an estimate it needs its interval.
Where two credible sources disagree, resist the temptation to pick one. The disagreement is usually definitional and the examination of it is often more informative than either figure, because it shows the reader what the measure actually is.
| Question | Why it matters | Where the answer is |
|---|---|---|
| What is the definition? | Decides who is counted | Publisher methodology note |
| Count or estimate? | Determines whether an interval applies | Series documentation |
| Which denominator? | Governs comparability of places | Population or caseload source |
| Adjusted or not? | Changes short-term direction | Series title and notes |
| Which vintage? | Revisions can look like change | Publication date |
| Any definitional break? | Splits a series in two | Revision history |
| How large is each unit? | Small denominators produce extreme rates | Denominator column of the same table |
Working checklist
- Read the definition before quoting the number.
- Ask who the measure excludes.
- Publish counts alongside percentages.
- Report the interval for any sample-based figure.
- Keep one basis of comparison in a story.
- State the vintage of the series used.
- Give absolute figures with any relative risk.
- Match the verb to the strength of the evidence.
- Give the denominator beside every position in a ranking.
Questions
Why do two surveys of the same thing disagree?
Margins of error cover sampling variation only. Question wording, non-response, coverage of the frame and interview mode add further differences that no formula quantifies.
Is a statistically significant difference automatically newsworthy?
No. Significance concerns whether an effect is distinguishable from none; it says nothing about size. A tiny significant difference may not matter.
Can figures for two countries simply be compared?
Only where the definitions are harmonised. National measures of apparently identical concepts frequently differ, and the gap can exceed the difference being reported.