Skip to main content

Independent editorial reference · no accreditation and no qualification awarded · general information only, not legal or professional advice

Open Data DeskGalway · IE

Field 03 · Analysis

Statistical literacy

Averages, rates, percentages, uncertainty and correlation: the small set of ideas that prevents most published numerical errors.

Reviewed 12 June 2026 · sources dated in text · general information only

01Which average, and why it matters

The arithmetic mean is the default and often the wrong choice. For a skewed distribution — incomes, waiting times, company sizes, compensation payments — the mean is pulled by a small number of large values and describes nobody's experience. The median, the midpoint of the ordered values, usually describes the typical case better.

The honest practice is to report which average is used and, where they diverge substantially, to give both. A gap between mean and median is itself informative: it says the distribution has a long tail, which is frequently the actual story.

02Counts, rates and denominators

A raw count answers how many and almost never answers whether that is a lot. Ten incidents in a city of a million and ten in a town of four thousand describe different situations. Converting to a rate requires an appropriate denominator: the population at risk, the number of cases handled, the number of premises inspected.

Denominators need the same scrutiny as numerators. Which population, from which year, on which boundaries, including or excluding which groups. A rate built on a mismatched denominator is more misleading than the count it replaced, because it looks like an analysis.

03Percentages, percentage points and base effects

Percentage change is computed as the new value minus the old, divided by the old, times one hundred. Two consequences are regularly missed. First, a fall and a rise of the same percentage do not cancel: a value falling from 100 to 80 loses 20 per cent, and returning to 100 requires a rise of 25 per cent. Second, a change measured from a small base produces a large percentage that means very little: two cases becoming six is a 200 per cent increase and four cases.

Percentage points are a different unit. A share moving from 40 per cent to 44 per cent has risen by four percentage points, which is a ten per cent increase in the share. Both statements are correct and they are not interchangeable; using them loosely is one of the easiest ways to overstate a finding.

04Uncertainty in sample-based figures

Any figure from a survey is an estimate of a population value, and it carries a margin of error determined by the sample size, the observed proportion and the confidence level chosen. Reporting the point estimate alone implies a precision the method cannot deliver.

The practical rule follows directly: when two figures differ by less than their combined margins, the honest statement is that the data do not establish a difference. Small subgroup breakdowns need special care, because the effective sample within a subgroup can be a fraction of the headline sample and the interval correspondingly wide.

05Seasonality, revisions and trend

Many official series move with the calendar. Comparing one month with the previous month in an unadjusted series mixes the seasonal pattern with any real change, which is why publishers offer seasonally adjusted versions and year-on-year comparisons. Choose one basis, state it, and do not switch between them inside a story.

Revisions are normal. Early estimates are refined as more returns arrive, so an apparent change between two publications may reflect the revision rather than the world. A responsible chart labels which vintage of the data it uses.

06Correlation, causation and confounding

Two series moving together is a starting point for inquiry, not a conclusion. Alternative explanations must be stated: a common third factor, a shared trend over time, reverse causation, selection in how the data was collected, or coincidence in a short series.

The language of a report should match the strength of the evidence. "Rose alongside" is defensible where "caused" is not, and a sentence that quietly promotes an association into a mechanism is the most common failure in numerical journalism.

07When the mix changes, the average moves on its own

An overall rate can rise while every group within it falls, and it can fall while every group rises. This is not a trick of presentation; it follows from the arithmetic whenever the composition of the population changes. If a population ages and the condition being measured is commoner in older people, the overall rate climbs without anything happening to any individual group's risk.

The consequence for reporting is that an aggregate change is never self-explanatory. Before describing a movement, break the figure down by the dimension most likely to have shifted — age, region, size of body, type of case — and check whether the group rates moved in the same direction as the total. Where they did not, the honest sentence describes the change in composition, because that is what happened.

Publishers handle this by standardising: recalculating the rate as it would be if the population structure had stayed fixed, which is why age-standardised figures exist alongside crude ones. Standardised and crude rates answer different questions, both are legitimate, and mixing them within one comparison produces a difference that belongs to the method rather than the world.

08Risk expressed so a reader can use it

Relative risk without absolute numbers is close to meaningless for a reader. A doubling of a risk that runs from one in a million to two in a million is not comparable to a doubling from one in fifty. Reporting both the relative change and the absolute figures, with the same denominator, lets the reader judge the size of the effect.

09Presenting numbers honestly

Precision should not exceed the method: a survey estimate does not warrant a decimal place. Every figure needs its unit, period, source and, where it exists, its uncertainty. Where two credible sources disagree, the disagreement should be preserved and the definitions examined, rather than resolved by choosing the more convenient figure.

Statistical choices and what they change
ChoiceEffect on the findingWhat to state
Mean or medianTypical value can shift substantiallyWhich measure, and both if they diverge
Count or rateComparability between placesThe denominator and its source year
Per cent or percentage pointsApparent size of a changeThe unit, explicitly
Confidence levelWidth of the intervalLevel and sample size
Adjusted or unadjustedDirection of short-term changeThe basis of comparison
Data vintageWhether a change is realPublication date of the series used
Crude or standardisedWhether composition is held constantWhich basis, and the reference structure

Checks before publishing

  • Say which average is being reported.
  • Give the denominator behind every rate.
  • Distinguish per cent from percentage points.
  • Report uncertainty for any sample-based figure.
  • Keep one basis of comparison throughout a story.
  • State alternative explanations for an association.
  • Check whether group rates moved with the total before describing a change.

Questions

When should the median be used instead of the mean?

Whenever the distribution is skewed, as with incomes or waiting times. Where the two differ substantially, report both.

Is a 200 per cent increase always significant?

No. From a small base it can mean a handful of cases. Always publish the underlying counts alongside the percentage.

Can two survey figures be compared directly?

Only if their difference exceeds their combined margins of error. Otherwise the data do not establish a difference.

Why can a total rate rise when every group's rate falls?

Because the composition of the population changed and the total is a weighted mixture of the groups. Break the figure down before describing the movement, and say so where the shift is in the mix.