Skip to main content

Independent editorial reference · no accreditation and no qualification awarded · general information only, not legal or professional advice

Open Data DeskGalway · IE

Terminology

Terminology of data reporting

Short, neutral definitions of the words that recur in data journalism. They are here to help read a document or a dataset note, not to replace a technical source.

Reviewed 12 June 2026 · sources dated in text · general information only

A great deal of confusion in reporting on numbers is vocabulary rather than mathematics. The same word carries different meanings depending on who uses it: a rate can be a proportion, a change over time or a charge; significance means one thing in statistics and another in ordinary speech; open data is used both for anything downloadable and for material released under a specific licence.

The definitions below describe how each term is generally used in official documentation and technical literature, and note where common usage diverges. They have no legal force: where a term appears in legislation, the instrument itself governs its meaning and a qualified professional should be consulted on any specific question.

Open data
Material published for reuse under a licence that permits it, usually with attribution. Availability for download does not by itself make data open.
Dataset metadata
The description accompanying a file: publisher, period, update frequency, licence, geographic level. It defines the limits of any claim the file can support.
Machine-readable
Structured so that software can parse it reliably. A table inside a PDF is readable by people and not, in this sense, by machines.
Denominator
The population or caseload a count is divided by to produce a rate. Its period, boundaries and definition must match the numerator.
Rate
A count expressed relative to a denominator, allowing comparison between places or periods of different size.
Mean
The arithmetic average. Pulled by extreme values, so it describes a skewed distribution poorly.
Median
The middle value of an ordered set. Usually the better description of a typical case where the distribution is skewed.
Percentage point
The arithmetic difference between two percentages. A move from 40 to 44 per cent is four percentage points and a ten per cent increase.
Base effect
The distortion produced when a percentage change is calculated from a small starting value, making a handful of cases look dramatic.
Margin of error
The interval around an estimate from a sample, given a confidence level. It covers sampling variation only, not question wording or non-response.
Confidence level
The proportion of similarly constructed intervals that would contain the true value; commonly 95 per cent, which corresponds to a z value of 1.96.
Statistical significance
A statement about whether an observed difference is distinguishable from no difference. It says nothing about whether the difference is large enough to matter.
Seasonal adjustment
Removal of the recurring calendar pattern from a series so that short-term changes can be compared.
Revision
A later, refined version of an earlier estimate. A difference between two publications may be the revision rather than a change in the world.
Vintage
The publication version of a series used in an analysis. Charts should state which vintage they draw on.
Suppression
Withholding a value, typically a small count, to protect identifiable individuals. Summing the visible rows understates the true total.
Sentinel value
A placeholder such as minus one or 9999 used to mark absence. Treated as a real number it corrupts every calculation it enters.
Near-duplicate
Two records describing the same entity with different spelling, punctuation or formatting. Resolving them requires a documented rule and manual review.
Join
Combining two tables on a shared key. Row counts before and after, and the number of unmatched records, must be checked.
Choropleth
A map shading areas by value. The choice of class boundaries changes the appearance more than the data does.
Provenance
The origin and history of a document or file: who produced it, when, and how it reached you. Established before content is analysed.
Corroboration
Agreement between sources that do not depend on each other. Multiple reposts of one original are a single source.
Beneficial owner
The natural person who ultimately owns or controls an entity. Registers of beneficial ownership can be incomplete or structured below reporting thresholds.
Register of charges
The public record of assets a company has pledged to lenders. It shows that security exists without disclosing the terms.
Reproducibility
The property of an analysis that another person, given the same sources and recorded steps, can rebuild the same figure.
Data minimisation
Holding only the personal data a purpose requires. A continuing duty in journalistic processing, not removed by the applicable exemptions.
Weighting
An adjustment applied to survey responses so that the sample composition matches known population totals. It corrects who is in the sample, not what they were asked or who declined to answer.
Non-response bias
The distortion that arises when the people who do not answer differ systematically from those who do. Increasing the sample size does not reduce it.
Unit of analysis
What each row of a table describes: a person, an incident, a premises, a contract, a year. Two files that look comparable often count different units.
Definitional break
The point in a series at which a category, a threshold or a collection method changed. Values either side of it are two series unless the publisher has restated the earlier ones.
Harmonised series
Figures compiled to a shared definition so that countries or regions can be compared. National measures of an apparently identical concept are usually not comparable.
Provisional estimate
A first published value that will be refined as later returns arrive. The publisher marks it as provisional; charts frequently draw it identically to a final figure.
Index with a base year
A series rescaled so that one chosen period equals one hundred. It makes relative movement legible and hides the level, so the base period must always be stated.
Statistical disclosure control
The techniques applied before release — suppression, rounding, banding, top-coding — so that individuals cannot be identified. They alter published values by design.
Redaction
The removal of specified content from a released document, usually third-party personal data. What was removed, and the ground cited for removing it, is itself information.
Public interest test
The weighing exercise applied under several access-to-records exemptions and in editorial harm assessment. A judgement to be recorded, not a threshold that can be computed.
Chain of custody
The documented history of an item of evidence from acquisition to publication: who obtained it, when, from where, and what has happened to it since.

Terms to use with care

Some widely used expressions have no stable technical content and are better avoided in reporting, or accompanied by an explanation. "The data shows" attributes agency to a file and usually precedes an inference rather than a record. "Correlation" used loosely implies a measured relationship where only a visual impression exists. "Anonymised" suggests a finality that re-identification research does not support, and "pseudonymised" is the more accurate description of most released datasets.

"Statistically significant" causes particular trouble because ordinary usage reads it as important. A difference can be significant and trivial, or substantial and unestablished, and a sentence that leaves the reader to guess which is meant is doing no work.

The material on this site prefers longer formulations that can be contested with evidence. "Recorded inspections fell by a fifth between 2019 and 2023 in the returns of eighteen of twenty-six authorities" can be checked and disputed; "inspections have collapsed" cannot.

The same word in two registers

Several entries above carry one meaning in statistical documentation and a different one in the language of access law or data protection, and a sentence that slides between the two is hard for a reader to catch. A record, in an access regime, is material a body already holds in some form, which is why a request for a record succeeds where a request for an explanation does not; in a dataset note the same word usually means a row.

Personal data covers information relating to a person who can be identified directly or indirectly, which is wider than a name and wider than what most people would call private. A suppressed cell in a statistical table exists because of that width: the value is withheld not because it is secret but because publishing it alongside the surrounding rows would make someone identifiable.

Public is the word that causes most trouble in open-source work. Publicly accessible describes what can be reached without a credential; published describes what a body has released deliberately for reuse; and neither settles whether compiling the material is proportionate. Keeping the three apart in a draft is the practical difference between describing what was available and asserting that it was fair game.

How these entries were built

Each definition was compared with prevailing usage in official statistical documentation, in the methodology notes of the relevant publishers and in peer-reviewed literature. Where common usage differs from technical usage, the entry says so. Where a word is used differently across jurisdictions, the entry describes the general sense rather than adopting one national definition.

The entries carry no legal weight. For any formal process, a filing or an assessment, the legislation in force and the relevant official guidance govern, and a qualified professional should be consulted.