Skip to main content

Independent editorial reference · no accreditation and no qualification awarded · general information only, not legal or professional advice

Open Data DeskGalway · IE

Method 01

From open data to a published story

A working sequence that takes a public dataset from download to a checked, published finding, with the decisions and failure points named at each stage.

Reviewed 12 June 2026 · sources dated in text · general information only

01Begin with a question the data could answer

Data reporting fails most often at the start, when a dataset is downloaded because it is available rather than because it answers something. An interesting file is not a story, and the effort spent looking for one inside it is usually wasted. The productive order is the opposite: state a question in a sentence, then ask what record would have to exist for it to be answerable.

A usable question specifies a subject, a measure, a population and a period. "Are inspections falling?" is not yet usable. "How many food-premises inspections did each local authority record in each year from 2018 to 2024, and how does that compare with the number of registered premises?" names the unit, the measure, the denominator and the range, and it can be tested against a catalogue in an afternoon.

The question should also be falsifiable. If both a rise and a fall would be written up as a story, the analysis has no discipline and will find whatever the analyst expects. Deciding in advance what result would mean there is no story is the cheapest protection against motivated reasoning.

02Establish what the source can support

Before any analysis, read the metadata: publisher, period, update frequency, last modification, geographic level, licence and methodology notes. The purpose is to establish the limits of the claim the file can support, which is a different exercise from checking that the file opens.

Three questions decide most of it. Is the population complete, or does the file cover only bodies that chose to report? Are the definitions stable across the whole period, or was a category redefined mid-series? Is the smallest available unit small enough for the comparison you intend, or will you be forced to a level of aggregation that hides the effect?

If the answers rule out the original question, that is a successful outcome for an hour's work. Reframing at this point costs nothing; discovering the same limitation after building an analysis costs a week and creates pressure to publish something anyway.

03Preserve the original and work in layers

Store the downloaded file untouched, with its access date, source URL and licence recorded beside it. Every subsequent step reads from that copy and writes elsewhere. This single rule prevents the most common irrecoverable failure in data work, which is an analysis resting on a file that was edited by hand in ways nobody can now reconstruct.

Keep raw, processed and output separated, and regenerate the lower layers rather than patching them. When an error is found late — and on a real story it will be — it is corrected once at its source and flows through to every table and chart, instead of being fixed in three places with one missed.

04Clean with a written record of decisions

Cleaning is where the analysis is quietly decided. Restructure human-formatted layouts into one header row and one row per observation; check types so identifiers keep leading zeros and dates are not reinterpreted; align units once and record the conversion factor; resolve near-duplicate names with a documented rule and manual review of the borderline cases.

Missing values need a meaning before they need a treatment. A blank, a zero, a negative sentinel and a suppression code used for small counts are four different things, and only the publisher can usually confirm which is which. Where counts are suppressed for privacy, the visible total is lower than the real one, and the published figure must say so.

Every decision taken by judgement rather than by rule belongs in the record, with the number of rows it affected. That record is what allows a colleague to reproduce your table, and what allows a correction to be traced to a step rather than announced as a general apology.

05Analyse the simplest sound version first

The first analysis should be the plainest one that could answer the question: counts by category and period, then a rate with an appropriate denominator, then a comparison. Sophistication added before the simple version is understood tends to hide errors rather than reveal patterns.

Choose the average with the shape of the distribution in mind, and report which one is used. Distinguish per cent from percentage points. Keep one basis of comparison — adjusted or unadjusted, year-on-year or period-on-period — and do not switch inside a story. Where figures come from a sample, carry the margin of error with them rather than quoting point estimates.

Then attack your own result. Recompute the headline figure by a different route; check whether it survives a change of period, denominator or exclusion rule; look for the boundary change, the reporting-system change or the one-off event that would explain it more boringly. A finding that only exists under one specification is not a finding.

06Estimate the uncertainty you are carrying

Where a figure comes from a survey rather than a full count, the margin of error determines what can be said about it. The tool below applies the standard expression for a proportion, and its main use is to show how quickly precision degrades in subgroups: a headline sample that supports a confident national figure often cannot support a claim about a region or an age band within it.

The practical rule it enforces is that two estimates differing by less than their combined intervals do not establish a difference. Reporting such a gap as a change is the most common numerical error in coverage of survey data, and it is entirely avoidable.

Where the calculator is

The margin-of-error calculator sits in the guide on reading statistics, together with the formula and the note on what it does not cover.

07Take the numbers to people

An analysis is a hypothesis about the world until people who know the field have tested it. Put the figures to the publisher of the data, to practitioners who generate the records, and to the bodies the finding concerns, with enough specificity that they can engage: the exact figures, the definitions used, the period and the method.

Expect the reply to change something. Publishers explain that a category was recorded differently in one region; practitioners point out that a fall reflects a change in what gets logged rather than in what happens; a body identifies a duplicate. Each of those is the analysis improving before publication rather than after.

This is also where the human material enters. A figure describes a pattern; a person describes what the pattern is like to live inside. The story is usually the two together, with the number establishing scale and the account establishing meaning.

08Write the finding so it cannot be over-read

The sentence carrying the main number should state the measure, the population, the period and the source: records of a specific kind, held by a named body, for a stated range. Vague formulations like "cases have soared" invite the reader to supply a magnitude the data does not support.

Language should track the strength of the evidence. "Rose alongside" where two series moved together; "is consistent with" where a mechanism is plausible but untested; "caused" only where the design supports it, which for administrative data is almost never. Alternative explanations belong in the piece, not in a defensive note afterwards.

Charts must agree with the text. Bar baselines at zero, denominators stated, missing and suppressed values marked, the classification method named on any map, and the underlying values available so a reader can check rather than trust.

09Publish the method with the story

A methodology note should give the sources with access dates, the period, the unit of analysis, the cleaning and matching rules, the calculations, the known limitations and a route for corrections. Where the licence allows, publish the processed data and the code; where it does not, say what cannot be released and why.

The note is not a technical appendix for specialists. It is the part of the story that lets a sceptical reader, a criticised body or a competitor check the work, and its existence is what distinguishes a documented finding from an assertion with a chart attached.

10A workable sequence

The shape matters more than the durations, which vary with the size of the dataset and the availability of the people involved. What should not compress is the space between having a result and publishing it: the re-derivation and the responses are where errors are caught.

  • Day 1Write the question in one sentence and define what would mean there is no story.
  • Day 1Read the metadata; decide whether the source can support the question.
  • Day 2Store the raw file with its access date; set up the three layers.
  • Days 2–4Clean with a written decision record; count rows through every join.
  • Days 4–6Run the simplest sound analysis; attempt to break your own result.
  • Days 6–8Put figures to the publisher, practitioners and the bodies concerned.
  • Days 8–9Independent re-derivation of the headline figure by a colleague.
  • Day 10Draft with the method note; check every chart against the text.

11Where these projects usually go wrong

The recurring failures are few and predictable. A dataset downloaded without a question. An absence of published data reported as an absence of the underlying activity. A rate built on a mismatched denominator. A comparison across a definitional or boundary change. Suppressed small counts summed as though they were zero. A percentage from a base of three. A single specification, chosen because it was the most striking.

Each of these is caught by a step above, and each survives when the sequence is skipped under deadline. The purpose of a written method is not procedural neatness; it is that the checks still happen on the day when there is no time for them.

Stages, output and the check that closes each one
StageOutputClosing check
QuestionOne testable sentenceA null result would be reportable
Source assessmentLimits of the claimPopulation and definitions stable
PreservationUntouched raw copyAccess date and licence recorded
CleaningAnalysable table plus logRow counts explained at every join
AnalysisFigures with uncertaintyResult survives a changed specification
Human checkResponses on the recordPublisher and subjects had specifics
PublicationStory plus method noteColleague re-derived the headline figure

Working checklist

  • State the question before opening the file.
  • Decide what result would mean there is no story.
  • Keep the raw download untouched and dated.
  • Log every judgement made while cleaning.
  • Give the denominator behind every rate.
  • Try to break your own finding before defending it.
  • Put exact figures and definitions to the publisher.
  • Have someone else re-derive the headline number.
  • Publish sources, method and limitations with the story.

Questions

How long does a data story take?

Anything from days to months depending on the source. What should not be compressed is the interval between having a result and publishing it, because that is where independent re-derivation and responses catch errors.

What if the data does not answer the question?

Reframe or stop. Discovering the limit in the first hour is a good outcome; discovering it after building an analysis creates pressure to publish something the data cannot support.

Is the methodology note necessary for a short piece?

Some form of it always is. Even three sentences naming the source, the period and the main limitation lets a reader judge the figure instead of taking it on trust.