Data provenance — Data, 14–17 years
Data provenance is the record of where data came from and what happened to it before a conclusion was made. It lets someone check, repeat or correct an analysis.
Idea
Data provenance is the trail behind a result: who collected the data, when, with which instrument, and what changes were made. It separates the original observations from cleaning, filtering, joining and calculating. A result is easier to trust when this trail can be followed.
Why
Without a record of the steps, a surprising result cannot be checked and an error may be repeated. Provenance became important as datasets grew across laboratories, companies and public agencies, with many people editing or combining them. The problem it solves is not only secrecy; it is being unable to tell how a number was made.
Worked example
A researcher reports that a city’s average temperature was 18 °C. The record shows: 100 sensor readings, two faulty sensors removed, all Fahrenheit readings changed using (F − 32) × 5/9, and the mean calculated from the remaining 98 values. Another researcher can repeat those steps, check the removed sensors and see exactly where a disagreement begins.
Common trap
People often think a polished final table is enough because it looks clearer than raw data. That is reasonable: tidy presentation helps readers. But if the cleaning rules and discarded cases are hidden, a neat table can conceal important choices; clarity needs both the result and its history.
Use
Scientists, journalists and public services keep provenance records for published figures and datasets. In a weather app, the trail can show which station supplied a reading and whether it was corrected. In a company, it helps explain why a report changed after a new file was added.
Keep exploring
Other languages
Loading MyLeoNes™…