My eight-step process
Reference · the words data comes wrapped in

What everything is called

The obstacle nobody warns you about is that one thing has four names. You search for “raw data”, find nothing, and conclude it does not exist, when the repository you were on calls it disaggregated and the next one calls it microdata.

So this is a phrasebook rather than a dictionary. Each entry gives the other names the same thing goes by, and says why it matters to an investigation. The steps do the teaching; this is here for the moment you are staring at a download page.

What a dataset is

The words for the shape of a file, which is the thing you need to know before anything else.

Indicator
Also calledvariableseriesmetricparametermeasure

One measured quantity, tracked over time or across places.

Why it mattersThis is the word most repositories organise themselves by, and the thing your research question has to name. Searching for a topic finds articles; searching for an indicator finds data.

Disaggregated
Also calledstation-levelmicrodataunit-recordraw recordsgranular

Broken down to the smallest unit the collector recorded, before any averaging.

Why it mattersThis is the form you want. A national annual mean has already had the variation you were going to analyse taken out of it, and no amount of processing puts it back.

Unit of analysis
Also calledobservationrecordone row

What a single row of your file represents.

Why it mattersIf you cannot finish the sentence "one row is...", you do not yet know what you are comparing. It is the first of step 4's four tests, and the fastest way to find out a dataset is not the one you thought.

Granularity
Also calledresolutionlevel of detailspatial resolutiontemporal resolution

How finely the data is cut, in space or in time.

Why it mattersMonthly against annual, station against country. A question at a finer grain than the data cannot be answered, and this is the mismatch that kills the most investigations before they start.

Coverage
Also calledextentscopecompleteness

Which places and which years are actually in the file.

Why it mattersRarely what the title implies. A dataset called global has holes, and where the holes are is usually not random, which makes coverage a finding as well as a constraint.

Time series
Also calledlongitudinalpanelcross-section

One thing measured repeatedly over time. A cross-section is many things measured once; a panel is many things measured repeatedly.

Why it mattersIt decides which question you can ask. Change over time needs a series; a comparison between places needs a cross-section; comparing change between places needs a panel, which is the rarest and the most powerful.

How the number was made

Every value is the output of a process. These are the processes, and they are not equally trustworthy.

In situ
Also calledground-basedfield measurementobserved

Measured directly at the place, by an instrument or a person.

Why it mattersThe strongest kind, and still has a detection limit, a calibration and somebody's judgement in it.

Modelled
Also calledestimatedderivedreanalysisgridded

Calculated from other measurements rather than observed.

Why it mattersYou are then analysing a model's output, which is a legitimate thing to do and a different claim from analysing an observation. Say which you have.

Remotely sensed
Also calledsatellite-derivedearth observationclassified imagery

Inferred from what a sensor detected, usually from orbit.

Why it mattersUnbeatable coverage, and it classifies rather than sees. Every classified product has a confusion it is known for, and its documentation usually names it.

Imputed
Also calledinterpolatedgap-filledestimated valuemodelled fill

A value that exists because the gap did, filled in by a rule.

Why it mattersIt is not a measurement, and analysing it as one overstates how much you know. Good datasets flag imputed values in a separate column, and the first thing to do is find out whether yours does.

Provisional
Also calledpreliminaryunvalidatedsubject to revision

Published quickly and expected to change.

Why it mattersThe figure you download today may not be the figure published in six months, which is exactly why your method has to record the date you accessed it.

Release
Also calledversionvintageeditionlast updated

Which issue of a dataset you have.

Why it mattersTwo people can download the same indicator a year apart and get different numbers. Recording the release is what makes your method repeatable rather than approximately repeatable.

Detection limit
Also calledlimit of quantificationbelow reporting limitcensored

The smallest amount the method can distinguish from nothing.

Why it mattersValues at the limit are often recorded as zero, which is a different claim from none. It changes what a mean means, and noticing it is a limitation worth marks.

What you do to it

The operations that turn a download into evidence. Naming the one you used is half of justifying it.

Normalise
Also calledper capitaper unit areastandardisescale

Divide out a difference you are not interested in.

Why it mattersComparing two countries' total emissions mostly compares their populations. Per capita compares something you meant to. It is the commonest and cheapest form of control on this route.

Baseline
Also calledreference periodindex yearanomaly

The fixed point that change is measured against. An anomaly is a departure from it.

Why it mattersClimate data is very often published as anomalies rather than values, and the baseline is chosen by the publisher. Two datasets on different baselines cannot be compared until you say so.

Stock and flow
Also calledlevel and rateamount and change

A stock is how much there is; a flow is how fast it changes.

Why it mattersForest area is a stock, deforestation rate is a flow, and a question that mixes them up cannot be answered. It is also the systems vocabulary from step 1 arriving in your spreadsheet.

Harmonise
Also calledreconcilealignmake comparable

Put two sources onto the same units, categories and dates before combining them.

Why it mattersThe moment you use a second dataset this becomes most of the work, and skipping it produces a join that looks fine and compares nothing.

Join
Also calledmergematchlink on a keylookup

Bring two datasets together on something they share, usually a date or a place.

Why it mattersThe single best demonstration of independent processing available on this route, and nobody does it by accident. Rainfall joined to almost anything else is the example worth copying.

Inclusion criteria
Also calledselection rulefilterexclusion criteriasubset

The rule that decides which rows you keep.

Why it mattersOn this route it is both your variable control and your integrity test. Fixed in advance and reported with what it removed, it is a method; applied after you can see the pattern, it is choosing your result.

Balanced panel
Also calledmatched samplepaired comparisoncommon set

Only the units present in every period you compare.

Why it mattersIt removes the objection that your early and late groups are different places. Cheap to do, and it converts the strongest criticism of a before-and-after comparison into a sentence saying you tested it.

How it misleads you

The named failures. A limitation you can name is worth more than three you gesture at.

Proxy
Also calledindicator ofstands in forsurrogate

Something measurable that you are treating as a stand-in for something you cannot measure.

Why it mattersAlmost every environmental indicator is one. Naming what yours is a proxy for, and where the two come apart, is one of the most reliable ways into the top band of Criterion F (Evaluation).

Ecological fallacy
Also calledaggregation biasthe aggregate is not the individual

Assuming that what is true of a group is true of the things inside it.

Why it mattersA national average can hide severe local damage; a site mean can hide the storm peaks that cause the problem. It is the characteristic error of working with somebody else's summaries.

Confounder
Also calledlurking variablethird variablecommon cause

Something that moves both of the things you are comparing.

Why it mattersYou controlled nothing physically, so this is the main threat to any relationship you find. Naming your two most likely confounders is worth more than any hedge about correlation and causation.

Selection bias
Also calledsampling biasnot missing at randomsurvivorship

The data you have is not a fair sample of the thing you care about.

Why it mattersMonitoring stations go where somebody expected a problem, or where access was easy. Missing data is rarely missing at random, and where the gaps are is often a finding in itself.

Data dredging
Also calledp-hackingfishingcherry-picking

Trying enough combinations that something comes out significant by chance.

Why it mattersWith forty columns and a spreadsheet this is trivially easy and invisible in the finished report. The protection is deciding your question and your inclusion rule before you look, and saying that you did.

Temporal misalignment
Also calleddifferent reporting periodslagmismatched frequency

Two series that describe the same period on different clocks.

Why it mattersA census every five years against monitoring every month. A financial year against a calendar year. If your sources are not on the same clock, saying so is part of the analysis rather than an apology for it.

Finding it, and crediting it

What the parts of a repository are called, and what a citation of data has to carry.

Repository
Also calledportaldata hubcataloguedata store

A place that publishes datasets, usually with a search and a licence.

Why it mattersThe word to use when searching. "Air quality repository" finds the archive; "air quality data" finds articles about air quality data.

Metadata
Also calledcodebookdata dictionarydocumentationmethodology notereadme

The document explaining what each column is, how it was measured, and by whom.

Why it mattersThis route's equivalent of calibrating an instrument, and it is where most of your evaluation material comes from. Fifteen minutes there pays for itself twice.

API
Also calledquery serviceendpointweb serviceREST

A way of asking a repository for exactly the rows you want, as a web address.

Why it mattersLess frightening than it sounds and better for you than a download button: the query URL *is* your extraction protocol, and a reader can re-run it rather than reconstructing it.

Licence
Also calledterms of useconditions of useCC BYODbLEtalabopen access

What you are permitted to do with the data, and what you must do in return.

Why it mattersOpen almost never means unattributed. Most licences require the source to be named, and naming it is the easy half of the ethics section on this route.

DOI
Also calleddigital object identifierpersistent identifier

A permanent address for a specific version of a dataset or paper.

Why it mattersIf your download offers one, use it. It is the strongest form of traceability there is, because it survives the publisher reorganising their website.

Accessed date
Also calledretrieveddownloaded onas of

The day you took your copy.

Why it mattersData gets revised. Without this a reader cannot tell whether they are looking at the numbers you had, and it is the one part of a data citation people leave out.

Using these to search

Put the word for the shape next to the word for the subject. “Air quality repository” finds the archive; “air quality data” finds articles about air quality data. “Station-level” or “disaggregated” next to your topic is the fastest way past a page of national totals.

What everything is called · ESS IA guide | Revise