Secondary data routeEvery step is showing the version for data somebody else collected. Switch to primary data
Criterion C: Method · 4 marks500 words · suggested

Find and select your data

You arrive with

A research question, a place, and a tension. Possibly also a dataset, if you found one before you finished the question. That order is fine on this route and often unavoidable: what the data can answer shapes what it is sensible to ask.

You leave with
  • A dataset you have opened, counted and understood
  • A protocol precise enough to rebuild your exact file
  • An inclusion rule you fixed before you looked
  • A licence, a credit and a date accessed

Your method is a set of instructions for getting the same file back.

Nobody handed you a river, so nothing about apparatus applies. What this criterion wants from you instead is an account precise enough that a stranger could go to the same source, apply the same filters, and end up with your dataset row for row. That is a harder standard than it sounds and a much easier one to hit, because everything you need to write down is on the screen at the moment you press download.

The same two questions as the fieldwork route. Both must still be yes.

Could someone rebuild your file?

Not find similar data. Rebuild yours: same source, same version, same filters, same rows. If a reader would have to ask you which dataset you meant, the answer is no.

Is there enough data?

Enough to answer the question you asked, and enough for the test you intend to run. On this route the amount is something you chose, so it is something you have to justify rather than something that happened to you.

The way this route fails, and it fails here

An investigation that only reviews what other people have written is explicitly not a repeatable method. Not a weaker one: not one. It fails the first gate outright and lands Criterion C (Method) in the 1 to 2 band however good the argument is.

The line between the two is not how much you read. It is whether you took unprocessed numbers and did something to them yourself. Quoting a conclusion from a paper is a literature review. Downloading that paper's dataset and reworking it is an investigation.

Everything below is how we suggest you actually do it.

Four tests, before you spend a day on it

2 min

Most secondary investigations are decided in the first twenty minutes, by whether the thing you found is usable. These four questions take a couple of minutes each and they save you from the failure that cannot be recovered: discovering at the end that your dataset could never have answered your question.

Is it numbers?
Not a picture of numbers

A chart on a page, a figure in a paper, a PDF of a table. None of those are data yet. You need values you can put in a spreadsheet and change. If the only way to get them is to read them off an axis, keep looking.

Can you say what one row is?
Out loud, in a sentence

One station in one year. One country in one month. If you cannot finish that sentence, you do not yet know what you would be comparing, and neither will your reader.

Does it cover your years and your places?
Both, not one

A strategy introduced in 2015 needs data either side of 2015. A question about a river needs stations on that river. Check before you write the question, not after.

Can you have it in ten minutes?
Today, not in a fortnight

Some of the best environmental archives need an account and a rights request that takes two weeks to approve. That is fine for a researcher and fatal for you. If it cannot be on your screen this morning, it is not your dataset.

The fourth one is the one people argue with, and it is the one I would hold hardest. An investigation you cannot start is worth less than a smaller one you can finish, and the good archive will still be there next year.

Open it before you trust it

2 min

A downloaded file arrives looking authoritative. It has a government logo somewhere behind it, a licence, a last-updated date. None of that tells you what is actually inside, and the gap between the two is where secondary investigations quietly fail.

So before anything else: how many rows are there really, what is one row, which years are missing, and does the current year look complete? Ten minutes with a spreadsheet, and it changes what you can honestly claim.

What to check, in this order
The row count, then the count of things you actually have. They are not always the same number.
One row, named. Write the sentence down; you will need it in your method anyway.
The years present. Sort them and look for gaps. A missing year is not usually announced.
The current year. Often present and incomplete, which makes it look like a collapse.
The extremes. A value that is suspiciously round, or repeated at the top of the range, is often a reporting ceiling rather than a measurement.
The zeros. Sometimes they mean none. Sometimes they mean below the detection limit, which is a different claim.
Thirty-seven thousand rows, seven hundred observationsGeneva rivers study

The Geneva bacteriology file downloads as 37,009 rows, which feels like a serious dataset. It holds 757. Every station-year is repeated about fifty times, identical apart from an internal object id that nobody would look twice at.

The means are unaffected, because duplicating a value fifty times does not move an average. The sample size is destroyed. A t-test on the file as it downloads has an n fifty times too large, so it returns a significant result whatever is in the data, and the p-value is meaningless in a way that is invisible in the output.

It took one line of code to notice and about a minute to fix. It would have invalidated every statistical claim in the report.

For your own investigation

Before you analyse anything, count the distinct things your file describes and compare that with its row count. If the two disagree, find out why. A published dataset from a public authority is not a clean dataset, and nothing in the download says otherwise.

The mistake this protects you from

Duplicated rows inflate your sample size without changing your averages, which is the worst combination: everything looks right and the statistics are wrong. It is the one error on this route that a reader cannot detect from your report, which is exactly why saying you checked is worth doing.

Could a stranger rebuild your file?

2 min

This is what repeatable means on this route, and it is a higher bar than it sounds. Not could someone find similar data. Could someone follow your written method and end up with the identical file, without asking you a single question.

It is also easier to hit than the fieldwork version, because everything you need to record is on your screen at the moment you download. Write it down then. Reconstructing it three weeks later is how the date accessed becomes a guess.

Record thisBecause without it
The publisher and the dataset's own namea reader searches the wrong catalogue
The version, release or last-updated datethe file changes under you and nobody can tell
The indicator or layer codeone publisher has forty things with similar names
Every filter, query or search term you appliedthe same page returns a different file
The date you downloaded itthe numbers may since have been revised
The file format and how many rows arrivedyour reader cannot tell whether they got the same thing

Where the data comes through a query rather than a download button, you are in luck: paste the query. A URL that returns the data is the most repeatable method anyone in this course will ever write, because a reader re-runs it rather than reconstructing it.

The test, in one sentence

Hand your method to somebody in another class. If they can produce your exact file, you have passed the first gate. If they come back and ask you which of the four datasets you meant, you have not.

Past tense here too, exactly as on the fieldwork route. This is a record of what you did, including the dataset you tried first and abandoned. That abandoned attempt is worth a sentence: it is evidence of selection rather than of luck, and step 7 will want it.

Every number was made by somebody, somehow

2 min

The commonest weakness in a secondary investigation is treating the values as facts that were simply lying around. They are the output of a process: someone chose a method, an instrument, a sampling frequency and a set of definitions, usually for a purpose that was not yours.

Reading the methodology or metadata page is this route's equivalent of calibrating an instrument, and it is where most of your Criterion F (Evaluation) material comes from. Fifteen minutes there pays for itself twice.

How the number was madeWhat it means for your claim
Measured directly, on siteThe strongest, and still has an instrument and a detection limit
Reported by an organisation or a countrySomeone had a reason to report it, and sometimes a reason not to
Modelled or estimated from other variablesYou are analysing a model's output, not an observation
Remotely sensed from satelliteExcellent coverage, and it classifies rather than sees; plantations can read as forest
Imputed to fill a gapThe value exists because the gap did, not because anyone measured it

Two more that catch people. Values get revised: the figure you downloaded in September may not be the figure published in March, which is why the date accessed matters. And a definition can change mid-series, so a step in your graph may be a change in what was being counted rather than a change in the world.

What the Geneva figure actually isGeneva rivers study

Not a measurement. Each row is an annual mean, calculated from a number of separate sampling visits that the file records alongside it. Some years rest on twelve visits and some on one.

That matters twice. A mean from a single visit is not comparable with a mean from twelve, so the two cannot sit in the same analysis without saying so. And an annual mean of a bacterial count systematically hides the thing that causes the problem, because contamination arrives in storm peaks and a scheduled fortnightly visit mostly misses them.

For your own investigation

Find out what one value is an average of, and over what. If the dataset records that for you, use it as a filter rather than ignoring it. If it does not, that absence is itself a limitation worth naming in step 7.

The CDN approach

Fix the rule before you look

2 min

On the fieldwork route you decide where the quadrats go before you know what is in them. Here, the equivalent decision is which rows to keep, and there is nothing physically stopping you making it after you have seen which version gives the nicer graph.

That is the integrity problem specific to this route, and it is worth naming plainly: choosing your countries, your years or your stations once you can see the relationship is the same act as throwing away the quadrats that disagreed with you. It is just quieter, and nobody can tell from the finished report.

So write the rule down first, apply it, and say in the method what it removed. The IB does not ask for this in these words; it asks for a repeatable method and controlled variables, and this is how you get both at once. I would go further and put the count in: started with 757, kept 646, and here is why.

Decided afterwards

I analysed the stations that showed a clear trend.

Unrepeatable, and the finding is a product of the choosing.

Decided in advance, and reported

I kept only station-years resting on at least ten sampling visits, which removed 111 of 757 observations, because a mean from one visit is not comparable with a mean from twelve.

A reader can re-run it and get your file back.

Rules worth considering
A minimum amount of underlying data behind each value
A fixed date window, chosen because of your strategy rather than your results
Only units present throughout, so early and late are the same set of places
A geographic boundary you can defend, such as one catchment or one country group
A habit worth stealing

Write your inclusion rule at the top of the spreadsheet, in a cell, before you make a single chart. It takes twenty seconds and it turns a good intention into a record. When step 7 asks what you would do differently, that cell is where the honest answer lives.

Your controls are arithmetic now

2 min

A fieldwork investigation holds things constant physically: the same time of day, the same slope aspect, the same observer. You have none of that. What you have instead is the ability to hold things constant mathematically, which is genuinely powerful and almost never used by students.

There are three moves, and naming which one you used is what turns a comparison into a controlled one.

Normalise
Divide the difference out

Per person, per square kilometre, per unit of output. Comparing two countries' total emissions mostly compares their populations; comparing per capita emissions compares something you meant to.

Restrict
Narrow until they match

Only inland cities, only sites below 200 m, only stations on the same river. You lose sample size and you gain a comparison where the confounding variable cannot vary.

Pair
Compare like with like

The same places early and late. Upstream against downstream on the same watercourse. Pairing removes everything that is constant within a pair, which is usually most of what worries you.

Why the pairing mattered hereGeneva rivers study

The headline was that Geneva's rivers got cleaner: median 40.0 CFU/ml in 1995 to 2000 against 17.4 in 2019 to 2024. But the monitoring network grew over those thirty years, from about nineteen stations to about thirty, so the two figures describe partly different sets of places.

Restricting to the 49 stations present in both periods gives 40.0 against 17.4, and the means move from 85.4 to 27.9. The improvement survives, and now it is an improvement at the same places rather than a change of subject.

Worth doing even when it confirms what you had: an unpaired comparison that happens to be right is still a comparison you cannot defend.

For your own investigation

Ask what else changed between your two groups besides the thing you are studying, then remove it by pairing or restricting rather than by hoping. Report both versions if they differ, because the difference between them is a finding.

How much data is enough, when you chose it

2 min

Here the question changes shape. On the fieldwork route, sample size is limited by how long you can stand in a field. Here it is a decision, which means it has to be justified rather than apologised for, in both directions.

The floor is set by the test you are going to run, exactly as on the fieldwork route. Decide the test first, then count backwards to how many rows you need to keep.

The test you plan to runRows it needs
A correlation10 pairs at least, 30 to be comfortable
A t-test10 or more per group
Chi-squared5 expected in every cell
The trap that is specific to this route

There is such a thing as too much. Run a correlation on ten thousand rows and almost anything comes out statistically significant, including relationships far too weak to mean anything. A p-value answers whether an effect exists, not whether it is big enough to care about.

So when your n runs into the thousands, quote the strength alongside the significance, r or R squared, and talk about the size of the effect rather than the smallness of the p. A moderator who sees p less than 0.001 on a correlation of 0.04 knows exactly what happened.

Bigger is not automatically better in the other direction either. Adding thirty more countries to reach a threshold, when twenty of them are not comparable with your original ten, buys a number and loses the argument. Sample size that comes from widening your inclusion rule has to be defended against the rule, not against the total.

What was enough here, and what was notGeneva rivers study

Seven hundred and fifty-seven station-years sounds enormous. For the comparison that actually got made, 1995 to 2000 against 2019 to 2024 on stations present in both, it came down to 49 places and 157 observations, which is comfortable for the test and nowhere near thousands.

Everything else was excluded by rules set in advance, and saying so is the difference between a sample and a selection.

For your own investigation

Count the rows that actually enter your analysis, not the rows in the file. That smaller number is your real sample size, it is the one your statistics rest on, and it belongs in the method where a reader can see it.

Ethics, when nobody gets wet

1 min

There is no risk assessment to write here and no consent form to hand anybody, which tempts students to skip the section entirely. Two things belong in it instead, and one of them is the most serious integrity question on this route.

The first is ordinary and quick. Say whose data it is, under what licence you may use it, and credit them properly. Open does not mean unattributed, most licences require the source to be named, and a dataset gets cited like any other source with the addition of the date you accessed it.

The second is the one that matters. Selective use of data is misconduct, not untidiness. Quietly dropping the years that spoil the pattern, or presenting a filtered subset as though it were the whole, is the same offence as inventing a reading. The protection is the previous beat: a rule decided in advance, applied to everything, and reported with what it removed.

What a dataset citation carries
Who published it, and the dataset's own title
The version, release or last-updated date
The URL you actually used
The date you downloaded it
The licence, where one is stated
One more, if people are in your data

Some open datasets describe individuals rather than places: health records, survey microdata, individual incomes. Those come with conditions, and the conditions are not optional because the file downloaded easily. If a dataset would let you identify a person, it does not belong in a school investigation.

Using AI at this stepLevel 3 · Targeted AI

It can explain what an indicator measures, what a licence permits, or how to write the spreadsheet formula that turns your raw column into the one you need. Ask it as many times as you like.

It cannot give you the numbers, or tell you a dataset exists. Both failures are silent here in a way they never are in a field: a chatbot will recite a plausible figure and invent a plausible repository, and neither announces itself the way a missing quadrat would. Every value comes out of a file you downloaded, and every dataset is one you have opened.

What this level means

Ready for step 5?

0 of 14

The first two are the criterion itself. Numbers 17 and 18 are the two that separate a defensible secondary investigation from a plausible one, and neither takes more than a few minutes.

Next: step 5, treat your data

You have a file and a rule for what is in it. Step 5 is where you turn its columns into something that answers your question, and where the good news arrives: tables, graphs and calculations all sit outside the word count. Decide your statistical test before you start processing, because it sets how much data you needed to keep.

The Geneva rivers investigation used on the secondary-data route of this guide is the author’s own analysis of a published cantonal dataset. Every figure quoted comes from the Canton of Geneva’s open data, downloaded on 10 September 2026; the duplication, the missing year and the varying campaign counts described in the guide are as found on that date and may since have been corrected. The selection rules, the analysis and the conclusions are the author’s and not the canton’s.