Secondary data routeYou are on the secondary data route: someone else collected it, you analyse it.Change on the start page
Criterion A: Research question and inquiry · 4 marks · shared with step 2500 words · suggestedSecondary data

Find your issue, and check it is measured

You arrive with

Nothing, or a vague sense of what interests you. That is the normal starting position and it is fine.

You leave with
  • An environmental issue you can state in a sentence
  • Proof that somebody measures it: one file, opened, not a page saying it exists
  • A rough sketch of the system involved
  • A note of who publishes those numbers, and where they live

You are not choosing your dataset here. You are finding out whether anybody has one.

Most advice tells you to choose a topic. The trouble is an issue that nothing measures: fine to care about, impossible to investigate. So this step ends with a file open, which is a check rather than a commitment. Step 4 is where a file becomes the dataset you can defend.

You cannot leave this step without all three

An issue
Not a topic

Something that is happening, to something, because of something. Water quality is a topic. Thirty years of investment in a sewer network, and whether the bacteria in the rivers noticed, is an issue.

Proof it is measured
One file, open in front of you

Numbers you can change in a spreadsheet, covering the years and the places your issue is about. Not a page saying the data exists. The file. You are not choosing your final dataset yet, and you may open three before one survives: that work is step 4.

A system
What flows through it

What enters, what leaves, what accumulates, what loops back. It carries weight here, because nothing about how you got your numbers shows that you understand them.

Everything below is how we suggest you actually do it.

The CDN approach

Choose the issue and the data at the same time

1 min

Choose a place and a file, and neither one without the other. Doing them in sequence is what costs students a fortnight.

Most environmental issues are not measured with the depth you need, near the place you care about, for long enough to see anything. Choose an issue first and look for data afterwards, and you can lose a week finding that out.

So work at both ends at once. Have an issue you can state in a sentence, and before you settle on it have a file open that describes it. If either one refuses to arrive, change the other rather than pushing on.

Issue first, then hunting

I want to investigate microplastics in the Rhône. Several hours later: a handful of one-off studies, a news article, and no series anyone can download.

The issue was fine. Nothing measures it in a form you can use.

Both at once

I want to investigate whether Geneva's rivers got cleaner. Twenty minutes later: thirty years of bacterial counts, station by station, one click, no account.

The issue survived contact with what exists.

This is not letting a dataset decide what you care about: the catalogue is organised by issue, so choosing a file is still choosing an issue. Just do not fall in love with a question before you know anyone has measured it.

Issues that are in the news

Twenty-two real arguments to start from

Each one is happening now, somebody is doing something about it, and opening one says what you could actually put in a spreadsheet. None of it is proof that the data exists: that is what the gallery is for.

In the news over the last
Loading headlines…
Water pollution
Rivers and lakes
Biodiversity
Air quality
Climate and weather
Ice and snow
Land and forests
Food and farming
Energy and emissions
People and development
Headlines from the Guardian.Now find out whether anyone measures it

The issue has to reach an ecosystem

1 min

Humans are part of the environment, so an investigation involving people is not off-syllabus. What Criterion A (Research question and inquiry) asks is that you identify an environmental issue and its impact on natural systems, and that people are not the whole of it.

The IB's own example of what does not work is the correlation between GDP and electricity usage: economics, with no environmental issue in it.

So apply one test before you commit to a pair of columns: does at least one side of your comparison describe something living, or a physical part of the environment? If both sides describe people and money, you have a social science investigation with an environmental word in the title.

Two columns about people

GDP per capita against carbon dioxide emissions, for sixty countries.

A real relationship, and nothing in it is an ecosystem.

One column about the environment

Nitrate concentration in a river, before and after the fertiliser rules in its catchment changed.

People caused it. The thing measured is the water.

The rescue is usually small. Keep the human driver you were interested in, and put an environmental measurement on the other side of the comparison. The driver explains why the pattern exists; the environmental column is what makes it an ESS investigation.

Does anyone measure this?

2 min

Three questions, and they take about ten minutes between them. They are what sinks investigations on this route, and all three can be answered before you have committed to anything.

At what depth?
The unit you want to compare

You want stations and they publish national totals. You want months and they publish decades. A national annual figure has already had the variation you were going to analyse averaged out of it.

For which years?
Both sides of your strategy

A strategy introduced in 2015 needs data before and after 2015. A series that begins in 2018 cannot tell you what it changed, however good it is.

Where you care about?
Not somewhere like it

Plenty of things are measured beautifully somewhere else. If the monitoring stops at a border and your issue does not, that is a limitation you will be writing about for the rest of the report.

If all three answers are yes, you have an investigation. If one is no, it is usually easier to move the question a little than to find different data: the same file will often answer a neighbouring question perfectly well.

Secondary data is not second best

The mark scheme says data may be primary or secondary, the criteria are identical, and this route reaches the same top band.

What it is not is a literature review. Summarise what other people concluded, rather than working something out from their numbers, and the method criterion caps at the bottom band.

What everything is called
A phrasebook for finding data. One thing has four names across four repositories, so you search for "raw data", find nothing, and conclude it does not exist. Each entry gives the other names and says why it matters.
Open the lexicon
Datasets that have been opened and counted
The catalogue, searchable by issue and ordered from Geneva outwards to global. Each entry says what one row is, and the counted ones say what went wrong when somebody actually downloaded them.
Browse the datasets

You can go bigger, and it costs something

1 min

The real gift of this route is scale. One download can hold fifty years, or fifteen cities, or a whole river basin, and questions that would take a research group a decade to measure are open to you this afternoon.

The cost arrives at step 3. A strategy is always somebody's, operating somewhere, and the bigger your question gets the harder it is to find one that acts on all of it. Ask about deforestation across five countries and the honest answer to "whose plan is this?" is that there are five of them, or none.

Match the scale of your question to the scale a strategy operates at: national law, a river-basin agreement, a cantonal programme, a city policy. Sit inside one rather than straddling several.

A question at this scaleWants a strategy of this kind
One catchment, one cantonA municipal or cantonal programme
A river that crosses a borderA treaty or a basin-wide agreement
One country over decadesNational legislation
Fifteen countries comparedSomething international, and it will be vaguer

Smaller is usually the better trade. A question about one canton has a named plan, a named authority and people who have publicly disagreed about it, which is three of the four things Criterion B (Strategy) is asking for.

Sketch the system

1 min

The job is to show that your variable sits inside something larger, with things entering, leaving, accumulating and looping back.

Your numbers arrived without you, so nothing about how you got them shows that you understand the thing they describe. A systems diagram is the cheapest proof, and diagrams sit outside the word count.

The system here, in one lineGeneva rivers study

Rain falls on roofs and roads and runs into sewers. Where the sewer carries foul and surface water together, heavy rain overloads it and some of the contents reach a river. The rivers carry it to the lake. Sunlight, dilution and time reduce it on the way, which is why the receiving water can be clean while the tributaries are not.

Drawn as a labelled flow with the sewer network at the top, that is one figure and no words at all. It also makes the strategy legible, because separating foul water from surface water acts on exactly the arrow the diagram puts at the top.

For your own investigation

Draw the path your measured quantity takes to get to the place it is measured. On this route that path is the thing you did not see for yourself, and putting it in a diagram is what turns a downloaded column into an environmental system.

Background, and the document nobody reads

1 min

The rule: every paragraph is there because your question needs it, and a general essay about the subject followed by an unrelated question costs marks on both strands at once.

One source almost nobody cites, and should: your dataset's methodology page or codebook, which says how the numbers were produced, by whom and with what definitions. Reading it is where most of your eventual evaluation comes from.

What the documentation was worth hereGeneva rivers study

The Geneva dataset's own record says the bacteriological method is the one normally used for surveillance of bathing water, and the file carries a column recording how many sampling visits each annual figure rests on.

Two sentences of documentation, and between them they decided the whole method: they are why the values are comparable with bathing-water thresholds at all, and why station-years resting on a single visit had to be filtered out before anything could be compared.

For your own investigation

Read your dataset's documentation before you write your background, and cite it. It is the one source that is definitely about your data rather than about your topic, and the questions it answers are the ones step 7 will ask you.

Using AI at this stepLevel 1 · AI Planning

It can help you get your bearings on an unfamiliar issue and suggest what a dataset of the kind you want is usually called, which is often the only thing standing between you and finding it.

It cannot tell you that a dataset exists. It names plausible repositories with complete confidence. Every dataset you mention is one you have opened.

What this level means

Ready for step 2?

Secondary data checklist0 of 14

The one about having opened the data, rather than found a page saying it exists, decides whether the next seven steps are possible at all. It is a feasibility check, not your method: counting the file, writing the protocol and fixing an inclusion rule are step 4. The three about a natural system are what stops a pair of downloaded columns being a social science investigation.

Next: step 2, write your research question

Step 2 turns your issue into something your file can answer. On this route the sentence has to name the indicator and its units, because those are things you chose from a list rather than things you decided to go and measure.

The Geneva rivers investigation used on the secondary-data route of this guide is the author’s own analysis of a published cantonal dataset. Every figure quoted comes from the Canton of Geneva’s open data, downloaded on 10 September 2026; the duplication, the missing year and the varying campaign counts described in the guide are as found on that date and may since have been corrected. The selection rules, the analysis and the conclusions are the author’s and not the canton’s.