Find your issue, and check it is measured
Nothing, or a vague sense of what interests you. That is the normal starting position and it is fine.
- An environmental issue you can state in a sentence
- Proof that somebody measures it: one file, opened, not a page saying it exists
- A rough sketch of the system involved
- A note of who publishes those numbers, and where they live
You are not choosing your dataset here. You are finding out whether anybody has one.
Most advice tells you to choose a topic. The trouble is an issue that nothing measures: fine to care about, impossible to investigate. So this step ends with a file open, which is a check rather than a commitment. Step 4 is where a file becomes the dataset you can defend.
You cannot leave this step without all three
Something that is happening, to something, because of something. Water quality is a topic. Thirty years of investment in a sewer network, and whether the bacteria in the rivers noticed, is an issue.
Numbers you can change in a spreadsheet, covering the years and the places your issue is about. Not a page saying the data exists. The file. You are not choosing your final dataset yet, and you may open three before one survives: that work is step 4.
What enters, what leaves, what accumulates, what loops back. It carries weight here, because nothing about how you got your numbers shows that you understand them.
Everything below is how we suggest you actually do it.
Choose the issue and the data at the same time
Choose a place and a file, and neither one without the other. Doing them in sequence is what costs students a fortnight.
Most environmental issues are not measured with the depth you need, near the place you care about, for long enough to see anything. Choose an issue first and look for data afterwards, and you can lose a week finding that out.
So work at both ends at once. Have an issue you can state in a sentence, and before you settle on it have a file open that describes it. If either one refuses to arrive, change the other rather than pushing on.
I want to investigate microplastics in the Rhône. Several hours later: a handful of one-off studies, a news article, and no series anyone can download.
The issue was fine. Nothing measures it in a form you can use.
I want to investigate whether Geneva's rivers got cleaner. Twenty minutes later: thirty years of bacterial counts, station by station, one click, no account.
The issue survived contact with what exists.
This is not letting a dataset decide what you care about: the catalogue is organised by issue, so choosing a file is still choosing an issue. Just do not fall in love with a question before you know anyone has measured it.
Twenty-two real arguments to start from
Each one is happening now, somebody is doing something about it, and opening one says what you could actually put in a spreadsheet. None of it is proof that the data exists: that is what the gallery is for.
The issue has to reach an ecosystem
Humans are part of the environment, so an investigation involving people is not off-syllabus. What Criterion A (Research question and inquiry) asks is that you identify an environmental issue and its impact on natural systems, and that people are not the whole of it.
The IB's own example of what does not work is the correlation between GDP and electricity usage: economics, with no environmental issue in it.
So apply one test before you commit to a pair of columns: does at least one side of your comparison describe something living, or a physical part of the environment? If both sides describe people and money, you have a social science investigation with an environmental word in the title.
GDP per capita against carbon dioxide emissions, for sixty countries.
A real relationship, and nothing in it is an ecosystem.
Nitrate concentration in a river, before and after the fertiliser rules in its catchment changed.
People caused it. The thing measured is the water.
The rescue is usually small. Keep the human driver you were interested in, and put an environmental measurement on the other side of the comparison. The driver explains why the pattern exists; the environmental column is what makes it an ESS investigation.
Does anyone measure this?
Three questions, and they take about ten minutes between them. They are what sinks investigations on this route, and all three can be answered before you have committed to anything.
You want stations and they publish national totals. You want months and they publish decades. A national annual figure has already had the variation you were going to analyse averaged out of it.
A strategy introduced in 2015 needs data before and after 2015. A series that begins in 2018 cannot tell you what it changed, however good it is.
Plenty of things are measured beautifully somewhere else. If the monitoring stops at a border and your issue does not, that is a limitation you will be writing about for the rest of the report.
If all three answers are yes, you have an investigation. If one is no, it is usually easier to move the question a little than to find different data: the same file will often answer a neighbouring question perfectly well.
The mark scheme says data may be primary or secondary, the criteria are identical, and this route reaches the same top band.
What it is not is a literature review. Summarise what other people concluded, rather than working something out from their numbers, and the method criterion caps at the bottom band.
You can go bigger, and it costs something
The real gift of this route is scale. One download can hold fifty years, or fifteen cities, or a whole river basin, and questions that would take a research group a decade to measure are open to you this afternoon.
The cost arrives at step 3. A strategy is always somebody's, operating somewhere, and the bigger your question gets the harder it is to find one that acts on all of it. Ask about deforestation across five countries and the honest answer to "whose plan is this?" is that there are five of them, or none.
Match the scale of your question to the scale a strategy operates at: national law, a river-basin agreement, a cantonal programme, a city policy. Sit inside one rather than straddling several.
| A question at this scale | Wants a strategy of this kind |
|---|---|
| One catchment, one canton | A municipal or cantonal programme |
| A river that crosses a border | A treaty or a basin-wide agreement |
| One country over decades | National legislation |
| Fifteen countries compared | Something international, and it will be vaguer |
Smaller is usually the better trade. A question about one canton has a named plan, a named authority and people who have publicly disagreed about it, which is three of the four things Criterion B (Strategy) is asking for.
Sketch the system
The job is to show that your variable sits inside something larger, with things entering, leaving, accumulating and looping back.
Your numbers arrived without you, so nothing about how you got them shows that you understand the thing they describe. A systems diagram is the cheapest proof, and diagrams sit outside the word count.
Rain falls on roofs and roads and runs into sewers. Where the sewer carries foul and surface water together, heavy rain overloads it and some of the contents reach a river. The rivers carry it to the lake. Sunlight, dilution and time reduce it on the way, which is why the receiving water can be clean while the tributaries are not.
Drawn as a labelled flow with the sewer network at the top, that is one figure and no words at all. It also makes the strategy legible, because separating foul water from surface water acts on exactly the arrow the diagram puts at the top.
Draw the path your measured quantity takes to get to the place it is measured. On this route that path is the thing you did not see for yourself, and putting it in a diagram is what turns a downloaded column into an environmental system.
Background, and the document nobody reads
The rule: every paragraph is there because your question needs it, and a general essay about the subject followed by an unrelated question costs marks on both strands at once.
One source almost nobody cites, and should: your dataset's methodology page or codebook, which says how the numbers were produced, by whom and with what definitions. Reading it is where most of your eventual evaluation comes from.
The Geneva dataset's own record says the bacteriological method is the one normally used for surveillance of bathing water, and the file carries a column recording how many sampling visits each annual figure rests on.
Two sentences of documentation, and between them they decided the whole method: they are why the values are comparable with bathing-water thresholds at all, and why station-years resting on a single visit had to be filtered out before anything could be compared.
Read your dataset's documentation before you write your background, and cite it. It is the one source that is definitely about your data rather than about your topic, and the questions it answers are the ones step 7 will ask you.
It can help you get your bearings on an unfamiliar issue and suggest what a dataset of the kind you want is usually called, which is often the only thing standing between you and finding it.
It cannot tell you that a dataset exists. It names plausible repositories with complete confidence. Every dataset you mention is one you have opened.
Ready for step 2?
Secondary data checklist0 of 14The one about having opened the data, rather than found a page saying it exists, decides whether the next seven steps are possible at all. It is a feasibility check, not your method: counting the file, writing the protocol and fixing an inclusion rule are step 4. The three about a natural system are what stops a pair of downloaded columns being a social science investigation.
Step 2 turns your issue into something your file can answer. On this route the sentence has to name the indicator and its units, because those are things you chose from a list rather than things you decided to go and measure.
The Geneva rivers investigation used on the secondary-data route of this guide is the author’s own analysis of a published cantonal dataset. Every figure quoted comes from the Canton of Geneva’s open data, downloaded on 10 September 2026; the duplication, the missing year and the varying campaign counts described in the guide are as found on that date and may since have been corrected. The selection rules, the analysis and the conclusions are the author’s and not the canton’s.
