Find your issue and your data
Nothing, or a vague sense of what interests you. That is the normal starting position and it is fine. What is different on this route is that you should expect to leave with a file open, not just an idea.
- An environmental issue you can state in a sentence
- A dataset you have opened, that describes it
- A rough sketch of the system involved
- The dataset's own documentation, read and cited
Most advice tells you to choose a topic. On this route, choose a place and a file, in the same afternoon.
The fieldwork route says choose somewhere you can stand. Here the binding constraint is different: most environmental issues are not measured at the grain you need, near the place you care about, for long enough to see anything. Working at both ends at once is what stops you discovering that two days in.
You cannot leave this step without all three
Something that is happening, to something, because of something. Water quality is a topic. Thirty years of investment in a sewer network, and whether the bacteria in the rivers noticed, is an issue.
Numbers you can change in a spreadsheet, covering the years and the places your issue is about. Not a page saying the data exists. The file.
What enters, what leaves, what accumulates, what loops back. It matters more here than on the fieldwork route, because nothing about how you got your numbers shows that you understand them.
Everything below is how we suggest you actually do it.
Choose the issue and the data in the same afternoon
The fieldwork route says choose a place before you choose a method, and means it. This route needs a second half of that sentence, and it is the one piece of advice here that will save you the most time.
The trouble with choosing an issue first and looking for data afterwards is that most environmental issues are not measured. Not measured at the grain you need, not measured near the place you care about, not measured for long enough to see the thing you want to see. You can spend two days finding that out.
So work at both ends at once. Have an issue you can state in a sentence, and inside the same afternoon have a file open that describes it. If either one refuses to arrive, change the other rather than pushing on.
I want to investigate microplastics in the Rhône. Two days later: a handful of one-off studies, a news article, and no series anyone can download.
The issue was fine. Nothing measures it in a form you can use.
I want to investigate whether Geneva's rivers got cleaner. Twenty minutes later: thirty years of bacterial counts, station by station, one click, no account.
The issue survived contact with what exists.
This is not the same as letting a dataset decide what you care about. The gallery is organised by environmental issue rather than by publisher precisely so you choose an issue while you choose a file. What you must not do is fall in love with a question before you know whether anyone has ever measured it.
Does anyone measure this?
Three questions, and they take about ten minutes between them. On the fieldwork route the equivalent question is whether you can reach the site repeatedly; this is the version that sinks investigations here.
You want stations and they publish national totals. You want months and they publish decades. A national annual figure has already had the variation you were going to analyse averaged out of it.
A strategy introduced in 2015 needs data before and after 2015. A series that begins in 2018 cannot tell you what it changed, however good it is.
Plenty of things are measured beautifully somewhere else. If the monitoring stops at a border and your issue does not, that is a limitation you will be writing about for the rest of the report.
If all three answers are yes, you have an investigation. If one is no, it is usually easier to move the question a little than to find different data: the same file will often answer a neighbouring question perfectly well.
Worth saying plainly, because students assume otherwise. The mark scheme states that data may be primary or secondary, the criteria are identical, and this route reaches the same top band.
What it is not is a literature review. If your report ends up summarising what other people concluded rather than working out something yourself from their numbers, the method criterion caps at the bottom band. Step 4 is where that line gets drawn properly.
You can go bigger, and it costs something
The real gift of this route is scale. A morning with a quadrat frame gets you one hillside; a download gets you fifty years, or fifteen cities, or a whole river basin. Questions that were simply unavailable to you are now open.
The cost arrives at step 3. A strategy is always somebody's, operating somewhere, and the bigger your question gets the harder it is to find one that acts on all of it. Ask about deforestation across five countries and the honest answer to "whose plan is this?" is that there are five of them, or none.
So match the scale of the question to the scale a strategy actually operates at. National law, a river-basin agreement, a cantonal programme, a city policy: each of those has a natural extent, and your question wants to sit inside one of them rather than straddle several.
| A question at this scale | Wants a strategy of this kind |
|---|---|
| One catchment, one canton | A municipal or cantonal programme |
| A river that crosses a border | A treaty or a basin-wide agreement |
| One country over decades | National legislation |
| Fifteen countries compared | Something international, and it will be vaguer |
Smaller is usually the better trade. A question about one canton has a named plan, a named authority and people who have publicly disagreed about it, which is three of the four things Criterion B (Strategy) is asking for.
Sketch the system
Exactly as on the fieldwork route, and for the same reason: show that your variable sits inside something larger, with things entering, leaving, accumulating and looping back.
It matters more here, not less. Your numbers arrived without you, so nothing about how you got them tells a reader that you understand the thing they describe. A systems diagram is the cheapest way to prove you do, and diagrams sit outside the word count while two hundred words of systems prose do not.
Rain falls on roofs and roads and runs into sewers. Where the sewer carries foul and surface water together, heavy rain overloads it and some of the contents reach a river. The rivers carry it to the lake. Sunlight, dilution and time reduce it on the way, which is why the receiving water can be clean while the tributaries are not.
Drawn as a labelled flow with the sewer network at the top, that is one figure and no words at all. It also makes the strategy legible, because separating foul water from surface water acts on exactly the arrow the diagram puts at the top.
Draw the path your measured quantity takes to get to the place it is measured. On this route that path is the thing you did not see for yourself, and putting it in a diagram is what turns a downloaded column into an environmental system.
Background, and the document nobody reads
The same rule as the fieldwork route: every paragraph is there because your question needs it, and a general essay about the subject followed by an unrelated question costs marks on both strands at once.
This route adds one source that students almost never cite and should. Your dataset has a methodology page, a metadata record or a codebook, saying how the numbers were produced, by whom, how often and with what definitions. It is background research about the thing you are actually going to analyse, it is traceable, and reading it is where most of your eventual evaluation comes from.
The Geneva dataset's own record says the bacteriological method is the one normally used for surveillance of bathing water, and the file carries a column recording how many sampling visits each annual figure rests on.
Two sentences of documentation, and between them they decided the whole method: they are why the values are comparable with bathing-water thresholds at all, and why station-years resting on a single visit had to be filtered out before anything could be compared.
Read your dataset's documentation before you write your background, and cite it. It is the one source that is definitely about your data rather than about your topic, and the questions it answers are the ones step 7 will ask you.
It can help you get your bearings on an unfamiliar issue and suggest what a dataset of the kind you want is usually called, which is often the only thing standing between you and finding it.
It cannot tell you that a dataset exists. It will name plausible repositories and plausible indicators with complete confidence and no way for you to tell which are real. Every dataset you mention is one you have opened, and every source you cite is one you read.
Ready for step 2?
0 of 12The last six are this route's own. Number 9 is the one that decides whether the next seven steps are possible at all.
Step 2 turns your issue into something your file can answer. On this route the sentence has to name the indicator and its units, because those are things you chose from a list rather than things you decided to go and measure.
The Geneva rivers investigation used on the secondary-data route of this guide is the author’s own analysis of a published cantonal dataset. Every figure quoted comes from the Canton of Geneva’s open data, downloaded on 10 September 2026; the duplication, the missing year and the varying campaign counts described in the guide are as found on that date and may since have been corrected. The selection rules, the analysis and the conclusions are the author’s and not the canton’s.
