Secondary data routeEvery step is showing the version for data somebody else collected. Switch to primary data
Criterion E: Analysis and conclusion · 6 marks550 words · suggested

Analysis and conclusion

You arrive with

Processed data, figures you built yourself, and a statistical result. No interpretation yet, because that is the six marks on this page.

You leave with
  • Every pattern described and then explained through the system
  • Two or three ways your data source could mislead, each with a direction
  • A conclusion in numbers that reaches exactly as far as your rows
  • A return to the issue and the strategy

Six marks for saying what your numbers mean, and one extra job this route cannot skip.

The ladder is describe, explain, reconnect, exactly as on the fieldwork route. What this route has to do as well is be honest about its instrument: you controlled nothing, you measured nothing, and the collection process that produced your values had its own purposes and its own failures. Naming those precisely is where most of the marks on this page actually are.

A strong analysis does three things, in order.

Describe
What the data shows

Name the pattern precisely. Which direction, how strong, and whether the statistics say it is real.

On its own, this is the lower band.

Explain
Why it shows that

What process produced the pattern? This is where your abiotic measurements earn their place, because they are the mechanism.

This is what moves you up.

Reconnect
What it means for the issue

Take the answer back to the environmental issue and the strategy you wrote about in step 3.

This is where the work stays ESS.

Everything below is how we suggest you actually do it.

Four things, one narrative

1 min

Bias, reliability, validity and uncertainty all have to be here. The common mistake is to give each one its own labelled paragraph, which makes them easy for an examiner to find and impossible to read.

Weave them instead. Take one trend, connect it to your question, then work through how far you trust it and why. It reads as thinking rather than as a checklist, and it sets up step 7 without you having to change gear.

What each one asks
Reliability, If someone repeated this without you there, would they get similar results? Weather between sampling days, drifting equipment, sites that changed.
Validity, Are you measuring what you think you are? Species richness is not species evenness, and calling either one 'biodiversity' hides the difference.
Uncertainty, Your instrument's precision, against the size of the difference you found. ±4% matters when the gap is 5%.
Bias, Who chose the sites, who estimated the values, and what would have made them consistently wrong in one direction.

Describe, explain, reconnect

2 min

The same ladder as the fieldwork route, and this route slips off the bottom rung more easily. A downloaded series describes itself: the line goes down, the difference is significant, the correlation is 0.6. It is very easy to write two pages that report a file back to a reader.

Explaining means naming the process that produced the pattern, in the environmental system you drew in step 1. Not that the numbers fell, but what happened, to what, that made them fall. That is where your systems diagram stops being decoration.

You also have a harder job than the fieldwork route on causation, and you should say so rather than hoping nobody notices. You controlled nothing physically. Everything you have is association, and the honest move is to name the mechanism you think is operating, then name the other things that could have produced the same pattern.

Describing, then explainingGeneva rivers study

Describing: median E. coli at the same 49 stations fell from 40.0 to 17.4 CFU/ml between 1995 to 2000 and 2019 to 2024.

Explaining: separating foul water from surface water reduces the number of occasions when heavy rain pushes untreated sewage into a watercourse. Fewer overflow events means fewer large inputs of faecal bacteria, and because the indicator is an annual mean of scheduled visits, removing the rare very large values moves the mean a long way.

And then the honesty: thirty years also brought stricter agricultural rules, changes in land use, and a different laboratory method more than once. The sewer programme is a candidate explanation, not a demonstrated cause, and the data cannot separate it from the others.

For your own investigation

Write the mechanism as a chain of physical events, then immediately name what else could have produced the same numbers. On this route that second sentence is not weakness, it is the thing that separates analysis from a press release.

Five ways a dataset misleads you

3 min

The beat above asks you to discuss bias, reliability, validity and uncertainty, and on this route those words point somewhere specific. Your instrument was somebody else's collection process, and these are its five characteristic failures. Naming one precisely, and saying which way it would push your result, is worth more than a paragraph about human error.

Proxy validity
Does it measure what you think?

E. coli is an indicator of faecal contamination, not a measure of harm. GDP per capita is not environmental stewardship. Ask what your indicator actually is, and what you are quietly claiming it stands for.

Aggregation
The ecological fallacy

A national figure can hide severe local damage; a site mean can hide the storm peaks that cause the problem. A pattern true of the aggregate need not be true of anything inside it.

Temporal misalignment
Different clocks

A census every five years against monitoring every month. An annual mean against a strategy that started in June. If your two series are not on the same clock, saying so is part of the analysis.

Reporting bias
Who produced it, and why

Figures submitted by the party being judged by them. Under-reporting where reporting is costly, absence where monitoring is absent. Missing data is rarely missing at random.

Classification error
What the instrument confuses

Remote sensing reads plantations as forest. An automated category groups things you would separate. Every classified dataset has a confusion it is known for, and its documentation usually says so.

The test for whether you have done this properly: can you say which direction each one would push your result? "Reporting bias may exist" is the bottom band. "Countries with weaker monitoring report lower values, which would flatten the relationship I found rather than create it" is the top one, because it tells a reader what to do with your conclusion.

Three of the five, in this datasetGeneva rivers study

Aggregation: each value is an annual mean of about twelve scheduled visits, and contamination arrives in storm peaks. The dataset systematically under-samples the mechanism it is measuring, which means the real difference between the periods is probably larger than the one measured, not smaller.

Proxy validity: E. coli indicates faecal contamination. It does not distinguish human sewage from agricultural runoff, and the strategy under discussion acts on one of those and not the other.

Temporal misalignment: the sewer programme has no single start date. It is a rolling network upgrade over thirty years, so there is no before and after to test, only two windows chosen by the student.

For your own investigation

Pick the two or three that genuinely apply to your data rather than listing all five, and for each one say which way it moves your answer. A limitation with a direction is analysis; a limitation without one is a disclaimer.

A conclusion that claims exactly what you found

2 min

Answer your research question, in numbers, in a couple of sentences. Nothing may appear here that was not in the analysis above it, and if you stated a hypothesis you must say whether the data supports it.

Then the discipline this route needs most: your conclusion reaches exactly as far as the rows you kept. Not the whole country, if you kept one canton. Not the present, if your series ends in 2024. Not all rivers, if you analysed the ones somebody chose to monitor.

Claims more than it can

Geneva's investment in its sewer network has successfully cleaned up its rivers.

Three claims, and the data supports part of one.

Claims what was found

Median E. coli at 49 monitoring stations fell from 40.0 to 17.4 CFU/ml between 1995 to 2000 and 2019 to 2024. The improvement is not uniform: 15 of the 49 stations were worse in the later period. The sewer programme is one candidate explanation among several that the data cannot separate.

A reader knows exactly what has and has not been shown.

Three conclusions this data cannot support

"The rivers are clean." Nothing here measured cleanliness. It measured one indicator organism, at scheduled intervals, at monitored places.

"The strategy worked." The study compared two periods. It did not compare places with the strategy against places without it, which is the comparison that would license the claim.

"Water quality is improving." True of the median and false at nearly a third of the stations, which is not a detail: it is a different finding.

Then travel back

1 min

The last thing your conclusion should do is return to where step 1 started: the environmental issue, and the strategy people are arguing about. This is the reconnection that makes it an ESS investigation rather than a piece of data analysis, and on this route it is the only place where the two halves of your report touch.

It does not need many words. What did your numbers add to the argument you described in step 3? Whose position do they support, whose do they complicate, and what would each side say about them?

What the numbers did to the argumentGeneva rivers study

The overall improvement supports the canton's case that thirty years of network investment was worth the money. The 15 worsening stations complicate it, and several of those sit downstream of infrastructure on the French side, which is an argument for the cross-border agreements rather than against the cantonal ones.

Neither side gets to claim the finding whole, which is the most useful thing a school investigation can say about a real disagreement.

For your own investigation

Say what your result does to each position you described in step 3, including the one you expected to lose. A finding that inconveniences everybody slightly is usually the honest one, and it is far more interesting to read than a verdict.

Using AI at this stepLevel 0 · No AI

It can do nothing here. If a general idea needs clarifying, do that before you open this section, not inside it.

It cannot interpret your results. Explaining your own pattern is what the six marks are for, and a tool that cannot see your file will produce a paragraph that sounds like analysis and contains none.

What this level means

Ready for step 7?

0 of 15

The last four belong to this route. Number 13 is the one that separates the top band from the middle: a limitation with a direction is analysis, one without is a disclaimer.

Next: step 7, evaluation

You have just written about bias, reliability, validity and uncertainty. Step 7 takes the same material and does something different with it: not how far your data can be trusted, but what you would change and what it would fix. The bridge is already built.

The Geneva rivers investigation used on the secondary-data route of this guide is the author’s own analysis of a published cantonal dataset. Every figure quoted comes from the Canton of Geneva’s open data, downloaded on 10 September 2026; the duplication, the missing year and the varying campaign counts described in the guide are as found on that date and may since have been corrected. The selection rules, the analysis and the conclusions are the author’s and not the canton’s.