Secondary data routeEvery step is showing the version for data somebody else collected. Switch to primary data
Criterion F: Evaluation · 6 marks600 words · suggested

Evaluation

You arrive with

A conclusion, and the discussion of how your data source could mislead that you wrote in step 6. Most of the raw material for this section is already on the page; what changes is what you do with it.

You leave with
  • Specific weaknesses of your data, ranked by how much they mattered
  • Your own selection rule, evaluated
  • An improvement for each, naming something real
  • Questions your investigation opened but could not close

Six marks for saying what was wrong with data you did not collect.

Which sounds like an odd thing to be marked on, until you notice how much of it was your decision. You chose the source, the indicator, the years, the places and the rule that kept some rows and dropped others. Every one of those is a methodological choice, and this is where you weigh them.

The bands turn on whether your weaknesses are generic or specific. The guide even defines generic for you: general to many methodologies, and not specifically relevant to the one you used. Which means the fix is mechanical. Take each sentence and ask whether it could have been written by someone who had never opened your file.

Generic · 1–2

The data may not be completely accurate.

Specific · 5–6

Each value is an annual mean of about twelve scheduled visits, so the dataset under-samples storm events, which is when contamination actually arrives. That flattens the difference between my two periods rather than creating it, so the improvement I found is likely an underestimate.

Generic · 1–2

More data would have improved reliability.

Specific · 5–6

The monitoring network grew from 19 stations to about 30 across the period, so my early and late groups describe partly different places. Restricting to the 49 present in both changed the median from 18.2 to 17.4, which is small enough to say the growth is not driving the result.

Generic · 1–2

The dataset may contain errors.

Specific · 5–6

The file publishes every station-year about fifty times over. I removed the duplicates before analysis, which left 757 observations rather than 37,009. Had I not, every significance test in this report would have been meaningless while looking entirely normal.

Everything below is how we suggest you actually do it.

Three strands, one chain

2 min

The commonest structural mistake is to write three separate lists: weaknesses here, improvements there, questions at the end, none of them speaking to each other. The descriptors join them up explicitly. Improvements must address the limitations you identified. Questions must bear on your conclusion.

Write it as a chain and the marks follow the structure. Take your two or three most significant weaknesses and give each one a full run: what it was, what it did to your answer, what you would change, and what that change would buy you.

The weakness

Specific to your method. And what it did to your conclusion, not just that it existed.

The improvement

Addressing that named weakness. Realistic for a student with a school's equipment.

What remains

A question with a different focus from your original, and a bearing on your answer.

If your data came from people

Questionnaires carry three built-in weaknesses the IB expects you to know about, and naming them precisely beats gesturing at "possible bias". Participation was voluntary, so the people who answered are not a fair sample of the people you asked; strong opinions turn up more often than mild ones. Respondents tend to avoid the two ends of a Likert scale, which squeezes your data toward the middle. And a long or complicated survey produces fatigue, so the later answers are worse than the early ones.

Each of these earns marks only when you connect it to your conclusion: which of your results would move, and in which direction, if the missing voices had answered?

The weaknesses are in the data, not in your hands

2 min

The fieldwork route evaluates a method somebody carried out. You did not carry one out, so the temptation here is to write about the analysis instead, and there is not much to say: a spreadsheet does not misread a scale.

What you evaluate instead is how the data was built: who chose the indicator, how finely it was measured, which places and years were left out, and which rows you yourself decided to keep. There is more to say here than on the fieldwork route, not less, because you controlled none of it. Step 6 asked you to discuss those weaknesses. This asks you to weigh them, rank them, and say what you would do about them.

Where the weakness livesThe question it answers
The indicatorDoes it really stand for the thing your question is about?
The resolutionIs it fine enough, in space and time, to show what you claimed?
The coverageWhich places and years are missing, and are they missing at random?
The collectionWho produced it, by what method, and with what reason to be careful or careless?
Your own selectionWhat did your inclusion rule remove, and could that have made the pattern?

The last row is the one students leave out and the one that is most obviously yours. You chose the rows. That choice is part of the method and it belongs in the evaluation of it.

Generic · 1 to 2

The data may not be completely accurate, and more data would have improved reliability.

Specific · 5 to 6

Each value is an annual mean of about twelve scheduled visits, so the dataset systematically under-samples storm events, which is when contamination actually arrives. The effect is one-directional: it flattens the difference between the two periods rather than creating it, so the improvement I found is likely to be an underestimate.

Test the limitation instead of asserting it

2 min

This is the single best move available on this route, and it is easier here than in fieldwork, because you still have all your data and re-running an analysis costs minutes.

Take the limitation you think matters most and rerun under a different assumption. Change the inclusion rule. Drop the years you were least sure about. Use the median instead of the mean. Then say what happened. "I tested this and the conclusion held" is what "evaluates" means as opposed to "describes", and it is worth more than any amount of careful hedging.

The result cuts both ways and both are useful. If the conclusion survives, you have evidence it is robust. If it does not, you have found something genuinely important, and reporting that honestly is worth far more than a finding that quietly depended on one choice.

A sensitivity test that passedGeneva rivers study

The monitoring network grew from about 19 stations to about 30 over the period, so the obvious objection is that the improvement is an artefact: different places, not better water.

Test it. Restrict to the 49 stations present in both windows and rerun. The medians go 40.0 to 17.4 and the means 85.4 to 27.9, against 40.0 to 18.2 and 89.1 to 27.9 for the unrestricted set. Practically identical, so the changing roster is not driving the result.

Ten minutes in a spreadsheet, and it converted the strongest objection to the study into a sentence saying the objection was tested and does not hold.

For your own investigation

Write down the strongest objection somebody could make to your result, then spend ten minutes trying to make it true. Reporting that you tried is worth more than any assurance that you thought about it.

Improvements that name something

1 min

An improvement that is only stated earns the bottom band. Say what it would achieve and you are in the top one, and that is two extra clauses per improvement.

This route has one characteristic failure here, and examiners know it: "use more data". More countries, more years, a bigger dataset. It is not an improvement, it is a wish, and it addresses no weakness you actually named. A real improvement names a specific dataset, a finer resolution, or a second source that would settle something.

Keep them feasible for a student. A different open dataset, a finer time step, a second indicator to cross-check with, a shorter and better-matched window. Those are things you could do next week.

ImprovementWhat it would buy
Join daily rainfall from the national weather serviceTurns "scheduled sampling misses storms" from an assertion into a measurement
Add the physico-chemistry file for the same stationsSeparates sewage sources from agricultural ones, which E. coli alone cannot
Restrict to stations with a full unbroken seriesRemoves the changing-roster objection by design rather than by testing
Find the dates each sector was converted to separate sewersGives a real before and after instead of two windows you chose

Notice that three of those four are things you could start this afternoon, and the fourth is a request to an authority that publishes most of its planning documents. That is the level to aim for.

Questions you opened but could not close

1 min

These have to point somewhere new. "What if I had more data?" is an improvement wearing a disguise, and the guidance names it as exactly that.

A real one has a different focus from your original question, arises from something you actually found, and bears on how far your conclusion reaches. The best ones come out of the part of your result that did not fit.

Two that go somewhereGeneva rivers study

Fifteen of the 49 stations got worse, and several sit downstream of infrastructure across the border. Is the pattern of worsening stations geographic, and does it follow the boundary of who is responsible for the network?

The lake bathing station has read zero for nearly thirty years while its tributaries carry hundreds of times more. At what point between the river mouth and the shore does that happen, and does it depend on the season?

Both came out of results the study could not explain, and both could be answered with data that already exists.

For your own investigation

Look at the part of your result that did not fit your explanation, and ask what would have to be true for it to make sense. On this route the answer is often another dataset, which makes it a question somebody could actually go and settle.

Using AI at this stepLevel 1 · AI Planning

It can be argued with once your own list exists. Ask it to attack a limitation you have already named and quantified, and see whether it holds.

It cannot generate the limitations. Ask a tool what is wrong with a dataset it has never opened and you get the five things wrong with all datasets, which is the definition of generic and the bottom band. The marks here are for what you found by looking at your file.

What this level means

Ready for step 8?

0 of 13

Number 12 is the one nobody thinks of: you chose which rows to keep, so that choice is part of the method and belongs in the evaluation of it.

Next: step 8, write it up

Every section now exists in some form. Step 8 is assembly, referencing and the word count, and it is where the budgets you have been carrying since step 1 finally get added up. Aim for 2,900, not 3,000, and you will find out why.

The Geneva rivers investigation used on the secondary-data route of this guide is the author’s own analysis of a published cantonal dataset. Every figure quoted comes from the Canton of Geneva’s open data, downloaded on 10 September 2026; the duplication, the missing year and the varying campaign counts described in the guide are as found on that date and may since have been corrected. The selection rules, the analysis and the conclusions are the author’s and not the canton’s.