Evaluation
A conclusion, and the discussion of how your data source could mislead that you wrote in step 6. Most of the raw material for this section is already on the page; what changes is what you do with it.
- Specific weaknesses of your data, ranked by how much they mattered
- Your own selection rule, evaluated
- An improvement for each, naming something real
- Questions your investigation opened but could not close
Six marks for saying what was wrong with data you did not collect.
Which sounds like an odd thing to be marked on, until you notice how much of it was your decision. You chose the source, the indicator, the years, the places and the rule that kept some rows and dropped others. Every one of those is a methodological choice, and this is where you weigh them.
The bands turn on whether your weaknesses are generic or specific. The test is mechanical: could this sentence have been written by someone who never opened your file?
The data may not be completely accurate.
Each value is an annual mean of about twelve scheduled visits, so the dataset under-samples storm events, which is when contamination actually arrives. That flattens the difference between my two periods rather than creating it, so the improvement I found is likely an underestimate.
More data would have improved reliability.
The monitoring network grew from 19 stations to about 30 across the period, so my early and late groups describe partly different places. Restricting to the 49 present in both changed the median from 18.2 to 17.4, which is small enough to say the growth is not driving the result.
The dataset may contain errors.
The file publishes every station-year about fifty times over. I removed the duplicates before analysis, which left 757 observations rather than 37,009. Had I not, every significance test in this report would have been meaningless while looking entirely normal.
Everything below is how we suggest you actually do it.
Three strands, one chain
The commonest structural mistake is to write three separate lists: weaknesses here, improvements there, questions at the end, none of them speaking to each other. The descriptors join them up explicitly. Improvements must address the limitations you identified. Questions must bear on your conclusion.
Write it as a chain and the marks follow the structure. Take your two or three most significant weaknesses and give each one a full run: what it was, what it did to your answer, what you would change, and what that change would buy you.
Two things do not belong here, both left over from the old course: the strengths of your method, and an application of your findings. If your template has headings for either, delete them; between them they can eat two hundred words the three marked strands needed.
Specific to your method. And what it did to your conclusion, not just that it existed.
Addressing that named weakness. Realistic for a student with a school's equipment.
A question with a different focus from your original, and a bearing on your answer.
Two shapes cost marks. A table of limitations against improvements looks organised, reads as a list, and scores lower than extended writing. And bullet points of three words each, which is what a section written after the word count ran out looks like.
The opposite failure is being simply wordy. Six hundred words, two or three weaknesses, each given a proper run. Plan the space, because this section is worth as much as your analysis and students most often arrive at it with nothing left.
Questionnaires carry three built-in weaknesses the IB expects you to know about, and naming them precisely beats gesturing at "possible bias". Participation was voluntary, so the people who answered are not a fair sample of the people you asked; strong opinions turn up more often than mild ones. Respondents tend to avoid the two ends of a Likert scale, which squeezes your data toward the middle. And a long or complicated survey produces fatigue, so the later answers are worse than the early ones.
Each of these earns marks only when you connect it to your conclusion: which of your results would move, and in which direction, if the missing voices had answered?
The weaknesses are in the data, not in your hands
This criterion evaluates a method, and the method that produced your numbers was carried out by somebody else. So the temptation is to write about your analysis instead, where there is very little to say: a spreadsheet does not misread a scale.
What you evaluate instead is how the data was built: who chose the indicator, how finely it was measured, which places and years were left out, and which rows you yourself decided to keep. Step 6 discussed those weaknesses; this weighs them, ranks them, and says what you would do about them.
| Where the weakness lives | The question it answers |
|---|---|
| The indicator | Does it really stand for the thing your question is about? |
| The resolution | Is it fine enough, in space and time, to show what you claimed? |
| The coverage | Which places and years are missing, and are they missing at random? |
| The collection | Who produced it, by what method, and with what reason to be careful or careless? |
| Your own selection | What did your inclusion rule remove, and could that have made the pattern? |
The last row is the one students leave out and the one that is most obviously yours. You chose the rows. That choice is part of the method and it belongs in the evaluation of it.
And put a number on the uncertainty, read off the documentation rather than an instrument. How many significant figures? Is there a detection limit? Rounded to what? A series reported to the nearest whole unit cannot support a difference of half a unit, and saying that in a sentence is worth more than a paragraph about data quality in general.
The data may not be completely accurate, and more data would have improved reliability.
Each value is an annual mean of about twelve scheduled visits, so the dataset systematically under-samples storm events, which is when contamination actually arrives. The effect is one-directional: it flattens the difference between the two periods rather than creating it, so the improvement I found is likely to be an underestimate.
Test the limitation instead of asserting it
This is the single best move available on this route: you still have all your data, and re-running an analysis costs minutes.
Take the limitation you think matters most and rerun under a different assumption. Change the inclusion rule. Drop the years you were least sure about. Use the median instead of the mean. Then say what happened. "I tested this and the conclusion held" is what "evaluates" means as opposed to "describes", and it is worth more than any amount of careful hedging.
The result cuts both ways and both are useful. If the conclusion survives, you have evidence it is robust. If it does not, you have found something genuinely important, and reporting that honestly is worth far more than a finding that quietly depended on one choice.
The monitoring network grew from about 19 stations to about 30 over the period, so the obvious objection is that the improvement is an artefact: different places, not better water.
Test it. Restrict to the 49 stations present in both windows and rerun. The medians go 40.0 to 17.4 and the means 85.4 to 27.9, against 40.0 to 18.2 and 89.1 to 27.9 for the unrestricted set. Practically identical, so the changing roster is not driving the result.
Ten minutes in a spreadsheet, and it converted the strongest objection to the study into a sentence saying the objection was tested and does not hold.
Write down the strongest objection somebody could make to your result, then spend ten minutes trying to make it true. Reporting that you tried is worth more than any assurance that you thought about it.
Improvements that name something
An improvement that is only stated earns the bottom band. Say what it would achieve and you are in the top one, and that is two extra clauses per improvement.
This route has one characteristic failure here, and examiners know it: "use more data". More countries, more years, a bigger dataset. It is not an improvement, it is a wish, and it addresses no weakness you actually named. A real improvement names a specific dataset, a finer resolution, or a second source that would settle something.
Keep them feasible for a student. A different open dataset, a finer time step, a second indicator to cross-check with, a shorter and better-matched window. Those are all things you could actually go and do.
| Improvement | What it would buy |
|---|---|
| Join daily rainfall from the national weather service | Turns "scheduled sampling misses storms" from an assertion into a measurement |
| Add the physico-chemistry file for the same stations | Separates sewage sources from agricultural ones, which E. coli alone cannot |
| Restrict to stations with a full unbroken series | Removes the changing-roster objection by design rather than by testing |
| Find the dates each sector was converted to separate sewers | Gives a real before and after instead of two windows you chose |
Notice that three of those four are things you could start straight away, and the fourth is a request to an authority that publishes most of its planning documents. That is the level to aim for.
Questions you opened but could not close
These have to point somewhere new. "What if I had more data?" is an improvement wearing a disguise, and the guidance names it as exactly that.
A real one has a different focus from your original question, arises from something you actually found, and bears on how far your conclusion reaches. The best ones come out of the part of your result that did not fit.
Then two mechanical things that are free marks. Give them their own subheading, called Unresolved questions, and write them as questions, with question marks. A moderator cannot award a strand they cannot find, and headings do not count against your word limit.
Fifteen of the 49 stations got worse, and several sit downstream of infrastructure across the border. Is the pattern of worsening stations geographic, and does it follow the boundary of who is responsible for the network?
The lake bathing station has read zero for nearly thirty years while its tributaries carry hundreds of times more. At what point between the river mouth and the shore does that happen, and does it depend on the season?
Both came out of results the study could not explain, and both could be answered with data that already exists.
Look at the part of your result that did not fit your explanation, and ask what would have to be true for it to make sense. On this route the answer is often another dataset, which makes it a question somebody could actually go and settle.
It can be argued with once your own list exists. Ask it to attack a limitation you have already named and quantified, and see whether it holds.
It cannot generate the limitations. Ask a tool what is wrong with a dataset it has never opened and you get the five things wrong with all datasets, which is the definition of generic and the bottom band. The marks here are for what you found by looking at your file.
Ready for step 8?
Secondary data checklist0 of 17The one about evaluating your own inclusion rule is the one nobody thinks of: you chose which rows to keep, so that choice is part of the method and belongs in the evaluation of it.
Every section now exists in some form. Step 8 is assembly, referencing and the word count, and it is where the budgets you have been carrying since step 1 finally get added up. Aim for 2,900, not 3,000, and you will find out why.
The Geneva rivers investigation used on the secondary-data route of this guide is the author’s own analysis of a published cantonal dataset. Every figure quoted comes from the Canton of Geneva’s open data, downloaded on 10 September 2026; the duplication, the missing year and the varying campaign counts described in the guide are as found on that date and may since have been corrected. The selection rules, the analysis and the conclusions are the author’s and not the canton’s.
