Criterion F: Evaluation · 6 marks600 words · suggested

Evaluation

You arrive with

A conclusion, and the discussion of bias, reliability, validity and uncertainty you wrote in step 6. Most of the raw material for this section is already on the page. What changes is what you do with it.

You leave with
  • Specific weaknesses, ranked by how much they mattered
  • An improvement for each, with what it would fix
  • Questions your investigation opened but could not close
  • A finished report, bar the writing up

Six marks for saying what was wrong with your own work.

The same as your analysis is worth, and half as much again as your strategy or your method. Students tend to treat this section as an apology and hurry through it. It is the best return on effort in the entire report, and the only place where being hard on yourself is worth marks.

The bands turn on whether your weaknesses are generic or specific. The guide even defines generic for you: general to many methodologies, and not specifically relevant to the one you used. Which means the fix is mechanical. Take each sentence and ask whether it could have been written by someone who had never seen your investigation.

Generic · 1–2

Human error may have affected the results.

Specific · 5–6

Percentage cover was estimated by eye by one observer. Any consistent tendency to over-read moss would shift Simpson's index the same way at both sites, so the comparison survives it better than the absolute values do.

Generic · 1–2

More data would have improved reliability.

Specific · 5–6

Fifteen quadrats per site clears the threshold for a t-test but not the thirty conventionally wanted before quoting a standard error of the mean. The difference between the sites is therefore better supported than the precision of either site mean.

Generic · 1–2

The weather could have affected the results.

Specific · 5–6

Sampling happened on one day in spring, but the mechanism under investigation operates in winter, through snow compaction. The vegetation was measured months after the disturbance it responds to.

Everything below is how we suggest you actually do it.

Three strands, one chain

1 min

The commonest structural mistake is to write three separate lists: weaknesses here, improvements there, questions at the end, none of them speaking to each other. The descriptors join them up explicitly. Improvements must address the limitations you identified. Questions must bear on your conclusion.

Write it as a chain and the marks follow the structure. Take your two or three most significant weaknesses and give each one a full run: what it was, what it did to your answer, what you would change, and what that change would buy you.

The weakness

Specific to your method. And what it did to your conclusion, not just that it existed.

The improvement

Addressing that named weakness. Realistic for a student with a school's equipment.

What remains

A question with a different focus from your original, and a bearing on your answer.

Quantify it if you possibly can

2 min

Most evaluations assert that a limitation was serious. Very few test it. If you can take your own data and calculate what would have happened under a different assumption, you have moved from claiming a weakness matters to showing whether it does.

This is what "evaluates" means as opposed to "describes", and it is worth more than another paragraph of careful prose. It is also, usefully, something a spreadsheet can do in ten minutes.

Testing a limitation instead of asserting itski piste study

Mosses were recorded as a single group. Off-piste that group is 52% of all cover, which looks like it could be doing serious damage to a diversity index built on evenness.

So test it. Assume the 52% was really two species in equal shares, 26% each, put those two figures into the index in place of the single block, and recalculate. That raises off-piste diversity from 2.18 to 3.24, which overturns the headline result. Alarming, until you notice the piste's feather moss is equally a lumped group at 45%. Split both, as consistent identification would, and the conclusion holds in every scenario tested.

So the real limitation is not lumping. It is uneven lumping between sites, and on that measure the two sites are comparable. The conclusion stands, with a caveat that can now be stated precisely rather than vaguely.

For your own investigation

Take your biggest limitation and recalculate under a different assumption before you write about it. A spreadsheet can do it in ten minutes, and "I tested this and the conclusion held" is what "evaluates" means; it is worth more than any amount of careful hedging.

Improvements, evaluated

1 min

An improvement that is only stated earns the bottom band. Say what it would achieve and you are in the top one. Two extra clauses per improvement is the whole difference.

Keep them feasible. A second observer, a hand lens, an extra site, a return visit in a different season: these are things a student can actually do. Satellite imagery and a three-year monitoring programme are not improvements, they are daydreams, and examiners read them as such.

ImprovementWhat it would buy
Two observers estimating cover independentlyTurns an assumed bias into a measured one
A hand lens, mosses split to speciesTests the headline result rather than caveating it
A second piste and forest pair elsewhereSeparates management from this particular hillside

Questions you opened but could not close

1 min

These have to point somewhere new. "What if I collected more data?" is not an unresolved question, it is an improvement wearing a disguise, and the guidance names it as exactly that.

A real one has a different focus from your original question, arises from something you actually found, and bears on how far your conclusion reaches. The best ones point back at your environmental issue, which is also the last opportunity in the report for the two halves of your work to touch.

Two that go somewhereski piste study

Is the evenness advantage on the piste permanent, or a stage that an ungroomed slope passes through on its way back to forest? The strategy assumes a destination; this study only saw one moment.

Would summer grazing alone hold the sward open without winter grooming? The two are confounded here, and separating them would tell the graziers and the operator whether their interests actually align.

For your own investigation

A question worth writing down changes what your conclusion would mean, and someone could actually go and answer it. "More research is needed" fails both tests; questions that come out of your own confounded variables pass them.

Using AI at this stepLevel 1 · AI Planning

It can be argued with once your own list exists, ask it to attack a limitation you have already named and quantified, and see whether it holds.

It cannot generate the limitations. Ask a tool what was wrong with a study it cannot see and you get precisely the generic answers this criterion puts in the bottom band. The marks here are for weaknesses only someone who was there could name.

What this level means

Ready for step 8?

0 of 10

Eight of these are the criterion itself. It is the most prescriptive one in the rubric, which makes it the easiest to score well on.

Next: step 8, write it up

Every section now exists in some form. Step 8 is assembly, referencing and the word count, and it is where the budgets you have been carrying since step 1 finally get added up. Aim for 2,900, not 3,000, and you will find out why.

The ski piste investigation used throughout this guide is a teaching reconstruction. The location, sampling design, plant records and soil data come from real fieldwork at Mijoux in the French Jura. Canopy cover measurements were added to illustrate good practice and were not part of the original fieldwork. Photographs are the author’s own, taken at the site.