Design your method
A research question, a place, and a tension. That last one is more useful here than it looks: knowing what each side of the argument would want to see is how you work out what is worth measuring.
- A method a stranger could follow
- Enough planned data to answer your question
- Your variables, named and justified
- Safety and ethics specific to what you are doing
Your method is an instruction manual, written afterwards.
Two things at once, and students usually get one of them. It has to be complete enough for someone else to follow, and it has to be a record of what you actually did rather than a plan for what you intend to do. Anything you changed in the field belongs in it.
This criterion is two questions. Both must be yes.
Could a third party replicate your investigation from what you wrote, with you nowhere near them? If any step needs you standing there to explain it, the answer is no.
Does the method generate enough to answer the question you asked? Not enough data here hurts you again in steps 5 and 6, when there is nothing to find a pattern in.
Fail either and you are in the 1–2 band, however well the other one is done. There is no partial credit and no ladder of command terms here. It is unusually blunt for an IB criterion, and unusually easy to check.
Everything below is how we suggest you actually do it.
Write it as a record, not a plan
The guidance is explicit that the method should not be a proposal but an account of what was done. That means past tense throughout, and it means including the things that went differently from how you imagined them.
If you ran a pilot and changed your interval afterwards, say so. If the third site was inaccessible and you moved it, say so. Those are not admissions of failure; they are what makes the account true, and step 7 will thank you for them.
Standard protocols should be cited rather than copied out. Nobody needs the Winkler method reproduced in full; a reference to it is enough.

That caption is the model, and it is doing more work than it looks. A figure number your text can point at, what the picture shows, where it was taken and how high, which way you were facing, when, and whose photograph it is. All of it came off the phone: coordinates, altitude and bearing are stored in the image, and your photo library will show them.
Credit your own photographs, author's own photograph is the phrase, and cite anyone else's like any other source. A photograph nobody can locate or date is decoration, and decoration earns nothing. Note also what the picture cannot tell you: the interval between quadrats, and how the tape was oriented. Photographs support a written method, they do not replace one.
Plus correct scientific names for any organism. Vaccinium myrtillus, not bilberry.
For the grid reference and the map, use the national mapping service rather than a screenshot of Google Maps. Both of these give you coordinates for any point you click, a topographic base map worth reproducing, and historical layers if your issue is about how a place has changed.
- Border
- Orientation
- Legend
- Title
- Scale bar
And credit the basemap; that is not one of the five letters, and it is a formal requirement of its own.
How much data is enough
The mark scheme says "sufficient" and leaves you to work out what that means. It is not a mystery, though, because the statistics you will run in step 5 have their own thresholds, and those set the floor.
| How to lay the sampling out | Rule of thumb |
|---|---|
| Along a gradient | at least 5 intervals |
| Repeats at each point | at least 3; the IB's benchmark is 5 |
| Within each zone | 3 minimum, 5 better |
| Survey responses | at least 30 per group you compare |
| Readings your chosen statistic wants | Per group |
|---|---|
| For a standard deviation | 5 or more |
| For a t-test | 10 or more |
| Before quoting a standard error | 30 or more |
The IB's own rule of thumb is five values of the independent variable, five repeats at each: twenty-five data points. It is written with laboratory work in mind, and the guidance itself concedes that a transect under time pressure may collect less, but if yours does, say so in the report, because five-by-five is the default a moderator arrives with. And for questionnaires, thirty means thirty per group you compare, not thirty in total: three age groups means ninety responses.
Decide which test you will run before you go out. Discovering afterwards that you have eight readings and needed ten is a bad afternoon.
Fifteen quadrats on the ski piste and fifteen in the forest beside it. The question asks whether the two differ, so the statistic that answers it is a t-test, and fifteen clears the threshold of ten comfortably. That is why the comparison holds up.
It does not clear thirty, the number conventionally wanted before quoting a standard error of the mean. Standard error is what would let you say how precisely each site's own mean is known, and put honest error bars on it. So fifteen is enough to be confident the two sites differ, and not enough to be confident about either site's exact value.
Working that out beforehand would have meant sampling more, or planning around it. Found afterwards, it became an evaluation point in step 7 instead; the second-best outcome, because a specific, quantified limitation still earns marks there, but it is a weakness you are explaining rather than one you avoided.
Decide which statistic answers your question before you go out, then count backwards to how many readings it needs. Doing it in that order is the difference between a limitation you avoided and one you have to write about.
Name the rule that places your samples
Most methods describe where the quadrats went and never say what decided it. "We spread them out evenly across the site" is a description of an outcome, not a rule someone else could follow, and justifying your sampling is hard when the choice has no name. There are only three names to learn, and the same three work whether you are placing quadrats or choosing people to survey.
| Placing quadrats | Choosing respondents | |
|---|---|---|
| Random | Lay a numbered grid over the site and let a random number generator pick the squares. Your eye never chooses. | Chance picks the people: house numbers or list positions from a generator, never whoever looks approachable. |
| Systematic | A fixed interval, applied without exception: a quadrat every five metres along the transect. | A fixed rule, applied without exception: every fifth person to pass a point, whoever they are. |
| Stratified | The sample mirrors the site. If a third of the slope is bare ground, a third of the quadrats go on bare ground. | The sample mirrors the population. If a third of visitors are families, a third of your respondents are families. |
Each buys something and costs something. Random removes your judgement from the choice, but on a small site it can land three quadrats in one corner. Systematic cannot cluster, but the interval can fall into step with a pattern in the ground: every five metres on ridged terrain samples the ridges. Stratified mirrors the site best, and demands you know its proportions first, which is one of the things a pilot is for. They also combine: a transect placed at random with quadrats at a fixed interval along it is random and systematic doing different jobs in one design.
One sentence in your method settles it: the rule, and why it fits this site or these people. That sentence is the justification the checklist asks for. Without the name, all you can write is that you spread things out.
Sometimes the honest description is opportunistic: you took the readings you could get, where access allowed. That is not a disqualification, it is a limitation, and it belongs named in the method and weighed in step 7 rather than dressed up as one of the three.
A questionnaire is an instrument too
Questionnaires look like the easy method: no probes to calibrate, no species to identify. Then the answers come back and half are unusable, because two people read one question two different ways, an open question produced thirty answers with nothing to count in them, and nobody remembers which responses were the practice run. Every one of those is a writing-stage failure, and every one is avoidable.
How many responses is enough? Thirty per group you compare is the floor from the thresholds earlier on this page. The RGS guide adds a stopping rule I have borrowed since: collect your thirty, do a rough count of the pattern, then collect fifteen or twenty more. If the picture holds, it was stable and you have enough. If it moved, it was never a conclusion yet, and stopping at thirty would have published the noise.
Decide before the first stranger: do you read the questions aloud, or hand the sheet over? Aloud is friendlier and lets you rephrase for anyone who needs it, but nobody holds eight options in their head, so pair it with a printed card of the choices they can point at.
And write for anyone. A question that needs the vocabulary of your course to answer will only measure who has taken your course.
The other half of asking people is consent, which has its own section below: what the form says, and why a blank copy goes in the appendix rather than the completed ones.
Measure the conditions, not just the thing
The IB does not require this. It asks for data appropriate to your question and leaves the rest to you. We ask for it anyway, and the reason is worth understanding rather than just complying with.
Measure only your organisms and you can report that two places differ. Measure the conditions as well and you can explain why they differ, which turns a correlation into a mechanism.
That mechanism is also the bridge back to step 3. A strategy acts on conditions: clearing, draining, grazing, building. If you have measured the conditions, you can say what the strategy would do. If you have not, your fieldwork and your argument sit side by side and never touch.
What is living there, and how much of it
The conditions that decide what can live there
The condition that actually mattered here was how much light reached the ground, because that is what grooming changes and what the plants respond to. The obvious instrument for it is a lux meter.
So light was measured with one: 12,000 lux on the piste against 5,000 in the forest. Two readings, taken at different moments under moving cloud. That is a number you can quote in a sentence and nothing more; there is nothing to correlate against diversity, and a cloud passing at the wrong moment could have reversed it.
The better approach was to measure canopy cover instead: the same underlying thing, but stable while you walk between plots. One reading per quadrat gives fifteen per site rather than one, so it can be plotted and correlated, and that correlation, rho = −0.77, turned out to be the strongest single piece of evidence in the whole study.
Measure the stable proxy for the condition rather than the condition at the moment you happened to be standing there. Ask of every abiotic variable: can I take one reading per quadrat, and will it still mean the same thing an hour later? If not, find the version that will.
Design the bias out while you still can
Step 6 will ask you to discuss bias and step 7 to evaluate it, and both reward honesty about it. But most of the bias in a school investigation is settled here, in the design, while it is still removable. It hides in a short list of places, and each one is worth thirty seconds against your plan.
One judge for anything judged. Wherever a value is estimated rather than measured, percentage cover, an impact score, the same person estimates it at every site. Splitting the sites between friends is faster, and it means the difference between your sites is partly the difference between your eyes.
A photograph at every point, facing the same way. Fix the rule before you start: one photo per sampling point, on the same bearing. It turns the photo record from a gallery of the interesting bits into the cheapest systematic dataset of the day.
What you cannot design out, write down in the field while it is in front of you. A bias named in your notes becomes step 6's discussion and step 7's evaluation; one remembered a month later becomes a guess.
The day itself
You will get less time in the field than you expect, and no second visit. The difference between a comfortable day and a wasted one is almost entirely what you settled before you left.
Which test you will run, and therefore how many readings you need. Where the sampling points go and how you will locate them. What each person is doing. Decisions made standing in the wind get made badly.
A blank notebook invites gaps. Print a table with a row per quadrat and a column per variable, so an empty cell is visibly a job rather than something you discover a week later.
The site from both ends, the apparatus in use, anything unusual, and any quadrat that surprises you. Photographs are outside the word count, and they are the only way to answer a question you have not thought of yet.
Do a pilot in the first ten minutes. Take one full set of readings at one point, then stop and look at it. Is the interval sensible? Is the variable actually varying? Is a reading taking three minutes when you have budgeted one? Changing your design after one point costs ten minutes; discovering the problem at the end costs the investigation. Whatever you change, write down that you changed it and why, step 7 pays for exactly that.
Record as you go, and record the awkward things too. The quadrat you moved because it landed on a path, the reading you repeated, the point you could not reach. Those notes are what makes the method an account rather than a plan, and they are where your specific limitations come from. Nobody reconstructs them afterwards.
One morning at Mijoux: thirty quadrats across two sites, percentage cover species by species, plus canopy cover, soil penetration depth and moisture at each one. The drone and ground photographs on this guide were taken twenty minutes apart on the same visit.
What it did not produce was a second chance. The lux-meter readings were abandoned as unusable, and the sample size turned out to be comfortable for a t-test and short for a standard error, both discovered afterwards, both now evaluation points rather than fixes.
Assume one visit. Before you go, write the list of numbers you must come home with; on the day, take a pilot set early and check it against that list. The things you cannot fix later are the sample size and the variables you did not think to measure.
Safety and ethics, about your investigation
No marks ride on a risk assessment any more, which has had an odd side effect: students either drop it entirely or paste in generic lab rules about tying back long hair. Neither helps.
What belongs here is what was actually risky about this work. Steep or uneven ground. Ticks. Working near water. Handling samples. And on the ethics side, what you did to leave the site as you found it, plus a consent form if people were involved.
One page is plenty. Say in the method that consent was obtained and how, then put a blank copy of the form in an appendix; that is the one thing an appendix is genuinely for. Never the completed ones: they carry other people's information, and an appendix is not read anyway.
Writing it in the future tense. "I will place five quadrats along each transect." It reads as a plan you might not have carried out, it hides everything you changed along the way, and it makes the account harder to trust. The fix takes ten minutes: go through and put the whole thing in the past tense, then add the two or three things that did not go as intended.
It can explain a technique you have not met before, what a Winkler titration measures, how a quadrat frame is used, as many times and as many ways as you need.
It cannot design the sampling. Where the quadrats go, how many, and at what interval are the decisions this criterion marks, and they depend on a place only you have stood in.
Ready for step 5?
0 of 12The first two are the criterion itself. The rest are how you get there.
Step 5 is where the good news arrives. Tables, graphs and calculations sit outside the word count entirely, so the six marks there cost you almost nothing to earn. Decide your statistical test now, before you go out, because it determines how much data you need to come back with.
The ski piste investigation used throughout this guide is a teaching reconstruction. The location, sampling design, plant records and soil data come from real fieldwork at Mijoux in the French Jura. Canopy cover measurements were added to illustrate good practice and were not part of the original fieldwork. Photographs are the author’s own, taken at the site.
