Criterion C: Quality and treatment of information collected · 6 marks · with step 5100 words · suggested

Your dataset

You arrive with

Every reading from the fieldwork day, and your sub-hypotheses. If you are missing readings, get them from the shared set now and note which ones are not your own.

You leave with
  • One clean spreadsheet, raw readings and calculated values kept apart
  • Discharge and hydraulic radius calculated for every site
  • A decision about which sites and variables you are using, and why
  • Any anomalous readings identified, and kept

This is where a shared dataset becomes your dataset.

Everyone walked away from the river with the same readings. What you do next, which sites you keep, which variables you use, what you calculate, what you do about the number that looks wrong, is a set of decisions nobody else makes for you, and Criterion C (Quality and treatment of information collected) marks the result.

QuantityHowUnitsWhat it tells you
Cross-sectional areaSum of segment areas: each depth × the interval between readings. Or mean depth × occupied widthHow much channel there is to carry water
DischargeCross-sectional area × mean velocitym³ s⁻¹ (cumecs)The volume passing a point each second, the headline downstream variable
Hydraulic radiusCross-sectional area ÷ wetted perimetermChannel efficiency. Higher means proportionally less water in contact with the bed and banks, so less friction
Why segments beat mean depth

Dividing the channel into subsections and summing their areas follows the real shape of the bed. A single mean depth × width treats the channel as a rectangle, which it never is, and the error is worst exactly where channels are most asymmetric, on a meander bend.

It is the same reason a gauging station divides its cross-section into verticals rather than taking one measurement in the middle.

Everything below is how we suggest you actually do it.

Use less data, better

3 min

You collected more than you need. That was deliberate; it is easier to collect eight variables once than to go back for the ninth, but it leaves you with a decision, and the instinct to use everything is the wrong one.

Pick the variables your sub-hypotheses actually need. Two or three variables analysed properly across nine sites will beat eight variables described in turn, because the top band of Criterion C (Quality and treatment of information collected) asks for data that is all directly relevant, and the top band of Criterion D (Written analysis) asks for analysis with no or only minor gaps in its supporting evidence. Both get harder the more you carry.

The same goes for sites. If a site was measured in obviously unrepresentative conditions, a stagnant backwater, a concrete culvert, a reach being worked on, you may leave it out. You must then say that you did, and why, in the report. Silently dropping a site is the thing you cannot do.

A cross-section calculator, for checking
Barcelona Field Studies Centre, “River Cross Section Calculator,” geographyfieldwork.com. Type in your ten depths and it draws the cross-section and returns area, wetted perimeter and hydraulic radius. Ignore the Manning’s n velocity option, that is a different method from the one we used.
Open the calculator
A ten-second check on wetted perimeter

Wetted perimeter is the measurement most likely to be wrong, because a tape laid along the bed lifts off between stones. The flowmeter manual (linked in step 3) gives a quick approximation that treats the channel as a rectangle: wetted perimeter ≈ width + (2 × mean depth).

It is not a better measurement than your tape, and it is not what you should report. It is a check. If your taped figure and this estimate are wildly apart at one site and close at the others, that site is worth re-reading before it goes into a hydraulic radius.

Use it second, never first

Do one site by hand, then put the same ten depths through the calculator. If the two agree, you have proved your spreadsheet formula works and you can trust it for the other eight. If they disagree, you have found a real error while it is still cheap to fix, and wetted perimeter is where it usually hides.

What you should not do is let it do the nine sites for you, or paste its diagrams into your write-up. Criterion C (Quality and treatment of information collected) is marking your processing and your presentation, and this is the one part of the investigation the fieldwork day already handed to everybody equally. Outsourcing it gives away the marks the shared dataset left you.

Everything, thinly

Eight variables, each with a scattergraph and a paragraph, none tied to a hypothesis. Eight figures and 850 words of analysis, about a hundred words each.

A hundred words is enough to describe a graph and nothing else. It reads as a tour of the data.

Three, properly

Two or three variables, each tied to a sub-hypothesis, each with the right graph, one statistical test where it genuinely fits, and roughly 650 words split between them, plus 100 for the anomalies and 100 drawing it together.

The same 850 words. Room to explain, and every figure is doing work.

One spreadsheet, raw and processed kept apart

3 min

Build one sheet with a row per site and columns for everything: site number, latitude, longitude, altitude, distance from source, then each measured variable, then each calculated one. Keep the raw readings and the values you have calculated visibly separate, different blocks, or a clear label, so that a reader can always see which numbers came out of the river and which came out of a formula.

Use formulas rather than typing calculated values in by hand. Not to save time, but because a formula is checkable and a typed number is not: if you find a depth reading was transcribed wrongly, everything downstream of it corrects itself.

Include the latitude and longitude columns even if you are not sure you will map anything. Step 5 needs them in decimal degrees, and adding them later means going back to the map for nine sites one at a time.

If any of your velocities came from a float rather than the flowmeter, apply the correction here rather than later: a float measures the fast surface water, so multiply by the surface-velocity coefficient you chose in step 3, commonly 0.85, to estimate the mean velocity of the cross-section, and record which value you used. Mixing corrected and uncorrected velocities in one column is the kind of error that survives all the way to the conclusion.

Raw data tables, calculations and appendices containing raw data do not count towards the word limit. Neither do tables of statistical or numerical data anywhere in the report. This is the cheapest evidence in the whole IA.

The class dataset
Every group's readings for all 9 sites, in one sheet: location, channel characteristics, cross-section depths, flowmeter and float velocities, bedload. 5 groups measured the same sites, so cross-check yours against the others before you trust an odd number.
Open the shared sheet
A row per site should carry
Site number, latitude, longitude, altitude
Distance from source, in metres, your x-axis for everything
The measured variables you are using, with units in the header
Cross-sectional area, discharge, hydraulic radius
A notes column for anything odd about that site
Keep the spread, not just the mean

Wherever you took repeat readings, the flowmeter's three trials at each position, the three float runs, the thirty stones, the spread across them is data, not noise to be averaged away. Carry the highest and lowest into the sheet beside the mean; it costs one column.

It pays twice. In the analysis, a site whose velocity trials disagree wildly is a fact about turbulent flow at that site, worth a sentence. In the evaluation, a mean quoted with its range is a limitation you can quantify, where a mean quoted alone leaves you writing about "human error".

The CDN approach

The reading that looks wrong is the most valuable one you have

2 min

Somewhere in your dataset there is a site that refuses to behave: a velocity that drops when it should rise, a channel that narrows downstream, a bed of angular pebbles a long way from the source. The instinct is to assume you measured it wrong and quietly lose it.

Do the opposite. The top band of Criterion D (Written analysis) specifically asks for outliers and anomalies, if present, to be explained and linked to the question, the hypotheses, the theory, the fieldwork location and the methods used. The band below only lists them. An anomaly you can explain is worth more marks than a clean trend, because explaining it is the thing the band above is made of.

And on this river you have an unusually good chance of explaining one. Reaches of the Asse were straightened and confined long ago, and since 2019 they have been the subject of a scheme to widen and restore them. A site in an engineered reach has a channel shape that somebody chose. That is not noise in your data; it is a second, human, control on the variables you are measuring.

Where to look firstour river

Three kinds of place are worth checking against your site list before you decide anything is an error.

Engineered or restored reaches. The flood-protection and renaturation scheme covers about 4.5 km inside Nyon, and is being delivered in stages, so check what has actually been built at your site. A widened or reprofiled channel can be wider, shallower or slower than distance from the source alone would predict.

Tributary junctions. The Calèves was reconnected to the Asse in 2019, having been diverted away decades earlier. Discharge steps up below a confluence rather than rising smoothly.

Structures and local conditions. Bridges, weirs and culverts fix the channel's width and bed. A pool sampled as though it were a riffle gives a slack velocity that has nothing to do with where it is on the river.

For your own investigation

Before you call a reading an error, check the site against a map and your own photographs. A number that is wrong because the tape slipped belongs in the evaluation; a number that is unexpected because somebody rebuilt the channel belongs in the analysis, and is worth far more.

Decide about a statistical test now, not later

6 min

A statistical test is optional. Nowhere does the subject guide say you must run one. Criterion C (Quality and treatment of information collected) says your techniques “may include statistical tests”, and its bands never mention statistics at all. Criterion B (Method(s) of investigation) says you “may describe statistical tests if appropriate”. Criterion D (Written analysis) asks for “descriptive and statistical techniques (if appropriate to the question formulated)”. Every one of those is conditional.

Be straight about the other half, though. Statistical techniques are named in the top two bands of Criterion D (Written analysis), and for a question about how a variable changes downstream a correlation is an obvious candidate for “appropriate”. So the honest position is not “tests do not matter”. It is that the guide asks you to judge whether one fits, and then to be right about it.

If you do run one, the natural choice here is Spearman’s rank correlation coefficient. You are testing whether one variable changes with distance from the source, which is a question about the strength and direction of a relationship between two ranked variables, and that is exactly what Spearman’s is for.

Decide now rather than after you have made your graphs, for two reasons. It tells you what your null hypothesis has to say, and it tells you whether you have enough sites for the result to mean anything. Choosing a test to fit a result you have already seen is the wrong way round, and it shows.

The output is a coefficient between −1 and +1, and a significance level. Both matter: a strong-looking correlation across nine sites can still fail to reach significance, and reporting the coefficient without saying whether it is significant leaves the reader unable to judge it.

Nine sites is not many

With nine pairs of values, a Spearman's coefficient has to be fairly large before it is significant at the 95% level. That is a genuine limitation of the design, not a mistake you made.

Say so, in the analysis, at the point where you report the result. "rs = 0.47, which does not reach significance at the 95% level with n = 9" is a stronger sentence than a coefficient reported alone, and it is exactly the kind of honesty Criterion F (Evaluation) rewards later.

Check your critical value in a table rather than assuming one, and say which you used. The threshold depends on your number of pairs and on whether you are testing in one direction or two, and since step 1 had you write directional hypotheses, that choice is one you should be able to justify.

“Appropriate to the data” is the actual question

The 5–6 band of Criterion D (Written analysis) already asks for techniques “that are appropriate to the data and the question formulated”. That phrase is doing the work, and our dataset makes it a real question rather than a formality.

Nine sites is a small n. Spearman’s needs a coefficient around 0.6 before it is significant at the 95% level with n = 9, so a relationship you can see on the graph can still fail the test. Examiner guidance for river investigations puts the comfortable minimum nearer 15 data points, with some textbooks accepting 10. Nine sites is below both, and that is a fact about the design rather than something you did wrong.

Our sites are not independent of each other. Sites 5 and 6 are 40 m apart and sites 8 and 9 are 200 m apart, so two of your nine points are nearly duplicates. Water is also carried downstream, so a site’s discharge is partly the site above it. Correlation assumes independent observations, and a chain of points down one river is not quite that.

Where your numbers are actually large is inside a site, not along the river. You measured thirty stones at each site, so a comparison of bedload between two sites rests on sixty observations rather than nine. If you want a test with real power behind it, that is where to look instead; step 6 sets out how to run it so that it is valid.

None of this makes a test wrong. It makes it a choice you have to defend, which is exactly what the descriptor asks. Run it, report it honestly including a non-significant result, and say what its limits are; that combination reads as judgement. Running it and reporting only the half that worked does not.

A bad test is worse than no test

Criterion D (Written analysis) puts it in as many words. Its 3–4 band reads: the written analysis includes descriptive techniques that are appropriate… Any statistical techniques used either are not relevant to the question formulated or contain errors. A test that is wrong, or that answers a question you never asked, does not sit there earning nothing. It pulls you into that band, below work that used no test at all.

So the decision is not “should I add a test to look rigorous?” It is: can I name why this test suits this data, run it correctly, interpret the result, and say what it means for a hypothesis I actually wrote? Four yeses and it is worth doing. Any no and you are better off without it, spending the words on explanation instead.

One test, understood

Running three different statistical tests does not demonstrate three times the skill. It usually demonstrates that none of them was chosen for a reason.

One test, applied to the relationships your hypotheses name, with its result interpreted properly, is what the descriptors ask for.

Using AI at this stepLevel 3 · Targeted AI

It can explain what a statistic does, what a spreadsheet formula means, or why an equation is built the way it is; the questions you would otherwise queue up to ask us.

It cannot do the arithmetic. Paste a table in and ask for the means and you will get numbers that look right and are not checkable, in a dataset you share with the class. An error here propagates silently into every graph and every sentence of your analysis.

What this level means

Ready for step 5?

0 of 11

This is the last step where a mistake is cheap. An error in the spreadsheet now becomes an error in every graph, every statistic and every sentence of the analysis.

Next, showing it

Step 5 is the other half of Criterion C (Quality and treatment of information collected): choosing presentation techniques that suit each variable, and making them well enough that the pattern is visible before a word of analysis is read.

Anything marked “our river” is specific to the fieldwork we do together on the River Asse at Nyon: the sites, the equipment and the way we collect the data. Site coordinates and altitudes are from our own fieldwork records and the exact list can change from year to year, so check them against the sheet you are given on the day. Photographs are the author’s own, taken at the river.