Criterion D: Treatment of data · 6 marks200 words · suggested

Treat your data

You arrive with

A notebook or spreadsheet of raw readings, and ideally a decision already made about which statistical test you are going to run. If you have not chosen one yet, do it before you start processing.

You leave with
  • Raw data, tabulated
  • Processed data that shows the pattern
  • A statistical result, correctly calculated
  • No interpretation whatsoever. That is next.

Six marks, and almost none of the evidence costs you a word.

This is the best-value section in the whole report, and most students never notice. Everything that earns marks here sits outside the 3,000: the tables, the graphs, the calculations. What costs words is describing them, and describing them earns nothing.

Does not count

Data tables, raw and processed. Every graph and chart. Equations, formulae and calculations. Diagrams and photographs. Citations and the bibliography.

Spend freely. Make them excellent.

Does count

The sentence saying what each table shows. Why you chose that statistical test. Your null and alternative hypotheses. What you did about outliers, and why.

About 200 words in total. That is all you need.

One caution that catches people out: appendices are not read by the examiner. Anything that matters has to be inside the report. Since tables do not count against you wherever they sit, put them in the body where they will actually be seen.

Everything below is how we suggest you actually do it.

Raw, then processed, then stop

1 min

Raw data goes in tables. Not lists, not prose, not a photograph of your field notebook. Separate tables for raw and processed are perfectly acceptable and usually clearer.

Then process it into something that shows the pattern. A well-made graph of raw data is not processing, and it is not analysis either. Something has to be calculated: an index, a mean with its spread, a correlation.

You do not need to show a worked example for standard processing. Nobody needs to see how you calculated a mean. Show the working where the calculation is unusual or where a reader would otherwise have to guess.

What that looked like hereski piste study

Percentage cover per species per quadrat is the raw data: thirty rows of it. Simpson's reciprocal diversity index calculated per quadrat is the processed data, it runs from 1, meaning a single species holds everything, up to the number of species present when they are perfectly even, so it rewards evenness as well as richness. One worked example of the calculation is shown, because the formula is not something a reader can assume.

The means, standard deviations and t-tests come next. All of it in tables, none of it costing a word.

For your own investigation

Put the numbers in tables and show the formula once, with a single worked example. Every sentence that walks the reader through a table is a sentence taken from a section where words are the only evidence you have.

Choosing the graph and the test

3 min

The graph type is a real decision, not a default. Comparing two groups wants a bar chart with error bars. Looking for a relationship between two continuous variables wants a scatter plot with a line of best fit. Picking the wrong one hides the very pattern you are trying to show.

Where you can, run an inferential test and report it properly: your null and alternative hypotheses, the statistic, the degrees of freedom, and the probability.

The two hypotheses are one sentence each, written before you test. The null says there is nothing going on, "there is no difference in Simpson's index between the piste and the forest", and the alternative says there is, without predicting which way unless you have a reason to. You are not trying to prove the alternative; you are seeing whether the data makes the null hard to believe.

Error bars are worth understanding rather than decorating with. Say which ones you plotted, standard deviation, standard error or a 95% confidence interval, because they answer different questions and a reader cannot tell by looking. And treat them as a picture of spread rather than a significance test: bars that overlap can still hide a real difference, and only the test settles it.

No, you do not show the working

You do not need to reproduce the arithmetic, the formula, or the rank-by-rank calculation. Nobody is marking your ability to add up, and every line of it costs words you need elsewhere.

What has to be visible is the decision trail: which test, why that test for that data, the value of the statistic, n or the degrees of freedom, the critical value or p, the significance level, and the verdict on the null hypothesis. Justifying the choice is where the credit is, because that is a judgement; the arithmetic is not.

The tables are free, since tables of numerical data sit outside the word count. Put a rank table or a chi-squared table of observed and expected values in the body rather than an appendix, where a reader can see the test was set up correctly.

One trap worth naming: chi-squared must be run on raw counts, never percentages or means.

Social Science Statistics
Free online calculators for every test below, t-test, Mann-Whitney U, Spearman's rank, chi-squared, ANOVA, with the working shown step by step.
Open the calculators

A calculator will run any test on any numbers you give it. It has no idea whether that test was the right one, which is the part you are marked on.

Which test, and how to justify it

4 min

Students usually approach this backwards, hunting for the name of a test and then checking whether their data fits it. Start at the other end. What kind of claim are you trying to make? That question alone narrows six tests to two, and the rest is checking your data qualifies.

A difference
Between groups

Two sites, three management regimes, cut against uncut. You are asking whether the groups sit in different places, comparing averages, or the overlap between two spreads.

A relationship
Between two measurements

As one thing rises, does the other rise or fall? Both variables measured on the same quadrat or transect point, so they come in pairs.

An association
Across categories

Counts falling into classes, species by substrate, responses by age group. Tallies of things, not measurements of them.

One piece of vocabulary first, because the table turns on it. Interval data is measured on a real scale with units, percentage cover, centimetres, degrees. Ordinal data is ranked but not measured, so the gaps between values mean nothing: a 1-to-5 trampling score, or shelter judged from none to heavy. Categorical data is counted into named classes: species by substrate, answers by age group. Percentage cover is interval; a score you assigned by eye is ordinal; a tally is categorical.

TestThe question it answersWhat it needsCheck before you use it
t-testDo two groups differ, on average?Interval data, around 10+ per groupEach group roughly normal, plot a histogram and look
Mann-Whitney UDo two groups differ, when the data is skewed or ranked?Interval or ordinal, 5+ per groupNo normality needed; the two distributions should be a similar shape
ANOVADo three or more groups differ?Interval data, around 10+ per groupRoughly normal; a significant result still does not say which pair differs
Pearson's rDo two measurements rise and fall together in a straight line?Interval data, 10+ pairs, 30 is betterThe scatter plot looks linear, and no single outlier is driving it
Spearman's rankDo they move together, when the relationship bends?Interval or ordinal, 7+ pairsConsistently one direction; it tolerates curves and outliers that Pearson will not
Chi-squaredAre counts spread across categories differently from what you would expect?Raw counts in categoriesEvery expected value is 5 or more, and they are counts, never percentages or densities

Two habits turn a lucky choice into a defensible one. Plot a histogram of each group before you assume normality; it takes a minute, and it is the whole difference between a t-test you can justify and one you cannot. If it comes out lopsided, that is not a failure: it is the reason you reach for Mann-Whitney instead, and saying so is worth a mark.

Decide whether your question is one-tailed or two-tailed before you run anything. Predicting only that two sites differ is two-tailed. Predicting that the piste specifically holds more diversity is one-tailed. The calculators ask, most fieldwork questions are two-tailed, and picking the one-tailed option because it gives a smaller probability is the kind of thing a moderator notices.

Two tests, two jobsski piste study

The investigation ran both, because it asked both kinds of question. A t-test compared Simpson's index between the two sites, a difference question, interval data, fifteen quadrats in each group, and returned t = 4.26, p = 0.0002.

Spearman's rank then took canopy cover against diversity across all thirty quadrats, which is a relationship question rather than a difference one, and returned rho = −0.77. Spearman rather than Pearson because the scatter bends rather than running straight, and Spearman does not mind.

For your own investigation

Name the test and its justification in the same sentence: what you were asking, what kind of data you had, and which condition you checked. "A t-test was used because the two sites are independent groups of interval data, and a histogram of each showed no strong skew" is one sentence, and it is the whole of the checklist item below.

Outliers, and the small things that cost marks

1 min

Systematically deleting awkward points is not acceptable. If a value has to go, it needs a full justification, and "it looked wrong" is not one. In almost every case the better move is to keep it and explain it, because an explained outlier is evidence of thinking and a deleted one is evidence of nothing.

Then the quiet mark-losers: a table without a title, an axis without units, decimal places implying precision your instrument never had. None of these are hard. All of them are noticed.

An outlier worth keepingski piste study

One forest quadrat scored 4.25 for diversity when the others sat between 1.5 and 2.6. Tempting to call it an error and drop it.

It turned out to be a gap in the canopy, and once canopy cover was measured it stopped being an anomaly and became the clearest single piece of evidence for what was driving the whole pattern. Deleting it would have thrown away the best result in the study.

For your own investigation

Never drop an outlier you cannot explain. Go back to your field notes and photographs and work out what was different about that spot first; a point that disagrees with the others often knows something the others do not.

Using AI at this stepLevel 3 · Targeted AI

It can explain what a test does, what its conditions mean and how to read its output; the thing you would otherwise queue up to ask a teacher.

It cannot choose the test, run your numbers or say what the result means. Choosing is what you have to justify below, and the meaning belongs in step 6, where it is worth six marks.

What this level means

Ready for step 6?

0 of 9

The last one is the hardest to obey and the most valuable.

Next: step 6, analysis and conclusion

Everything you have just been told not to do is what step 6 is for. The patterns you can see in your graphs, what is causing them, and what it means for your research question: all of that goes there, where six marks are waiting for it.

The ski piste investigation used throughout this guide is a teaching reconstruction. The location, sampling design, plant records and soil data come from real fieldwork at Mijoux in the French Jura. Canopy cover measurements were added to illustrate good practice and were not part of the original fieldwork. Photographs are the author’s own, taken at the site.