Secondary data routeEvery step is showing the version for data somebody else collected. Switch to primary data
Criterion D: Treatment of data · 6 marks200 words · suggested

Treat your data

You arrive with

A downloaded file, cleaned, with an inclusion rule already applied and written down. And ideally a decision about which statistical test you are going to run; if not, make it before you start processing.

You leave with
  • Raw data tabulated, with the full set in an appendix if it is long
  • At least one quantity the file did not contain
  • Your own figures, built from your own values
  • A statistical result, correctly calculated. No interpretation.

Six marks, almost none of the evidence costs a word, and one rule that is specific to this route.

Everything that earns marks here sits outside the 3,000: the tables, the graphs, the calculations. What costs words is describing them, and describing them earns nothing. The rule that belongs to this route: the figures have to be yours. A repository that draws a beautiful chart for you is offering you the one thing you cannot use.

Does not count

Data tables, raw and processed. Every graph and chart. Equations, formulae and calculations. Diagrams and photographs. Citations and the bibliography.

Spend freely. Make them excellent.

Does count

The sentence saying what each table shows. Why you chose that statistical test. Your null and alternative hypotheses. What you did about outliers, and why.

About 200 words in total. That is all you need.

One caution that catches people out: appendices are not read by the examiner. Anything that matters has to be inside the report. Since tables do not count against you wherever they sit, put them in the body where they will actually be seen.

Everything below is how we suggest you actually do it.

Choosing the graph and the test

3 min

The graph type is a real decision, not a default. Comparing two groups wants a bar chart with error bars. Looking for a relationship between two continuous variables wants a scatter plot with a line of best fit. Picking the wrong one hides the very pattern you are trying to show.

Where you can, run an inferential test and report it properly: your null and alternative hypotheses, the statistic, the degrees of freedom, and the probability.

The two hypotheses are one sentence each, written before you test. The null says there is nothing going on, "there is no difference between the two groups", and the alternative says there is, without predicting which way unless you have a reason to. You are not trying to prove the alternative; you are seeing whether the data makes the null hard to believe.

Error bars are worth understanding rather than decorating with. Say which ones you plotted, standard deviation, standard error or a 95% confidence interval, because they answer different questions and a reader cannot tell by looking. And treat them as a picture of spread rather than a significance test: bars that overlap can still hide a real difference, and only the test settles it.

No, you do not show the working

You do not need to reproduce the arithmetic, the formula, or the rank-by-rank calculation. Nobody is marking your ability to add up, and every line of it costs words you need elsewhere.

What has to be visible is the decision trail: which test, why that test for that data, the value of the statistic, n or the degrees of freedom, the critical value or p, the significance level, and the verdict on the null hypothesis. Justifying the choice is where the credit is, because that is a judgement; the arithmetic is not.

The tables are free, since tables of numerical data sit outside the word count. Put a rank table or a chi-squared table of observed and expected values in the body rather than an appendix, where a reader can see the test was set up correctly.

One trap worth naming: chi-squared must be run on raw counts, never percentages or means.

Social Science Statistics
Free online calculators for every test below, t-test, Mann-Whitney U, Spearman's rank, chi-squared, ANOVA, with the working shown step by step.
Open the calculators

A calculator will run any test on any numbers you give it. It has no idea whether that test was the right one, which is the part you are marked on.

A trend line is a claim

A line of best fit says the relationship has a shape, so it needs justifying like any other claim. Straight is not the default: fit what theory or the scatter itself suggests, and your graphing tool will happily fit a curve where one belongs.

Two lines are never defensible: one drawn through a scatter whose correlation is close to zero, and one drawn across categories that have no order, sites and species have no slope to find. Plain dot-to-dot joining is acceptable, and with a handful of points it is often the honest choice.

If you do fit a line, add : the share of the variation your line accounts for, read as a percentage. It is not the correlation coefficient in different clothes. r measures how strongly two variables move together; R² measures how well this particular line fits, and quoting one as the other is a real error.

Never the chart they already made

2 min

Most of the repositories you will use are generous with graphics. Our World in Data will draw you a beautiful line for any indicator; a cantonal portal will map its own monitoring stations; a UN agency will publish the figure from its own report. Every one of those is a trap.

Pasting somebody else's chart into your report demonstrates nothing. It shows no processing, no decisions and no understanding, and it is the single commonest way a secondary investigation loses marks in Criterion D (Treatment of data). The marks here are for what you did to the numbers.

So take the download, not the picture. Every repository worth using has a button that gives you the values behind the chart, and building your own figure from those values is both the requirement and, once you have the file, about five minutes' work.

Earns nothing

A screenshot of the publisher's own graph, captioned and cited.

Correctly attributed, and still evidence of no work.

Earns the marks

Your own chart, from the downloaded values, showing the subset you selected and a quantity you derived, with your own axes and units.

Same underlying data, entirely different piece of evidence.

One exception, handled properly

A published figure can appear in your background, in step 1, as a source like any other: cited, credited, and doing the job of explaining the issue.

What it cannot do is appear in your data section as though it were your processing. If it is somebody else's figure it is a reference, and references live where references live.

Raw, then processed, then stop

2 min

Show the raw data in a table, and where it runs to hundreds of rows show a structured sample in the body with the full set in an appendix. That is the one thing an appendix is genuinely for; nothing that matters should be in there, because appendices are not read.

Then process it into something the file did not contain. This is the part students on this route most often skip, because the data already looks finished. A downloaded column is raw material however tidy it looks, and pasting it into a table is not treatment.

Processing that counts
Normalising: per person, per square kilometre, per unit of output.
Change: a percentage change, a rate per year, a difference from a baseline.
Summarising: a mean or median with its spread, by group or by period.
Indexing: combining columns into a single measure you define and justify.
Joining: bringing a second dataset alongside on a shared key, usually a date or a place.
What was raw and what was processed hereGeneva rivers study

Raw: one row per station per year, the mean E. coli in CFU/ml, and the count of sampling campaigns behind it. Straight out of the file, once the fifty-fold duplication was removed.

Processed: the median and mean for each period, restricted to the stations present in both, and a per-station change from the first period to the last. That last column is the one the file does not contain, and it is the one that produced the finding that 15 of 49 stations got worse.

For your own investigation

Derive at least one quantity your download does not contain, and make it the one your conclusion turns on. If every number in your report could be read straight off the original file, you have tabulated somebody else's data rather than treated your own.

Gaps, zeros and ceilings

2 min

Fieldwork data is missing when you did not collect it, and you know exactly why. Downloaded data is missing for reasons nobody wrote down, and the file rarely distinguishes between them.

Three things to look for, and to say what you did about. A gap: a year or a station simply absent, which your chart must not quietly bridge. A zero: sometimes genuinely none, and often a value below the detection limit, which is a different claim and changes what a mean means. A ceiling: a value that recurs at exactly the top of the range, which is usually the reporting limit rather than a measurement.

None of these has one right answer. All of them have one right behaviour: notice it, decide, and say so in a sentence. An unexplained gap is a hole in your method; an explained one is a limitation, which is worth marks in step 7.

All three, in one fileGeneva rivers study

2007 is absent entirely, so the chart uses bars and leaves the slot empty rather than drawing a line across it. Zeros and values of 0.8 recur, which points to a detection limit rather than sterile water, so the analysis uses medians and treats the low end as "at or below the limit". And one value sits at exactly 1000, alone at the top of the range, which reads as a ceiling and was kept but flagged.

None of that took longer than ten minutes to notice. All of it would have been invisible in the finished report if nobody had looked.

For your own investigation

Sort your values and look at both ends before you calculate anything. Repeated round numbers at the top, repeated zeros at the bottom, and missing years in the middle are the three things a downloaded file will not tell you about itself.

Which test, and how to justify it

4 min

Identical to the fieldwork route, because the statistics do not care where the numbers came from. Start from the claim you are trying to make, not from the name of a test.

A difference
Between groups

Two periods, three catchments, countries above and below a threshold. You are asking whether the groups sit in different places.

A relationship
Between two measurements

As one rises, does the other rise or fall? Both measured on the same unit, so they come in pairs. This is where a second dataset joined on date or place earns its keep.

An association
Across categories

Counts falling into classes: stations by quality band, countries by income group. Tallies of things, not measurements of them.

One piece of vocabulary, because the table turns on it. Interval data is measured on a real scale with units. Ordinal data is ranked but not measured, so the gaps mean nothing: a five-point quality band, a stanine. Categorical data is counted into named classes. A downloaded file will happily give you all three in adjacent columns, and the column heading will not tell you which is which.

TestThe question it answersWhat it needsCheck before you use it
t-testDo two groups differ, on average?Interval data, around 10+ per groupEach group roughly normal, plot a histogram and look
Mann-Whitney UDo two groups differ, when the data is skewed or ranked?Interval or ordinal, 5+ per groupNo normality needed; the two distributions should be a similar shape
ANOVADo three or more groups differ?Interval data, around 10+ per groupRoughly normal; a significant result still does not say which pair differs
Pearson's rDo two measurements rise and fall together in a straight line?Interval data, 10+ pairs, 30 is betterThe scatter plot looks linear, and no single outlier is driving it
Spearman's rankDo they move together, when the relationship bends?Interval or ordinal, 7+ pairsConsistently one direction; it tolerates curves and outliers that Pearson will not
Chi-squaredAre counts spread across categories differently from what you would expect?Raw counts in categoriesEvery expected value is 5 or more, and they are counts, never percentages or densities

Plot a histogram of each group before you assume normality. It takes a minute and it is the whole difference between a t-test you can justify and one you cannot. And decide whether your question is one-tailed or two-tailed before you run anything; most comparisons are two-tailed, and choosing the one-tailed option because it gives a smaller probability is the kind of thing a moderator notices.

One trap belongs to this route in particular. With a large downloaded dataset, significance is cheap: run a correlation on thousands of rows and almost anything comes out significant, including relationships far too weak to matter. Quote the strength as well, r or R squared, and talk about the size of the effect rather than the smallness of the p.

The test, and its justification, in one sentenceGeneva rivers study

The Geneva comparison is a difference question on interval data with 85 observations in one period and 72 in the other, so a t-test answers it. Because the values are strongly skewed, with a long tail of high counts, a Mann-Whitney U on the same two groups is the more defensible choice and was run alongside it.

Saying that, in a sentence, is the whole of the justification: what was being asked, what kind of data it was, which condition was checked, and what the check found.

For your own investigation

Name the test and its justification in the same sentence, and check the condition rather than assuming it. "The data are skewed, so Mann-Whitney rather than a t-test" is one clause and it is worth a mark.

Using AI at this stepLevel 3 · Targeted AI

It can explain what a test does and what its conditions mean, and help you write the spreadsheet formula that turns your raw column into the one you need. Both are genuinely useful and neither is the marked judgement.

It cannot choose the test, run your numbers or tell you what a value is. A tool asked for a figure will produce one that looks right, and on this route a fabricated number is indistinguishable from a real one in the finished report. Every value comes from your file.

What this level means

Ready for step 6?

0 of 13

The last four are this route's own. Number 10 is the one that most often costs marks.

Next: step 6, analysis and conclusion

Everything you have just been told not to do is what step 6 is for. The patterns you can see in your graphs, what is causing them, and what it means for your research question: all of that goes there, where six marks are waiting for it.

The Geneva rivers investigation used on the secondary-data route of this guide is the author’s own analysis of a published cantonal dataset. Every figure quoted comes from the Canton of Geneva’s open data, downloaded on 10 September 2026; the duplication, the missing year and the varying campaign counts described in the guide are as found on that date and may since have been corrected. The selection rules, the analysis and the conclusions are the author’s and not the canton’s.