Datasets you can actually have this morning
Every link on this page was checked, leads to numbers rather than to a chart or a report, and needs no account. That last one matters more than it sounds: some of the best environmental archives take a fortnight to grant access, which is fine for a researcher and fatal for you.
Bringing your own is better than taking one from here. This exists so that nobody spends a morning failing to find one.
Opened, counted, and written up
6 files, all downloaded and taken apart. Row counts are what is really in the file, not what the row count says; the traps are what actually bit somebody. They come from one publisher and describe one river system, which means they share a download pattern, a set of station codes and the three strategies below. Six genuinely different investigations off one morning of teaching.
Criterion B (Strategy) needs one real strategy and one explained tension. These are real, dated and sourced. None of them is the obvious winner, which is the useful part: the one with the tightest link to the data has no documented opposition, and the one with the best-evidenced disagreement aims at river shape rather than at water chemistry.
The French State, the Rhône-Alpes region, the Canton of Geneva, the Haute-Savoie department, the Agence de l'eau Rhône-Méditerranée-Corse and basin users, with operational lead shared between the Communauté de communes du Genevois and the canton.
The tensionThe instrument exists because the goals needed reconciling across a border: Geneva legislates and pays for its own network, while part of what arrives comes from French territory it cannot legislate for. Several of the stations that got worse between the two periods are cross-border ones.
SourceThe canton sets the PGEE, which decides sector by sector whether the system is combined or separate. The secondary network is owned by the communes: more than 1,300 km of foul and surface-water sewers and 28 pumping stations. Since 1 January 2015 the FIA mutualises the cost across all of them.
The tensionWho pays, and on what basis. The FIA is funded by a one-off connection charge plus two annual ones: the canton and the communes pay on the impermeable public road surface connected to the network, and property owners pay on their drinking-water consumption. Mutualising across every commune means the bill and the benefit do not land in the same place. Note that I could not find documented opposition to it, so a student choosing this one has to find sourced positions rather than assert the disagreement.
SourceThe Grand Conseil, which in April 1997 added seven articles on renaturation to the cantonal water law of 5 July 1961 and created the fund. It is fed mainly by the hydraulic royalties paid by SIG and the Société des Forces Motrices de Chancy-Pougny, by pumping taxes, by federal subsidies and by donations, at around 11.8 million francs a year with a floor of 6 million.
The tensionThe best-documented of the three. Farmers in the sector judged the project's land footprint oversized, one description calling it science fiction, and argued that agricultural realities had not been taken into account and that farmland is the farmer's principal working tool and inheritance. Part of the land at stake sits in the surfaces d'assolement, the protected cantonal quota of the best arable land. An economic and a cultural claim against an environmental one, which is a tension along three of the five lines the criterion names.
SourceE. coli in Geneva's rivers and at the lake shore
Faecal contamination of the watercourses draining into Lake Geneva, and whether thirty years of investment in wastewater infrastructure shows up in the water.
- 1The file has 37,009 rows and 757 observations. Every station-year is repeated about fifty times, identical apart from the object id. Run a test on the file as downloaded and your sample size is wrong by a factor of fifty, so everything comes out significant.
- 22007 is missing from the dataset entirely.
- 3The number of campaigns behind each annual mean ranges from 1 to 12. A mean from a single visit is not comparable with a mean from twelve, so filter before you compare.
- 4The current year is present but incomplete, with a single campaign behind it.
- 5The station roster grows over time, from 19 stations in 1995 to around 30 now. Comparing early years with late years compares different sets of places unless you restrict to the stations present in both.
- 6One value sits at exactly 1000, which looks like a reporting ceiling rather than a measurement.
- 7Zeros and values of 0.8 recur, which suggests a detection limit rather than genuinely no bacteria.
- 8The only lake bathing station, Pâquis, reads zero in 24 of its 28 years. The receiving water is clean; the story is in the tributaries, and a study aimed at the lake shore has nothing to analyse.
- How has mean annual E. coli at Geneva's river monitoring stations changed between 1995 and 2024?
- A comparison of E. coli between stations on predominantly agricultural catchments and predominantly urban ones.
- Do stations downstream of a wastewater treatment plant differ from stations upstream of one?
Nutrients and major ions in Geneva's rivers, 1969 onwards
Nutrient enrichment of the watercourses draining into Lake Geneva, across the whole period in which phosphate left detergents, treatment plants were built and farming practice changed.
- 1Missing values are recorded as −99, not as blanks, and nothing in the file says so. Twenty parameters carry them. Average silica without removing them and you get a large negative concentration, which is impossible and which a spreadsheet will calculate without complaint.
- 2Silica is −99 in 856 of 950 rows, so that column is 90% missing rather than 90% present. Check how much of a column is real before you build a question on it.
- 31973, 1989 and 1990 are absent from the series entirely.
- 4The network grew from 6 stations in 1969 to about 30 now, so early and late years describe very different sets of places. Any before-and-after comparison needs restricting to the stations present in both.
- 5Station names are not unique. Twenty-two different stations are called "Embouchure", because every river has a mouth. Group by CODEMESURE, never by name, or you will silently merge twenty rivers into one.
- 6There are no unit columns here, unlike the physico-chemistry file. The units are in the documentation, and you have to go and read it.
- How has annual mean phosphate at Geneva's river monitoring stations changed between 1969 and 2025?
- A comparison of nitrate between stations on predominantly agricultural catchments and predominantly urban ones.
- Which parameter most often downgrades a station's grade, and has that changed over fifty years?
Metals in Geneva's rivers
Contamination of watercourses by metals from road runoff, industry and the urban surface, and whether the places that exceed the standards are the places you would expect.
- 1The same −99 convention as the major-elements file, and worse. **21% of mercury values are −99.** The mean of the column as it downloads is −20.555; with those values removed it is 0.403. A fiftyfold error, in the wrong direction, producing a negative concentration that cannot exist.
- 2Gadolinium and vanadium are −99 in about two thirds of rows, so those columns are mostly absent rather than mostly present.
- 3Copper is the downgrading parameter in 255 of 683 station-years, which makes it the obvious variable and also the one everybody else in your class will pick.
- 4Thirty metal columns is an invitation to try relationships until one comes out significant. Decide which metal your question is about before you open the file, and say in your method that you did.
- 5Station names are not unique. Group by CODEMESURE.
- 6These are small numbers with real detection limits, so a value of exactly zero is likelier to mean "below the limit" than "none present".
- A comparison of copper concentrations between stations on urban catchments and rural ones.
- How has zinc at Geneva's river monitoring stations changed between 1995 and 2025?
- Which metal most often determines a station's grade, and does that differ between river types?
Diatom algae in Geneva's rivers
Whether the microscopic algae living on the river bed report the same water quality the chemistry does, and what it means when they disagree.
- 1The file publishes 24,759 rows and holds 503 observations. Every station-year is repeated, usually 50 times but sometimes 28, 32, 33 or 34, identical apart from an internal id. The duplication is not even consistent, so checking one group and assuming the rest match will still leave you wrong.
- 2Most scores rest on **one or two campaigns**, not twelve. That is far less underlying sampling than the chemistry files, so a single year's score is a noisier number than it looks.
- 3The index is a made number rather than a measured one: it summarises a whole community into one value using a published formula. Read what it weights before you build a conclusion on it.
- 4Because the scores come already sorted into five bands, it is tempting to analyse the bands. They are ordinal, so a mean of them means nothing. Use the numeric index and keep the classes for description.
- 5Station names are not unique. Group by CODEMESURE.
- A comparison of the diatom index between stations upstream and downstream of urban areas.
- Do the diatom index and the chemistry agree about which stations are the worst?
- How has the diatom index at Geneva's river monitoring stations changed since 1998?
Benthic invertebrates in Geneva's rivers
Whether the insect larvae, worms and crustaceans living on the river bed recovered as the water got cleaner, and where they did not.
- 1The cleanest file in the family: 554 rows and 554 observations, no duplication at all. Which is itself the lesson, because two of its siblings are duplicated fiftyfold and nothing on the outside distinguishes them. Every file has to be checked on its own.
- 2**The series stops in 2022.** It sits alongside files running to 2025, so a comparison across the family silently ends three years early unless you notice.
- 3Scores rest on between 1 and 4 samples, and 46 of them rest on a single sample. Filter, or say why you did not.
- 4The index has a ceiling of 20, so improvement at an already-good station is compressed. A station moving from 17 to 18 is not comparable with one moving from 4 to 5.
- 5Station names are not unique. Group by CODEMESURE.
- How has the biological index at Geneva's river monitoring stations changed between 1995 and 2022?
- A comparison of invertebrate index scores between renatured reaches and channelised ones.
- Does the number of taxa tell the same story as the index score?
Full physico-chemistry of Geneva's rivers, 1962 to 2017
The chemistry of the watercourses draining into Lake Geneva across the whole period of post-war industrialisation, sewage treatment and agricultural change.
- 1**It stops in 2017.** It is the biggest and most detailed file in the family, which makes it look like the flagship, and it is the discontinued one. The major-elements file carries most of the same parameters through to 2025. Work out which one your question needs before you commit a day to this.
- 2Missing values are marked differently here from the rest of the family: small negative numbers such as −0.005 or −0.05, rather than −99. Same publisher, same subject, two conventions, and no warning in either file.
- 3Discharge is present in only 502 of 836 rows, so a question that normalises by flow loses 40% of the data before it starts.
- 4About 120 columns, half of them unit columns. The width is a genuine hazard: it makes trying relationships until one is significant almost effortless.
- 5Station names are not unique. Group by CODEMESURE.
- 6The overlap with the major-elements file is large but not exact. If you use both, harmonise the parameter names and check the units, because only one of the two carries unit columns.
- How did nitrate at Geneva's river monitoring stations change between 1962 and 2017?
- A comparison of dissolved oxygen before and after a named wastewater treatment plant opened.
- Which parameters improved over the period, and which did not improve at all?
Places to start looking
One rung shallower. Every link below was confirmed to resolve and to lead to numbers, and none of them has been opened and counted. Applying step 4’s four tests to one of these and finding it fails is a good morning’s work, not a broken promise. Nearest first, because you are likelier to have something to say about a river you have walked along.
Geneva
Fish community assessments by station under the Swiss monitoring scheme, 2009 to 2025. Only 92 rows across 55 stations, which is too few for most statistical tests, so treat it as a companion to one of the vetted files rather than the basis of an investigation.
Switzerland
Temperature, precipitation, sunshine and more, per station, with historical files going back decades and a separate station metadata file carrying altitude and coordinates. The best partner dataset in this list: joining rainfall to almost anything else on date is the clearest demonstration of independent processing you can make.
Everything the Confederation, the cantons and the communes publish, in one searchable place, including all the Geneva water datasets above. Search by keyword and filter the format to CSV, or you will spend the morning opening map layers.
Discharge, water level and water temperature at gauging stations across Switzerland, including the Rhône and the lake itself. Excellent data with an access catch: anything older than the recent window has to be requested, which is free but not immediate, so check before you build a question on it.
The Alps
Every length-change observation held for Swiss glaciers, some series running back into the nineteenth century, plus seasonal mass balance for selected glaciers. About as clean a climate-signal dataset as exists, and long enough that the question becomes which decade changed rather than whether anything did.
Ground temperatures, borehole profiles and rock glacier movement across the Swiss Alps. A harder starting point than GLAMOS and a genuinely current issue, since thawing permafrost is what destabilises high mountain infrastructure.
Snow, forest, avalanche and landscape datasets published by WSL and SLF researchers. Research-grade rather than official statistics, so read the documentation first, and the reward is data nobody else in your class will have.
France
Over 200 million analyses across more than 20,000 stations covering the whole of France, including nitrates, pesticides and metals. The query URL is your extraction protocol, which makes this one of the most repeatable methods a student can write, and the French bank of Léman is in it.
The French equivalent of opendata.swiss, and the place to find départemental and communal data for the areas across the border. Useful when a Geneva question needs its French half.
Europe
Measured concentrations from monitoring stations across Europe, with the historical archive alongside recent years. Comparing cities means thinking hard about what else differs between them, which is exactly the controlling work step 4 asks for.
Bathing water quality, land cover, emissions, waste and protected areas, compiled to a common standard across member states. Comparable definitions across countries is the thing this buys you and the thing most global datasets cannot promise.
Reanalysis climate data: temperature, precipitation, sea ice, at grid points across the world and back several decades. Powerful and the steepest learning curve in this list, so a good choice only if you already know your way around a spreadsheet.
Iceland
Monthly and annual values for temperature, precipitation and more, from 1961 onwards, station by station. A North Atlantic climate series to set against an Alpine one, which is a comparison very few students think to make.
Terminus positions for thirty to forty glaciers, measured annually by volunteers since 1930, which is one of the longest citizen-science environmental records anywhere. Worth trying, and worth having a second choice ready.
Global
Energy, emissions, food, land use and biodiversity, compiled from original sources into long comparable series. Take the CSV and never the chart: the picture is theirs, the analysis has to be yours, and cite the original source as well as them.
Development indicators alongside environmental ones: forest area, freshwater withdrawal, access to clean cooking fuel, national income. The obvious place to go when your question pairs an environmental variable with a socio-economic one.
Food production, livestock, fertiliser and pesticide use, and agricultural emissions, country by country and year by year. Fertiliser application per hectare against water quality is a classic ESS pairing that this file makes possible.
Air quality measurements from monitoring stations worldwide, including places with no national portal of their own. Coverage is uneven by design, which is itself worth writing about honestly in your evaluation.
Annual tree cover loss, fire alerts and carbon emissions from forest change, derived from satellite imagery. Read how the classification works before you build on it: plantations and native forest can look alike from orbit, which is the single most quotable limitation in this list.
Species occurrence records with coordinates and dates, pooled from museums, surveys and citizen science. Records where people looked rather than where species are, so anything about abundance needs care, and every download comes with its own DOI to cite.
Things that look like data and are not
Step 4 teaches four tests. This is the shortest way to show why they exist: each of these is a real published resource a student would reasonably click on, and none can carry an investigation.
It is a status map, not a series. Thirty-five rows, one per beach, each carrying a word such as “Bonne” and the date it was last refreshed. There is no number to process and no history to compare, and you cannot tell that from its title.
Four tests, and they take a couple of minutes each. Is it numbers rather than a picture of numbers? Can you say what one row is? Does it cover your years and your places? And can you have it on screen in the next ten minutes?
Step 4 · Find and select your data