The A Level Maths Large Data Set Explained, Board by Board
· Webrich Software · 6 min read
Every A Level Maths student studies a Large Data Set (LDS) — a real, published spreadsheet chosen by the exam board. It’s part of the statistics content, and exam questions are written on the assumption that you have worked with it, not just read about it. Students who have only looked at the data once often lose marks on questions that seem easy: interpreting a code, spotting a missing value, or explaining why a conclusion can’t be generalised.
This guide explains what each board’s data set contains and how to revise it efficiently. It’s written for the 2026–27 academic year; boards can and do change their data sets, so always check the current version on your board’s website.
What the Large Data Set is for
The aim is to make statistics feel like real data analysis. Instead of tidy textbook numbers, you work with a large spreadsheet that has:
- many variables, some numerical and some categorical
- codes that stand for categories, such as a number representing a fuel type
- missing values and odd entries
- context — where and when the data was collected, and by whom
In the exam, LDS questions usually test whether you understand that context. Typical questions ask you to:
- explain what a variable or code means, or state its units
- say why a sample from the data might not represent a wider population
- clean the data — decide what to do with a missing value or an outlier
- interpret a chart or summary statistic in context
- spot that a claim doesn’t match what you know about the data
Key idea: you don’t need to remember numbers from the LDS. You need to remember what kind of data it is — and that is something you only learn by using it.
Edexcel: weather data
Edexcel’s LDS is Met Office weather data for 1 May to 31 October in 1987 and 2015. It covers five UK weather stations — Camborne, Heathrow, Hurn, Leeming and Leuchars — and three overseas stations: Beijing, Jacksonville and Perth.
Things to know:
- Perth is in the southern hemisphere, so its May–October data runs through late autumn, winter and spring. That matters if a question compares it with UK summer weather.
- Variables include daily mean temperature, total rainfall, total sunshine, mean windspeed, maximum gust, humidity, cloud cover, visibility and pressure — but not every station records every variable.
- Some values are codes: “tr” means a trace of rain (less than 0.05 mm) and “n/a” means not available.
- Windspeeds are in knots, with Beaufort categories (for example, Light up to 10 knots, Fresh 17 to 21). You should know 1 knot ≈ 1.15 mph.
Revise it in the app’s Edexcel weather data section — the stations and variables subtopic is free, with 15 questions.
AQA: cars data
AQA’s LDS is a Department for Transport extract of cars first registered in one week of June 2002 or June 2016, from the five most-registered makes (BMW, Ford, Toyota, Vauxhall and Volkswagen), with keepers in London, the North West or the South West.
Things to know:
- Many variables are codes: the propulsion type, body type and keeper title are numbers that stand for categories. Count codes — never average them.
- Mass includes an allowance for the driver, and engine size is in cm³.
- There are several emissions variables (CO₂, CO, NOx, particulates and hydrocarbons), measured in g/km — some are missing for some cars.
- Many 2016 cars have company keepers, which hides the sex of the driver. That’s a classic “why can’t you conclude…” exam point.
Revise it in the AQA cars data section; cars data variables is free.
OCR A: census data
OCR A uses census data for England and Wales from 2001 and 2011, arranged by local authority, each with its region and a nine-character code.
Things to know:
- There are two themes: method of travel to work and age structure, each for both census years.
- The first three characters of the code give the type of authority — for example, London boroughs and metropolitan boroughs have different prefixes.
- People are counted where they live, not where they work — a frequent source of exam questions.
- Age groups are unequal classes, so histograms need frequency density (frequency ÷ class width), and “90 and over” is an open class that needs an assumed upper boundary before you estimate a mean.
Revise it in the OCR A census data section; census data variables is free.
MEI: a data set that changes
MEI (OCR B) rotates its Large Data Set, so which one you study depends on when you sit your exams.
Health survey data — A level exams in 2027
For A level exams in June 2027, MEI uses a sample of 200 people from the US National Health and Nutrition Examination Survey (NHANES, 2003–04). Variables include sex, age, weight, height, BMI, several body measurements, pulse and blood pressure readings.
Things to know: BMI = weight (kg) ÷ height (m)², the blood pressure averages follow a specific rule about which readings are used, and #N/A marks missing values. It’s a sample of US residents in 2003–04, so be careful generalising to the UK today. Revise it in the health survey section; health survey variables is free.
Countries data — A level exams from 2028
The next MEI data set covers countries and territories: population, birth and death rates per 1000, median age, CO₂ emissions, GDP per capita, health measures, mobile subscribers and life expectancy over time. It’s used for AS exams in 2027 and A level exams from 2028.
You’ll need rates and densities: births = population × rate ÷ 1000, density = population ÷ area. Revise it in the countries data section; countries data variables is free.
The board-by-board summary
| Board | Data set | Watch out for |
|---|---|---|
| Edexcel | Met Office weather, 1987 and 2015, eight stations | Perth’s seasons, “tr” and “n/a”, knots and Beaufort |
| AQA | DfT cars, registered June 2002 and June 2016 | Codes are categories; company keepers; missing emissions |
| OCR A | England and Wales census, 2001 and 2011 | Counted where people live; unequal and open age classes |
| MEI (2027) | NHANES health survey sample, 2003–04 | US sample; BMI; blood pressure averaging; #N/A |
| MEI (2028) | Countries data | Rates per 1000; per-person and density calculations |
How to revise the LDS
- Open the spreadsheet. Download it from your board’s website, sort it, filter it and draw a few charts. Twenty minutes with the real file is worth hours of reading about it.
- Make a one-page fact sheet. Source, years, locations or groups, every variable with its units, and every code. This is the knowledge exam questions test.
- Learn the quirks. Missing values, codes, unusual units, open classes, odd stations. Examiners write questions around exactly these.
- Practise “in context” answers. “The mean is higher” scores less than “the mean daily rainfall at Camborne was higher in 1987”. Always name the variable and the place, year or group.
- Practise cleaning decisions. For each kind of missing or extreme value, decide whether you would remove it, keep it or investigate — and say why.
Exam tip: when a question says “using your knowledge of the large data set”, the mark is almost always for a specific fact about the data — a code, a unit, a location, a missing variable — not for a general statistical point.
Practise with your board’s data
In the A Level Maths app, each board has its own Large Data Set section with three subtopics — the variables, cleaning the data, and interpreting it — plus revision notes and a Large Data Set mock test. It appears automatically once you choose your board. Browse every section on the topics page, or start with your board’s free questions above.
Frequently asked questions
Do I need to memorise the Large Data Set?
No. You need to be familiar with it — its variables, units, codes, where the data comes from and its quirks — so that you can answer questions about it quickly and spot when a claim doesn't fit the data. You will not be asked to recall individual values.
Will I have the Large Data Set in the exam?
Not the full spreadsheet. Questions give you any values you need, often a small extract, and expect you to bring your familiarity with the context.