Skip to content

Edexcel GCSE Maths revision notes

Statistics

Section S
6 specification points

Notes and three levels of exam-style practice for each registered specification point in this section.

Checked against Edexcel 1MA1 section S

Checked against Edexcel 1MA1 section S. Review basis: the qualification registry sourced from the Pearson Edexcel Level 1/Level 2 GCSE (9-1) in Mathematics (1MA1) specification; registry verification recorded 9 July 2026.

How this checking works

In the exam: Formula sheet provided · Paper 1 non-calculator

Loading your tier…

Open the printable pack
S1

Infer properties of populations or distributions from a sample, while knowing the limitations of sampling

Notes
Worked answers & exam appearances →
Evidence from your answers: none yet
Your confidence:

A self-report of how sure you feel. It does not measure mastery. Evidence from your answers reaches secure after the latest Tier 2/3 attempt is correct, with three correct distinct drills across at least two dates and two practice sources.

Explanation

  • A population is the whole group being studied; a sample is the smaller group from which data are collected.
  • If the sample is representative, its proportion or mean can be used to estimate a population value.
  • Scale a sample proportion by the population size when estimating a count.
  • The result is an estimate, not a certainty.
  • Small samples, convenience sampling, under-coverage and non-response can make the sample unrepresentative and the inference unreliable.
Worked example

In a random sample of 7575 residents, 2727 support a proposal. Estimate how many of the town's 12501250 residents support it and state the assumption needed.

  1. 1.Sample proportion =2775=0.36=\dfrac{27}{75}=0.36.
  2. 2.Estimated population count =0.36×1250=450=0.36\times1250=450.
  3. 3.The estimate assumes that the sample is representative of the town's residents.

Answer: About 450450 residents, provided the sample is representative.

Common mistakes

  • Don't fall into the trap of using the sample frequency as the population estimate without scaling.
  • Don't fall into the trap of assuming a large but biased sample must be representative.

Exam tip

State a limitation in context, explaining how it could make the sample unrepresentative.

Tier 1 · Easy

ORIGINAL

1

In a random sample of 5050 library users, 3030 prefer later opening. Estimate how many of the library's 10001000 users prefer later opening.

(2)

(Total for Question 1 is 2 marks)

Tier 2 · Standard

ORIGINAL

1

A random sample of 8080 items from a production run of 800800 contains 1818 items with a surface mark. Estimate the number in the whole run with a surface mark and state one limitation of the estimate.

(3)

(Total for Question 1 is 3 marks)

Tier 3 · Hard

ORIGINAL

1

A service invites a random sample of 450450 customers to answer a survey. Only 270270 reply, and 189189 of the replies support a change. Use the replies to estimate the number of supporters among all 1200012000 customers, then explain a serious limitation.

(4)

(Total for Question 1 is 4 marks)

Your progress and exam materials
S2

Interpret and construct frequency tables, bar charts, pie charts and pictograms (categorical data), vertical line charts (ungrouped discrete numerical data), tables and line graphs (time series)

Notes
Worked answers & exam appearances →
Evidence from your answers: none yet
Your confidence:

A self-report of how sure you feel. It does not measure mastery. Evidence from your answers reaches secure after the latest Tier 2/3 attempt is correct, with three correct distinct drills across at least two dates and two practice sources.

Explanation

  • Match the display to the data. Use separated bars for categories, a vertical line chart for ungrouped discrete numbers, and joined points in time order for a time series.
  • A pictogram needs a clear key.
  • For a pie chart, sector angle=frequencytotal×360\text{sector angle}=\dfrac{\text{frequency}}{\text{total}}\times360^\circ.
  • Any graph needs clear labels, an even scale and accurate plotting.
  • Bar widths and gaps should be consistent, and unrelated categories should not be joined by lines.
Worked example

A 126126^\circ sector of a pie chart represents 4242 students. Find the total number of students.

  1. 1.The sector represents the fraction 126360\dfrac{126}{360} of the total.
  2. 2.Total =42÷126360=42×360126=120=42\div\dfrac{126}{360}=42\times\dfrac{360}{126}=120.

Answer: 120120 students.

Common mistakes

  • Don't fall into the trap of using frequency directly as an angle instead of scaling to 360360^\circ.
  • Don't fall into the trap of joining categorical bars or plotting a time series out of order.

Exam tip

After calculating pie-chart sectors, check that all sector angles add to 360360^\circ.

Tier 1 · Easy

ORIGINAL

1

The data are 2,3,2,5,4,3,2,4,5,22,3,2,5,4,3,2,4,5,2. Construct a frequency table for the values 22, 33, 44 and 55.

(2)

(Total for Question 1 is 2 marks)

Tier 2 · Standard

ORIGINAL

1

In a survey of 240240 journeys, 5454 are made by bicycle. Work out the angle of the bicycle sector in a pie chart.

(2)

(Total for Question 1 is 2 marks)

Tier 3 · Hard

ORIGINAL

1

Quarterly sales are Q1: 320320, Q2: 410410, Q3: 390390, Q4: 520520, and the following Q1: 350350. State the five points to plot on a time-series line graph. Work out the percentage change from the first Q1 to the following Q1 and explain why comparing Q4 directly with the following Q1 could mislead.

(4)

(Total for Question 1 is 4 marks)

S3

Construct and interpret diagrams for grouped discrete and continuous data, i.e. histograms with equal and unequal class intervals and cumulative frequency graphs [Higher only]

Notes
Worked answers & exam appearances →
Evidence from your answers: none yet
Your confidence:

A self-report of how sure you feel. It does not measure mastery. Evidence from your answers reaches secure after the latest Tier 2/3 attempt is correct, with three correct distinct drills across at least two dates and two practice sources.

Explanation

  • Higher tier only. In a histogram, bar area represents frequency and bar height is frequency density: frequency density=frequencyclass width\text{frequency density}=\dfrac{\text{frequency}}{\text{class width}}.
  • Therefore $\text{frequency}=\text{frequency density}\times\text{class width}$.
  • Use the class boundaries for each width, draw touching bars, and label both axes.
  • For a cumulative frequency graph, plot the lower boundary with cumulative frequency 00, then each upper class boundary against its running total.
  • Estimate the median at n/2n/2, the lower quartile at n/4n/4 and the upper quartile at 3n/43n/4.
Worked example

The classes 10<x2510<x\leq25 and 25<x3525<x\leq35 are adjacent. The first has frequency 3030; the second has histogram height 1.41.4. Find the first bar height, the second frequency and the cumulative-frequency points.

  1. 1.First width =15=15, so its height is 30÷15=230\div15=2.
  2. 2.Second width =10=10, so its frequency is 1.4×10=141.4\times10=14.
  3. 3.The cumulative-frequency points are (10,0)(10,0), (25,30)(25,30) and (35,44)(35,44).

Answer: First height 22; second frequency 1414; cumulative points (10,0)(10,0), (25,30)(25,30) and (35,44)(35,44).

Common mistakes

  • Don't fall into the trap of using frequency as the bar height instead of dividing by class width.
  • Don't fall into the trap of plotting cumulative frequencies at class midpoints instead of upper boundaries.

Exam tip

Write the frequency-density formula beside the histogram before calculating any height or area.

Tier 1 · Easy

ORIGINAL

1

A histogram class is 15<x2015<x\leq20 and has frequency 2020. Work out the frequency density.

(2)

(Total for Question 1 is 2 marks)

Tier 2 · Standard

ORIGINAL

1

Grouped data have classes 0<x100<x\leq10, 10<x2510<x\leq25 and 25<x4025<x\leq40 with frequencies 1212, 3030 and 1818. Calculate the three frequency densities and identify the tallest histogram bar.

(3)

(Total for Question 1 is 3 marks)

Tier 3 · Hard

ORIGINAL

1

A cumulative frequency graph for 8080 values passes through (10,8)(10,8), (20,26)(20,26), (30,50)(30,50), (40,70)(40,70) and (50,80)(50,80). Use linear interpolation between the given points to estimate the median and the interquartile range.

(5)

(Total for Question 1 is 5 marks)

S4

Interpret, analyse and compare data-set distributions via graphs (incl. box plots), central tendency (median, mean, mode, modal class) and spread (range, outliers, quartiles, inter-quartile range)

Notes
Worked answers & exam appearances →
Evidence from your answers: none yet
Your confidence:

A self-report of how sure you feel. It does not measure mastery. Evidence from your answers reaches secure after the latest Tier 2/3 attempt is correct, with three correct distinct drills across at least two dates and two practice sources.

Explanation

  • A measure of central tendency describes a typical value: the mean uses every value, the median is the ordered middle, and the mode or modal class is most frequent.
  • Spread describes variation: range is maximum minus minimum.
  • At Higher tier, box plots also show quartiles and the interquartile range Q3Q1Q_3-Q_1.
  • To compare distributions, make one contextual statement about centre and one about spread.
  • Outliers can strongly affect the mean and range, so the median and interquartile range may be more representative.
Worked example

Delivery service A has median time 3838 minutes and range 2222 minutes. Service B has median time 4343 minutes and range 1212 minutes. Compare the distributions.

  1. 1.A has the lower median, so its deliveries are typically quicker.
  2. 2.B has the smaller range, so its delivery times are more consistent.

Answer: Service A is typically quicker, but service B has more consistent delivery times.

Common mistakes

  • Don't fall into the trap of giving comparison figures without stating what they mean in context.
  • Don't fall into the trap of finding the median before putting raw data in order.

Exam tip

A full comparison usually needs both a typical-value statement and a spread statement.

Tier 1 · Easy

ORIGINAL

1

For the data 4,6,6,9,104,6,6,9,10, work out the mean and the range.

(2)

(Total for Question 1 is 2 marks)

Tier 2 · Standard

ORIGINAL

1

Data set A has median 2424 and range 3333. Data set B has median 2525 and range 2929. Compare the two distributions.

(2)

(Total for Question 1 is 2 marks)

Tier 3 · Hard

ORIGINAL

1

The data are 12,13,13,14,15,15,16,5212,13,13,14,15,15,16,52. Work out the mean, median and range. Decide which of the mean or median better describes a typical value for these data. Give a reason.

(4)

(Total for Question 1 is 4 marks)

S5

Apply statistics to describe a population

Notes
Worked answers & exam appearances →
Evidence from your answers: none yet
Your confidence:

A self-report of how sure you feel. It does not measure mastery. Evidence from your answers reaches secure after the latest Tier 2/3 attempt is correct, with three correct distinct drills across at least two dates and two practice sources.

Explanation

  • Statistics can summarise a population or estimate its features from representative samples.
  • When groups have different sizes, combine their means using a weighted mean: add each group size multiplied by its mean, then divide by the total population.
  • When different subgroups have different sample proportions, estimate each subgroup separately before adding.
  • A simple average of group means is only valid when the groups are the same size.
Worked example

A town has 900900 northern and 600600 southern residents. In representative samples, 1212 of 4040 northern residents and 2020 of 5050 southern residents cycle to work. Estimate the town total and percentage.

  1. 1.North estimate =900×1240=270=900\times\dfrac{12}{40}=270.
  2. 2.South estimate =600×2050=240=600\times\dfrac{20}{50}=240.
  3. 3.Total =270+240=510=270+240=510; percentage =5101500×100=34%=\dfrac{510}{1500}\times100=34\%.

Answer: About 510510 residents, or 34%34\% of the town.

Common mistakes

  • Don't fall into the trap of averaging two group means without using the group sizes.
  • Don't fall into the trap of applying one subgroup's sample proportion to the whole population.

Exam tip

Turn each mean back into a total first; combine totals, then divide once.

Tier 1 · Easy

ORIGINAL

1

A representative sample of parcels has mean mass 2.4kg2.4\,\text{kg}. Estimate the total mass of 250250 parcels in the population.

(2)

(Total for Question 1 is 2 marks)

Tier 2 · Standard

ORIGINAL

1

A population contains 120120 junior members with mean attendance 6.56.5 sessions and 8080 senior members with mean attendance 88 sessions. Work out the mean attendance for the whole population.

(3)

(Total for Question 1 is 3 marks)

Tier 3 · Hard

ORIGINAL

1

A town has 18001800 residents in the north and 12001200 in the south. In representative samples, 1515 of 6060 northern residents and 2020 of 5050 southern residents cycle to work. Estimate the total number and percentage of the town's residents who cycle to work.

(5)

(Total for Question 1 is 5 marks)

S6

Use and interpret scatter graphs of bivariate data; recognise correlation, know it does not indicate causation; draw estimated lines of best fit; make predictions; interpolate/extrapolate with caution

Notes
Worked answers & exam appearances →
Evidence from your answers: none yet
Your confidence:

A self-report of how sure you feel. It does not measure mastery. Evidence from your answers reaches secure after the latest Tier 2/3 attempt is correct, with three correct distinct drills across at least two dates and two practice sources.

Explanation

  • A scatter graph displays paired values.
  • An upward pattern shows positive correlation, a downward pattern negative correlation, and no clear pattern no correlation.
  • Draw a line of best fit through the centre of the points with a roughly balanced spread above and below.
  • Use it for estimates: interpolation stays within the observed data range, while extrapolation goes beyond it and is less reliable.
  • Correlation does not prove causation because another variable, reverse causation or coincidence may explain the relationship.
An illustrative scatter graph with an upward trend and a straight line of best fit passing through the centre of the points.
Worked example

A scatter graph compares weekly revision time with test score for revision times from 11 to 88 hours. A sensible line of best fit gives a score of about 7272 at 66 hours and about 8484 when extended to 1010 hours. Interpret both estimates.

  1. 1.66 hours is inside the observed range, so 7272 is an interpolation and is reasonably reliable.
  2. 2.1010 hours is outside the observed range, so 8484 is an extrapolation and is less reliable because the trend may not continue.
  3. 3.The positive correlation does not prove that extra revision alone caused the higher scores; prior attainment could affect both variables.

Answer: 7272 is the more reliable interpolation; 8484 is a less reliable extrapolation, and the graph does not establish causation.

Common mistakes

  • Don't fall into the trap of saying that correlation proves one variable causes the other.
  • Don't fall into the trap of using a line of best fit far outside the observed range without warning.

Exam tip

Answer the command precisely: for 'describe the relationship', write 'as xx increases, yy tends to increase/decrease', not just 'positive/negative'. For a prediction, identify interpolation or extrapolation and comment on reliability.

Tier 1 · Easy

ORIGINAL

1

A scatter graph shows that as daily temperature increases, ice-cream sales usually increase. State the type of correlation and explain why the graph alone does not prove that temperature is the only cause of higher sales.

(2)

(Total for Question 1 is 2 marks)

Tier 2 · Standard

ORIGINAL

1

For data with observed xx-values from 55 to 2020, a line of best fit is y=1.8x+6y=1.8x+6. Estimate yy when x=14x=14. Explain why using the line at x=35x=35 is less reliable.

(3)

(Total for Question 1 is 3 marks)

Tier 3 · Hard

ORIGINAL

1

A scatter graph compares a puppy's age aa months with mass mm kg for ages from 22 to 1212 months. Its estimated line of best fit is m=0.42a+1.8m=0.42a+1.8. Estimate the mass at 99 months and at 1616 months. Comment on the reliability of both estimates and on whether the graph proves that age alone causes the change in mass.

(5)

(Total for Question 1 is 5 marks)

Want help turning these notes into marks?

Bring a tricky specification point or a recent answer, and we can work through the method and exam wording together.