Skip to content

Edexcel A-level Maths revision notes

Data presentation and interpretation

Section S2
Year 1
Year 1: this is the AS subject content the exam board publishes, which is what most schools teach in Year 12.
4 specification points

Notes and three levels of exam-style practice for each registered specification point in this section.

Checked against Edexcel 9MA0 section S2

Checked against Edexcel 9MA0 section S2. Review basis: the qualification registry sourced from the Pearson Edexcel Level 3 Advanced GCE in Mathematics (9MA0) specification; registry verification recorded 11 July 2026.

How this checking works

In the exam: Formulae booklet provided · calculator allowed in every paper

Open the printable pack
S2.1

Interpret diagrams for single-variable data, including understanding that area in a histogram represents frequency; connect to probability distributions.

Notes
Worked answers & exam appearances →
Evidence from your answers: none yet
Your confidence:

A self-report of how sure you feel. It does not measure mastery. Evidence from your answers reaches secure after the latest Tier 2/3 attempt is correct, with three correct distinct drills across at least two dates and two practice sources.

Explanation

  • In a histogram, frequency is proportional to bar area and frequency density is frequencyclass width\frac{\text{frequency}}{\text{class width}}; unequal class widths make bar height alone misleading.
  • Frequency polygons show class patterns, box plots summarise centre and spread, and cumulative frequency diagrams support estimates of medians, quartiles and counts below a value.
  • For a continuous probability density histogram, total area is 11 and the area above an interval is the probability of an observation in that interval.
  • Read class boundaries and axis scales before calculating.
  • A common error is to use frequency density as though it were frequency.
In a histogram, bar area represents frequency, so height is frequency density.
Worked example

A cumulative frequency diagram represents 8080 observations. The cumulative frequencies at x=10x=10 and x=20x=20 are 1818 and 5454 respectively. Estimate the number of observations satisfying 10<x2010<x\leq20.

  1. 1.The cumulative frequency at 2020 counts all 5454 observations up to that value, including the 1818 already counted up to 1010.
  2. 2.Subtract to isolate the interval: 5418=3654-18=36.

Answer: 3636

Common mistakes

  • Don't use bar height rather than class width times frequency density to compare histogram frequencies.
  • Don't read cumulative frequencies directly as class frequencies instead of subtracting the boundary totals.

Exam tip

For cumulative-frequency intervals, subtract the cumulative totals at the two boundaries and respect endpoint inequalities.

Tier 1 · Easy

ORIGINAL

1.

A histogram class is 12t<1712\leq t<17 and has frequency density 3.63.6. Find the frequency in this class.

(2)

(Total for Question 1 is 2 marks)

Tier 2 · Standard

ORIGINAL

1.

In a histogram, the class 10x<1510\leq x<15 has frequency density 44, and the class 15x<2515\leq x<25 has frequency density 2.52.5. Find the frequency in each class and the total frequency represented by these two bars.

(4)

(Total for Question 1 is 4 marks)

Tier 3 · Hard

ORIGINAL

1.

The probability density histogram for a continuous random variable XX has constant heights 0.120.12 on 0x<30\leq x<3, 0.080.08 on 3x<83\leq x<8, and 0.120.12 on 8x108\leq x\leq10. Verify that it defines a probability distribution and find P(1<X<6)P(1<X<6).

(5)

(Total for Question 1 is 5 marks)

Your progress and exam materials

This section: Evidence from your answers: 0/4 secureYour confidence: 0 self-rated secureTracker status: 0/4 secure, 0 shaky, 4 unseen

Overall: Evidence from your answers: 0/89 secureYour confidence: 0 self-rated secureTracker status: 0/89 secure, 0 shaky, 89 unseen

Progress is saved on this device for guests and accounts right now; cross-device account sync is not live yet.

Answer conventions

Follow the wording on the question and its mark scheme. awrt means an appropriately rounded value is accepted; an exact answer must stay as a fraction, surd, logarithm or multiple of π when required, and a rounded decimal may be disallowed. Include requested units and forms. A cso tag protects that accuracy mark, while earlier method marks follow the question-specific dependencies.

S2.2

Interpret scatter diagrams and regression lines for bivariate data, including recognising distinct sections of the population (regression calculations excluded); interpret correlation informally; correlation does not imply causation.

Notes
Worked answers & exam appearances →
Evidence from your answers: none yet
Your confidence:

A self-report of how sure you feel. It does not measure mastery. Evidence from your answers reaches secure after the latest Tier 2/3 attempt is correct, with three correct distinct drills across at least two dates and two practice sources.

Explanation

  • A scatter diagram shows paired bivariate data; describe the direction and strength of its association and note any outliers or clusters.
  • A regression line estimates the mean response for a given explanatory-variable value and is most defensible within the observed data range.
  • Distinct clusters may represent different sections of the population, so one overall correlation or regression line can hide different within-group patterns.
  • Correlation measures association, not causation; a lurking variable, reverse causation or coincidence may explain the observed relationship.
  • For y=axny=ax^n, plot logy\log y against logx\log x to get gradient nn and intercept loga\log a; for y=kbxy=kb^x, plot logy\log y against xx to get gradient logb\log b and intercept logk\log k.
Worked example

A scatter diagram of journey distance against journey time contains one cluster for bicycles and a separate cluster for cars. Explain why fitting one regression line to all journeys may be misleading.

  1. 1.Identify vehicle type as a grouping variable.
  2. 2.Since the relationship between distance and time is likely to differ between bicycles and cars, pooling the groups can produce a line that fits neither group well.
  3. 3.Analyse the sections separately.

Answer: The two clusters represent distinct sections of the population with different travel speeds.; A single line may mainly reflect the gap between the clusters rather than either within-group relationship.; Separate regression lines or separate analyses would be more informative.

Common mistakes

  • Don't extrapolate a regression line far beyond the observed data range.
  • Don't interpret an overall regression line despite distinct clusters representing different subpopulations.

Exam tip

Before using a regression line, inspect clusters and explain whether one relationship is credible for the whole population.

Tier 1 · Easy

ORIGINAL

1.

A scatter diagram of daily ice-cream sales against temperature shows a strong positive correlation. Interpret this and explain why it does not prove that higher temperature causes every increase in sales.

(2)

(Total for Question 1 is 2 marks)

Tier 2 · Standard

ORIGINAL

1.

Across a year, ice-cream sales and the number of reported sunburn cases have strong positive correlation. Explain why this does not establish that buying ice cream causes sunburn.

(3)

(Total for Question 1 is 3 marks)

Tier 3 · Hard

ORIGINAL

1.

For trees aged between 33 and 1818 years, a regression line of trunk diameter dd cm on age aa years is d=1.8a+4.2d=1.8a+4.2. A 1212-year-old tree has diameter 2929 cm. Interpret the difference between the observed and predicted values and critique using the line to predict the diameter of a 4040-year-old tree.

(4)

(Total for Question 1 is 4 marks)

S2.3

Interpret measures of central tendency and variation, extending to standard deviation; be able to calculate standard deviation, including from summary statistics.

Notes
Worked answers & exam appearances →
Evidence from your answers: none yet
Your confidence:

A self-report of how sure you feel. It does not measure mastery. Evidence from your answers reaches secure after the latest Tier 2/3 attempt is correct, with three correct distinct drills across at least two dates and two practice sources.

Explanation

  • The mean uses every value, while the median is resistant to extremes; choose a measure that suits the distribution and context.
  • Range and interquartile range measure spread using endpoints or quartiles; standard deviation measures typical spread about the mean using all observations.
  • For nn data values, use σ=x2n(xn)2\sigma=\sqrt{\frac{\sum x^2}{n}-\left(\frac{\sum x}{n}\right)^2} unless a different convention is stated.
  • When groups are combined, add nn, x\sum x and x2\sum x^2 before recalculating; averaging separate standard deviations is not valid.
  • Coding data by y=(xa)/by=(x-a)/b simplifies calculation, after which transform the statistics back; grouped-data percentiles require linear interpolation within the relevant class.
Worked example

For 2020 observations, x=310\sum x=310 and x2=5020\sum x^2=5020. Calculate the population standard deviation to 33 significant figures.

  1. 1.The mean is 310/20=15.5310/20=15.5.
  2. 2.Hence σ=5020/2015.52=251240.25=10.75=3.278\sigma=\sqrt{5020/20-15.5^2}=\sqrt{251-240.25}=\sqrt{10.75}=3.278\ldots, so the standard deviation is 3.283.28.

Answer: 3.283.28

Common mistakes

  • Don't average two group means without weighting them by their different sample sizes.
  • Don't use Sxx/nS_{xx}/n as the standard deviation without taking the square root.

Exam tip

Substitute the stated divisor into the variance formula, keep the square root until the end and check the result is non-negative.

Tier 1 · Easy

ORIGINAL

1.

Calculate the mean and population standard deviation of 4,7,7,8,94,7,7,8,9. Give the standard deviation to 33 significant figures.

(3)

(Total for Question 1 is 3 marks)

Tier 2 · Standard

ORIGINAL

1.

A data set xx has mean 1212 and standard deviation 33. A new variable is defined by y=5+2xy=5+2x. Find the mean and standard deviation of yy.

(3)

(Total for Question 1 is 3 marks)

Tier 3 · Hard

ORIGINAL

1.

Group A has 1212 values with mean 1818 and x2=3996\sum x^2=3996. Group B has 88 values with mean 2424 and x2=4736\sum x^2=4736. Find the mean and population standard deviation of all 2020 values.

(5)

(Total for Question 1 is 5 marks)

S2.4

Recognise and interpret possible outliers in data sets and statistical diagrams; select or critique data presentation techniques in context; clean data, including dealing with missing data, errors and outliers.

Notes
Worked answers & exam appearances →
Evidence from your answers: none yet
Your confidence:

A self-report of how sure you feel. It does not measure mastery. Evidence from your answers reaches secure after the latest Tier 2/3 attempt is correct, with three correct distinct drills across at least two dates and two practice sources.

Explanation

  • A common outlier rule flags values below Q11.5IQRQ_1-1.5\operatorname{IQR} or above Q3+1.5IQRQ_3+1.5\operatorname{IQR}, but context should guide the final decision.
  • Investigate a suspicious value against the original record before correcting or removing it; an unusual valid observation is not automatically an error.
  • Handle missing data transparently: record how many values are missing, avoid inventing unsupported values and consider whether missingness could bias conclusions.
  • Choose displays to suit the data and purpose: histograms for continuous grouped data, box plots for comparing distributions, and scatter diagrams for paired variables.
Worked example

A table of package masses contains one blank entry and one value 482482 among values near 48.248.2 grams. Describe a defensible way to clean these two entries before analysis.

  1. 1.First distinguish a data-entry error from a genuine extreme value by consulting the source.
  2. 2.Do not silently divide 482482 by 1010.
  3. 3.The blank contains no observed value, so omit it from calculations unless a justified imputation rule has been chosen, and document the decision so its possible bias is visible.

Answer: Check the original measurement record for both entries.; Correct 482482 to 48.248.2 only if the source confirms a decimal-point error; otherwise retain and flag it or exclude it with a stated reason.; Treat the blank as missing rather than replacing it without evidence, and report the reduced sample size or justified imputation method.

Common mistakes

  • Don't replace every missing value with zero, changing both the centre and spread of the data.
  • Don't delete an unusual value automatically instead of investigating whether it is an error or genuine observation.

Exam tip

Document separate rules for missing values, transcription errors and plausible outliers before recalculating summaries.

Tier 1 · Easy

ORIGINAL

1.

For a data set, Q1=14Q_1=14 and Q3=22Q_3=22. Use the 1.5IQR1.5\operatorname{IQR} rule to determine whether the value 3636 is a possible outlier.

(2)

(Total for Question 1 is 2 marks)

Tier 2 · Standard

ORIGINAL

1.

For a data set, the lower quartile is 1818 and the upper quartile is 3030. Use the 1.5×IQR1.5\times\operatorname{IQR} rule to decide whether a value of 5050 is an outlier, and state what should be done before removing it.

(4)

(Total for Question 1 is 4 marks)

Tier 3 · Hard

ORIGINAL

1.

A data set of 4040 readings was summarised as x=504\sum x=504 and x2=6856.9\sum x^2=6856.9. One reading was entered as 3131 but the source record confirms it should be 1313. Calculate the corrected mean and population standard deviation, and state the likely effect of the error on the original spread.

(5)

(Total for Question 1 is 5 marks)

Want help turning these notes into marks?

Bring a tricky specification point or a recent answer, and we can work through the method and exam wording together.