S2 Data presentation and interpretation — revision question pack

4 specification points · notes, questions, answers and worked methods

Checked against Edexcel 9MA0 section S2. Review basis: the qualification registry sourced from the Pearson Edexcel Level 3 Advanced GCE in Mathematics (9MA0) specification; registry verification recorded 11 July 2026.

How this checking works

S2.1 · Interpret diagrams for single-variable data, including understanding that area in a histogram represents frequency; connect to probability distributions.

Explanation

  • In a histogram, frequency is proportional to bar area and frequency density is frequencyclass width\frac{\text{frequency}}{\text{class width}}; unequal class widths make bar height alone misleading.
  • Frequency polygons show class patterns, box plots summarise centre and spread, and cumulative frequency diagrams support estimates of medians, quartiles and counts below a value.
  • For a continuous probability density histogram, total area is 11 and the area above an interval is the probability of an observation in that interval.
  • Read class boundaries and axis scales before calculating.
  • A common error is to use frequency density as though it were frequency.
In a histogram, bar area represents frequency, so height is frequency density.

Worked example

A cumulative frequency diagram represents 8080 observations. The cumulative frequencies at x=10x=10 and x=20x=20 are 1818 and 5454 respectively. Estimate the number of observations satisfying 10<x2010<x\leq20.

  1. 1.The cumulative frequency at 2020 counts all 5454 observations up to that value, including the 1818 already counted up to 1010.
  2. 2.Subtract to isolate the interval: 5418=3654-18=36.

Answer: 3636

Common mistakes

  • Don't use bar height rather than class width times frequency density to compare histogram frequencies.
  • Don't read cumulative frequencies directly as class frequencies instead of subtracting the boundary totals.

Exam tip

For cumulative-frequency intervals, subtract the cumulative totals at the two boundaries and respect endpoint inequalities.

Tier 1 · Easy

  1. 1.

    A histogram class is 12t<1712\leq t<17 and has frequency density 3.63.6. Find the frequency in this class.

    (2)

    (Total for Question 1 is 2 marks)

  2. 2.

    Two bars in a histogram represent equal frequencies. The first class has width 66 and frequency density 44. The second class has width 99. Find its frequency density.

    (2)

    (Total for Question 2 is 2 marks)

Tier 2 · Standard

  1. 1.

    In a histogram, the class 10x<1510\leq x<15 has frequency density 44, and the class 15x<2515\leq x<25 has frequency density 2.52.5. Find the frequency in each class and the total frequency represented by these two bars.

    (4)

    (Total for Question 1 is 4 marks)

  2. 2.

    On a cumulative frequency diagram for 8080 observations, the graph segment from (30,22)(30,22) to (42,58)(42,58) is a straight line. Estimate the median and the 6565th percentile.

    (4)

    (Total for Question 2 is 4 marks)

  3. 3.

    The vertical scale is missing from a histogram. The class 0x<50\leq x<5 has frequency 3030 and a bar height of 88 grid units. The class 5x<115\leq x<11 has a bar height of 44 grid units. Find the frequency density represented by one grid unit and the frequency in the second class.

    (4)

    (Total for Question 3 is 4 marks)

Tier 3 · Hard

  1. 1.

    The probability density histogram for a continuous random variable XX has constant heights 0.120.12 on 0x<30\leq x<3, 0.080.08 on 3x<83\leq x<8, and 0.120.12 on 8x108\leq x\leq10. Verify that it defines a probability distribution and find P(1<X<6)P(1<X<6).

    (5)

    (Total for Question 1 is 5 marks)

  2. 2.

    A histogram of 6060 observations has frequency densities 3.23.2, 2.02.0 and kk in the classes 0x<50\leq x<5, 5x<125\leq x<12 and 12x<2012\leq x<20 respectively. Find kk. Hence find the probability that a randomly selected observation from the data is at least 1010.

    (6)

    (Total for Question 2 is 6 marks)

  3. 3.

    A histogram has frequency densities 22, 55 and 33 in the classes 0x<50\leq x<5, 5x<95\leq x<9 and 9x<159\leq x<15 respectively. Estimate the median and the interquartile range, assuming observations are distributed uniformly within each class.

    (6)

    (Total for Question 3 is 6 marks)

  4. 4.

    Two histograms use the same class boundaries and both vertical axes show frequency density. In the class 10x<1510\leq x<15, sample A has 8080 observations in total and frequency density 44, while sample B has 120120 observations in total and frequency density 55. Calculate the frequency and the percentage of each sample in this class. Hence explain why comparing the two bar heights alone gives the wrong conclusion about which sample has the larger proportion in the class.

    (5)

    (Total for Question 4 is 5 marks)

  5. 5.

    A histogram of 7070 observations has frequency densities 33, 55, 22 and 44 in the classes 0x<40\leq x<4, 4x<104\leq x<10, 10x<a10\leq x<a and ax<20a\leq x<20, respectively, where 10<a<2010<a<20. Find aa. Assuming observations are distributed uniformly within each class, estimate the median and find the probability that a randomly selected observation is at least 1414.

    (7)

    (Total for Question 5 is 7 marks)

S2.2 · Interpret scatter diagrams and regression lines for bivariate data, including recognising distinct sections of the population (regression calculations excluded); interpret correlation informally; correlation does not imply causation.

Explanation

  • A scatter diagram shows paired bivariate data; describe the direction and strength of its association and note any outliers or clusters.
  • A regression line estimates the mean response for a given explanatory-variable value and is most defensible within the observed data range.
  • Distinct clusters may represent different sections of the population, so one overall correlation or regression line can hide different within-group patterns.
  • Correlation measures association, not causation; a lurking variable, reverse causation or coincidence may explain the observed relationship.
  • For y=axny=ax^n, plot logy\log y against logx\log x to get gradient nn and intercept loga\log a; for y=kbxy=kb^x, plot logy\log y against xx to get gradient logb\log b and intercept logk\log k.

Worked example

A scatter diagram of journey distance against journey time contains one cluster for bicycles and a separate cluster for cars. Explain why fitting one regression line to all journeys may be misleading.

  1. 1.Identify vehicle type as a grouping variable.
  2. 2.Since the relationship between distance and time is likely to differ between bicycles and cars, pooling the groups can produce a line that fits neither group well.
  3. 3.Analyse the sections separately.

Answer: The two clusters represent distinct sections of the population with different travel speeds.; A single line may mainly reflect the gap between the clusters rather than either within-group relationship.; Separate regression lines or separate analyses would be more informative.

Common mistakes

  • Don't extrapolate a regression line far beyond the observed data range.
  • Don't interpret an overall regression line despite distinct clusters representing different subpopulations.

Exam tip

Before using a regression line, inspect clusters and explain whether one relationship is credible for the whole population.

Tier 1 · Easy

  1. 1.

    A scatter diagram of daily ice-cream sales against temperature shows a strong positive correlation. Interpret this and explain why it does not prove that higher temperature causes every increase in sales.

    (2)

    (Total for Question 1 is 2 marks)

  2. 2.

    A scatter diagram contains one compact cluster for electric bicycles and a separate compact cluster for conventional bicycles. State why one regression line for all the points may be misleading.

    (2)

    (Total for Question 2 is 2 marks)

Tier 2 · Standard

  1. 1.

    Across a year, ice-cream sales and the number of reported sunburn cases have strong positive correlation. Explain why this does not establish that buying ice cream causes sunburn.

    (3)

    (Total for Question 1 is 3 marks)

  2. 2.

    For recordings made at wind speeds between 55 and 2525 km h1^{-1}, the regression line of sound level ss dB on wind speed ww km h1^{-1} is s=681.4ws=68-1.4w. Find the predicted sound level when w=18w=18 and interpret the gradient in context.

    (3)

    (Total for Question 2 is 3 marks)

  3. 3.

    A scatter diagram plots advertising cost in pounds against sales income in pounds, and its fitted line has gradient 0.400.40. The sales income is then converted from pounds to pence, without changing any observations. State the new gradient and describe the effect of this conversion on the direction and strength of the correlation.

    (3)

    (Total for Question 3 is 3 marks)

Tier 3 · Hard

  1. 1.

    For trees aged between 33 and 1818 years, a regression line of trunk diameter dd cm on age aa years is d=1.8a+4.2d=1.8a+4.2. A 1212-year-old tree has diameter 2929 cm. Interpret the difference between the observed and predicted values and critique using the line to predict the diameter of a 4040-year-old tree.

    (4)

    (Total for Question 1 is 4 marks)

  2. 2.

    A scatter diagram compares hours of optional practice with assessment score for two courses. Students on the advanced course form a cluster with both higher practice hours and higher scores than students on the introductory course. Within each course cluster, score tends to decrease slightly as practice hours increase, but the pooled data show positive correlation. Explain why the pooled correlation and a single regression line could be misleading, and why the diagram does not show that extra practice lowers scores.

    (5)

    (Total for Question 2 is 5 marks)

  3. 3.

    For observations with 1x91\leq x\leq9, the fitted equation predicting yy from xx is y=2x1y=2x-1. After the observation (9,25)(9,25) is omitted, the equation becomes y=1.2x+1.0y=1.2x+1.0. Calculate the residual of (9,25)(9,25) from the second fit and compare the two predicted values at x=8x=8. Interpret what these results show, and state why the observation should not be deleted solely on this evidence.

    (5)

    (Total for Question 3 is 5 marks)

  4. 4.

    For observations with 2x202\leq x\leq20, a regression line of log10y\log_{10}y on log10x\log_{10}x is log10y=log102+1.5log10x\log_{10}y=\log_{10}2+1.5\log_{10}x. Obtain a model of the form y=axny=ax^n, and use it to estimate yy when x=8x=8. Comment on using the model when x=50x=50, and explain why a strong linear correlation on the transformed plot would not by itself establish a causal relationship.

    (6)

    (Total for Question 4 is 6 marks)

  5. 5.

    For a bivariate data set with xˉ=10\bar{x}=10 and yˉ=30\bar{y}=30, the fitted equation for predicting yy from xx is y=2x+10y=2x+10. The fitted equation for predicting xx from yy is x=0.4y2x=0.4y-2. Estimate yy when x=12x=12, stating which equation you use. Show that both equations pass through (xˉ,yˉ)(\bar{x},\bar{y}). Explain why rearranging the second equation gives a different estimate and is not the appropriate procedure here.

    (5)

    (Total for Question 5 is 5 marks)

S2.3 · Interpret measures of central tendency and variation, extending to standard deviation; be able to calculate standard deviation, including from summary statistics.

Explanation

  • The mean uses every value, while the median is resistant to extremes; choose a measure that suits the distribution and context.
  • Range and interquartile range measure spread using endpoints or quartiles; standard deviation measures typical spread about the mean using all observations.
  • For nn data values, use σ=x2n(xn)2\sigma=\sqrt{\frac{\sum x^2}{n}-\left(\frac{\sum x}{n}\right)^2} unless a different convention is stated.
  • When groups are combined, add nn, x\sum x and x2\sum x^2 before recalculating; averaging separate standard deviations is not valid.
  • Coding data by y=(xa)/by=(x-a)/b simplifies calculation, after which transform the statistics back; grouped-data percentiles require linear interpolation within the relevant class.

Worked example

For 2020 observations, x=310\sum x=310 and x2=5020\sum x^2=5020. Calculate the population standard deviation to 33 significant figures.

  1. 1.The mean is 310/20=15.5310/20=15.5.
  2. 2.Hence σ=5020/2015.52=251240.25=10.75=3.278\sigma=\sqrt{5020/20-15.5^2}=\sqrt{251-240.25}=\sqrt{10.75}=3.278\ldots, so the standard deviation is 3.283.28.

Answer: 3.283.28

Common mistakes

  • Don't average two group means without weighting them by their different sample sizes.
  • Don't use Sxx/nS_{xx}/n as the standard deviation without taking the square root.

Exam tip

Substitute the stated divisor into the variance formula, keep the square root until the end and check the result is non-negative.

Tier 1 · Easy

  1. 1.

    Calculate the mean and population standard deviation of 4,7,7,8,94,7,7,8,9. Give the standard deviation to 33 significant figures.

    (3)

    (Total for Question 1 is 3 marks)

  2. 2.

    Eight readings have mean 1414. A ninth reading, 1616, is added. Find the mean of all nine readings, giving your answer to 33 significant figures.

    (2)

    (Total for Question 2 is 2 marks)

Tier 2 · Standard

  1. 1.

    A data set xx has mean 1212 and standard deviation 33. A new variable is defined by y=5+2xy=5+2x. Find the mean and standard deviation of yy.

    (3)

    (Total for Question 1 is 3 marks)

  2. 2.

    Machine A fills 1010 containers with x=200\sum x=200 and x2=4040\sum x^2=4040. Machine B fills 88 containers with x=164\sum x=164 and x2=3400\sum x^2=3400. Calculate the mean and population standard deviation for each machine, and compare their consistency. Give standard deviations to 33 significant figures.

    (5)

    (Total for Question 2 is 5 marks)

  3. 3.

    A set of 1616 observations has mean 7.57.5 and population standard deviation 2.252.25. Find x\sum x and x2\sum x^2.

    (4)

    (Total for Question 3 is 4 marks)

Tier 3 · Hard

  1. 1.

    Group A has 1212 values with mean 1818 and x2=3996\sum x^2=3996. Group B has 88 values with mean 2424 and x2=4736\sum x^2=4736. Find the mean and population standard deviation of all 2020 values.

    (5)

    (Total for Question 1 is 5 marks)

  2. 2.

    For 2525 observations of xx, the coded variable y=(x50)/4y=(x-50)/4 satisfies y=30\sum y=30 and y2=100\sum y^2=100. Find the mean and population standard deviation of xx. One further observation, x=62x=62, is then added. Find the new mean and population standard deviation. Give all non-integer answers to 33 significant figures.

    (7)

    (Total for Question 2 is 7 marks)

  3. 3.

    Eight observations have mean 1010 and population standard deviation 33. Six of the observations are 5,10,11,12,12,125,10,11,12,12,12; the other two are uu and vv, where uvu\leq v. Find uu and vv.

    (6)

    (Total for Question 3 is 6 marks)

  4. 4.

    A grouped frequency table has classes 0x<100\leq x<10, 10x<2010\leq x<20, 20x<3020\leq x<30 and 30x<5030\leq x<50, with respective frequencies 33, 55, 88 and 44. Using class midpoints, estimate the mean and population standard deviation. Give the standard deviation to 33 significant figures and explain why both answers are estimates rather than exact statistics for the raw data.

    (6)

    (Total for Question 4 is 6 marks)

  5. 5.

    Two groups together contain 1010 observations with mean 16.416.4 and population variance 28.4428.44. Group A contains 44 observations with mean 1111 and population variance 55. Find the mean and population standard deviation of the 66 observations in group B. Use unrounded values in your working and give the standard deviation to 33 significant figures.

    (6)

    (Total for Question 5 is 6 marks)

S2.4 · Recognise and interpret possible outliers in data sets and statistical diagrams; select or critique data presentation techniques in context; clean data, including dealing with missing data, errors and outliers.

Explanation

  • A common outlier rule flags values below Q11.5IQRQ_1-1.5\operatorname{IQR} or above Q3+1.5IQRQ_3+1.5\operatorname{IQR}, but context should guide the final decision.
  • Investigate a suspicious value against the original record before correcting or removing it; an unusual valid observation is not automatically an error.
  • Handle missing data transparently: record how many values are missing, avoid inventing unsupported values and consider whether missingness could bias conclusions.
  • Choose displays to suit the data and purpose: histograms for continuous grouped data, box plots for comparing distributions, and scatter diagrams for paired variables.

Worked example

A table of package masses contains one blank entry and one value 482482 among values near 48.248.2 grams. Describe a defensible way to clean these two entries before analysis.

  1. 1.First distinguish a data-entry error from a genuine extreme value by consulting the source.
  2. 2.Do not silently divide 482482 by 1010.
  3. 3.The blank contains no observed value, so omit it from calculations unless a justified imputation rule has been chosen, and document the decision so its possible bias is visible.

Answer: Check the original measurement record for both entries.; Correct 482482 to 48.248.2 only if the source confirms a decimal-point error; otherwise retain and flag it or exclude it with a stated reason.; Treat the blank as missing rather than replacing it without evidence, and report the reduced sample size or justified imputation method.

Common mistakes

  • Don't replace every missing value with zero, changing both the centre and spread of the data.
  • Don't delete an unusual value automatically instead of investigating whether it is an error or genuine observation.

Exam tip

Document separate rules for missing values, transcription errors and plausible outliers before recalculating summaries.

Tier 1 · Easy

  1. 1.

    For a data set, Q1=14Q_1=14 and Q3=22Q_3=22. Use the 1.5IQR1.5\operatorname{IQR} rule to determine whether the value 3636 is a possible outlier.

    (2)

    (Total for Question 1 is 2 marks)

  2. 2.

    A researcher wants to compare the median and spread of journey times for two train operators. State a suitable diagram and give one reason for your choice.

    (2)

    (Total for Question 2 is 2 marks)

Tier 2 · Standard

  1. 1.

    For a data set, the lower quartile is 1818 and the upper quartile is 3030. Use the 1.5×IQR1.5\times\operatorname{IQR} rule to decide whether a value of 5050 is an outlier, and state what should be done before removing it.

    (4)

    (Total for Question 1 is 4 marks)

  2. 2.

    Two schools have different numbers of students. A report uses two pie charts to compare their students' continuous weekly screen times. Critique this presentation and suggest a more informative display for comparing the distributions.

    (4)

    (Total for Question 2 is 4 marks)

  3. 3.

    The ordered data are 4,6,7,8,9,10,11,12,14,15,294,6,7,8,9,10,11,12,14,15,29. Use the Edexcel discrete convention: round a quartile position up when it is not a whole number; when it is a whole number, average the value in that position and the next value. Find Q1Q_1, Q3Q_3 and the outlier fences, and identify any possible outlier.

    (4)

    (Total for Question 3 is 4 marks)

Tier 3 · Hard

  1. 1.

    A data set of 4040 readings was summarised as x=504\sum x=504 and x2=6856.9\sum x^2=6856.9. One reading was entered as 3131 but the source record confirms it should be 1313. Calculate the corrected mean and population standard deviation, and state the likely effect of the error on the original spread.

    (5)

    (Total for Question 1 is 5 marks)

  2. 2.

    Ten response times, in seconds, are 8,9,9,10,10,10,11,11,12,408,9,9,10,10,10,11,11,12,40. For these data, Q1=9Q_1=9 and Q3=11Q_3=11. Use the 1.5IQR1.5\operatorname{IQR} rule to identify any possible outlier. Calculate the mean and median, then recommend which measure better represents a typical response time. The source confirms that the value 4040 is a genuine response during a network outage; state whether it should automatically be deleted.

    (5)

    (Total for Question 2 is 5 marks)

  3. 3.

    The ordered data are 3,4,5,6,7,8,9,10,x,14,15,253,4,5,6,7,8,9,10,x,14,15,25, where xx is an integer and 10x1410\leq x\leq14. Use the Edexcel discrete convention: round a quartile position up when it is not a whole number; when it is a whole number, average the value in that position and the next value. Find all values of xx for which 2525 is a possible outlier under the 1.5IQR1.5\operatorname{IQR} rule. State how the value 2525 should be treated if its source record confirms that it is genuine.

    (6)

    (Total for Question 3 is 6 marks)

  4. 4.

    Fifteen ordered readings are 2,4,5,7,8,9,10,11,12,13,14,15,16,18,302,4,5,7,8,9,10,11,12,13,14,15,16,18,30. A sixteenth reading is missing from the file, but its source record shows only that it is an integer from 1010 to 1313 inclusive. Do not replace it by an average. Using the Edexcel discrete convention, show that Q1Q_1, Q3Q_3 and the upper outlier fence are the same for every possible value of the missing reading. Hence decide whether 3030 is always a possible outlier, and state what should be recorded before the data are used.

    (6)

    (Total for Question 4 is 6 marks)

  5. 5.

    A data extract lists measurements 41,42,43,44,45,46,47,47,48,49,50,51041,42,43,44,45,46,47,47,48,49,50,510 and one blank entry. The two values 4747 have the same record identifier, the source confirms that 510510 is a decimal-point error for 51.051.0, and the blank cannot be recovered. State how each issue should be handled. For the cleaned data, find the sample size, median, quartiles and any possible outliers using the Edexcel discrete convention and the 1.5IQR1.5\operatorname{IQR} rule.

    (7)

    (Total for Question 5 is 7 marks)

Answer key

Answers begin on a new printed page so the question pack can be completed without the solutions alongside it.

S2.1 · Interpret diagrams for single-variable data, including understanding that area in a histogram represents frequency; connect to probability distributions.

Tier 1 · Easy

Mark scheme for S2.1 Tier 1 · Easy
QuestionSchemeMarks
1
  • 1818
2
(2 marks)2
Notes
The class width is 1712=517-12=5. Hence frequency == class width ×\times frequency density =5×3.6=18=5\times3.6=18.
2
  • 83\dfrac{8}{3}
2
(2 marks)2
Notes
Equal frequencies give equal bar areas. The first bar has area 6(4)=246(4)=24, so the second density is 24/9=8/324/9=8/3.

Tier 2 · Standard

Mark scheme for S2.1 Tier 2 · Standard
QuestionSchemeMarks
1
  • Frequencies 2020 and 2525
  • Total frequency 4545
4
(4 marks)4
Notes
Histogram frequency equals bar area, so multiply frequency density by class width. The first class has width 55 and frequency 4(5)=204(5)=20. The second has width 1010 and frequency 2.5(10)=252.5(10)=25. Their total is 20+25=4520+25=45.
2
  • Median =36=36
  • 6565th percentile =40=40
4
(4 marks)4
Notes
The median has cumulative frequency 4040. This is 18/36=1/218/36=1/2 of the way from 2222 to 5858, so its value is 30+(1/2)(4230)=3630+(1/2)(42-30)=36. The 6565th percentile has cumulative frequency 0.65(80)=520.65(80)=52, which is 30/36=5/630/36=5/6 of the way along the segment, giving 30+(5/6)(12)=4030+(5/6)(12)=40.
3
  • One grid unit represents frequency density 0.750.75.
  • The frequency in 5x<115\leq x<11 is 1818.
4
(4 marks)4
Notes
The first class has width 55, so its frequency density is 30/5=630/5=6. A height of 88 grid units therefore represents density 66, giving 6/8=0.756/8=0.75 density units per grid unit. The second bar has density 4(0.75)=34(0.75)=3 and width 66, so its area and frequency are 6(3)=186(3)=18.

Tier 3 · Hard

Mark scheme for S2.1 Tier 3 · Hard
QuestionSchemeMarks
1
  • Total area =1=1.
  • P(1<X<6)=0.48P(1<X<6)=0.48.
5
(5 marks)5
Notes
The total area is 3(0.12)+5(0.08)+2(0.12)=0.36+0.40+0.24=13(0.12)+5(0.08)+2(0.12)=0.36+0.40+0.24=1, so the histogram defines a probability distribution. The area from 11 to 33 is 2(0.12)=0.242(0.12)=0.24, and the area from 33 to 66 is 3(0.08)=0.243(0.08)=0.24. Therefore P(1<X<6)=0.24+0.24=0.48P(1<X<6)=0.24+0.24=0.48.
2
  • k=3.75k=3.75
  • P(X10)=1730P(X\geq10)=\dfrac{17}{30}
6
(6 marks)6
Notes
The first two frequencies are 5(3.2)=165(3.2)=16 and 7(2.0)=147(2.0)=14. The last class therefore has frequency 601614=3060-16-14=30, so k=30/8=3.75k=30/8=3.75. From 1010 to 1212, the histogram area gives frequency 2(2.0)=42(2.0)=4; all 3030 observations in the final class also qualify. Hence the probability is (4+30)/60=17/30(4+30)/60=17/30.
3
  • Estimated median =7.8=7.8
  • Estimated lower quartile =5.4=5.4 and upper quartile =11=11
  • Estimated interquartile range =5.6=5.6
6
(6 marks)6
Notes
The class frequencies are 5(2)=105(2)=10, 4(5)=204(5)=20 and 6(3)=186(3)=18, giving total frequency 4848. The median is the 2424th value, so interpolation in the second class gives 5+241020(4)=7.85+\frac{24-10}{20}(4)=7.8. The lower quartile is the 1212th value, giving 5+121020(4)=5.45+\frac{12-10}{20}(4)=5.4. The upper quartile is the 3636th value, in the third class, giving 9+363018(6)=119+\frac{36-30}{18}(6)=11. Hence the estimated IQR is 115.4=5.611-5.4=5.6.
4
  • Sample A has frequency 5(4)=205(4)=20, which is 20/80=25%20/80=25\% of sample A.
  • Sample B has frequency 5(5)=255(5)=25, which is 25/120=20.8%25/120=20.8\% of sample B to 33 significant figures.
  • The bar for B is taller and contains more observations, but A has the larger proportion in the class.
  • Frequency-density heights from samples of different total sizes do not directly compare relative frequencies; percentage-frequency density or calculated proportions are needed.
5
(5 marks)5
Notes
Both class widths are 55, so the bar areas give frequencies 2020 and 2525. Divide each frequency by its own sample total: 20/80=0.2520/80=0.25 and 25/120=0.2083325/120=0.20833\ldots. The different denominators reverse the comparison made from raw bar heights.
5
  • a=16a=16
  • Estimated median =8.6=8.6
  • P(X14)=27P(X\geq14)=\dfrac27
7
(7 marks)7
Notes
The total frequency is 4(3)+6(5)+(a10)(2)+(20a)(4)=1022a4(3)+6(5)+(a-10)(2)+(20-a)(4)=102-2a. Equating this to 7070 gives a=16a=16. The first two class frequencies are 1212 and 3030, so the 3535th value lies in 4x<104\leq x<10 and is estimated by 4+351230(6)=8.64+\frac{35-12}{30}(6)=8.6. From 1414 to 1616 the estimated frequency is 2(2)=42(2)=4, and from 1616 to 2020 it is 4(4)=164(4)=16. Hence P(X14)=(4+16)/70=2/7P(X\geq14)=(4+16)/70=2/7.

S2.2 · Interpret scatter diagrams and regression lines for bivariate data, including recognising distinct sections of the population (regression calculations excluded); interpret correlation informally; correlation does not imply causation.

Tier 1 · Easy

Mark scheme for S2.2 Tier 1 · Easy
QuestionSchemeMarks
1
  • Higher temperatures tend to be associated with higher ice-cream sales.
  • Correlation alone does not establish causation; other variables such as sunshine, holidays or visitor numbers may affect both quantities.
2
(2 marks)2
Notes
Translate positive correlation as a tendency for larger values of one variable to occur with larger values of the other. Then distinguish association from a causal conclusion by identifying a plausible lurking variable.
2
  • The clusters represent distinct sections of the population.
  • One line may describe mainly the separation between bicycle types rather than the relationship within either type, so the groups should be considered separately.
2
(2 marks)2
Notes
The bicycle type is a grouping variable. Pooling two separated groups can create an overall pattern that does not represent the within-group relationship, so separate analyses are more informative.

Tier 2 · Standard

Mark scheme for S2.2 Tier 2 · Standard
QuestionSchemeMarks
1
  • Correlation does not imply causation.
  • A valid lurking variable is hot or sunny weather, season, holiday periods or visitor numbers, any of which could increase both ice-cream sales and sunburn cases.
  • Evidence controlling for plausible lurking variables, or an appropriately designed study, would be needed to support a causal claim.
3
(3 marks)3
Notes
The scatter pattern records association only. Temperature or sunshine can affect both variables: hot days encourage ice-cream purchases and increase exposure to ultraviolet radiation. This common cause can generate the correlation even if ice cream has no effect on sunburn.
2
  • Predicted sound level =42.8=42.8 dB.
  • Within the observed range, the model predicts a decrease of 1.41.4 dB in mean sound level for each increase of 11 km h1^{-1} in wind speed.
3
(3 marks)3
Notes
Substitute w=18w=18: s=681.4(18)=42.8s=68-1.4(18)=42.8. The gradient 1.4-1.4 is the fitted change in the response variable for a one-unit increase in the explanatory variable; it describes association within the data range.
3
  • New gradient =100(0.40)=40=100(0.40)=40 pence of sales income per pound of advertising cost.
  • The direction remains positive.
  • The strength of the correlation is unchanged because multiplying every response value by the same positive constant only changes its scale, not the scatter pattern.
3
(3 marks)3
Notes
Converting pounds to pence multiplies every response value, and hence the fitted change in response, by 100100. The gradient becomes 4040. This positive rescaling does not alter which points lie relatively above or below one another or how tightly they follow a linear pattern, so it changes neither the positive direction nor the strength of the correlation.

Tier 3 · Hard

Mark scheme for S2.2 Tier 3 · Hard
QuestionSchemeMarks
1
  • The predicted diameter at a=12a=12 is 25.825.8 cm, so the difference between the observed and predicted values is 2925.8=3.229-25.8=3.2 cm.
  • The tree is 3.23.2 cm thicker than the value predicted by the line.
  • A prediction at 4040 years is extrapolation far outside the observed age range and the linear relationship may not continue.
4
(4 marks)4
Notes
Substitute a=12a=12 to get the fitted value d=1.8(12)+4.2=25.8d=1.8(12)+4.2=25.8. Residual == observed - fitted =2925.8=3.2=29-25.8=3.2 cm. Since 4040 is well beyond the data range 33 to 1818, using the line there is unsupported extrapolation.
2
  • Course level creates two distinct sections of the population.
  • The pooled positive correlation may be caused mainly by the separation of the two clusters and need not describe either course.
  • A single regression line may fit neither within-course relationship well, so the courses should be analysed separately.
  • The slight negative within-course association does not prove causation; prior attainment, difficulty or students seeking help could affect both practice time and score.
5
(5 marks)5
Notes
The overall association mixes a between-course difference with two different within-course patterns. Course level is therefore a lurking grouping variable, and one pooled line conceals the structure. Even after separating the courses, an observational association cannot establish that practice causes lower scores because other variables can influence both quantities.
3
  • The second line predicts 11.811.8 at x=9x=9, so the residual of (9,25)(9,25) is 2511.8=13.225-11.8=13.2.
  • At x=8x=8, the full-data prediction is 1515 and the prediction after omission is 10.610.6, a difference of 4.44.4.
  • The observation has a large residual and substantially influences the fitted line.
  • Influence does not prove a recording error; the source and context should be checked before retaining, correcting or excluding it.
5
(5 marks)5
Notes
Using the line fitted without the observation isolates how unusual it is relative to the remaining pattern: 25[1.2(9)+1.0]=13.225-[1.2(9)+1.0]=13.2. The two fitted values at x=8x=8 are 2(8)1=152(8)-1=15 and 1.2(8)+1.0=10.61.2(8)+1.0=10.6. Their difference shows that the point changes the fitted relationship materially. Statistical influence is a reason to investigate, not automatic evidence that the value is invalid.
4
  • y=2x1.5y=2x^{1.5}, so a=2a=2 and n=1.5n=1.5.
  • When x=8x=8, y2(8)1.5=45.3y\approx2(8)^{1.5}=45.3 to 33 significant figures.
  • x=50x=50 is outside the observed range, so this would be unsupported extrapolation and the power relationship may not continue.
  • Correlation on the transformed plot is association only; lurking variables, reverse causation or the data-generating process must be considered before making a causal claim.
6
(6 marks)6
Notes
Taking powers of 1010 gives y=2x1.5y=2x^{1.5}. Substitution gives 283=45.2542\sqrt{8^3}=45.254\ldots. The fitted relationship was supported only for 2x202\leq x\leq20, and a regression pattern does not identify a causal mechanism.
5
  • Use the regression line of yy on xx: the estimate is y=2(12)+10=34y=2(12)+10=34.
  • At x=10x=10, the first line gives y=30y=30; at y=30y=30, the second gives x=10x=10, so both pass through (10,30)(10,30).
  • Rearranging x=0.4y2x=0.4y-2 would give y=35y=35 when x=12x=12.
  • The two regression lines minimise deviations in different variables. To predict yy from a given xx, use the regression of yy on xx, not the algebraic inverse of the regression of xx on yy.
5
(5 marks)5
Notes
The response being estimated is yy, so substitute x=12x=12 into y=2x+10y=2x+10. The mean point satisfies both supplied equations. Rearrangement of the other line gives (12+2)/0.4=35(12+2)/0.4=35, but that line was fitted to predict xx from yy and generally is not the inverse of the yy-on-xx line unless correlation is perfect.

S2.3 · Interpret measures of central tendency and variation, extending to standard deviation; be able to calculate standard deviation, including from summary statistics.

Tier 1 · Easy

Mark scheme for S2.3 Tier 1 · Easy
QuestionSchemeMarks
1
  • Mean =7=7.
  • Standard deviation =1.67=1.67 to 33 significant figures.
3
(3 marks)3
Notes
x=35\sum x=35 and x2=259\sum x^2=259, so xˉ=35/5=7\bar{x}=35/5=7. Then σ=259/572=2.8=1.673\sigma=\sqrt{259/5-7^2}=\sqrt{2.8}=1.673\ldots, giving 1.671.67.
2
  • 14.214.2 to 33 significant figures
2
(2 marks)2
Notes
The original total is 8(14)=1128(14)=112. After adding 1616, the total is 128128, so the new mean is 128/9=14.222=14.2128/9=14.222\ldots=14.2 to 33 significant figures.

Tier 2 · Standard

Mark scheme for S2.3 Tier 2 · Standard
QuestionSchemeMarks
1
  • Mean 2929
  • Standard deviation 66
3
(3 marks)3
Notes
Adding 55 shifts every value and hence adds 55 to the mean but does not change the spread. Multiplying by 22 doubles both the mean contribution and the standard deviation. Thus the new mean is 5+2(12)=295+2(12)=29 and the new standard deviation is 2(3)=6|2|(3)=6.
2
  • Machine A: mean 2020, standard deviation 2.002.00.
  • Machine B: mean 20.520.5, standard deviation 2.182.18.
  • Machine A is slightly more consistent because it has the smaller standard deviation.
5
(5 marks)5
Notes
For A, xˉ=200/10=20\bar{x}=200/10=20 and σA=4040/10202=2\sigma_A=\sqrt{4040/10-20^2}=2. For B, xˉ=164/8=20.5\bar{x}=164/8=20.5 and σB=3400/820.52=4.75=2.179\sigma_B=\sqrt{3400/8-20.5^2}=\sqrt{4.75}=2.179\ldots. Smaller standard deviation means less spread, so A is more consistent.
3
  • x=120\sum x=120
  • x2=981\sum x^2=981
4
(4 marks)4
Notes
The sum is nxˉ=16(7.5)=120n\bar{x}=16(7.5)=120. Rearranging σ2=x2/nxˉ2\sigma^2=\sum x^2/n-\bar{x}^2 gives x2=n(σ2+xˉ2)=16(2.252+7.52)=16(61.3125)=981\sum x^2=n(\sigma^2+\bar{x}^2)=16(2.25^2+7.5^2)=16(61.3125)=981.

Tier 3 · Hard

Mark scheme for S2.3 Tier 3 · Hard
QuestionSchemeMarks
1
  • Combined mean =20.4=20.4.
  • Combined standard deviation =4.52=4.52 to 33 significant figures.
5
(5 marks)5
Notes
The combined sum is 12(18)+8(24)=40812(18)+8(24)=408, so the combined mean is 408/20=20.4408/20=20.4. The combined sum of squares is 3996+4736=87323996+4736=8732. Therefore σ=8732/2020.42=436.6416.16=20.44=4.521\sigma=\sqrt{8732/20-20.4^2}=\sqrt{436.6-416.16}=\sqrt{20.44}=4.521\ldots, giving 4.524.52.
2
  • Original mean =54.8=54.8 and original standard deviation =6.40=6.40.
  • New mean =55.1=55.1 and new standard deviation =6.43=6.43.
7
(7 marks)7
Notes
For yy, the mean is 30/25=1.230/25=1.2 and the variance is 100/251.22=2.56100/25-1.2^2=2.56, so σy=1.6\sigma_y=1.6. Since x=4y+50x=4y+50, xˉ=4(1.2)+50=54.8\bar{x}=4(1.2)+50=54.8 and σx=4(1.6)=6.4\sigma_x=4(1.6)=6.4. Also x=1370\sum x=1370 and x2=76100\sum x^2=76100. After adding 6262, these become 14321432 and 7994479944 for 2626 values. Thus the new mean is 1432/26=55.07691432/26=55.0769\ldots and the new standard deviation is 79944/26(1432/26)2=6.4266\sqrt{79944/26-(1432/26)^2}=6.4266\ldots, giving 55.155.1 and 6.436.43.
3
  • u=5u=5 and v=13v=13
6
(6 marks)6
Notes
The total sum is 8(10)=808(10)=80, while the six known values sum to 6262, so u+v=18u+v=18. The total sum of squares is 8(32+102)=8728(3^2+10^2)=872. The known squares sum to 678678, so u2+v2=194u^2+v^2=194. Hence 2uv=(u+v)2(u2+v2)=324194=1302uv=(u+v)^2-(u^2+v^2)=324-194=130, giving uv=65uv=65. Thus uu and vv are the roots of t218t+65=0=(t5)(t13)t^2-18t+65=0=(t-5)(t-13); the stated order gives u=5u=5 and v=13v=13.
4
  • Estimated mean =22.5=22.5
  • Estimated population standard deviation =11.1=11.1 to 33 significant figures
  • The calculations replace every observation in a class by its midpoint, so the within-class positions of the raw observations are unknown.
6
(6 marks)6
Notes
Use midpoints 5,15,25,405,15,25,40. Then fx=3(5)+5(15)+8(25)+4(40)=450\sum fx=3(5)+5(15)+8(25)+4(40)=450, so xˉ=450/20=22.5\bar{x}=450/20=22.5. Also fx2=3(25)+5(225)+8(625)+4(1600)=12600\sum fx^2=3(25)+5(225)+8(625)+4(1600)=12600. Hence σ=12600/2022.52=123.75=11.124\sigma=\sqrt{12600/20-22.5^2}=\sqrt{123.75}=11.124\ldots, giving 11.111.1. Midpoint substitution loses the unknown variation within each class.
5
  • Mean of group B =20=20
  • Population standard deviation of group B =3.42=3.42 to 33 significant figures
6
(6 marks)6
Notes
The combined sum is 10(16.4)=16410(16.4)=164 and group A contributes 4(11)=444(11)=44, so group B has sum 120120 and mean 2020. From σ2=x2/nxˉ2\sigma^2=\sum x^2/n-\bar{x}^2, the combined sum of squares is 10(28.44+16.42)=297410(28.44+16.4^2)=2974, while group A contributes 4(5+112)=5044(5+11^2)=504. Thus group B has sum of squares 24702470, variance 2470/6202=35/32470/6-20^2=35/3, and standard deviation 35/3=3.41565=3.42\sqrt{35/3}=3.41565\ldots=3.42.

S2.4 · Recognise and interpret possible outliers in data sets and statistical diagrams; select or critique data presentation techniques in context; clean data, including dealing with missing data, errors and outliers.

Tier 1 · Easy

Mark scheme for S2.4 Tier 1 · Easy
QuestionSchemeMarks
1
  • The upper boundary is 3434, so 3636 is a possible outlier.
2
(2 marks)2
Notes
The interquartile range is 2214=822-14=8. The upper outlier boundary is 22+1.5(8)=3422+1.5(8)=34. Since 36>3436>34, it is flagged as a possible outlier.
2
  • Use box plots drawn on a common scale.
  • They display the median and interquartile range, so the centres and spreads of the two distributions can be compared directly.
2
(2 marks)2
Notes
Box plots summarise each distribution using its median, quartiles and extremes. Putting them on the same scale makes the requested comparison direct.

Tier 2 · Standard

Mark scheme for S2.4 Tier 2 · Standard
QuestionSchemeMarks
1
  • IQR=12\operatorname{IQR}=12 and the upper outlier boundary is 4848.
  • 5050 is an outlier.
  • Check the original record and context: correct it only if an error is confirmed, retain it if genuine, or exclude it only with a justified and documented reason.
4
(4 marks)4
Notes
The interquartile range is 3018=1230-18=12. The upper boundary is Q3+1.5IQR=30+18=48Q_3+1.5\operatorname{IQR}=30+18=48, so 50>4850>48 is flagged as an outlier. Being an outlier does not prove the value is wrong: check the source and measurement conditions, then document any correction or exclusion.
2
  • Pie charts require screen time to be grouped into categories and do not show the detailed shape or spread of the continuous data.
  • Unless the sample sizes are shown, equal-looking sectors can hide very different frequencies between the schools.
  • Use comparative histograms with common class boundaries and frequency density, or box plots on a common scale.
  • Histograms compare distribution shape; box plots compare median and spread.
4
(4 marks)4
Notes
The variable is continuous, while pie charts reduce it to proportions in chosen categories and can conceal both sample size and distributional detail. Common-scale histograms preserve shape; common-scale box plots give a concise comparison of centre and spread.
3
  • Q1=7Q_1=7 and Q3=14Q_3=14
  • IQR=7\operatorname{IQR}=7; the fences are 3.5-3.5 and 24.524.5.
  • 2929 is the only possible outlier.
4
(4 marks)4
Notes
Here n=11n=11. The lower-quartile position is 11/4=2.7511/4=2.75, rounded up to the third value, so Q1=7Q_1=7. The upper-quartile position is 33/4=8.2533/4=8.25, rounded up to the ninth value, so Q3=14Q_3=14. Thus the IQR is 77 and the fences are 71.5(7)=3.57-1.5(7)=-3.5 and 14+1.5(7)=24.514+1.5(7)=24.5. Only 2929 lies outside them.

Tier 3 · Hard

Mark scheme for S2.4 Tier 3 · Hard
QuestionSchemeMarks
1
  • Corrected mean =12.15=12.15.
  • Corrected standard deviation =2.00=2.00.
  • The error inflated the original spread.
5
(5 marks)5
Notes
Replace the contribution of 3131 by 1313: the corrected sum is 50431+13=486504-31+13=486, and the corrected sum of squares is 6856.9312+132=6064.96856.9-31^2+13^2=6064.9. Thus xˉ=486/40=12.15\bar{x}=486/40=12.15 and σ=6064.9/4012.152=4=2.00\sigma=\sqrt{6064.9/40-12.15^2}=\sqrt{4}=2.00. The erroneous value lay much farther from the centre, so it made the original spread larger.
2
  • IQR=2\operatorname{IQR}=2, with lower boundary 66 and upper boundary 1414, so 4040 is the only possible outlier.
  • Mean =13=13 seconds and median =10=10 seconds.
  • The median better represents a typical response time because it is resistant to the extreme value.
  • 4040 should not automatically be deleted; it is genuine and should be retained or analysed separately according to whether outage performance is part of the population of interest.
5
(5 marks)5
Notes
The IQR is 119=211-9=2. The lower fence is 91.5(2)=69-1.5(2)=6 and the upper fence is 11+1.5(2)=1411+1.5(2)=14, so only 4040 is flagged. The total is 130130, giving mean 1313, and the middle two values are both 1010, giving median 1010. The mean is pulled upward by the genuine outage value. An outlier rule flags a value for investigation; it does not by itself justify deletion.
3
  • Q1=(5+6)/2=5.5Q_1=(5+6)/2=5.5, Q3=(x+14)/2Q_3=(x+14)/2 and IQR=(x+3)/2\operatorname{IQR}=(x+3)/2.
  • The upper fence is x+142+1.5(x+32)=5x+374\dfrac{x+14}{2}+1.5\left(\dfrac{x+3}{2}\right)=\dfrac{5x+37}{4}.
  • 2525 is a possible outlier when 25>(5x+37)/425>(5x+37)/4, so x<63/5=12.6x<63/5=12.6.
  • The integer values are x{10,11,12}x\in\{10,11,12\}; at x=13x=13, the upper fence is 25.525.5, so 2525 is not outside it.
  • A genuine value should not be deleted automatically; retain it or analyse it separately according to the population and purpose, documenting the decision.
6
(6 marks)6
Notes
With n=12n=12, n/4=3n/4=3 and 3n/4=93n/4=9 are whole numbers. The quartiles are therefore the midpoints of the 33rd and 44th values, and of the 99th and 1010th values: Q1=5.5Q_1=5.5 and Q3=(x+14)/2Q_3=(x+14)/2. Hence IQR=(x+3)/2\operatorname{IQR}=(x+3)/2 and the upper fence is (5x+37)/4(5x+37)/4. The strict inequality 25>(5x+37)/425>(5x+37)/4 gives x<12.6x<12.6, so the stated integer range gives 10,11,1210,11,12. A verified observation is not an error merely because it is unusual.
4
  • For n=16n=16, Q1Q_1 is the mean of the 44th and 55th values, so Q1=(7+8)/2=7.5Q_1=(7+8)/2=7.5.
  • For every allowed missing value, the 1212th and 1313th values are 1414 and 1515, so Q3=14.5Q_3=14.5.
  • The IQR is 77 and the upper fence is 14.5+1.5(7)=2514.5+1.5(7)=25.
  • 30>2530>25, so 3030 is a possible outlier for every allowed value of the missing reading.
  • Record the value as unresolved within the verified range, document that no unsupported imputation was made, and check the original source before any final correction or exclusion.
6
(6 marks)6
Notes
Since 16/4=416/4=4 and 3(16)/4=123(16)/4=12 are whole positions, average positions 44 and 55, then positions 1212 and 1313. Inserting any integer from 1010 to 1313 changes only the middle ordering: the relevant pairs remain 7,87,8 and 14,1514,15. The quartiles and fence are therefore invariant, so the outlier decision for 3030 is robust even though the missing reading itself remains unresolved.
5
  • Remove one of the duplicate 4747 records, correct 510510 to 51.051.0 from the confirmed source, and treat the unrecoverable blank as missing rather than as zero.
  • The cleaned ordered data are 41,42,43,44,45,46,47,48,49,50,5141,42,43,44,45,46,47,48,49,50,51, so n=11n=11 and the median is 4646.
  • Q1=43Q_1=43 and Q3=49Q_3=49.
  • The IQR is 66, giving fences 3434 and 5858; there are no possible outliers.
  • The cleaning decisions and reduced sample size should be documented.
7
(7 marks)7
Notes
A repeated identifier shows duplication, while the source record supplies evidence for the decimal correction. The blank supplies no observed value. After cleaning, the sixth value is the median. Since n/4=2.75n/4=2.75 and 3n/4=8.253n/4=8.25 are not whole numbers, round up to positions 33 and 99, giving quartiles 4343 and 4949. The resulting fences contain all 1111 values.