1.
(2)
(Total for Question 1 is 2 marks)
4 specification points · notes, questions, answers and worked methods
Checked against Edexcel 9MA0 section S2. Review basis: the qualification registry sourced from the Pearson Edexcel Level 3 Advanced GCE in Mathematics (9MA0) specification; registry verification recorded 11 July 2026.
Explanation
Worked example
A cumulative frequency diagram represents observations. The cumulative frequencies at and are and respectively. Estimate the number of observations satisfying .
Answer:
Common mistakes
Exam tip
For cumulative-frequency intervals, subtract the cumulative totals at the two boundaries and respect endpoint inequalities.
1.
(2)
(Total for Question 1 is 2 marks)
2.
(2)
(Total for Question 2 is 2 marks)
1.
(4)
(Total for Question 1 is 4 marks)
2.
(4)
(Total for Question 2 is 4 marks)
3.
(4)
(Total for Question 3 is 4 marks)
1.
(5)
(Total for Question 1 is 5 marks)
2.
(6)
(Total for Question 2 is 6 marks)
3.
(6)
(Total for Question 3 is 6 marks)
4.
(5)
(Total for Question 4 is 5 marks)
5.
(7)
(Total for Question 5 is 7 marks)
Explanation
Worked example
A scatter diagram of journey distance against journey time contains one cluster for bicycles and a separate cluster for cars. Explain why fitting one regression line to all journeys may be misleading.
Answer: The two clusters represent distinct sections of the population with different travel speeds.; A single line may mainly reflect the gap between the clusters rather than either within-group relationship.; Separate regression lines or separate analyses would be more informative.
Common mistakes
Exam tip
Before using a regression line, inspect clusters and explain whether one relationship is credible for the whole population.
1.
(2)
(Total for Question 1 is 2 marks)
2.
(2)
(Total for Question 2 is 2 marks)
1.
(3)
(Total for Question 1 is 3 marks)
2.
(3)
(Total for Question 2 is 3 marks)
3.
(3)
(Total for Question 3 is 3 marks)
1.
(4)
(Total for Question 1 is 4 marks)
2.
(5)
(Total for Question 2 is 5 marks)
3.
(5)
(Total for Question 3 is 5 marks)
4.
(6)
(Total for Question 4 is 6 marks)
5.
(5)
(Total for Question 5 is 5 marks)
Explanation
Worked example
For observations, and . Calculate the population standard deviation to significant figures.
Answer:
Common mistakes
Exam tip
Substitute the stated divisor into the variance formula, keep the square root until the end and check the result is non-negative.
1.
(3)
(Total for Question 1 is 3 marks)
2.
(2)
(Total for Question 2 is 2 marks)
1.
(3)
(Total for Question 1 is 3 marks)
2.
(5)
(Total for Question 2 is 5 marks)
3.
(4)
(Total for Question 3 is 4 marks)
1.
(5)
(Total for Question 1 is 5 marks)
2.
(7)
(Total for Question 2 is 7 marks)
3.
(6)
(Total for Question 3 is 6 marks)
4.
(6)
(Total for Question 4 is 6 marks)
5.
(6)
(Total for Question 5 is 6 marks)
Explanation
Worked example
A table of package masses contains one blank entry and one value among values near grams. Describe a defensible way to clean these two entries before analysis.
Answer: Check the original measurement record for both entries.; Correct to only if the source confirms a decimal-point error; otherwise retain and flag it or exclude it with a stated reason.; Treat the blank as missing rather than replacing it without evidence, and report the reduced sample size or justified imputation method.
Common mistakes
Exam tip
Document separate rules for missing values, transcription errors and plausible outliers before recalculating summaries.
1.
(2)
(Total for Question 1 is 2 marks)
2.
(2)
(Total for Question 2 is 2 marks)
1.
(4)
(Total for Question 1 is 4 marks)
2.
(4)
(Total for Question 2 is 4 marks)
3.
(4)
(Total for Question 3 is 4 marks)
1.
(5)
(Total for Question 1 is 5 marks)
2.
(5)
(Total for Question 2 is 5 marks)
3.
(6)
(Total for Question 3 is 6 marks)
4.
(6)
(Total for Question 4 is 6 marks)
5.
(7)
(Total for Question 5 is 7 marks)
Answers begin on a new printed page so the question pack can be completed without the solutions alongside it.
| Question | Scheme | Marks |
|---|---|---|
| 1 | 2 | |
| (2 marks) | 2 | |
| Notes | ||
| The class width is . Hence frequency class width frequency density . | ||
| 2 | 2 | |
| (2 marks) | 2 | |
| Notes | ||
| Equal frequencies give equal bar areas. The first bar has area , so the second density is . | ||
| Question | Scheme | Marks |
|---|---|---|
| 1 |
| 4 |
| (4 marks) | 4 | |
| Notes | ||
| Histogram frequency equals bar area, so multiply frequency density by class width. The first class has width and frequency . The second has width and frequency . Their total is . | ||
| 2 |
| 4 |
| (4 marks) | 4 | |
| Notes | ||
| The median has cumulative frequency . This is of the way from to , so its value is . The th percentile has cumulative frequency , which is of the way along the segment, giving . | ||
| 3 |
| 4 |
| (4 marks) | 4 | |
| Notes | ||
| The first class has width , so its frequency density is . A height of grid units therefore represents density , giving density units per grid unit. The second bar has density and width , so its area and frequency are . | ||
| Question | Scheme | Marks |
|---|---|---|
| 1 |
| 5 |
| (5 marks) | 5 | |
| Notes | ||
| The total area is , so the histogram defines a probability distribution. The area from to is , and the area from to is . Therefore . | ||
| 2 | 6 | |
| (6 marks) | 6 | |
| Notes | ||
| The first two frequencies are and . The last class therefore has frequency , so . From to , the histogram area gives frequency ; all observations in the final class also qualify. Hence the probability is . | ||
| 3 |
| 6 |
| (6 marks) | 6 | |
| Notes | ||
| The class frequencies are , and , giving total frequency . The median is the th value, so interpolation in the second class gives . The lower quartile is the th value, giving . The upper quartile is the th value, in the third class, giving . Hence the estimated IQR is . | ||
| 4 |
| 5 |
| (5 marks) | 5 | |
| Notes | ||
| Both class widths are , so the bar areas give frequencies and . Divide each frequency by its own sample total: and . The different denominators reverse the comparison made from raw bar heights. | ||
| 5 |
| 7 |
| (7 marks) | 7 | |
| Notes | ||
| The total frequency is . Equating this to gives . The first two class frequencies are and , so the th value lies in and is estimated by . From to the estimated frequency is , and from to it is . Hence . | ||
| Question | Scheme | Marks |
|---|---|---|
| 1 |
| 2 |
| (2 marks) | 2 | |
| Notes | ||
| Translate positive correlation as a tendency for larger values of one variable to occur with larger values of the other. Then distinguish association from a causal conclusion by identifying a plausible lurking variable. | ||
| 2 |
| 2 |
| (2 marks) | 2 | |
| Notes | ||
| The bicycle type is a grouping variable. Pooling two separated groups can create an overall pattern that does not represent the within-group relationship, so separate analyses are more informative. | ||
| Question | Scheme | Marks |
|---|---|---|
| 1 |
| 3 |
| (3 marks) | 3 | |
| Notes | ||
| The scatter pattern records association only. Temperature or sunshine can affect both variables: hot days encourage ice-cream purchases and increase exposure to ultraviolet radiation. This common cause can generate the correlation even if ice cream has no effect on sunburn. | ||
| 2 |
| 3 |
| (3 marks) | 3 | |
| Notes | ||
| Substitute : . The gradient is the fitted change in the response variable for a one-unit increase in the explanatory variable; it describes association within the data range. | ||
| 3 |
| 3 |
| (3 marks) | 3 | |
| Notes | ||
| Converting pounds to pence multiplies every response value, and hence the fitted change in response, by . The gradient becomes . This positive rescaling does not alter which points lie relatively above or below one another or how tightly they follow a linear pattern, so it changes neither the positive direction nor the strength of the correlation. | ||
| Question | Scheme | Marks |
|---|---|---|
| 1 |
| 4 |
| (4 marks) | 4 | |
| Notes | ||
| Substitute to get the fitted value . Residual observed fitted cm. Since is well beyond the data range to , using the line there is unsupported extrapolation. | ||
| 2 |
| 5 |
| (5 marks) | 5 | |
| Notes | ||
| The overall association mixes a between-course difference with two different within-course patterns. Course level is therefore a lurking grouping variable, and one pooled line conceals the structure. Even after separating the courses, an observational association cannot establish that practice causes lower scores because other variables can influence both quantities. | ||
| 3 |
| 5 |
| (5 marks) | 5 | |
| Notes | ||
| Using the line fitted without the observation isolates how unusual it is relative to the remaining pattern: . The two fitted values at are and . Their difference shows that the point changes the fitted relationship materially. Statistical influence is a reason to investigate, not automatic evidence that the value is invalid. | ||
| 4 |
| 6 |
| (6 marks) | 6 | |
| Notes | ||
| Taking powers of gives . Substitution gives . The fitted relationship was supported only for , and a regression pattern does not identify a causal mechanism. | ||
| 5 |
| 5 |
| (5 marks) | 5 | |
| Notes | ||
| The response being estimated is , so substitute into . The mean point satisfies both supplied equations. Rearrangement of the other line gives , but that line was fitted to predict from and generally is not the inverse of the -on- line unless correlation is perfect. | ||
| Question | Scheme | Marks |
|---|---|---|
| 1 |
| 3 |
| (3 marks) | 3 | |
| Notes | ||
| and , so . Then , giving . | ||
| 2 |
| 2 |
| (2 marks) | 2 | |
| Notes | ||
| The original total is . After adding , the total is , so the new mean is to significant figures. | ||
| Question | Scheme | Marks |
|---|---|---|
| 1 |
| 3 |
| (3 marks) | 3 | |
| Notes | ||
| Adding shifts every value and hence adds to the mean but does not change the spread. Multiplying by doubles both the mean contribution and the standard deviation. Thus the new mean is and the new standard deviation is . | ||
| 2 |
| 5 |
| (5 marks) | 5 | |
| Notes | ||
| For A, and . For B, and . Smaller standard deviation means less spread, so A is more consistent. | ||
| 3 | 4 | |
| (4 marks) | 4 | |
| Notes | ||
| The sum is . Rearranging gives . | ||
| Question | Scheme | Marks |
|---|---|---|
| 1 |
| 5 |
| (5 marks) | 5 | |
| Notes | ||
| The combined sum is , so the combined mean is . The combined sum of squares is . Therefore , giving . | ||
| 2 |
| 7 |
| (7 marks) | 7 | |
| Notes | ||
| For , the mean is and the variance is , so . Since , and . Also and . After adding , these become and for values. Thus the new mean is and the new standard deviation is , giving and . | ||
| 3 |
| 6 |
| (6 marks) | 6 | |
| Notes | ||
| The total sum is , while the six known values sum to , so . The total sum of squares is . The known squares sum to , so . Hence , giving . Thus and are the roots of ; the stated order gives and . | ||
| 4 |
| 6 |
| (6 marks) | 6 | |
| Notes | ||
| Use midpoints . Then , so . Also . Hence , giving . Midpoint substitution loses the unknown variation within each class. | ||
| 5 |
| 6 |
| (6 marks) | 6 | |
| Notes | ||
| The combined sum is and group A contributes , so group B has sum and mean . From , the combined sum of squares is , while group A contributes . Thus group B has sum of squares , variance , and standard deviation . | ||
| Question | Scheme | Marks |
|---|---|---|
| 1 |
| 2 |
| (2 marks) | 2 | |
| Notes | ||
| The interquartile range is . The upper outlier boundary is . Since , it is flagged as a possible outlier. | ||
| 2 |
| 2 |
| (2 marks) | 2 | |
| Notes | ||
| Box plots summarise each distribution using its median, quartiles and extremes. Putting them on the same scale makes the requested comparison direct. | ||
| Question | Scheme | Marks |
|---|---|---|
| 1 |
| 4 |
| (4 marks) | 4 | |
| Notes | ||
| The interquartile range is . The upper boundary is , so is flagged as an outlier. Being an outlier does not prove the value is wrong: check the source and measurement conditions, then document any correction or exclusion. | ||
| 2 |
| 4 |
| (4 marks) | 4 | |
| Notes | ||
| The variable is continuous, while pie charts reduce it to proportions in chosen categories and can conceal both sample size and distributional detail. Common-scale histograms preserve shape; common-scale box plots give a concise comparison of centre and spread. | ||
| 3 |
| 4 |
| (4 marks) | 4 | |
| Notes | ||
| Here . The lower-quartile position is , rounded up to the third value, so . The upper-quartile position is , rounded up to the ninth value, so . Thus the IQR is and the fences are and . Only lies outside them. | ||
| Question | Scheme | Marks |
|---|---|---|
| 1 |
| 5 |
| (5 marks) | 5 | |
| Notes | ||
| Replace the contribution of by : the corrected sum is , and the corrected sum of squares is . Thus and . The erroneous value lay much farther from the centre, so it made the original spread larger. | ||
| 2 |
| 5 |
| (5 marks) | 5 | |
| Notes | ||
| The IQR is . The lower fence is and the upper fence is , so only is flagged. The total is , giving mean , and the middle two values are both , giving median . The mean is pulled upward by the genuine outage value. An outlier rule flags a value for investigation; it does not by itself justify deletion. | ||
| 3 |
| 6 |
| (6 marks) | 6 | |
| Notes | ||
| With , and are whole numbers. The quartiles are therefore the midpoints of the rd and th values, and of the th and th values: and . Hence and the upper fence is . The strict inequality gives , so the stated integer range gives . A verified observation is not an error merely because it is unusual. | ||
| 4 |
| 6 |
| (6 marks) | 6 | |
| Notes | ||
| Since and are whole positions, average positions and , then positions and . Inserting any integer from to changes only the middle ordering: the relevant pairs remain and . The quartiles and fence are therefore invariant, so the outlier decision for is robust even though the missing reading itself remains unresolved. | ||
| 5 |
| 7 |
| (7 marks) | 7 | |
| Notes | ||
| A repeated identifier shows duplication, while the source record supplies evidence for the decimal correction. The blank supplies no observed value. After cleaning, the sixth value is the median. Since and are not whole numbers, round up to positions and , giving quartiles and . The resulting fences contain all values. | ||