Skip to content
S2.4

Recognise and interpret possible outliers in data sets and statistical diagrams; select or critique data presentation techniques in context; clean data, including dealing with missing data, errors and outliers.

Draft — not yet indexed

Outliers and cleaning data

Worked answers and methods for S2.4 on Edexcel A-level Maths 9MA0.

Explanation

  • A common outlier rule flags values below Q11.5IQRQ_1-1.5\operatorname{IQR} or above Q3+1.5IQRQ_3+1.5\operatorname{IQR}, but context should guide the final decision.
  • Investigate a suspicious value against the original record before correcting or removing it; an unusual valid observation is not automatically an error.
  • Handle missing data transparently: record how many values are missing, avoid inventing unsupported values and consider whether missingness could bias conclusions.
  • Choose displays to suit the data and purpose: histograms for continuous grouped data, box plots for comparing distributions, and scatter diagrams for paired variables.

Worked example

A table of package masses contains one blank entry and one value 482482 among values near 48.248.2 grams. Describe a defensible way to clean these two entries before analysis.

  1. 1.First distinguish a data-entry error from a genuine extreme value by consulting the source.
  2. 2.Do not silently divide 482482 by 1010.
  3. 3.The blank contains no observed value, so omit it from calculations unless a justified imputation rule has been chosen, and document the decision so its possible bias is visible.

Answer: Check the original measurement record for both entries.; Correct 482482 to 48.248.2 only if the source confirms a decimal-point error; otherwise retain and flag it or exclude it with a stated reason.; Treat the blank as missing rather than replacing it without evidence, and report the reduced sample size or justified imputation method.

Common mistakes

  • Don't replace every missing value with zero, changing both the centre and spread of the data.
  • Don't delete an unusual value automatically instead of investigating whether it is an error or genuine observation.

Exam tip

Document separate rules for missing values, transcription errors and plausible outliers before recalculating summaries.

Worked practice

Q1
Tier 1 · Easy

1.

For a data set, Q1=14Q_1=14 and Q3=22Q_3=22. Use the 1.5IQR1.5\operatorname{IQR} rule to determine whether the value 3636 is a possible outlier.

(2)

(Total for Question 1 is 2 marks)

Mark scheme

Mark scheme for question 1
QuestionSchemeMarks
1
  • The upper boundary is 3434, so 3636 is a possible outlier.
2
Notes
The interquartile range is 2214=822-14=8. The upper outlier boundary is 22+1.5(8)=3422+1.5(8)=34. Since 36>3436>34, it is flagged as a possible outlier.

(2 marks)

Q2
Tier 2 · Standard

2.

For a data set, the lower quartile is 1818 and the upper quartile is 3030. Use the 1.5×IQR1.5\times\operatorname{IQR} rule to decide whether a value of 5050 is an outlier, and state what should be done before removing it.

(4)

(Total for Question 2 is 4 marks)

Mark scheme

Mark scheme for question 2
QuestionSchemeMarks
2
  • IQR=12\operatorname{IQR}=12 and the upper outlier boundary is 4848.
  • 5050 is an outlier.
  • Check the original record and context: correct it only if an error is confirmed, retain it if genuine, or exclude it only with a justified and documented reason.
4
Notes
The interquartile range is 3018=1230-18=12. The upper boundary is Q3+1.5IQR=30+18=48Q_3+1.5\operatorname{IQR}=30+18=48, so 50>4850>48 is flagged as an outlier. Being an outlier does not prove the value is wrong: check the source and measurement conditions, then document any correction or exclusion.

(4 marks)

Q3
Tier 3 · Hard

3.

A data set of 4040 readings was summarised as x=504\sum x=504 and x2=6856.9\sum x^2=6856.9. One reading was entered as 3131 but the source record confirms it should be 1313. Calculate the corrected mean and population standard deviation, and state the likely effect of the error on the original spread.

(5)

(Total for Question 3 is 5 marks)

Mark scheme

Mark scheme for question 3
QuestionSchemeMarks
3
  • Corrected mean =12.15=12.15.
  • Corrected standard deviation =2.00=2.00.
  • The error inflated the original spread.
5
Notes
Replace the contribution of 3131 by 1313: the corrected sum is 50431+13=486504-31+13=486, and the corrected sum of squares is 6856.9312+132=6064.96856.9-31^2+13^2=6064.9. Thus xˉ=486/40=12.15\bar{x}=486/40=12.15 and σ=6064.9/4012.152=4=2.00\sigma=\sqrt{6064.9/40-12.15^2}=\sqrt{4}=2.00. The erroneous value lay much farther from the centre, so it made the original spread larger.

(5 marks)

Q4
Tier 1 · Easy

4.

A researcher wants to compare the median and spread of journey times for two train operators. State a suitable diagram and give one reason for your choice.

(2)

(Total for Question 4 is 2 marks)

Mark scheme

Mark scheme for question 4
QuestionSchemeMarks
4
  • Use box plots drawn on a common scale.
  • They display the median and interquartile range, so the centres and spreads of the two distributions can be compared directly.
2
Notes
Box plots summarise each distribution using its median, quartiles and extremes. Putting them on the same scale makes the requested comparison direct.

(2 marks)

Q5
Tier 2 · Standard

5.

Two schools have different numbers of students. A report uses two pie charts to compare their students' continuous weekly screen times. Critique this presentation and suggest a more informative display for comparing the distributions.

(4)

(Total for Question 5 is 4 marks)

Mark scheme

Mark scheme for question 5
QuestionSchemeMarks
5
  • Pie charts require screen time to be grouped into categories and do not show the detailed shape or spread of the continuous data.
  • Unless the sample sizes are shown, equal-looking sectors can hide very different frequencies between the schools.
  • Use comparative histograms with common class boundaries and frequency density, or box plots on a common scale.
  • Histograms compare distribution shape; box plots compare median and spread.
4
Notes
The variable is continuous, while pie charts reduce it to proportions in chosen categories and can conceal both sample size and distributional detail. Common-scale histograms preserve shape; common-scale box plots give a concise comparison of centre and spread.

(4 marks)

Q6
Tier 3 · Hard

6.

Ten response times, in seconds, are 8,9,9,10,10,10,11,11,12,408,9,9,10,10,10,11,11,12,40. For these data, Q1=9Q_1=9 and Q3=11Q_3=11. Use the 1.5IQR1.5\operatorname{IQR} rule to identify any possible outlier. Calculate the mean and median, then recommend which measure better represents a typical response time. The source confirms that the value 4040 is a genuine response during a network outage; state whether it should automatically be deleted.

(5)

(Total for Question 6 is 5 marks)

Mark scheme

Mark scheme for question 6
QuestionSchemeMarks
6
  • IQR=2\operatorname{IQR}=2, with lower boundary 66 and upper boundary 1414, so 4040 is the only possible outlier.
  • Mean =13=13 seconds and median =10=10 seconds.
  • The median better represents a typical response time because it is resistant to the extreme value.
  • 4040 should not automatically be deleted; it is genuine and should be retained or analysed separately according to whether outage performance is part of the population of interest.
5
Notes
The IQR is 119=211-9=2. The lower fence is 91.5(2)=69-1.5(2)=6 and the upper fence is 11+1.5(2)=1411+1.5(2)=14, so only 4040 is flagged. The total is 130130, giving mean 1313, and the middle two values are both 1010, giving median 1010. The mean is pulled upward by the genuine outage value. An outlier rule flags a value for investigation; it does not by itself justify deletion.

(5 marks)

Q7
Tier 2 · Standard

7.

The ordered data are 4,6,7,8,9,10,11,12,14,15,294,6,7,8,9,10,11,12,14,15,29. Use the Edexcel discrete convention: round a quartile position up when it is not a whole number; when it is a whole number, average the value in that position and the next value. Find Q1Q_1, Q3Q_3 and the outlier fences, and identify any possible outlier.

(4)

(Total for Question 7 is 4 marks)

Mark scheme

Mark scheme for question 7
QuestionSchemeMarks
7
  • Q1=7Q_1=7 and Q3=14Q_3=14
  • IQR=7\operatorname{IQR}=7; the fences are 3.5-3.5 and 24.524.5.
  • 2929 is the only possible outlier.
4
Notes
Here n=11n=11. The lower-quartile position is 11/4=2.7511/4=2.75, rounded up to the third value, so Q1=7Q_1=7. The upper-quartile position is 33/4=8.2533/4=8.25, rounded up to the ninth value, so Q3=14Q_3=14. Thus the IQR is 77 and the fences are 71.5(7)=3.57-1.5(7)=-3.5 and 14+1.5(7)=24.514+1.5(7)=24.5. Only 2929 lies outside them.

(4 marks)

Q8
Tier 3 · Hard

8.

The ordered data are 3,4,5,6,7,8,9,10,x,14,15,253,4,5,6,7,8,9,10,x,14,15,25, where xx is an integer and 10x1410\leq x\leq14. Use the Edexcel discrete convention: round a quartile position up when it is not a whole number; when it is a whole number, average the value in that position and the next value. Find all values of xx for which 2525 is a possible outlier under the 1.5IQR1.5\operatorname{IQR} rule. State how the value 2525 should be treated if its source record confirms that it is genuine.

(6)

(Total for Question 8 is 6 marks)

Mark scheme

Mark scheme for question 8
QuestionSchemeMarks
8
  • Q1=(5+6)/2=5.5Q_1=(5+6)/2=5.5, Q3=(x+14)/2Q_3=(x+14)/2 and IQR=(x+3)/2\operatorname{IQR}=(x+3)/2.
  • The upper fence is x+142+1.5(x+32)=5x+374\dfrac{x+14}{2}+1.5\left(\dfrac{x+3}{2}\right)=\dfrac{5x+37}{4}.
  • 2525 is a possible outlier when 25>(5x+37)/425>(5x+37)/4, so x<63/5=12.6x<63/5=12.6.
  • The integer values are x{10,11,12}x\in\{10,11,12\}; at x=13x=13, the upper fence is 25.525.5, so 2525 is not outside it.
  • A genuine value should not be deleted automatically; retain it or analyse it separately according to the population and purpose, documenting the decision.
6
Notes
With n=12n=12, n/4=3n/4=3 and 3n/4=93n/4=9 are whole numbers. The quartiles are therefore the midpoints of the 33rd and 44th values, and of the 99th and 1010th values: Q1=5.5Q_1=5.5 and Q3=(x+14)/2Q_3=(x+14)/2. Hence IQR=(x+3)/2\operatorname{IQR}=(x+3)/2 and the upper fence is (5x+37)/4(5x+37)/4. The strict inequality 25>(5x+37)/425>(5x+37)/4 gives x<12.6x<12.6, so the stated integer range gives 10,11,1210,11,12. A verified observation is not an error merely because it is unusual.

(6 marks)

Q9
Tier 3 · Hard

9.

Fifteen ordered readings are 2,4,5,7,8,9,10,11,12,13,14,15,16,18,302,4,5,7,8,9,10,11,12,13,14,15,16,18,30. A sixteenth reading is missing from the file, but its source record shows only that it is an integer from 1010 to 1313 inclusive. Do not replace it by an average. Using the Edexcel discrete convention, show that Q1Q_1, Q3Q_3 and the upper outlier fence are the same for every possible value of the missing reading. Hence decide whether 3030 is always a possible outlier, and state what should be recorded before the data are used.

(6)

(Total for Question 9 is 6 marks)

Mark scheme

Mark scheme for question 9
QuestionSchemeMarks
9
  • For n=16n=16, Q1Q_1 is the mean of the 44th and 55th values, so Q1=(7+8)/2=7.5Q_1=(7+8)/2=7.5.
  • For every allowed missing value, the 1212th and 1313th values are 1414 and 1515, so Q3=14.5Q_3=14.5.
  • The IQR is 77 and the upper fence is 14.5+1.5(7)=2514.5+1.5(7)=25.
  • 30>2530>25, so 3030 is a possible outlier for every allowed value of the missing reading.
  • Record the value as unresolved within the verified range, document that no unsupported imputation was made, and check the original source before any final correction or exclusion.
6
Notes
Since 16/4=416/4=4 and 3(16)/4=123(16)/4=12 are whole positions, average positions 44 and 55, then positions 1212 and 1313. Inserting any integer from 1010 to 1313 changes only the middle ordering: the relevant pairs remain 7,87,8 and 14,1514,15. The quartiles and fence are therefore invariant, so the outlier decision for 3030 is robust even though the missing reading itself remains unresolved.

(6 marks)

Q10
Tier 3 · Hard

10.

A data extract lists measurements 41,42,43,44,45,46,47,47,48,49,50,51041,42,43,44,45,46,47,47,48,49,50,510 and one blank entry. The two values 4747 have the same record identifier, the source confirms that 510510 is a decimal-point error for 51.051.0, and the blank cannot be recovered. State how each issue should be handled. For the cleaned data, find the sample size, median, quartiles and any possible outliers using the Edexcel discrete convention and the 1.5IQR1.5\operatorname{IQR} rule.

(7)

(Total for Question 10 is 7 marks)

Mark scheme

Mark scheme for question 10
QuestionSchemeMarks
10
  • Remove one of the duplicate 4747 records, correct 510510 to 51.051.0 from the confirmed source, and treat the unrecoverable blank as missing rather than as zero.
  • The cleaned ordered data are 41,42,43,44,45,46,47,48,49,50,5141,42,43,44,45,46,47,48,49,50,51, so n=11n=11 and the median is 4646.
  • Q1=43Q_1=43 and Q3=49Q_3=49.
  • The IQR is 66, giving fences 3434 and 5858; there are no possible outliers.
  • The cleaning decisions and reduced sample size should be documented.
7
Notes
A repeated identifier shows duplication, while the source record supplies evidence for the decimal correction. The blank supplies no observed value. After cleaning, the sixth value is the median. Since n/4=2.75n/4=2.75 and 3n/4=8.253n/4=8.25 are not whole numbers, round up to positions 33 and 99, giving quartiles 4343 and 4949. The resulting fences contain all 1111 values.

(7 marks)

Verified exam appearances

We have not yet indexed a verified real-paper appearance for S2.4. Browse the Edexcel A-level Maths 9MA0 past papers directly.

Other points in S2 Data presentation and interpretation

Want help turning this into marks?

Bring S2.4 or any tricky specification point, and we can work through the method and exam wording together.