Skip to content
P5

Understand that empirical unbiased samples tend towards theoretical probability distributions, with increasing sample size

Sampling and probability

Worked answers, methods and verified real exam appearances for P5 on Edexcel GCSE Maths 1MA1.

Explanation

  • An empirical distribution is based on observed relative frequencies; a theoretical distribution is predicted by a probability model.
  • With unbiased, independent trials, a larger sample usually gives more stable relative frequencies that are closer to the theoretical probabilities.
  • This is a long-run tendency, not a promise that every larger sample will be closer.
  • Increasing sample size reduces random variation, but it does not remove bias caused by an unfair device or poor sampling method.

Worked example

A fair coin gives 162162 heads in 300300 tosses and 10381038 heads in 20002000 tosses. Compare the two relative frequencies with 0.50.5.

  1. 1.For 300300 tosses: 162300=0.54\dfrac{162}{300}=0.54.
  2. 2.For 20002000 tosses: 10382000=0.519\dfrac{1038}{2000}=0.519.
  3. 3.0.5190.519 is closer to the theoretical probability 0.50.5 than 0.540.54.

Answer: The larger sample gives the closer relative frequency in this experiment.

Common mistakes

  • Don't fall into the trap of saying that a larger sample guarantees an exact theoretical distribution.
  • Don't fall into the trap of assuming more trials can correct a biased experiment.

Exam tip

When comparing experiments, calculate relative frequencies rather than comparing raw counts.

Worked practice

Q1
Tier 1 · Easy

1

A fair coin gives a relative frequency of heads of 0.640.64 after 2525 tosses. Explain what is likely to happen to the relative frequency as many more unbiased tosses are made.

(1)

(Total for Question 1 is 1 mark)

Mark scheme

Mark scheme for question 1
QuestionAnswerMarkMark scheme
1
  • It is likely to move closer to the theoretical probability 0.50.5, although it need not do so after every extra toss.
1Use the long-run tendency: for unbiased coin tosses the empirical proportion tends towards the theoretical probability 1/21/2 as the sample becomes large.
Q2
Tier 2 · Standard

2

A spinner is designed to land on green with probability 0.30.3. It lands on green 88 times in 2020 spins and 6363 times in 200200 spins. Compare the two empirical probabilities with the theoretical probability.

(3)

(Total for Question 2 is 3 marks)

Mark scheme

Mark scheme for question 2
QuestionAnswerMarkMark scheme
2
  • For 2020 spins, the relative frequency is 0.40.4.
  • For 200200 spins, the relative frequency is 0.3150.315.
  • The larger sample gives the value closer to 0.30.3.
3Calculate 8/20=0.48/20=0.4 and 63/200=0.31563/200=0.315. Their distances from 0.30.3 are 0.10.1 and 0.0150.015, so the empirical result from 200200 spins is closer to the theoretical probability.
Q3
Tier 3 · Hard

3

A fair die is tested. In the first 6060 rolls, a six appears 1616 times. After 600600 rolls in total, a six has appeared 108108 times. A student says the die must be biased because neither relative frequency equals 16\dfrac{1}{6}. Evaluate the student's claim.

(4)

(Total for Question 3 is 4 marks)

Mark scheme

Mark scheme for question 3
QuestionAnswerMarkMark scheme
3
  • The relative frequencies are 16600.267\dfrac{16}{60}\approx0.267 and 108600=0.18\dfrac{108}{600}=0.18.
  • The theoretical probability is 160.167\dfrac{1}{6}\approx0.167.
  • The larger-sample result is closer to the theoretical value, and an exact match is not expected, so the evidence does not establish bias.
4Compare each empirical probability with 1/61/6. The differences are about 0.1000.100 for the first 6060 rolls and 0.0130.013 after 600600 rolls. Random variation explains a non-zero difference, and the movement towards 1/61/6 as the sample grows is consistent with a fair die.
Q4
Tier 1 · Easy

4

Kai tosses an unbiased coin many times and recalculates the relative frequency of heads after every 1010 tosses. Give a reason why the first few relative frequencies are likely to fluctuate more than the later ones.

(1)

(Total for Question 4 is 1 mark)

Mark scheme

Mark scheme for question 4
QuestionAnswerMarkMark scheme
4
  • The early relative frequencies use small samples, so each new result has a larger effect. With more tosses, the relative frequency tends to become more stable and approach 0.50.5.
1For a small total, one extra head changes the fraction by a relatively large amount. As the denominator grows, one result changes the fraction less, so random variation has less effect.
Q5
Tier 2 · Standard

5

A spinner has theoretical probabilities 0.20.2, 0.30.3 and 0.50.5 for red, yellow and blue. In 5050 spins the frequencies are 88, 1919 and 2323. In 500500 spins the frequencies are 102102, 146146 and 252252. Work out the two empirical distributions and decide which is closer to the theoretical distribution.

(4)

(Total for Question 5 is 4 marks)

Mark scheme

Mark scheme for question 5
QuestionAnswerMarkMark scheme
5
  • 5050 spins: red 0.160.16, yellow 0.380.38, blue 0.460.46.
  • 500500 spins: red 0.2040.204, yellow 0.2920.292, blue 0.5040.504.
  • The 500500-spin distribution is closer to the theoretical distribution.
4Divide each frequency by its experiment total. For 5050 spins this gives 8/50=0.168/50=0.16, 19/50=0.3819/50=0.38 and 23/50=0.4623/50=0.46. For 500500 spins it gives 102/500=0.204102/500=0.204, 146/500=0.292146/500=0.292 and 252/500=0.504252/500=0.504. Every value in the larger sample is closer to its theoretical probability.
Q6
Tier 3 · Hard

6

A council wants to estimate the proportion of all teenagers who play chess online. It first asks 3030 teenage chess-club members and 1818 say yes. It later asks 900900 teenage chess-club members at several tournaments and 540540 say yes. Work out both empirical probabilities. Is the second result necessarily a reliable estimate for all teenagers because its sample is larger? Give a reason for your answer.

(4)

(Total for Question 6 is 4 marks)

Mark scheme

Mark scheme for question 6
QuestionAnswerMarkMark scheme
6
  • No. Both empirical probabilities are 0.60.6.
  • The second result is more stable for chess-club members, but the sample is biased towards teenagers interested in chess. A larger biased sample need not estimate the proportion for all teenagers reliably.
4The first empirical probability is 18/30=0.618/30=0.6 and the second is 540/900=0.6540/900=0.6. The larger result has less random variation for the sampled group, but every person sampled is a chess-club member, so increasing the sample size does not remove this selection bias.
Q7
Tier 2 · Standard

7

A fair spinner has probability 0.250.25 of landing on blue. It lands on blue 55 times in its first 2020 spins and 5656 times after 200200 spins in total. A student says a larger sample must always give a relative frequency closer to the theoretical probability. Work out both relative frequencies and decide whether these results support the student's claim. Give a reason for your answer.

(3)

(Total for Question 7 is 3 marks)

Mark scheme

Mark scheme for question 7
QuestionAnswerMarkMark scheme
7
  • No, the results do not support the claim. The relative frequencies are 0.250.25 and 0.280.28.
  • The larger-sample value is farther from 0.250.25. A larger unbiased sample tends to be more stable, but it is not guaranteed to be closer every time.
3Calculate 5/20=0.255/20=0.25 and 56/200=0.2856/200=0.28. The first value matches the theoretical probability, whereas the later value differs by 0.030.03. Long-run convergence is a tendency rather than a guarantee for every growing sample.
Q8
Tier 3 · Hard

8

A wheel has four sectors with theoretical probabilities 0.150.15, 0.250.25, 0.300.30 and 0.300.30 for red, blue, green and white. In 4040 spins the frequencies are 99, 88, 1313 and 1010. In 400400 spins the frequencies are 5858, 104104, 126126 and 112112. Work out both empirical distributions. For the 4040-spin distribution, write down the colour whose empirical probability differs most from its theoretical probability. State, with a reason, which of the two distributions you expect to be closer to the theoretical distribution.

(5)

(Total for Question 8 is 5 marks)

Mark scheme

Mark scheme for question 8
QuestionAnswerMarkMark scheme
8
  • 4040 spins: (0.225,0.20,0.325,0.25)(0.225,0.20,0.325,0.25); 400400 spins: (0.145,0.26,0.315,0.28)(0.145,0.26,0.315,0.28).
  • Red differs most in the 4040-spin distribution (difference 0.0750.075).
  • The 400400-spin distribution, because a larger number of trials gives relative frequencies closer to the theoretical probabilities.
5Divide the first four frequencies by 4040 and the second four by 400400. For 4040 spins the absolute differences are 0.075,0.05,0.025,0.050.075,0.05,0.025,0.05, so red differs most. For 400400 spins the differences are 0.005,0.01,0.015,0.020.005,0.01,0.015,0.02, all smaller, and the larger number of trials is expected to lie closer to the theoretical distribution.
Q9
Tier 3 · Hard

9

A spinner lands on red 3434 times in its first 160160 spins. It is then spun a further 8080 times. After all 240240 spins, the relative frequency of red is exactly 0.250.25. A maintenance warning is issued if the relative frequency of red during the further 8080 spins is greater than 0.300.30. Decide whether a warning should be issued.

(4)

(Total for Question 9 is 4 marks)

Mark scheme

Mark scheme for question 9
QuestionAnswerMarkMark scheme
9
  • 2626 of the further spins land on red.
  • Yes. Their relative frequency is 26/80=0.32526/80=0.325, which is greater than 0.300.30.
4A relative frequency of 0.250.25 after 240240 spins means there are 240×0.25=60240\times0.25=60 red results altogether. Therefore 6034=2660-34=26 of the further 8080 spins land on red. Since 26/80=0.325>0.3026/80=0.325>0.30, the maintenance warning should be issued.
Q10
Tier 3 · Hard

10

Two wheels are designed to land on red with probability 0.250.25. Wheel A lands on red 1717 times in its first 6060 spins and 156156 times after 600600 spins in total. Wheel B lands on red 1616 times in its first 6060 spins and 204204 times after 600600 spins in total. Work out the four relative frequencies. Which wheel gives stronger evidence that it may be biased? Explain your answer.

(5)

(Total for Question 10 is 5 marks)

Mark scheme

Mark scheme for question 10
QuestionAnswerMarkMark scheme
10
  • Wheel A: 0.2830.283 (to 3 d.p.) initially and 0.260.26 overall.
  • Wheel B: 0.2670.267 (to 3 d.p.) initially and 0.340.34 overall.
  • Wheel B gives stronger evidence that it may be biased, because its larger-sample relative frequency remains 0.090.09 above 0.250.25, whereas wheel A's is only 0.010.01 above it.
5Divide each red frequency by its number of spins. The early results are both fairly close to 0.250.25, but the 600600-spin results are more stable evidence. Wheel A moves closer to 0.250.25, while wheel B moves much farther away and remains there over the larger sample.

Verified exam appearances

SeriesPaperQuestionMarksCalculatorTierLinks
2022-111HQ104Non-calculatorHigherQPMS
2019-061FQ172Non-calculatorFoundationQPMS
2023-062HQ104AllowedHigherQPMS

Other points in P Probability

Want help turning this into marks?

Bring P5 or any tricky specification point, and we can work through the method and exam wording together.