S4 Statistical distributions — revision question pack

3 specification points · notes, questions, answers and worked methods

Checked against Edexcel 9MA0 section S4. Review basis: the qualification registry sourced from the Pearson Edexcel Level 3 Advanced GCE in Mathematics (9MA0) specification; registry verification recorded 11 July 2026.

How this checking works

S4.1 · Understand and use simple, discrete probability distributions (mean and variance of discrete random variables excluded), including the binomial distribution as a model; calculate probabilities using the binomial distribution.

Explanation

  • A discrete probability distribution lists possible values with probabilities between 00 and 11 whose total is 11; a discrete uniform distribution assigns equal probability to each value. Use XB(n,p)X\sim\operatorname{B}(n,p) for a fixed number nn of independent trials, each with two outcomes and constant success probability pp.
  • For a binomial variable, P(X=r)=(nr)pr(1p)nrP(X=r)=\binom{n}{r}p^r(1-p)^{n-r}; cumulative probabilities are often most efficiently found with a calculator.
  • Translate inequalities carefully: for integer-valued XX, P(X<r)=P(Xr1)P(X<r)=P(X\leq r-1) and P(Xr)=1P(Xr1)P(X\geq r)=1-P(X\leq r-1).
  • Use the cumulative binomial probability P(X4)=r=04(14r)(0.4)r(0.6)14r=0.279256P(X\leq4)=\sum_{r=0}^{4}\binom{14}{r}(0.4)^r(0.6)^{14-r}=0.279256\ldots.
  • Therefore P(X4)=0.2793P(X\leq4)=0.2793.

Worked example

Let XB(14,0.4)X\sim\operatorname{B}(14,0.4). Find P(X4)P(X\leq4) to 44 decimal places.

  1. 1.Use the cumulative binomial probability P(X4)=r=04(14r)(0.4)r(0.6)14r=0.279256P(X\leq4)=\sum_{r=0}^{4}\binom{14}{r}(0.4)^r(0.6)^{14-r}=0.279256\ldots.
  2. 2.Therefore P(X4)=0.2793P(X\leq4)=0.2793.

Answer: 0.27930.2793

Common mistakes

  • Don't use P(X<k)P(X<k) when the question asks for P(Xk)P(X\leq k).
  • Don't calculate a single binomial probability when the event is cumulative, or omit an endpoint from the sum.

Exam tip

Translate the inequality into the exact binomial range before using cumulative probabilities and round only at the end.

Tier 1 · Easy

  1. 1.

    A discrete random variable XX takes values 00, 11 and 22 with probabilities kk, 3k3k and 4k4k respectively. Find kk and P(X1)P(X\geq1).

    (3)

    (Total for Question 1 is 3 marks)

  2. 2.

    A discrete random variable YY takes values 3-3, 1-1, 22 and 55 with probabilities 0.100.10, 0.250.25, 0.400.40 and 0.250.25 respectively. Find P(Y>0)P(Y>0) and P(Y2)P(|Y|\leq2).

    (2)

    (Total for Question 2 is 2 marks)

Tier 2 · Standard

  1. 1.

    The random variable XX has a discrete uniform distribution on {2,0,2,4,6}\{-2,0,2,4,6\}. Write down its probability distribution, then find P(X<3)P(X<3) and P(X4)P(|X|\geq4).

    (3)

    (Total for Question 1 is 3 marks)

  2. 2.

    Independent trials each have success probability 0.200.20. Find the smallest number of trials needed for the probability of at least one success to exceed 0.950.95. State this probability for your minimum number of trials to 44 decimal places.

    (4)

    (Total for Question 2 is 4 marks)

  3. 3.

    Let XB(20,0.35)X\sim\operatorname{B}(20,0.35). Given P(X3)=0.0443756P(X\leq3)=0.0443756, P(X4)=0.1181966P(X\leq4)=0.1181966 and P(X8)=0.7623776P(X\leq8)=0.7623776, find P(4X8)P(4\leq X\leq8) and P(4<X8)P(4<X\leq8). Give each answer to 44 decimal places.

    (4)

    (Total for Question 3 is 4 marks)

Tier 3 · Hard

  1. 1.

    A binomial random variable XX has 1010 trials and satisfies P(X=0)=0.810P(X=0)=0.8^{10}. Find its success probability pp, then calculate P(X3)P(X\geq3) to 44 decimal places.

    (5)

    (Total for Question 1 is 5 marks)

  2. 2.

    Let XB(12,0.25)X\sim\operatorname{B}(12,0.25). Given that at least one success occurs, find the probability that at least three successes occur. Give your answer to 44 decimal places, using unrounded values in your working.

    (5)

    (Total for Question 2 is 5 marks)

  3. 3.

    Eight independent trials share a constant success probability pp, where 0<p<10<p<1, and XX denotes their total number of successes. Given that P(X=2)=3P(X=1)P(X=2)=3P(X=1), find pp. Hence find P(X2)P(X\geq2) to 44 decimal places.

    (6)

    (Total for Question 3 is 6 marks)

  4. 4.

    A binomial random variable XB(n,p)X\sim\operatorname{B}(n,p) satisfies P(X=1)=6P(X=0)P(X=1)=6P(X=0) and P(X=2)=2P(X=1)P(X=2)=2P(X=1), where n2n\geq2 and 0<p<10<p<1. Find nn and pp. Hence find the exact value of P(X2)P(X\geq2).

    (6)

    (Total for Question 4 is 6 marks)

  5. 5.

    Machine A produces 88 independently inspected components and machine B produces 77. Every component, regardless of machine, has the same probability 0.400.40 of being defective, and all inspections are independent. Let TT be the total number of defective components. State the distribution of TT and calculate P(T8)P(T\geq8) to 44 decimal places. Explain why the same binomial distribution for TT would not be valid if machine B instead had defect probability 0.200.20.

    (6)

    (Total for Question 5 is 6 marks)

S4.2 · Understand and use the Normal distribution as a model; find probabilities using the Normal distribution; link to histograms, mean, standard deviation, points of inflection and the binomial distribution.

Explanation

  • Write XN(μ,σ2)X\sim\operatorname{N}(\mu,\sigma^2), where μ\mu is the mean and σ\sigma is the standard deviation; standardise with Z=XμσZ=\frac{X-\mu}{\sigma}.
  • The Normal curve is continuous, symmetric about μ\mu, has total area 11, and has points of inflection at μσ\mu-\sigma and μ+σ\mu+\sigma.
  • A Normal model is plausible for a roughly symmetric, unimodal histogram with no strong outliers, but context and the variable's possible values also matter.
  • A binomial distribution may be approximated by a Normal distribution when nn is large and pp is close to 0.50.5; use a continuity correction.
A Normal density is symmetric about its mean; its inflection points occur at μσ\mu-\sigma and μ+σ\mu+\sigma, and shaded area represents probability.

Worked example

The lifetime LL of a component, in hours, is modelled by LN(72,82)L\sim\operatorname{N}(72,8^2). Find the lifetime exceeded by exactly 10%10\% of components, to 11 decimal place.

  1. 1.If P(L>l)=0.10P(L>l)=0.10, then P(Ll)=0.90P(L\leq l)=0.90.
  2. 2.The 0.900.90 standard Normal quantile is z=1.28155z=1.28155\ldots.
  3. 3.Hence l=72+8(1.28155)=82.252l=72+8(1.28155\ldots)=82.252\ldots, so the required lifetime is 82.382.3 hours.

Answer: 82.382.3 hours

Common mistakes

  • Don't use the variance as the denominator when standardising instead of the standard deviation.
  • Don't use the lower-tail quantile for a value exceeded by a stated percentage.

Exam tip

Convert an exceedance percentage to the corresponding lower-tail probability before standardising or using an inverse Normal calculation.

Tier 1 · Easy

  1. 1.

    A random variable has distribution XN(50,62)X\sim\operatorname{N}(50,6^2). Find P(X<56)P(X<56) to 44 decimal places.

    (2)

    (Total for Question 1 is 2 marks)

  2. 2.

    The random variable TT has distribution TN(72,92)T\sim\operatorname{N}(72,9^2). State the two values of TT at which the Normal density curve has points of inflection.

    (2)

    (Total for Question 2 is 2 marks)

Tier 2 · Standard

  1. 1.

    A random variable XX is modelled by XN(64,42)X\sim\operatorname{N}(64,4^2). Find P(58<X<70)P(58<X<70) to 44 decimal places.

    (4)

    (Total for Question 1 is 4 marks)

  2. 2.

    A variable XX is modelled by XN(μ,52)X\sim\operatorname{N}(\mu,5^2). Given that P(X<42)=0.10P(X<42)=0.10 and that the 0.100.10 standard Normal quantile is 1.2816-1.2816, find μ\mu to 33 significant figures.

    (4)

    (Total for Question 2 is 4 marks)

  3. 3.

    A variable XX is modelled by XN(μ,σ2)X\sim\operatorname{N}(\mu,\sigma^2). Given P(X<42)=0.20P(X<42)=0.20 and P(X<58)=0.80P(X<58)=0.80, find μ\mu and σ\sigma. Give σ\sigma to 33 significant figures. You may use the standard Normal quantiles 0.8416-0.8416 and 0.84160.8416.

    (5)

    (Total for Question 3 is 5 marks)

Tier 3 · Hard

  1. 1.

    Let XB(200,0.35)X\sim\operatorname{B}(200,0.35). Use a Normal approximation with a continuity correction to estimate P(X82)P(X\geq82), giving your answer to 44 decimal places.

    (5)

    (Total for Question 1 is 5 marks)

  2. 2.

    A variable XX is modelled by XN(30,σ2)X\sim\operatorname{N}(30,\sigma^2). Given that P(26<X<34)=0.80P(26<X<34)=0.80, find σ\sigma to 33 significant figures. Hence find P(X>36X>32)P(X>36\mid X>32) to 44 decimal places. Use the unrounded value of σ\sigma in your working. You may use the 0.900.90 standard Normal quantile 1.28161.2816.

    (6)

    (Total for Question 2 is 6 marks)

  3. 3.

    A process measurement XX is modelled by a Normal distribution whose density curve has points of inflection at 492492 and 508508. Measurements outside the interval 484X516484\leq X\leq516 are rejected. Find the probability that a measurement is accepted and the expected number rejected from 12001200 measurements. Given that a measurement is accepted, find the probability that it exceeds 505505. Give probabilities to 44 decimal places and the expected number to the nearest integer. Use unrounded values in your working.

    (7)

    (Total for Question 3 is 7 marks)

  4. 4.

    A variable XX is modelled by XN(μ,σ2)X\sim\operatorname{N}(\mu,\sigma^2). Given that P(X<44)=P(X>68)P(X<44)=P(X>68) and P(50<X<62)=0.60P(50<X<62)=0.60, find μ\mu and σ\sigma. You may use the 0.800.80 standard Normal quantile 0.84160.8416. Hence find the two values of XX at which the density curve has points of inflection, giving them to 33 significant figures. A histogram represents 900900 such measurements, with bar area equal to frequency. Estimate the total frequency between the points of inflection, and explain why agreement with this frequency alone would not establish that a Normal model is suitable.

    (7)

    (Total for Question 4 is 7 marks)

  5. 5.

    A variable XX is modelled by XN(μ,σ2)X\sim\operatorname{N}(\mu,\sigma^2). Given that P(X<40)=0.10P(X<40)=0.10 and P(X<55)=0.70P(X<55)=0.70, find μ\mu and σ\sigma to 33 significant figures. Hence find P(45<X<60)P(45<X<60) to 44 decimal places. Use unrounded values in your working. You may use the standard Normal quantiles 1.2816-1.2816 and 0.52440.5244.

    (7)

    (Total for Question 5 is 7 marks)

S4.3 · Select an appropriate probability distribution for a context, with appropriate reasoning, including recognising when the binomial or Normal model may not be appropriate.

Explanation

  • Select a distribution by matching its assumptions to the variable and data-generating process, not merely because its parameters can be estimated.
  • A binomial model needs a fixed number of trials, two outcomes per trial, independence and a constant success probability.
  • A Normal model is continuous and symmetric with unbounded tails, so it can be unsuitable for strongly skewed, bounded or discrete data.
  • Support a choice with contextual evidence such as histogram shape, stability over time and dependence; state how a failed assumption could affect predictions.

Worked example

The masses of loaves from a stable production line form a roughly symmetric, single-peaked histogram with no clear outliers. Explain why a Normal model may be suitable and why a binomial model is not.

  1. 1.Match the continuous measurement and bell-shaped empirical pattern to the features of a Normal distribution.
  2. 2.Reject the binomial model because the response is a measured mass, not a discrete success count.

Answer: Mass is continuous and the observed distribution is approximately symmetric and unimodal, supporting a Normal model.; A binomial model counts successes in a fixed number of two-outcome trials, so it does not model individual loaf masses.

Common mistakes

  • Don't use a binomial model when the success probability changes between trials.
  • Don't choose a distribution from the graph's shape alone and ignore the type of variable and trial assumptions.

Exam tip

Justify a model using the variable type, distribution shape and process assumptions, and name any condition that may fail.

Tier 1 · Easy

  1. 1.

    A manufacturer inspects 1212 independently chosen switches. Each switch has probability 0.040.04 of being faulty. State a suitable distribution for the number FF of faulty switches and give its parameters.

    (2)

    (Total for Question 1 is 2 marks)

  2. 2.

    Waiting times are non-negative and their histogram is strongly right-skewed. Give two reasons why a Normal model may be unsuitable.

    (2)

    (Total for Question 2 is 2 marks)

Tier 2 · Standard

  1. 1.

    A quality inspector selects 2020 components without replacement from a batch of only 5050 and records the number that are defective. Explain why a binomial model may be inappropriate and suggest a more suitable approach.

    (3)

    (Total for Question 1 is 3 marks)

  2. 2.

    A player takes 3030 shots in sequence. The probability of scoring decreases as the player becomes tired. Explain why a binomial model for the total number of goals may be inappropriate and suggest a refinement.

    (4)

    (Total for Question 2 is 4 marks)

  3. 3.

    A system examines 5050 independently selected messages, each with probability 0.080.08 of being spam, and records the number flagged. It also records file-download times, whose histogram is continuous, roughly symmetric and single-peaked with no clear outliers. Select a suitable distribution for each variable and justify both choices.

    (4)

    (Total for Question 3 is 4 marks)

Tier 3 · Hard

  1. 1.

    A technician proposes DB(500,p)D\sim\operatorname{B}(500,p) for the number of defective pixels on each screen. Defects tend to occur in neighbouring clusters, and pp varies between production shifts. Critique the model and suggest how the data should be used before choosing a replacement.

    (5)

    (Total for Question 1 is 5 marks)

  2. 2.

    A business wants to model the number of visits to its website in an hour. There is no fixed maximum number of visits, and the observed distribution is discrete, non-negative and strongly right-skewed. Explain why neither a binomial distribution nor a Normal distribution is well justified. State how the business should use its data before selecting another model.

    (5)

    (Total for Question 2 is 5 marks)

  3. 3.

    Recovery times in a study form a bimodal histogram. Every observation comes from either treatment A or treatment B, and the separate histogram for each treatment is continuous, approximately symmetric and single-peaked. Explain why one Normal distribution for all recovery times is inappropriate. Propose a more suitable modelling strategy and state two checks required before using it.

    (5)

    (Total for Question 3 is 5 marks)

  4. 4.

    A factory records the number XX of faulty items among 400400 independently produced items, each with constant fault probability 0.500.50. Select an exact distribution for XX and explain why a Normal approximation is reasonable. Use the approximation, with a continuity correction, to estimate P(190X210)P(190\leq X\leq210) to 44 decimal places. State one process change that would undermine both the exact binomial model and this approximation.

    (6)

    (Total for Question 4 is 6 marks)

  5. 5.

    A company proposes TN(5.25,2.52)T\sim\operatorname{N}(5.25,2.5^2) for a non-negative service time TT, measured in minutes. Calculate the probability that this model assigns to an impossible negative time and the expected number of such modelled values in 18001800 services. Give the probability to 44 decimal places and the expected number to the nearest integer. Assess the model and describe what evidence should be examined before choosing a replacement.

    (5)

    (Total for Question 5 is 5 marks)

Answer key

Answers begin on a new printed page so the question pack can be completed without the solutions alongside it.

S4.1 · Understand and use simple, discrete probability distributions (mean and variance of discrete random variables excluded), including the binomial distribution as a model; calculate probabilities using the binomial distribution.

Tier 1 · Easy

Mark scheme for S4.1 Tier 1 · Easy
QuestionSchemeMarks
1
  • k=18k=\frac18.
  • P(X1)=78P(X\geq1)=\frac78.
3
(3 marks)3
Notes
Probabilities sum to 11, so k+3k+4k=8k=1k+3k+4k=8k=1 and k=1/8k=1/8. Therefore P(X1)=3k+4k=7k=7/8P(X\geq1)=3k+4k=7k=7/8.
2
  • P(Y>0)=0.65P(Y>0)=0.65
  • P(Y2)=0.65P(|Y|\leq2)=0.65
2
(2 marks)2
Notes
The positive values are 22 and 55, giving 0.40+0.25=0.650.40+0.25=0.65. The values satisfying Y2|Y|\leq2 are 1-1 and 22, giving 0.25+0.40=0.650.25+0.40=0.65.

Tier 2 · Standard

Mark scheme for S4.1 Tier 2 · Standard
QuestionSchemeMarks
1
  • P(X=x)=15P(X=x)=\dfrac15 for each x{2,0,2,4,6}x\in\{-2,0,2,4,6\}
  • P(X<3)=35P(X<3)=\dfrac35
  • P(X4)=25P(|X|\geq4)=\dfrac25
3
(3 marks)3
Notes
There are five equally likely values, so each has probability 1/51/5. The event X<3X<3 contains 2,0,2-2,0,2, giving 3/53/5. The event X4|X|\geq4 contains 44 and 66, giving 2/52/5.
2
  • 1414 trials
  • P(X1)=0.9560P(X\geq1)=0.9560 when n=14n=14
4
(4 marks)4
Notes
For nn trials, P(X1)=10.8nP(X\geq1)=1-0.8^n. The requirement is 0.8n<0.050.8^n<0.05, so n>log(0.05)/log(0.8)=13.425n>\log(0.05)/\log(0.8)=13.425\ldots. The smallest integer is 1414, and 10.814=0.956019=0.95601-0.8^{14}=0.956019\ldots=0.9560.
3
  • P(4X8)=0.7180P(4\leq X\leq8)=0.7180
  • P(4<X8)=0.6442P(4<X\leq8)=0.6442
4
(4 marks)4
Notes
For the inclusive lower endpoint, subtract only values up to 33: P(4X8)=P(X8)P(X3)=0.76237760.0443756=0.7180020P(4\leq X\leq8)=P(X\leq8)-P(X\leq3)=0.7623776-0.0443756=0.7180020. For the strict lower endpoint, subtract values up to 44: P(4<X8)=0.76237760.1181966=0.6441810P(4<X\leq8)=0.7623776-0.1181966=0.6441810. Rounding gives 0.71800.7180 and 0.64420.6442.

Tier 3 · Hard

Mark scheme for S4.1 Tier 3 · Hard
QuestionSchemeMarks
1
  • p=0.2p=0.2.
  • P(X3)=0.3222P(X\geq3)=0.3222.
5
(5 marks)5
Notes
For XB(10,p)X\sim\operatorname{B}(10,p), P(X=0)=(1p)10P(X=0)=(1-p)^{10}. Hence (1p)10=0.810(1-p)^{10}=0.8^{10}, so p=0.2p=0.2. Then P(X3)=1P(X2)=1[0.810+10(0.2)(0.8)9+(102)(0.2)2(0.8)8]=0.322200P(X\geq3)=1-P(X\leq2)=1-[0.8^{10}+10(0.2)(0.8)^9+\binom{10}{2}(0.2)^2(0.8)^8]=0.322200\ldots, giving 0.32220.3222.
2
  • 0.62930.6293
5
(5 marks)5
Notes
Since X3X\geq3 implies X1X\geq1, the required conditional probability is P(X3)/P(X1)P(X\geq3)/P(X\geq1). Now P(X3)=1P(X2)=0.609324P(X\geq3)=1-P(X\leq2)=0.609324\ldots and P(X1)=1P(X=0)=10.7512=0.968323P(X\geq1)=1-P(X=0)=1-0.75^{12}=0.968323\ldots. Their ratio is 0.6292570.629257\ldots, giving 0.62930.6293.
3
  • p=613p=\dfrac6{13}
  • P(X2)=0.9445P(X\geq2)=0.9445
6
(6 marks)6
Notes
For XB(8,p)X\sim\operatorname{B}(8,p), P(X=2)=28p2(1p)6P(X=2)=28p^2(1-p)^6 and P(X=1)=8p(1p)7P(X=1)=8p(1-p)^7. Since 0<p<10<p<1, division gives 7p2(1p)=3\frac{7p}{2(1-p)}=3, so 7p=6(1p)7p=6(1-p) and p=6/13p=6/13. Therefore P(X2)=1P(X=0)P(X=1)=1(7/13)88(6/13)(7/13)7=0.944473=0.9445P(X\geq2)=1-P(X=0)-P(X=1)=1-(7/13)^8-8(6/13)(7/13)^7=0.944473\ldots=0.9445.
4
  • n=3n=3 and p=23p=\dfrac23
  • P(X2)=2027P(X\geq2)=\dfrac{20}{27}
6
(6 marks)6
Notes
Write q=1pq=1-p. The first ratio gives np/q=6np/q=6. Also P(X=2)/P(X=1)=(n1)p/(2q)=2P(X=2)/P(X=1)=(n-1)p/(2q)=2, so (n1)p/q=4(n-1)p/q=4. Subtracting gives p/q=2p/q=2, hence p=2/3p=2/3 and then n=3n=3. Therefore P(X2)=1P(X=0)P(X=1)=1(1/3)33(2/3)(1/3)2=20/27P(X\geq2)=1-P(X=0)-P(X=1)=1-(1/3)^3-3(2/3)(1/3)^2=20/27.
5
  • TB(15,0.40)T\sim\operatorname{B}(15,0.40)
  • P(T8)=0.2131P(T\geq8)=0.2131
  • If machine B had probability 0.200.20, the 1515 trials would not share one constant success probability, so their total would not have a binomial distribution with one parameter pp.
6
(6 marks)6
Notes
The two independent counts have the same success probability, so pooling their 8+7=158+7=15 trials gives TB(15,0.40)T\sim\operatorname{B}(15,0.40). Thus P(T8)=r=815(15r)(0.4)r(0.6)15r=0.213103=0.2131P(T\geq8)=\sum_{r=8}^{15}\binom{15}{r}(0.4)^r(0.6)^{15-r}=0.213103\ldots=0.2131. Different machine probabilities would violate the constant-pp condition even if all component outcomes remained independent.

S4.2 · Understand and use the Normal distribution as a model; find probabilities using the Normal distribution; link to histograms, mean, standard deviation, points of inflection and the binomial distribution.

Tier 1 · Easy

Mark scheme for S4.2 Tier 1 · Easy
QuestionSchemeMarks
1
  • 0.84130.8413
2
(2 marks)2
Notes
Standardise: z=(5650)/6=1z=(56-50)/6=1. Therefore P(X<56)=P(Z<1)=0.841344P(X<56)=P(Z<1)=0.841344\ldots, which is 0.84130.8413.
2
  • 6363 and 8181
2
(2 marks)2
Notes
A Normal density has points of inflection at μσ\mu-\sigma and μ+σ\mu+\sigma. Here these are 729=6372-9=63 and 72+9=8172+9=81.

Tier 2 · Standard

Mark scheme for S4.2 Tier 2 · Standard
QuestionSchemeMarks
1
  • 0.86640.8664
4
(4 marks)4
Notes
Standardising gives z=(5864)/4=1.5z=(58-64)/4=-1.5 and z=(7064)/4=1.5z=(70-64)/4=1.5. Hence P(58<X<70)=P(1.5<Z<1.5)=Φ(1.5)Φ(1.5)=0.933190.06681=0.8664P(58<X<70)=P(-1.5<Z<1.5)=\Phi(1.5)-\Phi(-1.5)=0.93319\ldots-0.06681\ldots=0.8664.
2
  • μ=48.4\mu=48.4 to 33 significant figures
4
(4 marks)4
Notes
Standardising the tenth percentile gives (42μ)/5=1.2816(42-\mu)/5=-1.2816. Hence 42μ=6.40842-\mu=-6.408, so μ=48.408\mu=48.408, which is 48.448.4 to 33 significant figures.
3
  • μ=50\mu=50
  • σ=9.51\sigma=9.51 to 33 significant figures
5
(5 marks)5
Notes
Standardising the two percentiles gives (42μ)/σ=0.8416(42-\mu)/\sigma=-0.8416 and (58μ)/σ=0.8416(58-\mu)/\sigma=0.8416. Adding the equations, or using symmetry, gives μ=50\mu=50. Then 8/σ=0.84168/\sigma=0.8416, so σ=8/0.8416=9.50570\sigma=8/0.8416=9.50570\ldots, giving 9.519.51 to 33 significant figures.

Tier 3 · Hard

Mark scheme for S4.2 Tier 3 · Hard
QuestionSchemeMarks
1
  • 0.04410.0441
5
(5 marks)5
Notes
Use YN(np,np(1p))=N(70,45.5)Y\sim\operatorname{N}(np,np(1-p))=\operatorname{N}(70,45.5). The continuity correction gives P(X82)P(Y>81.5)P(X\geq82)\approx P(Y>81.5). Thus z=(81.570)/45.5=1.7049z=(81.5-70)/\sqrt{45.5}=1.7049\ldots, so the upper-tail probability is 1Φ(1.7049)=0.04411-\Phi(1.7049\ldots)=0.0441.
2
  • σ=3.12\sigma=3.12
  • P(X>36X>32)=0.1046P(X>36\mid X>32)=0.1046
6
(6 marks)6
Notes
By symmetry, P(X<34)=0.90P(X<34)=0.90, so 4/σ=1.28164/\sigma=1.2816 and σ=3.1211\sigma=3.1211\ldots. Since X>36X>36 implies X>32X>32, the conditional probability is P(X>36)/P(X>32)P(X>36)/P(X>32). The corresponding standardised values are 1.92241.9224 and 0.64080.6408, giving tail probabilities 0.02727770.0272777\ldots and 0.2608260.260826\ldots. Their ratio is 0.1045820.104582\ldots, giving 0.10460.1046.
3
  • μ=500\mu=500 and σ=8\sigma=8
  • P(484X516)=0.9545P(484\leq X\leq516)=0.9545
  • Expected number rejected =55=55
  • P(X>505484X516)=0.2548P(X>505\mid 484\leq X\leq516)=0.2548
7
(7 marks)7
Notes
The inflection points are μσ\mu-\sigma and μ+σ\mu+\sigma, so their midpoint gives μ=500\mu=500 and half their separation gives σ=8\sigma=8. The acceptance bounds standardise to 2-2 and 22, hence the acceptance probability is Φ(2)Φ(2)=0.954499\Phi(2)-\Phi(-2)=0.954499\ldots. The expected rejected count is 1200[10.954499]=54.60031200[1-0.954499\ldots]=54.6003\ldots, giving 5555. Within the acceptance interval, exceeding 505505 means 505<X516505<X\leq516; the lower standardised value is 0.6250.625. Thus the conditional probability is [Φ(2)Φ(0.625)]/[Φ(2)Φ(2)]=0.254830=0.2548[\Phi(2)-\Phi(0.625)]/[\Phi(2)-\Phi(-2)]=0.254830\ldots=0.2548.
4
  • μ=56\mu=56 and σ=7.13\sigma=7.13 to 33 significant figures
  • The points of inflection are at X=48.9X=48.9 and X=63.1X=63.1 to 33 significant figures.
  • Estimated frequency between the points of inflection =614=614 to the nearest whole number.
  • Matching one central histogram area does not check the model's symmetry, unimodality, tails or possible outliers, so the full histogram and context must also be considered.
7
(7 marks)7
Notes
Equal opposite-tail probabilities in a Normal distribution place the mean halfway between 4444 and 6868, so μ=56\mu=56. The interval 50<X<6250<X<62 is therefore μ6<X<μ+6\mu-6<X<\mu+6. Its central probability is 0.600.60, leaving 0.200.20 in each tail, so 6/σ=0.84166/\sigma=0.8416 and σ=6/0.8416=7.12928\sigma=6/0.8416=7.12928\ldots. The points of inflection are μ±σ=48.8707\mu\pm\sigma=48.8707\ldots and 63.129363.1293\ldots. The probability between them is P(1<Z<1)=0.682689P(-1<Z<1)=0.682689\ldots, so the histogram area represents an estimated frequency 900(0.682689)=614.420900(0.682689\ldots)=614.420\ldots, giving 614614 to the nearest whole number. Agreement over only this interval is not a complete model check.
5
  • μ=50.6\mu=50.6 and σ=8.31\sigma=8.31 to 33 significant figures
  • P(45<X<60)=0.6216P(45<X<60)=0.6216
7
(7 marks)7
Notes
Standardising gives (40μ)/σ=1.2816(40-\mu)/\sigma=-1.2816 and (55μ)/σ=0.5244(55-\mu)/\sigma=0.5244. Subtraction gives 15/σ=1.806015/\sigma=1.8060, so σ=8.305647\sigma=8.305647\ldots and μ=40+1.2816σ=50.644518\mu=40+1.2816\sigma=50.644518\ldots. Therefore P(45<X<60)=Φ((60μ)/σ)Φ((45μ)/σ)=0.621622=0.6216P(45<X<60)=\Phi((60-\mu)/\sigma)-\Phi((45-\mu)/\sigma)=0.621622\ldots=0.6216.

S4.3 · Select an appropriate probability distribution for a context, with appropriate reasoning, including recognising when the binomial or Normal model may not be appropriate.

Tier 1 · Easy

Mark scheme for S4.3 Tier 1 · Easy
QuestionSchemeMarks
1
  • FB(12,0.04)F\sim\operatorname{B}(12,0.04)
2
(2 marks)2
Notes
There are 1212 fixed, independent trials, each switch is faulty or not faulty, and the fault probability is constant at 0.040.04. Therefore a binomial model with n=12n=12 and p=0.04p=0.04 is suitable.
2
  • A Normal distribution is symmetric, unlike the strongly right-skewed data.
  • A Normal distribution assigns positive probability to negative waiting times, which are impossible.
2
(2 marks)2
Notes
Compare both the observed shape and the possible values with the features of a Normal distribution. The mismatch in symmetry and support makes the model doubtful.

Tier 2 · Standard

Mark scheme for S4.3 Tier 2 · Standard
QuestionSchemeMarks
1
  • Sampling without replacement from a small batch makes the trials dependent and changes the defect probability after each selection.
  • If the total number of defective components in the batch is known, use a hypergeometric model; equivalently, calculate with conditional probabilities that update after each selection.
  • A binomial model would be an adequate approximation only when the sample size is small relative to the batch size, which is not true here because 20/50=0.420/50=0.4.
3
(3 marks)3
Notes
A binomial model requires a fixed success probability and independent trials. Removing 2020 of only 5050 components changes the composition appreciably, so these conditions fail. If the batch contains a known fixed number of defectives, the exact count distribution is hypergeometric; a probability tree with updated proportions is equivalent. Binomial can approximate sampling without replacement only when the sampling fraction is small, unlike the 40%40\% fraction here.
2
  • Although there is a fixed number of two-outcome trials, the success probability is not constant.
  • Fatigue may also make later outcomes dependent on the earlier workload or sequence.
  • A single binomial model is therefore inappropriate.
  • Use trial-specific or fatigue-stage scoring probabilities estimated from suitable data.
4
(4 marks)4
Notes
A binomial model requires independent trials with one constant probability of success. The stated fatigue mechanism violates the constant-probability condition and may introduce dependence. A refined model should allow the probability to change with shot number or fatigue stage.
3
  • For the number of spam messages, use SB(50,0.08)S\sim\operatorname{B}(50,0.08) because there is a fixed number of independent two-outcome trials with constant probability.
  • A Normal model may be suitable for the download times because the variable is continuous and the supplied histogram is approximately symmetric and unimodal without clear outliers.
4
(4 marks)4
Notes
Match the count to the binomial trial conditions: 5050 fixed selections, spam or not spam, independence and constant probability 0.080.08. The time variable is measured on a continuous scale, and its stated empirical shape matches the main features of a Normal density, so a Normal model is plausible rather than guaranteed.

Tier 3 · Hard

Mark scheme for S4.3 Tier 3 · Hard
QuestionSchemeMarks
1
  • Clustering violates independence between pixel outcomes.
  • Shift-to-shift variation violates the constant-pp assumption.
  • The binomial model is therefore likely to understate variation and tail probabilities.
  • Analyse screens by shift and inspect the empirical count distribution or cluster structure before selecting a model.
5
(5 marks)5
Notes
Although the number of pixels is fixed and each pixel is defective or not, two essential binomial assumptions fail. Neighbouring outcomes are dependent and different shifts have different probabilities. Both mechanisms create extra variation relative to a single binomial distribution, so compare separate-shift data and observed counts with candidate models rather than forcing one common pp.
2
  • A binomial distribution requires a fixed number of trials with two outcomes, but the number of possible visits is not a success count from a fixed number of trials.
  • A Normal distribution is continuous and symmetric with unbounded tails, whereas the observed counts are discrete, non-negative and strongly right-skewed.
  • The business should examine empirical frequencies across comparable hours and check whether conditions such as time of day change the distribution.
  • Any replacement model should be checked by comparing its predicted frequencies or tail probabilities with further observed data.
5
(5 marks)5
Notes
Match each candidate distribution to both the variable type and its generating process. The binomial trial structure is absent, while the Normal shape and support conflict with the data. Stratifying comparable hours and validating predicted against observed frequencies provides evidence for a replacement rather than choosing one only by name.
3
  • The combined bimodal shape is not well represented by one symmetric, single-peaked Normal distribution.
  • Treatment type creates two distinct sections of the population, so combining them hides different centres or spreads.
  • Model the two treatments separately, using a separate Normal distribution for each if its data support that choice.
  • Check each treatment's histogram for approximate symmetry, a single peak and influential outliers.
  • Check that observations are collected comparably and independently within each treatment, then validate each model against further data or predicted frequencies.
5
(5 marks)5
Notes
A mixture of two groups can be bimodal even when each group is individually close to Normal. One fitted Normal would place too much probability between the peaks and misrepresent both groups. Stratifying by the known treatment variable addresses the generating mechanism; the shape, support, independence and predictive fit of each proposed group model must still be checked.
4
  • XB(400,0.50)X\sim\operatorname{B}(400,0.50) exactly.
  • np=n(1p)=200np=n(1-p)=200, so a Normal approximation is reasonable: YN(200,100)Y\sim\operatorname{N}(200,100).
  • P(190X210)P(189.5<Y<210.5)=0.7063P(190\leq X\leq210)\approx P(189.5<Y<210.5)=0.7063.
  • A change causing fault probabilities to vary between items, or faults to occur in dependent clusters, would violate the binomial assumptions and invalidate the stated approximation.
6
(6 marks)6
Notes
The fixed number of independent two-outcome trials with constant pp gives the exact binomial model. Its mean and variance are 200200 and 100100, and both expected outcome counts are large. Applying continuity correction gives standardised bounds 1.05-1.05 and 1.051.05, so the estimate is Φ(1.05)Φ(1.05)=0.706282=0.7063\Phi(1.05)-\Phi(-1.05)=0.706282\ldots=0.7063.
5
  • P(T<0)=P(Z<2.1)=0.0179P(T<0)=P(Z<-2.1)=0.0179.
  • Expected number below zero =1800(0.017864)=32=1800(0.017864\ldots)=32 to the nearest integer.
  • Assigning about 1.8%1.8\% probability to impossible values is a material support mismatch, so the Normal model is doubtful unless the context or measurements have been misunderstood.
  • Inspect the empirical distribution for skewness, bounds, outliers and changes between service conditions, then compare candidate models with fresh observed frequencies or tail probabilities.
5
(5 marks)5
Notes
Standardising zero gives z=(05.25)/2.5=2.1z=(0-5.25)/2.5=-2.1, so the lower-tail probability is 0.0178640.017864\ldots. Multiplying by 18001800 gives 32.155932.1559\ldots, which rounds to 3232. Model selection must consider possible values as well as centre and spread, and should be validated against observed data rather than chosen from parameters alone.