S Statistics — revision question pack

6 specification points · notes, questions, answers and worked methods

Checked against Edexcel 1MA1 section S. Review basis: the qualification registry sourced from the Pearson Edexcel Level 1/Level 2 GCSE (9-1) in Mathematics (1MA1) specification; registry verification recorded 9 July 2026.

How this checking works

Loading your tier…

Answer ALL questions.

Write your answers in the spaces provided.

You must write down all the stages in your working.

S1 · Infer properties of populations or distributions from a sample, while knowing the limitations of sampling

Explanation

  • A population is the whole group being studied; a sample is the smaller group from which data are collected.
  • If the sample is representative, its proportion or mean can be used to estimate a population value.
  • Scale a sample proportion by the population size when estimating a count.
  • The result is an estimate, not a certainty.
  • Small samples, convenience sampling, under-coverage and non-response can make the sample unrepresentative and the inference unreliable.

Worked example

In a random sample of 7575 residents, 2727 support a proposal. Estimate how many of the town's 12501250 residents support it and state the assumption needed.

  1. 1.Sample proportion =2775=0.36=\dfrac{27}{75}=0.36.
  2. 2.Estimated population count =0.36×1250=450=0.36\times1250=450.
  3. 3.The estimate assumes that the sample is representative of the town's residents.

Answer: About 450450 residents, provided the sample is representative.

Common mistakes

  • Don't fall into the trap of using the sample frequency as the population estimate without scaling.
  • Don't fall into the trap of assuming a large but biased sample must be representative.

Exam tip

State a limitation in context, explaining how it could make the sample unrepresentative.

Tier 1 · Easy

  1. 1

    In a random sample of 5050 library users, 3030 prefer later opening. Estimate how many of the library's 10001000 users prefer later opening.

    (2)

    (Total for Question 1 is 2 marks)

  2. 2

    A council wants the views of all 60006000 residents about a cycle lane and surveys 8080 visitors to a sports centre. Write down the population and give one reason why the sample may be unrepresentative.

    (2)

    (Total for Question 2 is 2 marks)

Tier 2 · Standard

  1. 1

    A random sample of 8080 items from a production run of 800800 contains 1818 items with a surface mark. Estimate the number in the whole run with a surface mark and state one limitation of the estimate.

    (3)

    (Total for Question 1 is 3 marks)

  2. 2

    Two random samples estimate support for a new crossing. In a sample of 2525, 1414 people support it. In a sample of 200200, 104104 people support it. Use the more reliable sample to estimate how many of 30003000 residents support the crossing. Give a reason for your choice of sample.

    (3)

    (Total for Question 2 is 3 marks)

  3. 3

    A recycling team uses a sample of 160160 households to estimate that 10501050 of the district's 30003000 households use a food-waste collection. Work out how many households in the sample used the collection. Give one reason why the estimate of 10501050 may not equal the true district total.

    (3)

    (Total for Question 3 is 3 marks)

Tier 3 · Hard

  1. 1

    A service invites a random sample of 450450 customers to answer a survey. Only 270270 reply, and 189189 of the replies support a change. Use the replies to estimate the number of supporters among all 1200012000 customers, then explain a serious limitation.

    (4)

    (Total for Question 1 is 4 marks)

  2. 2

    In a representative sample of 120120 households, 7878 own at least one pet. In a separate representative sample of 9090 pet-owning households, 5454 own a dog. Use both samples to estimate the percentage of all households that own a dog. Explain why using 54/9054/90 as the estimate would be wrong.

    (4)

    (Total for Question 2 is 4 marks)

  3. 3

    A club emails 320320 randomly selected members. Of the 224224 who reply, 9898 intend to renew their membership. Without making an assumption about the members who did not reply, work out the least and greatest possible percentages of the 320320 sampled members who intend to renew. Explain why using 98/22498/224 to describe all club members may be unreliable.

    (4)

    (Total for Question 3 is 4 marks)

  4. 4

    A coach operator surveys 150150 passengers who booked their journeys using its app. Of these passengers, 9696 say the service was satisfactory. Use the sample to estimate how many of the operator's 32003200 weekly passengers are satisfied. Explain why this sampling method could make the estimate unreliable.

    (4)

    (Total for Question 4 is 4 marks)

  5. 5

    A leisure centre wants the views of all its adult members. Method A uses a random-number generator to select 120120 members from the complete membership list and arranges a time for each selected member to answer. Method B telephones members between 1010 am and 22 pm on a weekday until 120120 members have answered. State which method is more likely to give a representative sample. Explain why, referring to a group Method B is likely to exclude.

    (3)

    (Total for Question 5 is 3 marks)

S2 · Interpret and construct frequency tables, bar charts, pie charts and pictograms (categorical data), vertical line charts (ungrouped discrete numerical data), tables and line graphs (time series)

Explanation

  • Match the display to the data. Use separated bars for categories, a vertical line chart for ungrouped discrete numbers, and joined points in time order for a time series.
  • A pictogram needs a clear key.
  • For a pie chart, sector angle=frequencytotal×360\text{sector angle}=\dfrac{\text{frequency}}{\text{total}}\times360^\circ.
  • Any graph needs clear labels, an even scale and accurate plotting.
  • Bar widths and gaps should be consistent, and unrelated categories should not be joined by lines.

Worked example

A 126126^\circ sector of a pie chart represents 4242 students. Find the total number of students.

  1. 1.The sector represents the fraction 126360\dfrac{126}{360} of the total.
  2. 2.Total =42÷126360=42×360126=120=42\div\dfrac{126}{360}=42\times\dfrac{360}{126}=120.

Answer: 120120 students.

Common mistakes

  • Don't fall into the trap of using frequency directly as an angle instead of scaling to 360360^\circ.
  • Don't fall into the trap of joining categorical bars or plotting a time series out of order.

Exam tip

After calculating pie-chart sectors, check that all sector angles add to 360360^\circ.

Tier 1 · Easy

  1. 1

    The data are 2,3,2,5,4,3,2,4,5,22,3,2,5,4,3,2,4,5,2. Construct a frequency table for the values 22, 33, 44 and 55.

    (2)

    (Total for Question 1 is 2 marks)

  2. 2

    A vertical line chart shows the number of goals scored by a team in each match. The lines at 00, 11, 22 and 33 goals have frequencies 33, 88, 66 and 44 respectively. Write down the frequency for 22 goals and the modal number of goals.

    (2)

    (Total for Question 2 is 2 marks)

Tier 2 · Standard

  1. 1

    In a survey of 240240 journeys, 5454 are made by bicycle. Work out the angle of the bicycle sector in a pie chart.

    (2)

    (Total for Question 1 is 2 marks)

  2. 2

    A café records the number of hot drinks sold on four consecutive Mondays as 4242, 4646, 4444 and 5151. Write down the four points for a time-series graph, using Monday number on the horizontal axis, and describe the overall trend.

    (2)

    (Total for Question 2 is 2 marks)

  3. 3

    A survey of 4848 pupils records their favourite after-school activity. Sport has frequency 1414, music has frequency 1010 and drama has frequency 88. The remaining pupils choose art. A bar chart uses a scale of 1cm1\,\text{cm} for 22 pupils. Work out the frequency for art, write down the height of each of the four bars, and state whether the bars should touch.

    (4)

    (Total for Question 3 is 4 marks)

Tier 3 · Hard

  1. 1

    Quarterly sales are Q1: 320320, Q2: 410410, Q3: 390390, Q4: 520520, and the following Q1: 350350. State the five points to plot on a time-series line graph. Work out the percentage change from the first Q1 to the following Q1 and explain why comparing Q4 directly with the following Q1 could mislead.

    (4)

    (Total for Question 1 is 4 marks)

  2. 2

    A pie chart represents 300300 votes. Three sectors have angles 7272^\circ, 108108^\circ and 126126^\circ. Work out the angle of the fourth sector and the frequency represented by each of the four sectors.

    (4)

    (Total for Question 2 is 4 marks)

  3. 3

    A bar chart compares 6464 bookings on Monday with 7676 bookings on Tuesday. The vertical axis begins at 6060, so the visible parts of the bars have heights of 44 units and 1616 units. Sam says, ‘Tuesday had four times as many bookings as Monday.’ Is Sam correct? Work out the percentage increase from Monday to Tuesday and explain how the axis makes the chart misleading.

    (4)

    (Total for Question 3 is 4 marks)

  4. 4

    A time-series graph uses the number of months after January 2024 on the horizontal axis and the meter reading on the vertical axis. The readings are 120120 in January 2024, 150150 in April 2024, 132132 in October 2024 and 168168 in July 2025. Write down the four coordinates that should be plotted, in date order.

    (4)

    (Total for Question 4 is 4 marks)

  5. 5

    A pie chart represents the after-school clubs chosen by 240240 pupils. The school has exactly these four clubs: coding, drama, art and music. The coding sector has angle 7272^\circ and represents 4848 pupils. There are 3232 pupils who choose drama. The numbers who choose art and music are in the ratio 5:35 : 3. Work out the angles of the art, music and drama sectors.

    (4)

    (Total for Question 5 is 4 marks)

S3 · Construct and interpret diagrams for grouped discrete and continuous data, i.e. histograms with equal and unequal class intervals and cumulative frequency graphs [Higher only]

Explanation

  • Higher tier only. In a histogram, bar area represents frequency and bar height is frequency density: frequency density=frequencyclass width\text{frequency density}=\dfrac{\text{frequency}}{\text{class width}}.
  • Therefore $\text{frequency}=\text{frequency density}\times\text{class width}$.
  • Use the class boundaries for each width, draw touching bars, and label both axes.
  • For a cumulative frequency graph, plot the lower boundary with cumulative frequency 00, then each upper class boundary against its running total.
  • Estimate the median at n/2n/2, the lower quartile at n/4n/4 and the upper quartile at 3n/43n/4.

Worked example

The classes 10<x2510<x\leq25 and 25<x3525<x\leq35 are adjacent. The first has frequency 3030; the second has histogram height 1.41.4. Find the first bar height, the second frequency and the cumulative-frequency points.

  1. 1.First width =15=15, so its height is 30÷15=230\div15=2.
  2. 2.Second width =10=10, so its frequency is 1.4×10=141.4\times10=14.
  3. 3.The cumulative-frequency points are (10,0)(10,0), (25,30)(25,30) and (35,44)(35,44).

Answer: First height 22; second frequency 1414; cumulative points (10,0)(10,0), (25,30)(25,30) and (35,44)(35,44).

Common mistakes

  • Don't fall into the trap of using frequency as the bar height instead of dividing by class width.
  • Don't fall into the trap of plotting cumulative frequencies at class midpoints instead of upper boundaries.

Exam tip

Write the frequency-density formula beside the histogram before calculating any height or area.

Tier 1 · Easy

  1. 1

    A histogram class is 15<x2015<x\leq20 and has frequency 2020. Work out the frequency density.

    (2)

    (Total for Question 1 is 2 marks)

  2. 2

    A histogram class has width 88 and frequency density 2.52.5. Work out the frequency in the class.

    (1)

    (Total for Question 2 is 1 mark)

Tier 2 · Standard

  1. 1

    Grouped data have classes 0<x100<x\leq10, 10<x2510<x\leq25 and 25<x4025<x\leq40 with frequencies 1212, 3030 and 1818. Calculate the three frequency densities and identify the tallest histogram bar.

    (3)

    (Total for Question 1 is 3 marks)

  2. 2

    In a histogram, the class 5<x155<x\leq15 has frequency 2424 and its bar is 6cm6\,\text{cm} high. The bar for 15<x2115<x\leq21 is 7.5cm7.5\,\text{cm} high. Work out the frequency in the second class.

    (3)

    (Total for Question 2 is 3 marks)

  3. 3

    Three adjacent classes in a histogram have class widths 44, 66 and 1010. Their bars measure 6cm6\,\text{cm}, 5cm5\,\text{cm} and 3cm3\,\text{cm} high respectively. There are 8484 values altogether and there are no other classes. Work out the frequency in each class.

    (4)

    (Total for Question 3 is 4 marks)

Tier 3 · Hard

  1. 1

    A cumulative frequency graph for 8080 values passes through (10,8)(10,8), (20,26)(20,26), (30,50)(30,50), (40,70)(40,70) and (50,80)(50,80). Use linear interpolation between the given points to estimate the median and the interquartile range.

    (5)

    (Total for Question 1 is 5 marks)

  2. 2

    The class 12<xb12<x\leq b has frequency 3232 and frequency density 22. The next class is b<x36b<x\leq36 and has frequency 1818. The cumulative frequency at x=12x=12 is 1515. Work out bb, the cumulative frequencies at x=bx=b and x=36x=36, and the frequency density of the second class.

    (5)

    (Total for Question 2 is 5 marks)

  3. 3

    For 5050 values, a student plots the cumulative-frequency points (0,0)(0,0), (10,8)(10,8), (20,23)(20,23), (30,19)(30,19), (40,42)(40,42) and (50,50)(50,50). Explain why the point (30,19)(30,19) cannot be correct. The frequency in the class 20<x3020<x\leq30 is 1212. Work out the correct point at x=30x=30, the frequency in the class 30<x4030<x\leq40, and the class containing the median.

    (4)

    (Total for Question 3 is 4 marks)

  4. 4

    The histogram class 4<x164<x\leq16 has frequency 3636. The adjacent class 16<xb16<x\leq b has the same bar area and a bar twice as tall. The class b<x34b<x\leq34 has a bar one third as tall as the second bar. The bars use the same vertical scale. Work out bb and the total frequency in the three classes.

    (5)

    (Total for Question 4 is 5 marks)

  5. 5

    Higher only: A histogram has classes 0<x80<x\leq8, 8<x208<x\leq20, 20<x3220<x\leq32 and 32<x4232<x\leq42. Their frequency densities are 1.51.5, 2.52.5, 33 and 1.21.2 respectively. Assuming values are evenly spread within each class, estimate how many values are greater than 2626. Also estimate the value below which 70%70\% of the data lie.

    (5)

    (Total for Question 5 is 5 marks)

S4 · Interpret, analyse and compare data-set distributions via graphs (incl. box plots), central tendency (median, mean, mode, modal class) and spread (range, outliers, quartiles, inter-quartile range)

Explanation

  • A measure of central tendency describes a typical value: the mean uses every value, the median is the ordered middle, and the mode or modal class is most frequent.
  • Spread describes variation: range is maximum minus minimum.
  • At Higher tier, box plots also show quartiles and the interquartile range Q3Q1Q_3-Q_1.
  • To compare distributions, make one contextual statement about centre and one about spread.
  • Outliers can strongly affect the mean and range, so the median and interquartile range may be more representative.

Worked example

Delivery service A has median time 3838 minutes and range 2222 minutes. Service B has median time 4343 minutes and range 1212 minutes. Compare the distributions.

  1. 1.A has the lower median, so its deliveries are typically quicker.
  2. 2.B has the smaller range, so its delivery times are more consistent.

Answer: Service A is typically quicker, but service B has more consistent delivery times.

Common mistakes

  • Don't fall into the trap of giving comparison figures without stating what they mean in context.
  • Don't fall into the trap of finding the median before putting raw data in order.

Exam tip

A full comparison usually needs both a typical-value statement and a spread statement.

Tier 1 · Easy

  1. 1

    For the data 4,6,6,9,104,6,6,9,10, work out the mean and the range.

    (2)

    (Total for Question 1 is 2 marks)

  2. 2

    For the data 2,4,4,6,9,10,122,4,4,6,9,10,12, write down the mode and the median.

    (2)

    (Total for Question 2 is 2 marks)

Tier 2 · Standard

  1. 1

    Data set A has median 2424 and range 3333. Data set B has median 2525 and range 2929. Compare the two distributions.

    (2)

    (Total for Question 1 is 2 marks)

  2. 2

    Five readings have mean 1818. A sixth reading of 2424 is added. Work out the new mean and explain why it is greater than 1818.

    (3)

    (Total for Question 2 is 3 marks)

  3. 3

    A data set has mean 1212, median 1111 and range 99. Every value in the data set is doubled and then 33 is added. Work out the mean, median and range of the new data set.

    (3)

    (Total for Question 3 is 3 marks)

Tier 3 · Hard

  1. 1

    The data are 12,13,13,14,15,15,16,5212,13,13,14,15,15,16,52. Work out the mean, median and range. Decide which of the mean or median better describes a typical value for these data. Give a reason.

    (4)

    (Total for Question 1 is 4 marks)

  2. 2

    The ordered data are 6,9,11,x,18,206,9,11,x,18,20. Their mean is 1313. Work out xx, the median and the range.

    (4)

    (Total for Question 2 is 4 marks)

  3. 3

    Seven integer values are written in order of size: 3,7,a,11,b,16,203,7,a,11,b,16,20. Their mean is 1111 and their only mode is 77. Work out aa, bb and the range.

    (4)

    (Total for Question 3 is 4 marks)

  4. 4

    The values 44, 77, 1010 and 1313 have frequencies 22, 55, kk and 11 respectively. The mean is 88. Work out kk, the median, the mode and the range.

    (4)

    (Total for Question 4 is 4 marks)

  5. 5

    Higher only: Nine values are written in order: 4,7,9,12,15,18,x,26,314,7,9,12,15,18,x,26,31. The upper quartile is 2424. The median is excluded when the lower and upper halves are formed. Work out xx, the lower quartile and the interquartile range.

    (4)

    (Total for Question 5 is 4 marks)

S5 · Apply statistics to describe a population

Explanation

  • Statistics can summarise a population or estimate its features from representative samples.
  • When groups have different sizes, combine their means using a weighted mean: add each group size multiplied by its mean, then divide by the total population.
  • When different subgroups have different sample proportions, estimate each subgroup separately before adding.
  • A simple average of group means is only valid when the groups are the same size.

Worked example

A town has 900900 northern and 600600 southern residents. In representative samples, 1212 of 4040 northern residents and 2020 of 5050 southern residents cycle to work. Estimate the town total and percentage.

  1. 1.North estimate =900×1240=270=900\times\dfrac{12}{40}=270.
  2. 2.South estimate =600×2050=240=600\times\dfrac{20}{50}=240.
  3. 3.Total =270+240=510=270+240=510; percentage =5101500×100=34%=\dfrac{510}{1500}\times100=34\%.

Answer: About 510510 residents, or 34%34\% of the town.

Common mistakes

  • Don't fall into the trap of averaging two group means without using the group sizes.
  • Don't fall into the trap of applying one subgroup's sample proportion to the whole population.

Exam tip

Turn each mean back into a total first; combine totals, then divide once.

Tier 1 · Easy

  1. 1

    A representative sample of parcels has mean mass 2.4kg2.4\,\text{kg}. Estimate the total mass of 250250 parcels in the population.

    (2)

    (Total for Question 1 is 2 marks)

  2. 2

    The mean weekly electricity use for a representative sample of 6464 households is 118kWh118\,\text{kWh}. Estimate the mean weekly electricity use for all households in the village.

    (1)

    (Total for Question 2 is 1 mark)

Tier 2 · Standard

  1. 1

    A population contains 120120 junior members with mean attendance 6.56.5 sessions and 8080 senior members with mean attendance 88 sessions. Work out the mean attendance for the whole population.

    (3)

    (Total for Question 1 is 3 marks)

  2. 2

    A company has 1616 part-time employees whose mean working time is 1515 hours per week. The full-time employees have a mean working time of 3939 hours per week. The mean working time for all the employees is 2323 hours per week. Work out the number of full-time employees.

    (3)

    (Total for Question 2 is 3 marks)

  3. 3

    A school has 24002400 pupils. The mean number of days absent last term was 1.51.5 per pupil. Of the pupils, 25%25\% had no absence. Work out the mean number of days absent for the pupils who had at least one day absent.

    (3)

    (Total for Question 3 is 3 marks)

Tier 3 · Hard

  1. 1

    A town has 18001800 residents in the north and 12001200 in the south. In representative samples, 1515 of 6060 northern residents and 2020 of 5050 southern residents cycle to work. Estimate the total number and percentage of the town's residents who cycle to work.

    (5)

    (Total for Question 1 is 5 marks)

  2. 2

    A company has 280280 office employees and 420420 warehouse employees. Overall, 32%32\% of the employees work part-time. Of the office employees, 20%20\% work part-time. Work out the percentage of the warehouse employees who work part-time.

    (4)

    (Total for Question 2 is 4 marks)

  3. 3

    A biologist records the masses of fish from a lake in two random samples. Sample A has 4040 fish with a mean mass of 620620 g. Sample B has 6060 fish with a mean mass of 570570 g. Work out the mean mass of all 100100 sampled fish. Give one reason why this combined mean is a better estimate of the mean mass of fish in the lake than the mean of sample A alone. You must show all your working.

    (4)

    (Total for Question 3 is 4 marks)

  4. 4

    A survey estimates the number of oak trees affected by a disease in three woods. Wood A has 900900 oak trees; 2121 of a random sample of 3030 are affected. Wood B has 600600 oak trees; 1616 of a random sample of 2525 are affected. Wood C has 500500 oak trees; 1414 of a random sample of nn are affected. Using these samples gives an estimated total of 12641264 affected oak trees. Work out nn. State which wood has the lowest estimated proportion of affected oak trees.

    (4)

    (Total for Question 4 is 4 marks)

  5. 5

    A representative random sample of 8080 one-square-metre plots has a mean of 6.46.4 seedlings per plot. Use the sample to estimate the number of seedlings in a habitat of area 1250m21250\,\text{m}^2. A later census finds 76807680 seedlings. Work out the percentage error in the estimate as a percentage of the census figure, and state whether the estimate is too high or too low. Write the percentage to two decimal places.

    (4)

    (Total for Question 5 is 4 marks)

S6 · Use and interpret scatter graphs of bivariate data; recognise correlation, know it does not indicate causation; draw estimated lines of best fit; make predictions; interpolate/extrapolate with caution

Explanation

  • A scatter graph displays paired values.
  • An upward pattern shows positive correlation, a downward pattern negative correlation, and no clear pattern no correlation.
  • Draw a line of best fit through the centre of the points with a roughly balanced spread above and below.
  • Use it for estimates: interpolation stays within the observed data range, while extrapolation goes beyond it and is less reliable.
  • Correlation does not prove causation because another variable, reverse causation or coincidence may explain the relationship.
An illustrative scatter graph with an upward trend and a straight line of best fit passing through the centre of the points.

Worked example

A scatter graph compares weekly revision time with test score for revision times from 11 to 88 hours. A sensible line of best fit gives a score of about 7272 at 66 hours and about 8484 when extended to 1010 hours. Interpret both estimates.

  1. 1.66 hours is inside the observed range, so 7272 is an interpolation and is reasonably reliable.
  2. 2.1010 hours is outside the observed range, so 8484 is an extrapolation and is less reliable because the trend may not continue.
  3. 3.The positive correlation does not prove that extra revision alone caused the higher scores; prior attainment could affect both variables.

Answer: 7272 is the more reliable interpolation; 8484 is a less reliable extrapolation, and the graph does not establish causation.

Common mistakes

  • Don't fall into the trap of saying that correlation proves one variable causes the other.
  • Don't fall into the trap of using a line of best fit far outside the observed range without warning.

Exam tip

Answer the command precisely: for 'describe the relationship', write 'as xx increases, yy tends to increase/decrease', not just 'positive/negative'. For a prediction, identify interpolation or extrapolation and comment on reliability.

Tier 1 · Easy

  1. 1

    A scatter graph shows that as daily temperature increases, ice-cream sales usually increase. State the type of correlation and explain why the graph alone does not prove that temperature is the only cause of higher sales.

    (2)

    (Total for Question 1 is 2 marks)

  2. 2

    A scatter graph compares altitude with air temperature and shows a downward pattern. Write down the type of correlation. Give one reason why using the graph to estimate the air temperature at an altitude beyond the plotted points would be unreliable.

    (2)

    (Total for Question 2 is 2 marks)

Tier 2 · Standard

  1. 1

    For data with observed xx-values from 55 to 2020, a line of best fit is y=1.8x+6y=1.8x+6. Estimate yy when x=14x=14. Explain why using the line at x=35x=35 is less reliable.

    (3)

    (Total for Question 1 is 3 marks)

  2. 2

    A scatter graph contains 1212 points. A student's line of best fit runs across the full plotted range, but it is made from two straight sections joined at a corner and 1010 of the 1212 points lie above it. Write down two errors in the student's line of best fit.

    (2)

    (Total for Question 2 is 2 marks)

  3. 3

    A scatter graph uses wing length in centimetres on the horizontal axis and mass in grams on the vertical axis. One bird has wing length 180mm180\,\text{mm} and mass 75g75\,\text{g}. The line of best fit is y=3x+20y=3x+20. Write down the coordinates of this bird on the graph. Work out the value predicted by the line and how far the recorded mass is above or below that value.

    (3)

    (Total for Question 3 is 3 marks)

Tier 3 · Hard

  1. 1

    A scatter graph compares a puppy's age aa months with mass mm kg for ages from 22 to 1212 months. Its estimated line of best fit is m=0.42a+1.8m=0.42a+1.8. Estimate the mass at 99 months and at 1616 months. Comment on the reliability of both estimates and on whether the graph proves that age alone causes the change in mass.

    (5)

    (Total for Question 1 is 5 marks)

  2. 2

    For primary-school pupils, a scatter graph of shoe length against reading score shows strong positive correlation. Explain why this does not show that longer shoes cause better reading, suggest a variable that could affect both measurements, and explain why a prediction for an adult would be unreliable.

    (3)

    (Total for Question 2 is 3 marks)

  3. 3

    A scatter graph compares time since charging, xx hours, with battery level, y%y\%, for times from 44 to 1818 hours. An estimated line of best fit has gradient 3-3 and passes through (10,52)(10,52). Work out the equation of the line in the form y=mx+cy=mx+c. Use the line to estimate the time when the battery level is 40%40\%. State whether this estimate is interpolation or extrapolation.

    (4)

    (Total for Question 3 is 4 marks)

  4. 4

    For data with observed xx-values from 22 to 1414, a line of best fit is y=4.5x+12y=4.5x+12. A point has coordinates (8,54)(8,54). Work out how many units above or below the line's prediction this point lies. Another point has x=13x=13 and lies 1818 units below the line's prediction. Work out its observed yy-value. State which point lies farther from the line.

    (4)

    (Total for Question 4 is 4 marks)

  5. 5

    Two scatter graphs use the same observed xx-range, from 33 to 1212. Their estimated lines of best fit are y=5x+19y=5x+19 for group A and y=3x+35y=3x+35 for group B. Work out the value of xx at which the lines predict the same value of yy, and find this value of yy. At x=10x=10, state which group has the greater predicted value and by how much. State whether the prediction at x=10x=10 is interpolation or extrapolation.

    (4)

    (Total for Question 5 is 4 marks)

Answer key

Answers begin on a new printed page so the question pack can be completed without the solutions alongside it.

S1 · Infer properties of populations or distributions from a sample, while knowing the limitations of sampling

Tier 1 · Easy

Mark scheme for S1 Tier 1 · Easy
QuestionAnswerMarkMark scheme
1
  • 600600 users
2The sample proportion is 30/50=0.630/50=0.6. Apply this to the population: 0.6×1000=6000.6\times1000=600.
2
  • The population is all 60006000 residents.
  • Visitors to the sports centre may have different views from residents who do not use it (accept any valid reason that the sports-centre sample may be biased).
2The group whose views the council wants is all 60006000 residents, so that is the population. The sample comes from only one location and may over-represent people who use the sports centre, so it may not represent all residents.

Tier 2 · Standard

Mark scheme for S1 Tier 2 · Standard
QuestionAnswerMarkMark scheme
1
  • 180180 items.
  • The sample may differ from the population by chance, so the estimate need not equal the true number.
3The sample is one tenth of the production run, so scale 1818 by 1010 to get 180180. Because only a sample was inspected, sampling variation remains even if the selection was random.
2
  • 15601560 residents.
  • The sample of 200200 is more reliable because a larger random sample is likely to have less sampling variation.
3Use the larger random sample. Its support proportion is 104200=0.52\dfrac{104}{200}=0.52, so the estimate is 0.52×3000=15600.52\times3000=1560. A larger random sample is generally less affected by chance variation than a sample of 2525.
3
  • 5656 households.
  • Households that use the food-waste collection may be more or less likely to appear in the sample — for example, uptake may vary between streets or seasons — so the sample proportion may differ from the district's (accept any valid sampling limitation).
3The estimated population proportion is 1050/3000=0.351050/3000=0.35. Apply this proportion to the sample: 0.35×160=560.35\times160=56. Even a representative sample gives an estimate, because another sample could contain a different proportion.

Tier 3 · Hard

Mark scheme for S1 Tier 3 · Hard
QuestionAnswerMarkMark scheme
1
  • Estimated supporters =8400=8400.
  • Non-response bias may make the replies unrepresentative because the 180180 non-responders may have different views.
4Among replies, the support proportion is 189/270=0.7189/270=0.7, giving 0.7×12000=84000.7\times12000=8400. However, the estimate assumes responders and non-responders have similar opinions; the low response rate may break that assumption.
2
  • Estimated percentage =39%=39\%.
  • 54/9054/90 estimates the proportion of pet-owning households that own a dog, not the proportion of all households that own a dog.
4Estimate the proportion of households that own a pet as 78/120=0.6578/120=0.65. Estimate the proportion of pet-owning households that own a dog as 54/90=0.654/90=0.6. Therefore the estimated proportion of all households that own a dog is 0.65×0.6=0.390.65\times0.6=0.39, or 39%39\%. The second sample alone has pet-owning households as its population, so its 60%60\% cannot be applied to all households.
3
  • Least possible percentage =30.625%=30.625\% (or 30.6%30.6\% to 1 d.p.); greatest possible percentage =60.625%=60.625\% (or 60.6%60.6\% to 1 d.p.).
  • The members who replied may have different intentions from those who did not reply, so the replies may be unrepresentative of all club members.
4There are 320224=96320-224=96 non-responders. The least possible number intending to renew is 9898, giving 98/320×100=30.625%98/320\times100=30.625\%. The greatest is 98+96=19498+96=194, giving 194/320×100=60.625%194/320\times100=60.625\%. This wide range shows the possible effect of non-response bias.
4
  • Estimated number satisfied =2048=2048.
  • Passengers who book at a ticket counter or by telephone are excluded, and their experience or opinion of the service may differ from that of app users.
4The sample proportion satisfied is 96/150=0.6496/150=0.64, so the estimate is 0.64×3200=20480.64\times3200=2048. The sampling frame contains only app users and excludes passengers using other booking methods, so it may not represent all weekly passengers.
5
  • Method A is more likely to give a representative sample.
  • Method B is likely to exclude members who are at work, at college or otherwise unavailable during weekday daytime; their views may differ from those of members who are available then.
3Choose Method A first. Selecting at random from the complete membership list gives every adult member a chance of selection. Method B favours members who are available during weekday daytime and can omit working members, students and others who are away then, so it is more likely to be biased.

S2 · Interpret and construct frequency tables, bar charts, pie charts and pictograms (categorical data), vertical line charts (ungrouped discrete numerical data), tables and line graphs (time series)

Tier 1 · Easy

Mark scheme for S2 Tier 1 · Easy
QuestionAnswerMarkMark scheme
1
  • 2:42:4, 3:23:2, 4:24:2, 5:25:2.
2Tally each value once. The frequencies 4+2+2+2=104+2+2+2=10 match the number of data values.
2
  • The frequency for 22 goals is 66.
  • The modal number of goals is 11.
2Read the height of the line at 22 goals to get frequency 66. The greatest frequency is 88, at 11 goal, so the mode is 11.

Tier 2 · Standard

Mark scheme for S2 Tier 2 · Standard
QuestionAnswerMarkMark scheme
1
  • 8181^\circ
2The bicycle fraction is 54/24054/240. Multiply by a full turn: 54240×360=81\dfrac{54}{240}\times360^\circ=81^\circ.
2
  • (1,42)(1,42), (2,46)(2,46), (3,44)(3,44) and (4,51)(4,51).
  • The number of hot drinks sold shows an overall increase (also accept noting the fall from 4646 to 4444 between Mondays 22 and 33).
2Pair each Monday number with its sales value and plot the points in time order. Although sales fall from 4646 to 4444 between Mondays 22 and 33, they rise from 4242 to 5151 overall.
3
  • Art frequency =16=16.
  • Bar heights: sport 7cm7\,\text{cm}, music 5cm5\,\text{cm}, drama 4cm4\,\text{cm} and art 8cm8\,\text{cm}.
  • The bars should not touch because the activities are categories.
4The art frequency is 48(14+10+8)=1648-(14+10+8)=16. Divide each frequency by 22 to use the stated scale, giving heights 77, 55, 44 and 8cm8\,\text{cm}. A categorical bar chart has gaps between its bars.

Tier 3 · Hard

Mark scheme for S2 Tier 3 · Hard
QuestionAnswerMarkMark scheme
1
  • (Q1,320),(Q2,410),(Q3,390),(Q4,520),(next Q1,350)(\text{Q1},320),(\text{Q2},410),(\text{Q3},390),(\text{Q4},520),(\text{next Q1},350).
  • Percentage increase =9.375%=9.375\%.
  • Q4 and Q1 may have different seasonal patterns, so that comparison may not represent an underlying fall.
4Plot the values in time order and join consecutive points. Like-for-like Q1 change is 350320=30350-320=30, so the percentage change is 30/320×100=9.375%30/320\times100=9.375\%. Comparing different quarters confounds the change with possible seasonality.
2
  • Fourth angle =54=54^\circ.
  • The frequencies are 6060, 9090, 105105 and 4545, corresponding respectively to 7272^\circ, 108108^\circ, 126126^\circ and 5454^\circ.
4The missing angle is 360(72+108+126)=54360^\circ-(72^\circ+108^\circ+126^\circ)=54^\circ. Since 300/360=56300/360=\dfrac{5}{6}, multiply each angle by 56\dfrac{5}{6}: the frequencies are 6060, 9090, 105105 and 4545. They add to 300300.
3
  • No, Tuesday did not have four times as many bookings as Monday.
  • The percentage increase is 18.75%18.75\%.
  • Starting the axis at 6060 makes the visible heights 44 and 1616, exaggerating the difference; the actual values are 6464 and 7676.
4The increase is 7664=1276-64=12, so the percentage increase is 12/64×100=18.75%12/64\times100=18.75\%. The ratio of the actual values is 76/64=1.187576/64=1.1875, not 44. The fourfold visual ratio comes only from subtracting the truncated-axis baseline of 6060.
4
  • The coordinates are (0,120)(0,120), (3,150)(3,150), (9,132)(9,132) and (18,168)(18,168), in date order.
4The dates are 00, 33, 99 and 1818 months after January 2024. Pair each of these data values with its meter reading to give (0,120)(0,120), (3,150)(3,150), (9,132)(9,132) and (18,168)(18,168).
5
  • Art sector =150=150^\circ, music sector =90=90^\circ and drama sector =48=48^\circ.
4After the 4848 coding pupils and 3232 drama pupils, 2404832=160240-48-32=160 pupils remain. Split 160160 in the ratio 5:35 : 3 to get 100100 art pupils and 6060 music pupils. Each pupil represents 360/240=1.5360^\circ/240=1.5^\circ, so the three required angles are 150150^\circ, 9090^\circ and 4848^\circ.

S3 · Construct and interpret diagrams for grouped discrete and continuous data, i.e. histograms with equal and unequal class intervals and cumulative frequency graphs [Higher only]

Tier 1 · Easy

Mark scheme for S3 Tier 1 · Easy
QuestionAnswerMarkMark scheme
1
  • 44
2The class width is 2015=520-15=5. Frequency density is 20/5=420/5=4.
2
  • 2020
1Frequency is the area of the bar: frequency=class width×frequency density=8×2.5=20\text{frequency}=\text{class width}\times\text{frequency density}=8\times2.5=20.

Tier 2 · Standard

Mark scheme for S3 Tier 2 · Standard
QuestionAnswerMarkMark scheme
1
  • Frequency densities: 1.21.2, 22, 1.21.2.
  • The class 10<x2510<x\leq25 has the tallest bar.
3Divide each frequency by its class width: 12/10=1.212/10=1.2, 30/15=230/15=2 and 18/15=1.218/15=1.2. The greatest density, 22, gives the tallest bar.
2
  • 1818
3For 5<x155<x\leq15, the class width is 1010, so the frequency density is 24/10=2.424/10=2.4. A height of 6cm6\,\text{cm} represents density 2.42.4, so 1cm1\,\text{cm} represents 0.40.4. The second density is 7.5×0.4=37.5\times0.4=3, and its width is 66, giving frequency 3×6=183\times6=18.
3
  • The frequencies are 2424, 3030 and 3030, respectively.
4Histogram frequency is proportional to bar area. The relative areas are 4×6=244\times6=24, 6×5=306\times5=30 and 10×3=3010\times3=30, in the ratio 4:5:54 : 5 : 5. The 1414 ratio parts represent 8484 values, so each part represents 66. The frequencies are therefore 2424, 3030 and 3030.

Tier 3 · Hard

Mark scheme for S3 Tier 3 · Hard
QuestionAnswerMarkMark scheme
1
  • Median 25.8\approx25.8.
  • Lower quartile 16.7\approx16.7 and upper quartile =35=35.
  • Interquartile range 18.3\approx18.3.
5The median is the 4040th value, between cumulative frequencies 2626 and 5050: 20+40265026×1025.820+\dfrac{40-26}{50-26}\times10\approx25.8. The lower quartile is the 2020th value: 10+208268×1016.710+\dfrac{20-8}{26-8}\times10\approx16.7. The upper quartile is the 6060th value: 30+60507050×10=3530+\dfrac{60-50}{70-50}\times10=35. Hence IQR3516.7=18.3IQR\approx35-16.7=18.3.
2
  • b=28b=28.
  • Cumulative frequency =47=47 at x=28x=28 and 6565 at x=36x=36.
  • The second frequency density is 2.252.25.
5The first class width is 32/2=1632/2=16, so b=12+16=28b=12+16=28. Add the first frequency to get cumulative frequency 15+32=4715+32=47 at x=28x=28, then add 1818 to get 6565 at x=36x=36. The second width is 3628=836-28=8, so its frequency density is 18/8=2.2518/8=2.25.
3
  • A cumulative frequency cannot decrease as xx increases, but the plotted value falls from 2323 to 1919.
  • The correct point is (30,35)(30,35).
  • The frequency in 30<x4030<x\leq40 is 77.
  • The median is in the class 20<x3020<x\leq30.
4The cumulative frequency is already 2323 at x=20x=20, so it cannot be 1919 at the larger boundary x=30x=30. Add the class frequency 1212 to get 23+12=3523+12=35, giving (30,35)(30,35). The next class frequency is 4235=742-35=7. The 2525th and 2626th values lie after cumulative frequency 2323 and no later than cumulative frequency 3535, so the median lies in 20<x3020<x\leq30.
4
  • b=22b=22 and the total frequency is 9696.
5The first class has width 1212, so its frequency density is 36/12=336/12=3. The second bar has density 66 and, because it has the same area, frequency 3636. Its width is therefore 36/6=636/6=6, giving b=16+6=22b=16+6=22. The third density is 6/3=26/3=2 and its width is 3422=1234-22=12, so its frequency is 2424. The total is 36+36+24=9636+36+24=96.
5
  • Estimated number of values greater than 2626 is 3030.
  • 70%70\% of the data lie below an estimated value of 2727.
5The class frequencies are 8(1.5)=128(1.5)=12, 12(2.5)=3012(2.5)=30, 12(3)=3612(3)=36 and 10(1.2)=1210(1.2)=12, so there are 9090 values. Above 2626, the part of the third class has estimated frequency (3226)×3=18(32-26)\times3=18, and the final class contributes 1212, giving 3030. The 70%70\% point is the 0.70×90=630.70\times90=63rd value. The cumulative frequency is 4242 at x=20x=20, so interpolate 2121 values into the third class: 20+634236×12=2720+\dfrac{63-42}{36}\times12=27.

S4 · Interpret, analyse and compare data-set distributions via graphs (incl. box plots), central tendency (median, mean, mode, modal class) and spread (range, outliers, quartiles, inter-quartile range)

Tier 1 · Easy

Mark scheme for S4 Tier 1 · Easy
QuestionAnswerMarkMark scheme
1
  • Mean =7=7; range =6=6.
2The total is 4+6+6+9+10=354+6+6+9+10=35, so the mean is 35/5=735/5=7. The range is 104=610-4=6.
2
  • Mode =4=4; median =6=6.
244 occurs most often, so it is the mode. The data are ordered and the fourth of the seven values is 66, so the median is 66.

Tier 2 · Standard

Mark scheme for S4 Tier 2 · Standard
QuestionAnswerMarkMark scheme
1
  • B has the slightly higher median: 2525 compared with 2424.
  • B has the smaller range: 2929 compared with 3333, so its values are less spread out overall.
2Compare the typical values using the medians, then compare spread using the ranges. B has both the larger centre and the smaller spread.
2
  • New mean =19=19.
  • The added reading is greater than the original mean, so it raises the mean.
3The original total is 5×18=905\times18=90. The new total is 90+24=11490+24=114, so the new mean is 114/6=19114/6=19. Since 2424 is above the original mean, the mean increases.
3
  • Mean =27=27; median =25=25; range =18=18.
3The mean and median both undergo the same transformation as every value: 2×12+3=272\times12+3=27 and 2×11+3=252\times11+3=25. Doubling every value doubles the range, while adding the same amount to every value does not change the range, so the new range is 2×9=182\times9=18.

Tier 3 · Hard

Mark scheme for S4 Tier 3 · Hard
QuestionAnswerMarkMark scheme
1
  • Mean =18.75=18.75, median =14.5=14.5, range =40=40.
  • The median is more representative because 5252 is an outlier that pulls up the mean.
4The total is 150150, giving mean 150/8=18.75150/8=18.75. The median is (14+15)/2=14.5(14+15)/2=14.5 and the range is 5212=4052-12=40. The isolated value 5252 has a strong effect on the mean but not on the median, so the median better represents the main cluster.
2
  • x=14x=14; median =12.5=12.5; range =14=14.
4The total of all six values is 6×13=786\times13=78. The known values total 6+9+11+18+20=646+9+11+18+20=64, so x=7864=14x=78-64=14. The median is the mean of the third and fourth values: (11+14)/2=12.5(11+14)/2=12.5. The range is 206=1420-6=14.
3
  • a=7a=7, b=13b=13 and the range is 1717.
4For 77 to be the only mode, it must occur again, so a=7a=7. A mean of 1111 for seven values gives a total of 7777. The known values, including a=7a=7, total 6464, so b=7764=13b=77-64=13. This is consistent with the stated order. The range is 203=1720-3=17.
4
  • k=4k=4, median =7=7, mode =7=7 and range =9=9.
4The total frequency is 8+k8+k and the total of the values is 2(4)+5(7)+10k+13=56+10k2(4)+5(7)+10k+13=56+10k. Therefore 56+10k=8(8+k)56+10k=8(8+k), so 2k=82k=8 and k=4k=4. There are 1212 values; the sixth and seventh are both 77, so the median is 77. The greatest frequency is 55, so the mode is 77, and the range is 134=913-4=9.
5
  • x=22x=22, lower quartile =8=8 and interquartile range =16=16.
4After excluding the median 1515, the upper half is 18,x,26,3118,x,26,31. Its median is (x+26)/2=24(x+26)/2=24, so x=22x=22, which is consistent with the stated order. The lower half is 4,7,9,124,7,9,12, so the lower quartile is (7+9)/2=8(7+9)/2=8. The interquartile range is 248=1624-8=16.

S5 · Apply statistics to describe a population

Tier 1 · Easy

Mark scheme for S5 Tier 1 · Easy
QuestionAnswerMarkMark scheme
1
  • 600kg600\,\text{kg}
2Apply the sample mean to all 250250 parcels: 2.4×250=600kg2.4\times250=600\,\text{kg}.
2
  • 118kWh118\,\text{kWh}
1Use the sample mean as the estimate of the population mean. Because the sample is representative, the estimated mean for all households is 118kWh118\,\text{kWh}.

Tier 2 · Standard

Mark scheme for S5 Tier 2 · Standard
QuestionAnswerMarkMark scheme
1
  • 7.17.1 sessions
3The total attendances represented are 120×6.5=780120\times6.5=780 and 80×8=64080\times8=640. Divide their sum by all 200200 members: (780+640)/200=7.1(780+640)/200=7.1.
2
  • 88 full-time employees
3Let nn be the number of full-time employees. The total weekly hours give 16×15+39n=23(16+n)16\times15+39n=23(16+n). Therefore 240+39n=368+23n240+39n=368+23n, so 16n=12816n=128 and n=8n=8.
3
  • 22 days
3The total number of absence days is 2400×1.5=36002400\times1.5=3600. The number of pupils with at least one day absent is 75%75\% of 24002400, which is 18001800. Their mean is therefore 3600/1800=23600/1800=2 days.

Tier 3 · Hard

Mark scheme for S5 Tier 3 · Hard
QuestionAnswerMarkMark scheme
1
  • Estimated total =930=930 residents.
  • Estimated percentage =31%=31\%.
5For the north, estimate 1800×15/60=4501800\times15/60=450. For the south, estimate 1200×20/50=4801200\times20/50=480. The total is 450+480=930450+480=930 from a population of 30003000, so the percentage is 930/3000×100=31%930/3000\times100=31\%. Calculating by subgroup correctly accounts for their different sizes.
2
  • 40%40\%
4There are 280+420=700280+420=700 employees, so 0.32×700=2240.32\times700=224 work part-time. Of the office employees, 0.20×280=560.20\times280=56 work part-time. Therefore 22456=168224-56=168 warehouse employees work part-time, giving 168/420×100=40%168/420\times100=40\%.
3
  • 590590 g
  • The combined sample is larger, so its mean is less affected by sampling variation and is likely to be closer to the population mean.
4The total mass in sample A is 40×620=24,80040\times620=24{,}800 g and in sample B is 60×570=34,20060\times570=34{,}200 g. The combined total is 59,00059{,}000 g across 100100 fish, so the combined mean is 59,000÷100=59059{,}000\div100=590 g. A larger sample gives a more reliable estimate of the population mean.
4
  • n=28n=28.
  • Wood C has the lowest estimated proportion of affected oak trees.
4The estimates for woods A and B are 900×21/30=630900\times21/30=630 and 600×16/25=384600\times16/25=384. Wood C therefore contributes 1264630384=2501264-630-384=250 to the estimated total. Its estimated affected proportion is 250/500=0.5250/500=0.5, so 14/n=0.514/n=0.5 and n=28n=28. The estimated proportions for A, B and C are 0.70.7, 0.640.64 and 0.50.5, respectively, so Wood C has the lowest proportion.
5
  • Estimated number =8000=8000 seedlings.
  • The percentage error is 4.17%4.17\% (or 4.2%4.2\% to 1 d.p.) and the estimate is too high.
4The estimated density is 6.46.4 seedlings per square metre, giving 6.4×1250=80006.4\times1250=8000 seedlings. The estimate exceeds the census by 80007680=3208000-7680=320. Relative to the census, the percentage error is 320/7680×100=4.1666%4.17%320/7680\times100=4.1666\ldots\%\approx4.17\%, so the estimate is too high.

S6 · Use and interpret scatter graphs of bivariate data; recognise correlation, know it does not indicate causation; draw estimated lines of best fit; make predictions; interpolate/extrapolate with caution

Tier 1 · Easy

Mark scheme for S6 Tier 1 · Easy
QuestionAnswerMarkMark scheme
1
  • Positive correlation.
  • Correlation does not establish causation; another factor such as weekends or visitor numbers could affect sales.
2An upward association is positive correlation. Then identify a plausible confounding variable to show why the paired data alone cannot isolate a causal effect.
2
  • Negative correlation.
  • Beyond the plotted points the pattern may not continue, so the estimate is extrapolation and unreliable.
2A downward pattern from left to right is negative correlation. The graph only gives evidence for altitudes within the plotted range; assuming the same pattern beyond it is extrapolation, which may not hold.

Tier 2 · Standard

Mark scheme for S6 Tier 2 · Standard
QuestionAnswerMarkMark scheme
1
  • y=31.2y=31.2 when x=14x=14.
  • x=35x=35 is outside the observed range, so this would be extrapolation and the trend may not continue.
3Substitute x=14x=14: y=1.8(14)+6=25.2+6=31.2y=1.8(14)+6=25.2+6=31.2. Since 1414 lies inside the data range this is interpolation, whereas 3535 lies outside it.
2
  • It should be one straight line, not two sections joined at a corner.
  • It does not pass through the centre of the points, with a roughly equal number of points on each side.
2A line of best fit should be a single straight line through the centre of the scatter. The corner makes the student's line unsuitable, and having 1010 of 1212 points above it shows that the points are not balanced around the line.
3
  • The coordinates are (18,75)(18,75).
  • The predicted mass is 74g74\,\text{g}, so the recorded mass is 1g1\,\text{g} above the line.
3Convert 180mm180\,\text{mm} to 18cm18\,\text{cm}, so the point is (18,75)(18,75). Substituting x=18x=18 gives y=3(18)+20=74y=3(18)+20=74. Since 7574=175-74=1, the recorded mass is 1g1\,\text{g} above the line.

Tier 3 · Hard

Mark scheme for S6 Tier 3 · Hard
QuestionAnswerMarkMark scheme
1
  • At 99 months, estimated mass =5.58kg=5.58\,\text{kg}.
  • At 1616 months, estimated mass =8.52kg=8.52\,\text{kg}.
  • The 99-month estimate is interpolation and is more reliable; the 1616-month estimate is extrapolation.
  • The association does not prove age alone causes the change because factors such as breed or diet may also affect mass.
5Substitute into the line: for a=9a=9, m=0.42(9)+1.8=5.58m=0.42(9)+1.8=5.58; for a=16a=16, m=0.42(16)+1.8=8.52m=0.42(16)+1.8=8.52. The first age lies within 22 to 1212, while the second lies outside. The graph shows association only and does not control other variables.
2
  • Correlation does not prove that longer shoes cause better reading.
  • Age could affect both shoe length and reading score (accept another valid variable affecting both).
  • An adult is outside the range of the pupil data, so the prediction would be extrapolation and the trend may not continue.
3The graph shows association only. Older pupils tend to have both longer feet and more developed reading skills, so age is a confounding variable. An adult is outside the group and data range studied, making any estimate an unreliable extrapolation.
3
  • The line is y=3x+82y=-3x+82.
  • Estimated time =14=14 hours.
  • Interpolation, because 1414 hours is within the observed range from 44 to 1818 hours.
4Write y=3x+cy=-3x+c. Using (10,52)(10,52) gives 52=30+c52=-30+c, so c=82c=82 and the line is y=3x+82y=-3x+82. Set y=40y=40: 40=3x+8240=-3x+82, so 3x=423x=42 and x=14x=14. This lies inside the observed range, so it is interpolation.
4
  • The point (8,54)(8,54) lies 66 units above the line's prediction.
  • The observed value at x=13x=13 is 52.552.5.
  • The point with x=13x=13 lies farther from the line.
4At x=8x=8, the line predicts 4.5(8)+12=484.5(8)+12=48. Since 5448=654-48=6, the point is 66 units above the prediction. At x=13x=13, the line predicts 4.5(13)+12=70.54.5(13)+12=70.5, so a point 1818 units below it has observed value 70.518=52.570.5-18=52.5. Since 18>618>6, the second point is farther from the line.
5
  • The lines give the same prediction at x=8x=8, where y=59y=59.
  • At x=10x=10, group A has the greater predicted value by 44.
  • The predictions at x=10x=10 are interpolations.
4Set the predictions equal: 5x+19=3x+355x+19=3x+35, so 2x=162x=16 and x=8x=8. Substitution gives y=59y=59. At x=10x=10, group A predicts 6969 and group B predicts 6565, so A is greater by 44. Since 1010 lies within the observed range 33 to 1212, both predictions are interpolations.