Skip to content
S2.2

Interpret scatter diagrams and regression lines for bivariate data, including recognising distinct sections of the population (regression calculations excluded); interpret correlation informally; correlation does not imply causation.

Draft — not yet indexed

Scatter diagrams and regression lines

Worked answers and methods for S2.2 on Edexcel A-level Maths 9MA0.

Explanation

  • A scatter diagram shows paired bivariate data; describe the direction and strength of its association and note any outliers or clusters.
  • A regression line estimates the mean response for a given explanatory-variable value and is most defensible within the observed data range.
  • Distinct clusters may represent different sections of the population, so one overall correlation or regression line can hide different within-group patterns.
  • Correlation measures association, not causation; a lurking variable, reverse causation or coincidence may explain the observed relationship.
  • For y=axny=ax^n, plot logy\log y against logx\log x to get gradient nn and intercept loga\log a; for y=kbxy=kb^x, plot logy\log y against xx to get gradient logb\log b and intercept logk\log k.

Worked example

A scatter diagram of journey distance against journey time contains one cluster for bicycles and a separate cluster for cars. Explain why fitting one regression line to all journeys may be misleading.

  1. 1.Identify vehicle type as a grouping variable.
  2. 2.Since the relationship between distance and time is likely to differ between bicycles and cars, pooling the groups can produce a line that fits neither group well.
  3. 3.Analyse the sections separately.

Answer: The two clusters represent distinct sections of the population with different travel speeds.; A single line may mainly reflect the gap between the clusters rather than either within-group relationship.; Separate regression lines or separate analyses would be more informative.

Common mistakes

  • Don't extrapolate a regression line far beyond the observed data range.
  • Don't interpret an overall regression line despite distinct clusters representing different subpopulations.

Exam tip

Before using a regression line, inspect clusters and explain whether one relationship is credible for the whole population.

Worked practice

Q1
Tier 1 · Easy

1.

A scatter diagram of daily ice-cream sales against temperature shows a strong positive correlation. Interpret this and explain why it does not prove that higher temperature causes every increase in sales.

(2)

(Total for Question 1 is 2 marks)

Mark scheme

Mark scheme for question 1
QuestionSchemeMarks
1
  • Higher temperatures tend to be associated with higher ice-cream sales.
  • Correlation alone does not establish causation; other variables such as sunshine, holidays or visitor numbers may affect both quantities.
2
Notes
Translate positive correlation as a tendency for larger values of one variable to occur with larger values of the other. Then distinguish association from a causal conclusion by identifying a plausible lurking variable.

(2 marks)

Q2
Tier 2 · Standard

2.

Across a year, ice-cream sales and the number of reported sunburn cases have strong positive correlation. Explain why this does not establish that buying ice cream causes sunburn.

(3)

(Total for Question 2 is 3 marks)

Mark scheme

Mark scheme for question 2
QuestionSchemeMarks
2
  • Correlation does not imply causation.
  • A valid lurking variable is hot or sunny weather, season, holiday periods or visitor numbers, any of which could increase both ice-cream sales and sunburn cases.
  • Evidence controlling for plausible lurking variables, or an appropriately designed study, would be needed to support a causal claim.
3
Notes
The scatter pattern records association only. Temperature or sunshine can affect both variables: hot days encourage ice-cream purchases and increase exposure to ultraviolet radiation. This common cause can generate the correlation even if ice cream has no effect on sunburn.

(3 marks)

Q3
Tier 3 · Hard

3.

For trees aged between 33 and 1818 years, a regression line of trunk diameter dd cm on age aa years is d=1.8a+4.2d=1.8a+4.2. A 1212-year-old tree has diameter 2929 cm. Interpret the difference between the observed and predicted values and critique using the line to predict the diameter of a 4040-year-old tree.

(4)

(Total for Question 3 is 4 marks)

Mark scheme

Mark scheme for question 3
QuestionSchemeMarks
3
  • The predicted diameter at a=12a=12 is 25.825.8 cm, so the difference between the observed and predicted values is 2925.8=3.229-25.8=3.2 cm.
  • The tree is 3.23.2 cm thicker than the value predicted by the line.
  • A prediction at 4040 years is extrapolation far outside the observed age range and the linear relationship may not continue.
4
Notes
Substitute a=12a=12 to get the fitted value d=1.8(12)+4.2=25.8d=1.8(12)+4.2=25.8. Residual == observed - fitted =2925.8=3.2=29-25.8=3.2 cm. Since 4040 is well beyond the data range 33 to 1818, using the line there is unsupported extrapolation.

(4 marks)

Q4
Tier 1 · Easy

4.

A scatter diagram contains one compact cluster for electric bicycles and a separate compact cluster for conventional bicycles. State why one regression line for all the points may be misleading.

(2)

(Total for Question 4 is 2 marks)

Mark scheme

Mark scheme for question 4
QuestionSchemeMarks
4
  • The clusters represent distinct sections of the population.
  • One line may describe mainly the separation between bicycle types rather than the relationship within either type, so the groups should be considered separately.
2
Notes
The bicycle type is a grouping variable. Pooling two separated groups can create an overall pattern that does not represent the within-group relationship, so separate analyses are more informative.

(2 marks)

Q5
Tier 2 · Standard

5.

For recordings made at wind speeds between 55 and 2525 km h1^{-1}, the regression line of sound level ss dB on wind speed ww km h1^{-1} is s=681.4ws=68-1.4w. Find the predicted sound level when w=18w=18 and interpret the gradient in context.

(3)

(Total for Question 5 is 3 marks)

Mark scheme

Mark scheme for question 5
QuestionSchemeMarks
5
  • Predicted sound level =42.8=42.8 dB.
  • Within the observed range, the model predicts a decrease of 1.41.4 dB in mean sound level for each increase of 11 km h1^{-1} in wind speed.
3
Notes
Substitute w=18w=18: s=681.4(18)=42.8s=68-1.4(18)=42.8. The gradient 1.4-1.4 is the fitted change in the response variable for a one-unit increase in the explanatory variable; it describes association within the data range.

(3 marks)

Q6
Tier 3 · Hard

6.

A scatter diagram compares hours of optional practice with assessment score for two courses. Students on the advanced course form a cluster with both higher practice hours and higher scores than students on the introductory course. Within each course cluster, score tends to decrease slightly as practice hours increase, but the pooled data show positive correlation. Explain why the pooled correlation and a single regression line could be misleading, and why the diagram does not show that extra practice lowers scores.

(5)

(Total for Question 6 is 5 marks)

Mark scheme

Mark scheme for question 6
QuestionSchemeMarks
6
  • Course level creates two distinct sections of the population.
  • The pooled positive correlation may be caused mainly by the separation of the two clusters and need not describe either course.
  • A single regression line may fit neither within-course relationship well, so the courses should be analysed separately.
  • The slight negative within-course association does not prove causation; prior attainment, difficulty or students seeking help could affect both practice time and score.
5
Notes
The overall association mixes a between-course difference with two different within-course patterns. Course level is therefore a lurking grouping variable, and one pooled line conceals the structure. Even after separating the courses, an observational association cannot establish that practice causes lower scores because other variables can influence both quantities.

(5 marks)

Q7
Tier 2 · Standard

7.

A scatter diagram plots advertising cost in pounds against sales income in pounds, and its fitted line has gradient 0.400.40. The sales income is then converted from pounds to pence, without changing any observations. State the new gradient and describe the effect of this conversion on the direction and strength of the correlation.

(3)

(Total for Question 7 is 3 marks)

Mark scheme

Mark scheme for question 7
QuestionSchemeMarks
7
  • New gradient =100(0.40)=40=100(0.40)=40 pence of sales income per pound of advertising cost.
  • The direction remains positive.
  • The strength of the correlation is unchanged because multiplying every response value by the same positive constant only changes its scale, not the scatter pattern.
3
Notes
Converting pounds to pence multiplies every response value, and hence the fitted change in response, by 100100. The gradient becomes 4040. This positive rescaling does not alter which points lie relatively above or below one another or how tightly they follow a linear pattern, so it changes neither the positive direction nor the strength of the correlation.

(3 marks)

Q8
Tier 3 · Hard

8.

For observations with 1x91\leq x\leq9, the fitted equation predicting yy from xx is y=2x1y=2x-1. After the observation (9,25)(9,25) is omitted, the equation becomes y=1.2x+1.0y=1.2x+1.0. Calculate the residual of (9,25)(9,25) from the second fit and compare the two predicted values at x=8x=8. Interpret what these results show, and state why the observation should not be deleted solely on this evidence.

(5)

(Total for Question 8 is 5 marks)

Mark scheme

Mark scheme for question 8
QuestionSchemeMarks
8
  • The second line predicts 11.811.8 at x=9x=9, so the residual of (9,25)(9,25) is 2511.8=13.225-11.8=13.2.
  • At x=8x=8, the full-data prediction is 1515 and the prediction after omission is 10.610.6, a difference of 4.44.4.
  • The observation has a large residual and substantially influences the fitted line.
  • Influence does not prove a recording error; the source and context should be checked before retaining, correcting or excluding it.
5
Notes
Using the line fitted without the observation isolates how unusual it is relative to the remaining pattern: 25[1.2(9)+1.0]=13.225-[1.2(9)+1.0]=13.2. The two fitted values at x=8x=8 are 2(8)1=152(8)-1=15 and 1.2(8)+1.0=10.61.2(8)+1.0=10.6. Their difference shows that the point changes the fitted relationship materially. Statistical influence is a reason to investigate, not automatic evidence that the value is invalid.

(5 marks)

Q9
Tier 3 · Hard

9.

For observations with 2x202\leq x\leq20, a regression line of log10y\log_{10}y on log10x\log_{10}x is log10y=log102+1.5log10x\log_{10}y=\log_{10}2+1.5\log_{10}x. Obtain a model of the form y=axny=ax^n, and use it to estimate yy when x=8x=8. Comment on using the model when x=50x=50, and explain why a strong linear correlation on the transformed plot would not by itself establish a causal relationship.

(6)

(Total for Question 9 is 6 marks)

Mark scheme

Mark scheme for question 9
QuestionSchemeMarks
9
  • y=2x1.5y=2x^{1.5}, so a=2a=2 and n=1.5n=1.5.
  • When x=8x=8, y2(8)1.5=45.3y\approx2(8)^{1.5}=45.3 to 33 significant figures.
  • x=50x=50 is outside the observed range, so this would be unsupported extrapolation and the power relationship may not continue.
  • Correlation on the transformed plot is association only; lurking variables, reverse causation or the data-generating process must be considered before making a causal claim.
6
Notes
Taking powers of 1010 gives y=2x1.5y=2x^{1.5}. Substitution gives 283=45.2542\sqrt{8^3}=45.254\ldots. The fitted relationship was supported only for 2x202\leq x\leq20, and a regression pattern does not identify a causal mechanism.

(6 marks)

Q10
Tier 3 · Hard

10.

For a bivariate data set with xˉ=10\bar{x}=10 and yˉ=30\bar{y}=30, the fitted equation for predicting yy from xx is y=2x+10y=2x+10. The fitted equation for predicting xx from yy is x=0.4y2x=0.4y-2. Estimate yy when x=12x=12, stating which equation you use. Show that both equations pass through (xˉ,yˉ)(\bar{x},\bar{y}). Explain why rearranging the second equation gives a different estimate and is not the appropriate procedure here.

(5)

(Total for Question 10 is 5 marks)

Mark scheme

Mark scheme for question 10
QuestionSchemeMarks
10
  • Use the regression line of yy on xx: the estimate is y=2(12)+10=34y=2(12)+10=34.
  • At x=10x=10, the first line gives y=30y=30; at y=30y=30, the second gives x=10x=10, so both pass through (10,30)(10,30).
  • Rearranging x=0.4y2x=0.4y-2 would give y=35y=35 when x=12x=12.
  • The two regression lines minimise deviations in different variables. To predict yy from a given xx, use the regression of yy on xx, not the algebraic inverse of the regression of xx on yy.
5
Notes
The response being estimated is yy, so substitute x=12x=12 into y=2x+10y=2x+10. The mean point satisfies both supplied equations. Rearrangement of the other line gives (12+2)/0.4=35(12+2)/0.4=35, but that line was fitted to predict xx from yy and generally is not the inverse of the yy-on-xx line unless correlation is perfect.

(5 marks)

Verified exam appearances

We have not yet indexed a verified real-paper appearance for S2.2. Browse the Edexcel A-level Maths 9MA0 past papers directly.

Other points in S2 Data presentation and interpretation

Want help turning this into marks?

Bring S2.2 or any tricky specification point, and we can work through the method and exam wording together.