MyLeoNes™

Regression lines and residuals — Data, 14–17 years

A line summarises how one numerical variable tends to change with another, while residuals show what the line misses.

Idea

A regression line is a deliberately simple summary drawn through a cloud of paired data. Its slope describes the typical change in one variable for a one-unit increase in the other. A residual is the actual value minus the line’s predicted value.

Why use it?

A scatterplot can contain too many points to describe one by one. Regression gives a compact description and can make cautious predictions inside the range of the data. Residuals are important because they reveal whether the line is systematically wrong.

Worked example

Suppose a fitted line predicts exam score from study hours as score = 52 + 6h. For 4 hours, prediction = 52 + 6×4 = 76. If the student actually scores 81, the residual is 81 − 76 = 5. The student scored five points above the line’s prediction.

Common trap

The line is often mistaken for a rule that every individual must follow. It only describes an average tendency, so two people with the same x-value can have different outcomes. It is also risky to predict far beyond the observed range; that is extrapolation.

Where it appears

Businesses use regression to estimate sales from advertising or demand from price. Scientists may predict a measured quantity from another measurement. In both cases, the line is a planning aid, not a guarantee, and residuals help show how unreliable particular predictions may be.

Keep exploring

Other languages

Loading MyLeoNes™…