Regression lines and residuals — Data, 14–17 years
A line summarises how one numerical variable tends to change with another, while residuals show what the line misses.
Idea
A regression line is a deliberately simple summary drawn through a cloud of paired data. Its slope describes the typical change in one variable for a one-unit increase in the other. A residual is the actual value minus the line’s predicted value.
Why use it?
A scatterplot can contain too many points to describe one by one. Regression gives a compact description and can make cautious predictions inside the range of the data. Residuals are important because they reveal whether the line is systematically wrong.
Worked example
Suppose a fitted line predicts exam score from study hours as score = 52 + 6h. For 4 hours, prediction = 52 + 6×4 = 76. If the student actually scores 81, the residual is 81 − 76 = 5. The student scored five points above the line’s prediction.
Common trap
The line is often mistaken for a rule that every individual must follow. It only describes an average tendency, so two people with the same x-value can have different outcomes. It is also risky to predict far beyond the observed range; that is extrapolation.
Where it appears
Businesses use regression to estimate sales from advertising or demand from price. Scientists may predict a measured quantity from another measurement. In both cases, the line is a planning aid, not a guarantee, and residuals help show how unreliable particular predictions may be.
Keep exploring
Other languages
Loading MyLeoNes™…