← back to statistics

Correlation and Linear Regression

Statistics through experiments · 15 of 16 · CC BY-SA 4.0

A line compresses a relationship. Its residuals show what the compression leaves behind.

Suppose we record hours of practice and a quiz score for each student. We now have two numerical variables per person. A scatterplot puts practice on the horizontal axis and score on the vertical axis. Each dot keeps one person’s measurements together.

Before fitting a line, look for direction, curvature, clusters, and unusual points. Then try moving the line below. Can you reduce the vertical gaps between the dots and their predictions?

How close can your line get?

Each dot is an invented observation. Move the line to make the vertical gaps small. Those gaps are residuals: observed value minus predicted value.

Predict y from x · dashed vertical lines are residuals
Observations, fitted line, and residuals; exact values are in the table.00551010xy

Sum of squared residuals: 23.00. Best possible for these points: 4.73.

Adding a point keeps your current line. Predict how the best line will change, then fit again. Axes stay fixed; parts of a trial line outside the plot are clipped.

Read predictions and residuals
xObserved yPredicted yResidual
122.50-0.50
243.001.00
333.50-0.50
454.001.00
554.500.50
685.003.00
775.501.50
896.003.00

Inference for the least-squares line

Fitted slope: 0.940. 95% interval for the slope: [0.605, 1.276]. Two-sided p-value for zero slope: < 0.001. R² = 0.887.

These describe the best-fitting line for the current points, regardless of where you move your trial line. The t procedure uses 6 degrees of freedom and assumes a linear mean relationship with independent, normal, constant-variance errors. These invented points illustrate the calculation; they are not evidence from a sampled population.

The line you have built

ŷ = b₀ + b₁x. The hat on y marks a prediction. b₀ is the intercept; b₁ is the slope. Your current line is ŷ = 2.00 + (0.50)x. Least squares chooses the coefficients that minimize the sum of squared residuals, ∑(y − ŷ)².

A good fit does not establish causation. Extrapolating beyond observed x values needs additional justification.

A moment to think

One prediction is 3 units too high and another is 3 too low. Is the fit perfect because the errors cancel?

What the line says

The slope is the change in predicted outcome for a one-unit increase in the predictor. If hours predict quiz points, a slope of two means two predicted points per additional hour. It describes an association unless the design and causal assumptions justify more.

The intercept is the prediction when the predictor is zero. Sometimes that has a useful meaning; sometimes zero lies far outside the observed range. An intercept is still needed to position the line even when it has little direct interpretation.

A residual is observed outcome minus predicted outcome. Least squares chooses the slope and intercept that minimize the sum of squared residuals. Adding an unusual point can change that choice substantially. A point far out along the predictor axis has leverage; whether it strongly changes the fitted line also depends on where its outcome falls.

Association without units

Pearson’s correlation, r, summarizes the direction and strength of a linear relationship on a scale from −1 to 1. Positive values go with an upward tendency; negative values with a downward tendency. A value near zero does not rule out a strong curved relationship.

Correlation does not change when hours become minutes. The slope does: its units change. Neither a large correlation nor a steep slope can establish that practice caused the score differences. Motivation, prior knowledge, and access to help could affect both variables.

For an ordinary least-squares line with an intercept and one predictor, R² equals r². It is the fraction of the outcome’s sample variation around its mean accounted for by the fitted line. It is not the fraction of students predicted correctly, and it is not a fraction caused by the predictor.

A moment to think

Students who practice more tend to score higher in an observational survey. What does a positive slope establish?

The fitted slope is an estimate too

A different sample could produce a different slope. The activity displays a t interval and a two-sided test of zero population slope for the best-fitting line, independently of your trial line. Adding a point changes both the fit and its uncertainty.

The textbook t procedure assumes the mean relationship is linear, errors are independent and normally distributed, and their variance is constant across predictor values. Normality supports exact small-sample inference. Residual patterns can reveal curvature, changing spread, or unusual observations; independence is chiefly a question about data collection.

Plot residuals against the predictor or fitted values: a curve suggests a missed pattern, while a widening fan suggests changing variability. The activity’s dashed gaps show individual residuals; the deeper chapter develops the diagnostic view. Passing a visual check does not prove all assumptions true.

An idea to take with you

A fitted relationship, with room for uncertainty

ŷ = b₀ + b₁x
residual = y − ŷ
R² = 1 − Σ(y − ŷ)² / Σ(y − ȳ)²

The hat marks a prediction. b₀ positions the line; b₁ sets its slope. R² compares the remaining squared variation with the variation left by predicting the same mean for everyone.

For the classical simple-regression slope test, t = b₁/SE(b₁), with n − 2 degrees of freedom. We have estimated two line coefficients. This tests zero slope within the specified model; it does not test every possible relationship.

Return to the experiment ↑

A moment to think

Which is usually wider at the same predictor value: an interval for the population mean outcome or one for a new person’s outcome?

Extrapolation uses the line beyond the observed predictor range. A pattern seen over one to five hours need not continue to twenty hours. Good fit within the observed range does not supply evidence outside it.

Words you met

TermMeaning
CorrelationA unit-free summary of a linear association’s direction and strength.
SlopeThe fitted change in outcome per unit increase in predictor.
ResidualObserved outcome minus fitted prediction.
R²The fraction of sample outcome variation accounted for by the fitted model, with an intercept.
Prediction intervalAn interval for a new individual outcome, including individual variation.
Neighbors

Written by June Kim. The conversational approach was inspired by Danielle Navarro’s Learning Statistics with R (CC BY-SA 4.0). The prose, experiments, and questions here were created for this book. This chapter’s text and illustrations are also shared under CC BY-SA 4.0.