← back to statistics

Comparing Means with Student’s t

Statistics through experiments · 12 of 16 · CC BY-SA 4.0

Two groups of people and two measurements of the same people need different comparisons.

One group tries a new practice routine and another uses the usual routine. Their average scores differ by three points. Is three a clear signal, or small compared with the uncertainty in our estimate?

There is another way to collect scores: measure the same people before and after practice. A person who starts high may finish high under either routine. Keeping track of who is paired with whom can remove much of that person-to-person variation from the comparison.

The activity uses constructed measurements so you can isolate these ideas. First compare independent groups. Then switch to paired measurements and inspect the individual changes.

The design chooses the standard error

These are constructed quiz-score examples. Choose two independent groups of twelve people, or twelve people measured twice. The paired example has large differences between people but much smaller changes within a person.

Group A mean: 50.0. Group B mean: 53.0. Difference: 3.0 points.

95% interval for the mean difference · fixed −12 to 12 point scale
-10 points0 points10 points

Welch t: estimated difference 3.00, SE 1.33, df 21.9. 95% interval [0.25, 5.75]. Two-sided p-value 0.034 for a zero mean difference.

Read the measurements
Row (not a pair)AB
14852
25254
34749
45557
55055
64951
75353
84650
95158
105456
114547
125054

The paired method assumes independent people and approximately normal differences for this small sample. Welch’s method assumes independent groups and observations, with approximately normal populations for small samples; it allows unequal variances. Inspect data and design before using either method. Before-and-after changes alone do not establish causation.

Why a new distribution?

Earlier we used a normal model with known population standard deviation. Usually we do not know that spread: we estimate it from the sample. The standard error is now uncertain too, especially with few observations.

Student’s t distribution allows for that extra uncertainty. It is symmetric like the standard normal distribution, but has heavier tails. Its shape depends on degrees of freedom: the amount of independent information left after estimation. More degrees of freedom bring it closer to the normal distribution.

For a single sample mean, the statistic is t = (x̄ − μ₀)/(s/√n), where μ₀ is the proposed population mean. With independent, normally distributed observations it follows a t distribution with n − 1 degrees of freedom under the null. For larger samples the method can tolerate some non-normality, but extreme outliers and strong skew deserve attention.

Two independent groups

For separate groups, each sample mean adds uncertainty to the difference. Welch’s t procedure estimates those contributions separately and does not require equal population variances. Its degrees of freedom can be fractional; they are calculated from both groups’ sizes and variances.

Independent observations within and between groups are essential. Small samples need approximately normal populations for a reliable t approximation. Inspect the measurements and the design; a formula cannot detect classroom clustering or repair a biased sample.

A moment to think

The same twelve students take a quiz before and after practice. What should a paired analysis summarize?

Let each person be their own comparison

In a paired t procedure, subtract each before score from its matching after score. Now we have one sample of differences. Its mean is the estimated change; its standard deviation describes how much the changes vary between people.

A paired comparison can be much more precise when before and after scores move together. It is not automatically more precise for every dataset. The differences must be independent across pairs, and small-sample inference assumes their population is approximately normal. The original before and after scores need not each have the same spread.

Improvement after practice alone does not prove practice caused it. Familiarity with the quiz, outside study, or the passage of time could also matter. Pairing handles a dependence structure; a randomized control comparison handles a causal question.

An idea to take with you

Student’s t compares signal with estimated noise

t = (estimated difference − null difference) / SE

For independent groups, SE = √(s²A/nA + s²B/nB). For n matched pairs, SE = sd/√n, where sd is the standard deviation of the paired differences.

A 95% interval is the estimated difference ± t* × SE. The critical value t* comes from the appropriate t distribution. It is larger than 1.96 for finite degrees of freedom, reflecting the extra uncertainty in estimating spread.

Return to the experiment ↑

A moment to think

A 95% interval for a mean difference runs from −1 to 7 points. What is a fair reading?

Use the direction and units when reporting a result: “new routine minus usual routine, three points,” followed by its interval. A signed difference is easier to interpret than a t statistic alone.

Words you met

TermMeaning
Student’s tA family of heavier-tailed distributions used when estimating variability.
Degrees of freedomThe independent information available to estimate variability in the procedure.
Welch’s procedureA comparison of independent means that allows unequal population variances.
Paired dataMeasurements linked by a person or another meaningful match.
Null differenceThe difference proposed by the null, often zero.
Neighbors

Written by June Kim. The conversational approach was inspired by Danielle Navarro’s Learning Statistics with R (CC BY-SA 4.0). The prose, experiments, and questions here were created for this book. This chapter’s text and illustrations are also shared under CC BY-SA 4.0.