Comparing Means with Student’s t
Statistics through experiments · 12 of 16 · CC BY-SA 4.0
Two groups of people and two measurements of the same people need different comparisons.
One group tries a new practice routine and another uses the usual routine. Their average scores differ by three points. Is three a clear signal, or small compared with the uncertainty in our estimate?
There is another way to collect scores: measure the same people before and after practice. A person who starts high may finish high under either routine. Keeping track of who is paired with whom can remove much of that person-to-person variation from the comparison.
The activity uses constructed measurements so you can isolate these ideas. First compare independent groups. Then switch to paired measurements and inspect the individual changes.
The design chooses the standard error
These are constructed quiz-score examples. Choose two independent groups of twelve people, or twelve people measured twice. The paired example has large differences between people but much smaller changes within a person.
Group A mean: 50.0. Group B mean: 53.0. Difference: 3.0 points.
Welch t: estimated difference 3.00, SE 1.33, df 21.9. 95% interval [0.25, 5.75]. Two-sided p-value 0.034 for a zero mean difference.
Read the measurements
| Row (not a pair) | A | B |
|---|---|---|
| 1 | 48 | 52 |
| 2 | 52 | 54 |
| 3 | 47 | 49 |
| 4 | 55 | 57 |
| 5 | 50 | 55 |
| 6 | 49 | 51 |
| 7 | 53 | 53 |
| 8 | 46 | 50 |
| 9 | 51 | 58 |
| 10 | 54 | 56 |
| 11 | 45 | 47 |
| 12 | 50 | 54 |
The paired method assumes independent people and approximately normal differences for this small sample. Welch’s method assumes independent groups and observations, with approximately normal populations for small samples; it allows unequal variances. Inspect data and design before using either method. Before-and-after changes alone do not establish causation.
Why a new distribution?
Earlier we used a normal model with known population standard deviation. Usually we do not know that spread: we estimate it from the sample. The standard error is now uncertain too, especially with few observations.
Student’s t distribution allows for that extra uncertainty. It is symmetric like the standard normal distribution, but has heavier tails. Its shape depends on degrees of freedom: the amount of independent information left after estimation. More degrees of freedom bring it closer to the normal distribution.
For a single sample mean, the statistic is t = (x̄ − μ₀)/(s/√n), where μ₀ is the proposed population mean. With independent, normally distributed observations it follows a t distribution with n − 1 degrees of freedom under the null. For larger samples the method can tolerate some non-normality, but extreme outliers and strong skew deserve attention.
Two independent groups
For separate groups, each sample mean adds uncertainty to the difference. Welch’s t procedure estimates those contributions separately and does not require equal population variances. Its degrees of freedom can be fractional; they are calculated from both groups’ sizes and variances.
Independent observations within and between groups are essential. Small samples need approximately normal populations for a reliable t approximation. Inspect the measurements and the design; a formula cannot detect classroom clustering or repair a biased sample.
A moment to think
Let each person be their own comparison
In a paired t procedure, subtract each before score from its matching after score. Now we have one sample of differences. Its mean is the estimated change; its standard deviation describes how much the changes vary between people.
A paired comparison can be much more precise when before and after scores move together. It is not automatically more precise for every dataset. The differences must be independent across pairs, and small-sample inference assumes their population is approximately normal. The original before and after scores need not each have the same spread.
Improvement after practice alone does not prove practice caused it. Familiarity with the quiz, outside study, or the passage of time could also matter. Pairing handles a dependence structure; a randomized control comparison handles a causal question.
An idea to take with you
Student’s t compares signal with estimated noise
t = (estimated difference − null difference) / SE
For independent groups, SE = √(s²A/nA + s²B/nB). For n matched pairs, SE = sd/√n, where sd is the standard deviation of the paired differences.
A 95% interval is the estimated difference ± t* × SE. The critical value t* comes from the appropriate t distribution. It is larger than 1.96 for finite degrees of freedom, reflecting the extra uncertainty in estimating spread.
Return to the experiment ↑A moment to think
Use the direction and units when reporting a result: “new routine minus usual routine, three points,” followed by its interval. A signed difference is easier to interpret than a t statistic alone.
Words you met
| Term | Meaning |
|---|---|
| Student’s t | A family of heavier-tailed distributions used when estimating variability. |
| Degrees of freedom | The independent information available to estimate variability in the procedure. |
| Welch’s procedure | A comparison of independent means that allows unequal population variances. |
| Paired data | Measurements linked by a person or another meaningful match. |
| Null difference | The difference proposed by the null, often zero. |
Neighbors
- Worked procedures: Inference for Means.
- Return to sampling distributions and the CLT, or explore the Central Limit Theorem in the probability textbook.
Written by June Kim. The conversational approach was inspired by Danielle Navarro’s Learning Statistics with R (CC BY-SA 4.0). The prose, experiments, and questions here were created for this book. This chapter’s text and illustrations are also shared under CC BY-SA 4.0.