Chi-Square Tests and ANOVA
Statistics through experiments · 14 of 16 · CC BY-SA 4.0
The kind of measurement tells us which differences to compare.
A club offers three meeting times. Do members favor one, or are the choices equally popular? Elsewhere, students are randomly assigned to three groups that use different practice routines. Do their mean scores differ? Both questions involve three groups, but one records counts and the other records numerical measurements.
Below are two constructed examples. They are controls for exploring a calculation, not new random studies. Change the counts in the first. Then move the group means and the within-group spread in the second.
Counts or measurements?
These constructed examples let you change a pattern while holding other features fixed. They are not repeated random samples.
Three categories: count the choices
Ninety independently sampled people choose one of three meeting times. The null model gives each option probability one third, so each expected count is 30.
| Option | Observed | Expected under null |
|---|---|---|
| A | 30 | 30 |
| B | 30 | 30 |
| C | 30 | 30 |
Chi-square statistic 0.00, df = 2. Right-tail p-value 1.000.
Three groups: compare numerical outcomes
Now three independent groups each provide six quiz scores. Compare their mean differences with the variation within groups.
| Group | Mean | Observed scores |
|---|---|---|
| 1 | 50 | 45, 47, 49, 51, 53, 55 |
| 2 | 50 | 45, 47, 49, 51, 53, 55 |
| 3 | 50 | 45, 47, 49, 51, 53, 55 |
ANOVA F = 0.00, df = (2, 15). Right-tail p-value 1.000.
The chi-square approximation requires independent cases and adequate expected counts (all are 30 here). Classical one-way ANOVA assumes independent observations, approximately normal errors, and equal population variances. A small omnibus p-value says at least one null restriction fails; it does not identify every pair that differs.
Counts: compare observed with expected
If ninety independent members each choose one of three equally popular meeting times, the expected counts are thirty, thirty, and thirty. Actual counts need not match exactly. A chi-square goodness-of-fit test measures how far they depart from these prespecified proportions.
For each category, subtract the expected count from the observed count, square the gap, and divide by the expected count. Add the contributions. Bigger gaps produce a larger statistic. Squaring keeps opposite departures from canceling.
With three categories and fixed probabilities, two counts can vary freely; the third is fixed by the total. That gives two degrees of freedom. More generally it is k − 1 for k categories when no probabilities are estimated from the data. Estimating parameters can change that count.
The usual chi-square approximation needs independent observations and adequate expected counts. Having every expected count at least five is a common introductory screening rule. Sparse tables may need exact or simulation methods. Do not apply a chi-square test to percentages without their underlying counts.
A moment to think
Two categorical variables
Suppose we record both year of study and preferred meeting time. A two-way table counts each combination. A chi-square test of independence compares the observed table with the counts expected if those two variables were independent.
For each cell, the expected count is its row total × its column total / the overall total. The statistic uses the same squared-gap rule; degrees of freedom are (rows − 1)(columns − 1). A small p-value supports an association. It does not tell us which variable caused the other or which cells explain all of the difference.
Means: compare between with within
Now try the numerical scores. Move the group means apart while holding within-group spread fixed. Then hold the means fixed and increase the spread. The same mean differences become less compelling when individual scores are much more variable.
Analysis of variance, or ANOVA, compares variation between group means with variation within groups. Its null says all population means are equal. Its alternative says at least one differs. The F statistic is a ratio of those two variation estimates.
The classical one-way procedure assumes independent observations, approximately normal errors within each group, and equal population variances. Equal sample sizes help with some departures, but do not remove the assumptions. Strong variance differences may call for Welch’s ANOVA; clustered measurements need an analysis that respects the clusters.
An idea to take with you
Choose the comparison that matches the data
Counts: χ² = Σ(observed − expected)² / expected
Means: F = between-group mean square / within-group mean square
Both compare a pattern with what a null model can explain. Chi-square adds standardized count discrepancies. ANOVA compares variation estimates; under its null, their scales are comparable. A large statistic can be evidence against the corresponding null.
Return to the experiment ↑A moment to think
Plot the groups and report differences with uncertainty. A single omnibus p-value is an invitation to understand the pattern, not a substitute for describing it.
Words you met
| Term | Meaning |
|---|---|
| Goodness of fit | Agreement between observed category counts and a specified distribution. |
| Expected count | A count predicted under the null model. |
| Test of independence | A test for association between categorical variables. |
| ANOVA | A comparison of group means using between-group and within-group variation. |
| Omnibus test | A test of a joint claim that does not identify every individual difference. |
Neighbors
- Find chi-square examples in Inference for Proportions and ANOVA in Inference for Means.
- Review multiple testing before making many pairwise comparisons.
Written by June Kim. The conversational approach was inspired by Danielle Navarro’s Learning Statistics with R (CC BY-SA 4.0). The prose, experiments, and questions here were created for this book. This chapter’s text and illustrations are also shared under CC BY-SA 4.0.