Hypothesis Tests and p-Values
Statistics through experiments · 9 of 16 · CC BY-SA 4.0
An unusual result needs a comparison: unusual under which explanation?
Suppose you suspect a coin favors heads. Before flipping it, you decide to flip exactly twenty times and look for unusually many heads. You get fifteen.
A fair coin can do that. The question is how often it would produce a result at least as far in the direction we were looking for. Try repeating the twenty-flip experiment with a fair coin below.
A world where the coin really is fair
Our question, chosen before looking at the data: does this coin favor heads? Every simulated experiment flips a fair coin 20 independent times.
No fair-coin experiments recorded yet.
Dashed line: observed head count = 15.0.
Read the chart as a table
| Value range | Count |
|---|---|
| -0.5–0.5 | 0 |
| 0.5–1.5 | 0 |
| 1.5–2.5 | 0 |
| 2.5–3.5 | 0 |
| 3.5–4.5 | 0 |
| 4.5–5.5 | 0 |
| 5.5–6.5 | 0 |
| 6.5–7.5 | 0 |
| 7.5–8.5 | 0 |
| 8.5–9.5 | 0 |
| 9.5–10.5 | 0 |
| 10.5–11.5 | 0 |
| 11.5–12.5 | 0 |
| 12.5–13.5 | 0 |
| 13.5–14.5 | 0 |
| 14.5–15.5 | 0 |
| 15.5–16.5 | 0 |
| 16.5–17.5 | 0 |
| 17.5–18.5 | 0 |
| 18.5–19.5 | 0 |
| 19.5–20.5 | 0 |
Each range includes its lower end and excludes its upper end, except the last range, which includes both.
Amber bars are the outcomes counted: 15 heads or more. Changing the observed count reuses the same simulated fair-coin outcomes.
Check the exact probability
For a fair coin, the probability of 15 or more heads in 20 flips is 2.07%. The simulation estimates this value. If no simulated experiment reached the tail, the probability still need not be zero.
The latest 10,000 simulated experiments are retained. This is a one-sided test for an excess of heads; a question about bias in either direction would need both tails.
Give the explanation a fair trial
The null model is our comparison model: independent flips, each with a 50% chance of heads. The test statistic is the number of heads in twenty flips. Each simulated experiment gives one value of that statistic.
We count fifteen, sixteen, and every larger result through twenty. A result of eighteen heads would be even more evidence in the direction we specified, so it belongs in the comparison too.
The fraction of simulated counts at least fifteen estimates a p-value. The exact probability here is about 0.0207, or 2.07%. Simulation gets noisier when the events we count are rare. Seeing zero qualifying results in a short simulation does not make the true probability zero.
A moment to think
What should we conclude?
Fifteen or more heads is uncommon under this null model. That gives us evidence against it. It does not tell us the probability that the coin is fair: we calculated the probability of results assuming fairness.
Sometimes a study needs a decision rule. A significance threshold, often written α (“alpha”), is chosen before seeing the data. For example, reject the null when p ≤ 0.05.
Even a fair coin sometimes crosses that threshold. Such a rejection is a false positive. For a valid test, the threshold limits the long-run false-positive rate under the null; with discrete counts, the actual rate can be lower than the threshold.
The activity lets you change the hypothetical observation to understand the comparison. In an actual study, changing the question, stopping rule, or reported result after looking at the evidence can invalidate the advertised error rate.
You can now read this
A p-value is a conditional comparison
p-value = P(X ≥ 15 | fair, independent flips)
X counts heads in twenty flips. ≥ means “at least.” The vertical bar means “given”: calculate the probability under this model.
More generally, a p-value is the probability, under the specified null model, of a test statistic at least as extreme as the observed one according to the test’s comparison rule.
It is not the probability that the null is true, and it is not the size or practical importance of an effect.
Return to the experiment ↑A moment to think
Keep the question in view
A large p-value does not prove no effect. A small study may simply be too noisy to distinguish useful differences. A small p-value does not guarantee a useful effect either.
You now have a connected set of questions: What did we measure? Whom did we sample? How variable is our estimate? What would happen under a proposed explanation? Next, we will design a classroom experiment and use those questions to compare groups.
Words you met
| Term | Meaning |
|---|---|
| Null model | The probability model used as the reference for a test. |
| Test statistic | A numerical summary of the evidence relevant to the test. |
| p-value | A tail probability calculated under the null model. |
| Significance threshold | A preselected cutoff for the testing decision. |
| False positive | Rejecting the null when it is true. |
Neighbors
- See Foundations for Inference and Inference for Proportions for tests beyond this coin. Or follow the OpenIntro guide from its first chapter.
- The probability textbook’s Important Distributions gives the binomial formula behind the exact tail probability. Conditional Probability develops what it means to calculate a probability given an assumption.
Written by June Kim. The conversational approach was inspired by Danielle Navarro’s Learning Statistics with R (CC BY-SA 4.0). The prose, experiments, and questions here were created for this book. This chapter’s text and illustrations are also shared under CC BY-SA 4.0.