← back to statistics

Hypothesis Tests and p-Values

Statistics through experiments · 9 of 16 · CC BY-SA 4.0

An unusual result needs a comparison: unusual under which explanation?

Suppose you suspect a coin favors heads. Before flipping it, you decide to flip exactly twenty times and look for unusually many heads. You get fifteen.

A fair coin can do that. The question is how often it would produce a result at least as far in the direction we were looking for. Try repeating the twenty-flip experiment with a fair coin below.

A world where the coin really is fair

Our question, chosen before looking at the data: does this coin favor heads? Every simulated experiment flips a fair coin 20 independent times.

No fair-coin experiments recorded yet.

Outcomes under the fair-coin model
Outcomes under the fair-coin model. 0 values; exact bin counts are available below.012Frequency01020Heads in 20 flipsRun an experiment to add values.

Dashed line: observed head count = 15.0.

Read the chart as a table
Outcomes under the fair-coin model
Value rangeCount
-0.5–0.50
0.5–1.50
1.5–2.50
2.5–3.50
3.5–4.50
4.5–5.50
5.5–6.50
6.5–7.50
7.5–8.50
8.5–9.50
9.5–10.50
10.5–11.50
11.5–12.50
12.5–13.50
13.5–14.50
14.5–15.50
15.5–16.50
16.5–17.50
17.5–18.50
18.5–19.50
19.5–20.50

Each range includes its lower end and excludes its upper end, except the last range, which includes both.

Amber bars are the outcomes counted: 15 heads or more. Changing the observed count reuses the same simulated fair-coin outcomes.

Check the exact probability

For a fair coin, the probability of 15 or more heads in 20 flips is 2.07%. The simulation estimates this value. If no simulated experiment reached the tail, the probability still need not be zero.

The latest 10,000 simulated experiments are retained. This is a one-sided test for an excess of heads; a question about bias in either direction would need both tails.

Give the explanation a fair trial

The null model is our comparison model: independent flips, each with a 50% chance of heads. The test statistic is the number of heads in twenty flips. Each simulated experiment gives one value of that statistic.

We count fifteen, sixteen, and every larger result through twenty. A result of eighteen heads would be even more evidence in the direction we specified, so it belongs in the comparison too.

The fraction of simulated counts at least fifteen estimates a p-value. The exact probability here is about 0.0207, or 2.07%. Simulation gets noisier when the events we count are rare. Seeing zero qualifying results in a short simulation does not make the true probability zero.

A moment to think

Why do we count sixteen heads as well as exactly fifteen?

What should we conclude?

Fifteen or more heads is uncommon under this null model. That gives us evidence against it. It does not tell us the probability that the coin is fair: we calculated the probability of results assuming fairness.

Sometimes a study needs a decision rule. A significance threshold, often written α (“alpha”), is chosen before seeing the data. For example, reject the null when p ≤ 0.05.

Even a fair coin sometimes crosses that threshold. Such a rejection is a false positive. For a valid test, the threshold limits the long-run false-positive rate under the null; with discrete counts, the actual rate can be lower than the threshold.

The activity lets you change the hypothetical observation to understand the comparison. In an actual study, changing the question, stopping rule, or reported result after looking at the evidence can invalidate the advertised error rate.

You can now read this

A p-value is a conditional comparison

p-value = P(X ≥ 15 | fair, independent flips)

X counts heads in twenty flips. ≥ means “at least.” The vertical bar means “given”: calculate the probability under this model.

More generally, a p-value is the probability, under the specified null model, of a test statistic at least as extreme as the observed one according to the test’s comparison rule.

It is not the probability that the null is true, and it is not the size or practical importance of an effect.

Return to the experiment ↑

A moment to think

A properly designed A/B test reports p = 0.02. Which reading is justified?

Keep the question in view

A large p-value does not prove no effect. A small study may simply be too noisy to distinguish useful differences. A small p-value does not guarantee a useful effect either.

You now have a connected set of questions: What did we measure? Whom did we sample? How variable is our estimate? What would happen under a proposed explanation? Next, we will design a classroom experiment and use those questions to compare groups.

Words you met

TermMeaning
Null modelThe probability model used as the reference for a test.
Test statisticA numerical summary of the evidence relevant to the test.
p-valueA tail probability calculated under the null model.
Significance thresholdA preselected cutoff for the testing decision.
False positiveRejecting the null when it is true.
Neighbors

Written by June Kim. The conversational approach was inspired by Danielle Navarro’s Learning Statistics with R (CC BY-SA 4.0). The prose, experiments, and questions here were created for this book. This chapter’s text and illustrations are also shared under CC BY-SA 4.0.