← back to statistics

Errors, Power, and Multiple Testing

Statistics through experiments · 13 of 16 · CC BY-SA 4.0

A test can sound an alarm when nothing changed, or stay quiet when something did.

Imagine testing the same study procedure in many possible worlds. In one, the new lesson has no effect. In another, it adds five points on average. We get a fresh sample each time and apply exactly the same decision rule.

In an actual study we do not know which world we inhabit. In a simulation we can choose. That lets us count mistakes and learn what a significance threshold does.

One study can miss; many studies reveal the method

Each study compares two independent normal populations with known standard deviation 10 points. The test asks whether their means differ in either direction. You know the true difference because you set this simulated world.

0 studies. Detected effects: 0. Missed effects: 0.

The model’s power is 42.4%. Repeating a finite batch need not reproduce that percentage exactly.

Estimated differences across studies
Estimated differences across studies. 0 values; exact bin counts are available below.012Frequency-30540Difference (points)Run an experiment to add values.

Dashed line: true difference = 5.0.

Read the chart as a table
Estimated differences across studies
Value rangeCount
-30.0–-28.00
-28.0–-26.00
-26.0–-24.00
-24.0–-22.00
-22.0–-20.00
-20.0–-18.00
-18.0–-16.00
-16.0–-14.00
-14.0–-12.00
-12.0–-10.00
-10.0–-8.00
-8.0–-6.00
-6.0–-4.00
-4.0–-2.00
-2.0–0.00
0.0–2.00
2.0–4.00
4.0–6.00
6.0–8.00
8.0–10.00
10.0–12.00
12.0–14.00
14.0–16.00
16.0–18.00
18.0–20.00
20.0–22.00
22.0–24.00
24.0–26.00
26.0–28.00
28.0–30.00
30.0–32.00
32.0–34.00
34.0–36.00
36.0–38.00
38.0–40.00

Each range includes its lower end and excludes its upper end, except the last range, which includes both.

The calculation uses a two-sided z test because population spread is known in this model. Real mean comparisons usually estimate spread and use t methods. The latest 10,000 studies are retained. The chart covers −30 to 40; any rarer estimate outside that range still counts in the results above.

What if we test many questions?

At threshold 0.05, the chance of at least one false positive across 1 independent tests is 1 − (1 − 0.05)1 = 5.0%. Independence and true nulls are assumptions of this calculation.

A simple planned family-wise correction tests each of m questions at α/m (Bonferroni). This controls the probability of any false positive at most α for valid individual tests, even when tests are dependent; it can reduce power.

Two kinds of mistake

Start with a true difference of zero. Whenever the procedure rejects the no-difference null, it makes a Type I error, or false positive. The preselected threshold α controls that rejection probability under the null when the procedure’s assumptions hold. It is a property of repeated use, not the probability that this particular conclusion is false.

Now choose a real difference of five points. A test that fails to reject makes a Type II error for that alternative. Its probability is called β (“beta”). Power, 1 − β, is the probability of rejecting the null when that specified effect is real.

This activity assumes independent normal measurements with a known population standard deviation of ten points in each group. The two-sided z test therefore has an exact normal reference. Real studies commonly estimate variability and use t procedures; the same ideas about errors and power apply, but their calculations differ.

A moment to think

A planned study has 80% power to detect a five-point effect. What does that mean?

Make a small signal easier to hear

Hold the true effect fixed and increase the sample size. The estimates become less scattered, so more clear the rejection threshold. Larger effects are also easier to detect. Greater measurement variability makes detection harder.

Lowering α asks for stronger evidence before rejecting. With everything else fixed, this reduces false positives but also reduces power. There is no universal threshold that removes both kinds of mistake.

Plan around the smallest effect that would change a decision, not the largest effect you hope to see. Estimate plausible variability, allow for missing outcomes, and choose a sample size before collecting results. A study can have enough power to detect a trivial effect, yet answer the practical question poorly.

What if we keep looking?

Open the multiple-tests panel. With twenty independent tests of true null hypotheses, each at α = 0.05, the probability of at least one false positive is about 64%. None of the individual tests changed; we gave chance more opportunities to produce a headline.

A simple Bonferroni correction tests each of m prespecified hypotheses at α/m. It controls the probability of any false rejection at no more than α even when the tests are dependent, provided the individual p-values are valid. It can be conservative. Other methods answer related questions, such as controlling the expected fraction of false discoveries.

Repeatedly checking one accumulating dataset and stopping at the first small p-value creates a related problem. A fixed-sample test assumes its planned stopping rule. Sequential designs can allow interim looks, but require methods designed for them. Exploratory analysis remains useful; label it and seek new data for confirmation.

An idea to take with you

Power describes a procedure in a specified world

Type I error: reject a true null
Power = P(reject the null | specified real effect)
Type II error probability = 1 − power

Under m independent true-null tests each rejecting with probability α, P(at least one false positive) = 1 − (1 − α)m. Independence is needed for this equality; it is not needed for the Bonferroni bound.

Return to the experiment ↑

A moment to think

A small study reports p = 0.20 and a wide interval. What should we conclude?

After a study, the effect estimate and confidence interval usually communicate uncertainty more directly than plugging the observed effect into a new power calculation. Power earns its place most clearly when deciding what study to run.

Words you met

TermMeaning
Type I errorRejecting the null when it is true.
Type II errorFailing to reject the null under a specified alternative.
PowerThe probability of rejecting under a specified alternative.
MultiplicityHaving several opportunities to make a false discovery.
Practical importanceWhether the size of an effect matters for the decision at hand.
Neighbors

Written by June Kim. The conversational approach was inspired by Danielle Navarro’s Learning Statistics with R (CC BY-SA 4.0). The prose, experiments, and questions here were created for this book. This chapter’s text and illustrations are also shared under CC BY-SA 4.0.