← back to statistics

Comparing Proportions

Statistics through experiments · 11 of 16 · CC BY-SA 4.0

A difference between two percentages is an estimate. Shuffling helps us ask how surprising it is.

A website offers two versions of a sign-up page. Each visitor sees either A or B, assigned at random, and either signs up or does not. We want to know whether B changes the probability of signing up.

In this simulator you can see what a real experimenter cannot: the probabilities that generate the outcomes. Version A has a 30% sign-up probability. Choose B’s advantage, draw a study, and see how closely the observed difference follows it.

Two versions of a sign-up page

Each visitor receives exactly one randomly assigned version, A or B; the outcome is whether they sign up. This simulated world has sign-up probability 30% for A. You control B’s true difference.

No study yet. Predict whether one study’s observed difference will equal the true difference.

For the randomization test, keep every observed sign-up and non-sign-up fixed. Shuffle them between equally sized groups, as if assignment had no effect on anyone’s outcome. Count shuffled differences at least as far from zero as the observed difference, in either direction.

Differences after shuffling under no effect
Differences after shuffling under no effect. 0 values; exact bin counts are available below.012Frequency-1000100B − A (percentage points)Run an experiment to add values.
Read the chart as a table
Differences after shuffling under no effect
Value rangeCount
-100.0–-95.00
-95.0–-90.00
-90.0–-85.00
-85.0–-80.00
-80.0–-75.00
-75.0–-70.00
-70.0–-65.00
-65.0–-60.00
-60.0–-55.00
-55.0–-50.00
-50.0–-45.00
-45.0–-40.00
-40.0–-35.00
-35.0–-30.00
-30.0–-25.00
-25.0–-20.00
-20.0–-15.00
-15.0–-10.00
-10.0–-5.00
-5.0–0.00
0.0–5.00
5.0–10.00
10.0–15.00
15.0–20.00
20.0–25.00
25.0–30.00
30.0–35.00
35.0–40.00
40.0–45.00
45.0–50.00
50.0–55.00
55.0–60.00
60.0–65.00
65.0–70.00
70.0–75.00
75.0–80.00
80.0–85.00
85.0–90.00
90.0–95.00
95.0–100.00

Each range includes its lower end and excludes its upper end, except the last range, which includes both.

0 shuffled assignments.

At most the latest 10,000 shuffles are retained. The Monte Carlo calculation is (extreme shuffles + 1) / (shuffles + 1), so finite simulation never claims a zero p-value. A new study or setting clears the old reference distribution.

Keep three quantities separate

The population proportion is the underlying probability of success for a group. The sample proportion is successes divided by observed trials. A hat marks the estimate: p̂ is read “p-hat.” The effect estimate here is p̂B − p̂A.

If A gets 30% and B gets 40%, the difference is ten percentage points. The relative increase is ten divided by thirty, about 33%. Both describe the same change, but they answer different questions. A report should say which it uses.

For the activity’s model, visitors are independent, each appears once, and the success probability stays constant within each group. In practice, repeat visits, shared households, missing outcomes, or one visitor seeing both versions can require a different design or analysis.

A moment to think

A sign-up rate rises from 30% to 40%. What is the absolute difference?

Build a world where the labels do not matter

Under a sharp null hypothesis of no treatment effect for any visitor, each recorded response would be the same under A or B. We can keep the responses and repeatedly reassign the group labels, preserving the original group sizes. Every shuffle gives a difference that random assignment alone could produce.

That collection is a randomization distribution. It centers near zero. We compare our observed difference with it, counting shuffled differences at least as far from zero in either direction. This is a two-sided test: the alternative allows B to help or hurt.

Choose the direction before seeing the data. Choosing a favorable one-sided test afterward changes the error rate. Our simulated p-value includes a small correction: add one to both the extreme count and the number of shuffles. A finite simulation then cannot misleadingly report zero.

How much might the rate change?

The p-value asks about compatibility with the null. A confidence interval asks about the precision of the estimated difference. The activity shows a large-sample interval only when each group has at least ten successes and ten failures. That is a screening rule for this approximation, not a guarantee that any study passing it is sound.

Its standard error uses the two observed proportions separately. A normal-based test of equal proportions often instead pools them under the null. Our shuffle test and approximate interval also use different procedures, so their conclusions need not match exactly at a cutoff. For sparse counts, use an appropriate exact or simulation-based method rather than forcing this approximation.

An idea to take with you

Difference, uncertainty, and a null comparison

d̂ = p̂B − p̂A
SE ≈ √[p̂A(1 − p̂A)/nA + p̂B(1 − p̂B)/nB]
Approximate 95% interval: d̂ ± 1.96 SE

Each group contributes uncertainty. Independent contributions add as variances inside the square root. The interval gives a range of differences compatible with the data under the approximation; the shuffle p-value compares the observed difference with a specified no-effect model.

Return to the experiment ↑

A moment to think

A large experiment finds an increase of 0.2 percentage points with p = 0.001. What else matters?

Try setting the true advantage to zero and generating several studies. Some look more impressive than others. That variation is why a single low p-value is evidence within a model, not a certificate of truth.

Words you met

TermMeaning
Sample proportionThe fraction of observed trials with the outcome of interest.
Percentage pointOne hundredth on the probability scale; used for differences between percentages.
Randomization distributionStatistics from possible treatment assignments under a specified null.
Two-sided testA comparison that allows departures in either direction.
Effect sizeThe magnitude of a difference on a meaningful scale.
Neighbors

Written by June Kim. The conversational approach was inspired by Danielle Navarro’s Learning Statistics with R (CC BY-SA 4.0). The prose, experiments, and questions here were created for this book. This chapter’s text and illustrations are also shared under CC BY-SA 4.0.