Comparing Proportions
Statistics through experiments · 11 of 16 · CC BY-SA 4.0
A difference between two percentages is an estimate. Shuffling helps us ask how surprising it is.
A website offers two versions of a sign-up page. Each visitor sees either A or B, assigned at random, and either signs up or does not. We want to know whether B changes the probability of signing up.
In this simulator you can see what a real experimenter cannot: the probabilities that generate the outcomes. Version A has a 30% sign-up probability. Choose B’s advantage, draw a study, and see how closely the observed difference follows it.
Two versions of a sign-up page
Each visitor receives exactly one randomly assigned version, A or B; the outcome is whether they sign up. This simulated world has sign-up probability 30% for A. You control B’s true difference.
No study yet. Predict whether one study’s observed difference will equal the true difference.
For the randomization test, keep every observed sign-up and non-sign-up fixed. Shuffle them between equally sized groups, as if assignment had no effect on anyone’s outcome. Count shuffled differences at least as far from zero as the observed difference, in either direction.
Read the chart as a table
| Value range | Count |
|---|---|
| -100.0–-95.0 | 0 |
| -95.0–-90.0 | 0 |
| -90.0–-85.0 | 0 |
| -85.0–-80.0 | 0 |
| -80.0–-75.0 | 0 |
| -75.0–-70.0 | 0 |
| -70.0–-65.0 | 0 |
| -65.0–-60.0 | 0 |
| -60.0–-55.0 | 0 |
| -55.0–-50.0 | 0 |
| -50.0–-45.0 | 0 |
| -45.0–-40.0 | 0 |
| -40.0–-35.0 | 0 |
| -35.0–-30.0 | 0 |
| -30.0–-25.0 | 0 |
| -25.0–-20.0 | 0 |
| -20.0–-15.0 | 0 |
| -15.0–-10.0 | 0 |
| -10.0–-5.0 | 0 |
| -5.0–0.0 | 0 |
| 0.0–5.0 | 0 |
| 5.0–10.0 | 0 |
| 10.0–15.0 | 0 |
| 15.0–20.0 | 0 |
| 20.0–25.0 | 0 |
| 25.0–30.0 | 0 |
| 30.0–35.0 | 0 |
| 35.0–40.0 | 0 |
| 40.0–45.0 | 0 |
| 45.0–50.0 | 0 |
| 50.0–55.0 | 0 |
| 55.0–60.0 | 0 |
| 60.0–65.0 | 0 |
| 65.0–70.0 | 0 |
| 70.0–75.0 | 0 |
| 75.0–80.0 | 0 |
| 80.0–85.0 | 0 |
| 85.0–90.0 | 0 |
| 90.0–95.0 | 0 |
| 95.0–100.0 | 0 |
Each range includes its lower end and excludes its upper end, except the last range, which includes both.
0 shuffled assignments.
At most the latest 10,000 shuffles are retained. The Monte Carlo calculation is (extreme shuffles + 1) / (shuffles + 1), so finite simulation never claims a zero p-value. A new study or setting clears the old reference distribution.
Keep three quantities separate
The population proportion is the underlying probability of success for a group. The sample proportion is successes divided by observed trials. A hat marks the estimate: p̂ is read “p-hat.” The effect estimate here is p̂B − p̂A.
If A gets 30% and B gets 40%, the difference is ten percentage points. The relative increase is ten divided by thirty, about 33%. Both describe the same change, but they answer different questions. A report should say which it uses.
For the activity’s model, visitors are independent, each appears once, and the success probability stays constant within each group. In practice, repeat visits, shared households, missing outcomes, or one visitor seeing both versions can require a different design or analysis.
A moment to think
Build a world where the labels do not matter
Under a sharp null hypothesis of no treatment effect for any visitor, each recorded response would be the same under A or B. We can keep the responses and repeatedly reassign the group labels, preserving the original group sizes. Every shuffle gives a difference that random assignment alone could produce.
That collection is a randomization distribution. It centers near zero. We compare our observed difference with it, counting shuffled differences at least as far from zero in either direction. This is a two-sided test: the alternative allows B to help or hurt.
Choose the direction before seeing the data. Choosing a favorable one-sided test afterward changes the error rate. Our simulated p-value includes a small correction: add one to both the extreme count and the number of shuffles. A finite simulation then cannot misleadingly report zero.
How much might the rate change?
The p-value asks about compatibility with the null. A confidence interval asks about the precision of the estimated difference. The activity shows a large-sample interval only when each group has at least ten successes and ten failures. That is a screening rule for this approximation, not a guarantee that any study passing it is sound.
Its standard error uses the two observed proportions separately. A normal-based test of equal proportions often instead pools them under the null. Our shuffle test and approximate interval also use different procedures, so their conclusions need not match exactly at a cutoff. For sparse counts, use an appropriate exact or simulation-based method rather than forcing this approximation.
An idea to take with you
Difference, uncertainty, and a null comparison
d̂ = p̂B − p̂A
SE ≈ √[p̂A(1 − p̂A)/nA + p̂B(1 − p̂B)/nB]
Approximate 95% interval: d̂ ± 1.96 SE
Each group contributes uncertainty. Independent contributions add as variances inside the square root. The interval gives a range of differences compatible with the data under the approximation; the shuffle p-value compares the observed difference with a specified no-effect model.
Return to the experiment ↑A moment to think
Try setting the true advantage to zero and generating several studies. Some look more impressive than others. That variation is why a single low p-value is evidence within a model, not a certificate of truth.
Words you met
| Term | Meaning |
|---|---|
| Sample proportion | The fraction of observed trials with the outcome of interest. |
| Percentage point | One hundredth on the probability scale; used for differences between percentages. |
| Randomization distribution | Statistics from possible treatment assignments under a specified null. |
| Two-sided test | A comparison that allows departures in either direction. |
| Effect size | The magnitude of a difference on a meaningful scale. |
Neighbors
- Continue with formulas and examples in Inference for Proportions.
- For independent yes-or-no trials, see the binomial model in Important Distributions.
Written by June Kim. The conversational approach was inspired by Danielle Navarro’s Learning Statistics with R (CC BY-SA 4.0). The prose, experiments, and questions here were created for this book. This chapter’s text and illustrations are also shared under CC BY-SA 4.0.