Experimental Design
Statistics through experiments · 10 of 16 · CC BY-SA 4.0
The most useful statistical decision often happens before anyone collects a number.
A school has a new lesson and wants to know whether it improves learning. Twelve teachers volunteer, each with a class of twenty students. Half the classes will use the new lesson; half will continue with the usual lesson.
We could collect 240 test scores and calculate a difference. But first we need to decide what that difference could mean. Try designing the study below. Each choice changes the strength of the comparison.
Choose the design before seeing the scores
Twelve classrooms have 20 students each. Teachers will deliver either a new lesson or the usual lesson to their entire class for four weeks. Your question: does the new lesson change the mean score on a common final quiz?
Who gets the lesson?
The new lesson is a treatment. The usual lesson is the control: a comparison that tells us what might happen without the change. A control need not mean doing nothing. Here it means continuing ordinary teaching.
If teachers choose their own group, enthusiasm or prior achievement could differ between groups. Such a confounder affects the outcome and is entangled with the treatment. We may then mistake a difference between teachers or students for a lesson effect.
Random assignment gives the treatment allocation a known chance mechanism. It helps separate treatment from pre-existing differences on average across possible assignments. It does not guarantee perfectly matched groups in any one study.
Random sampling answers a different question: who enters the study? Random assignment helps a causal comparison among participants; representative sampling supports extending results to a population. Volunteer classrooms can support a randomized experiment without representing every school.
A moment to think
Count the independent opportunities to learn
Everyone in a classroom receives the same teacher-delivered lesson. The experimental unit is therefore a classroom: it is the unit independently assigned to treatment. Twenty students sharing a classroom are not twenty independent treatment assignments.
These groups are called clusters. Treating all 240 scores as independent can make uncertainty look too small. One simple analysis compares the twelve classroom means; a more detailed analysis models the clustering explicitly. Collecting more scores within one classroom helps describe that classroom, but adding independent classrooms supplies new treatment replications.
With only twelve classrooms, we may want similar classrooms on both sides. We can form six pairs using a pre-study measure, then randomly assign one classroom in each pair to each lesson. This is blocking. The analysis should respect those pairs, for example by examining six differences in classroom means.
Decide what success will mean
Choose a primary outcome before seeing results: the common quiz score after four weeks, for example. Decide how long to run the study, how to handle missing scores, and which exclusions are allowed. Recording this plan in advance is called preregistration. It makes the intended test distinguishable from later exploration.
Blinding hides group identity when that helps prevent expectations from influencing a measurement. Teachers will know which lesson they teach, but someone grading coded quiz papers may not need to know. Explain participation clearly, obtain appropriate permission, and protect students’ records.
An idea to take with you
Design determines what a difference can mean
Recruit participants → choose the experimental unit → randomize within the planned design → measure the prespecified outcome → analyze at the level of independent assignment.
Randomization, control, and replication support a fair comparison. Blocking can reduce unwanted variation. Blinding can reduce measurement bias. None replaces careful implementation or honest reporting.
Return to the experiment ↑A moment to think
A well-designed experiment can still give an uncertain answer. That is useful information: it tells us what the evidence can support. Next we will compare groups while keeping both the estimated difference and its uncertainty visible.
Words you met
| Term | Meaning |
|---|---|
| Treatment and control | The condition being studied and its planned comparison. |
| Confounder | A variable entangled with the explanatory variable that also influences the outcome. |
| Experimental unit | The smallest unit independently assigned to a treatment. |
| Blocking | Grouping similar units before randomizing within groups. |
| Preregistration | Recording the study’s questions and analysis decisions before examining outcomes. |
Neighbors
- Revisit sampling and assignment, or explore Introduction to Data.
- For the probability language behind independent events, see Conditional Probability.
Written by June Kim. The conversational approach was inspired by Danielle Navarro’s Learning Statistics with R (CC BY-SA 4.0). The prose, experiments, and questions here were created for this book. This chapter’s text and illustrations are also shared under CC BY-SA 4.0.