Reading a Statistical Study
Statistics through experiments · 16 of 16 · CC BY-SA 4.0
A useful conclusion keeps the design, the size of the effect, and the remaining uncertainty together.
You have met many statistical tools. Reading a study does not require using all of them at once. Start with a few ordinary questions: What was measured? Who took part? What was compared? How much did the groups differ, and how uncertain is that difference?
Let us return to the classroom study and practice making a claim small enough for its evidence. The following report is fictional. Its rounded numbers are supplied for interpretation, not generated by the earlier simulator.
A report lands on your desk
Twelve volunteer classrooms, each with twenty students, join a four-week study. Six classrooms are randomly assigned a new lesson and six the usual lesson. The primary outcome is the classroom’s mean score on a common quiz. Coded papers are graded without the grader knowing the lesson assignment.
The outcome, sample size, and analysis were recorded before results were examined. All 240 students completed the quiz. An independent-groups t analysis of the twelve classroom means estimates:
New lesson − usual lesson: +4 points
95% confidence interval: approximately [−1, 9] points
Two-sided p-value: approximately 0.10
The school had decided in advance that an improvement of at least three points would justify the extra preparation time.
First, identify the evidence
The measured variable is numerical: quiz score. The intervention is assigned to classrooms, so the independent treatment units are twelve classrooms. The analysis uses their means. Treating every pupil as an independent assignment would overstate what this design replicated.
Random assignment supports a causal comparison among these participating classrooms, provided the lessons were delivered as planned, classrooms did not interfere with one another, and the outcome model is adequate. Volunteering leaves a separate question about generalization to other schools, teachers, and students.
A moment to think
Read the range before choosing a headline
The point estimate favors the new lesson by four points. The interval includes a small disadvantage, zero difference, and benefits larger than the three-point threshold. The data remain too imprecise to distinguish these possibilities well.
At a preselected α of 0.05, this test does not reject the zero-difference null. That is not evidence that the lessons are equivalent. It also does not erase the positive point estimate. Both the estimate and its uncertainty belong in the report.
A p-value around 0.10 means that, under the specified zero-difference model and assumptions, test statistics at least this extreme in either direction have probability around 10%. It does not mean a 10% chance that the null is true or a 90% chance that the lesson works.
A moment to think
Write the conclusion yourself
Before opening the example below, try writing two sentences. Include the comparison, estimated difference, uncertainty, and the main limit on what you can claim.
Compare with a possible report
In this randomized study of twelve volunteer classrooms, the new lesson increased the estimated mean quiz score by four points relative to the usual lesson, with a 95% confidence interval from approximately −1 to 9 points and a two-sided p-value of approximately 0.10. The results remain compatible with both little benefit and an educationally worthwhile improvement; they do not establish equivalence, and generalization beyond these classrooms requires additional evidence.
What would help next?
More independent classrooms could narrow uncertainty if the study remains well run. A follow-up could use a prespecified blocking scheme, recruit across relevant settings, and track whether the lesson was delivered consistently. Planning power around the worthwhile three-point effect would help choose its size.
Missing scores would deserve attention too. If missingness differed by group or achievement, analyzing only complete responses could change the comparison. Record what is missing, why, and how the planned analysis addresses it. A larger dataset cannot automatically repair a missing-data mechanism.
An idea to take with you
A statistical claim has a scope
Claim = comparison + effect estimate + uncertainty + design + scope
This is a reporting guide, not an algebraic identity. Each part keeps the others honest. A precise estimate from a biased sample, a small p-value for a trivial effect, and a promising estimate with a wide interval tell different stories.
You can now ask what a study learned without demanding either certainty or a verdict of “nothing happened.”
Return to the experiment ↑A moment to think
The introductory route ends here. The OpenIntro guide below adds worked calculations and runnable examples. Return to an experiment whenever a formula starts to feel detached from the question it answers.
Words you met
| Term | Meaning |
|---|---|
| Point estimate | One best estimate of the quantity of interest. |
| Practical threshold | An effect size large enough to matter for a specified decision. |
| Generalization | Extending a conclusion beyond the participants and setting studied. |
| Equivalence | Similarity within a prespecified meaningful margin, requiring an appropriate analysis. |
| Replication | A new opportunity to examine a result with further independently collected evidence. |
Neighbors
- Continue with the nine-chapter OpenIntro reading guide.
- Keep the wider process in view with Scientific Method, or strengthen the mathematical foundations with Introduction to Probability.
Written by June Kim. The conversational approach was inspired by Danielle Navarro’s Learning Statistics with R (CC BY-SA 4.0). The prose, experiments, and questions here were created for this book. This chapter’s text and illustrations are also shared under CC BY-SA 4.0.