Center and Spread
Statistics through experiments · 3 of 16 · CC BY-SA 4.0
An average fits in a sentence. The people behind it rarely do.
Five people earn $40,000, $50,000, $60,000, $70,000, and $80,000 a year. If they pooled their salaries and shared the money equally, each would have $60,000. That equal share is their mean, one kind of average.
Now a person earning a million dollars joins them. Nobody else gets a raise. Before pressing the button, predict what happens to the average.
Five people, five salaries
Change a salary with its slider. The plot and summaries follow your changes. All amounts are fictional annual salaries.
Your mean: $300,000 shared across 5 people = $60,000 each.
Same mean, different spread
Try these two groups. Both have a mean and median of $60,000.
Population standard deviation for these 5 people: $14,142. Larger means more spread around the mean.
What did we measure?
Each person gives us an observation. Salary is a variable: something we record that can differ from person to person. We could also record their occupation.
Salary is numerical; differences and averages have meaning. Occupation is categorical: teacher, cook, engineer. We can count how many people belong to each category, but an “average occupation” makes no sense. A number used as a label, such as a postal code, is categorical too.
Two ways to find the middle
The mean listens to every dollar. That is useful when you care about a total divided among people, but an unusually large value can pull it far from most observations.
The median asks a different question: after sorting the salaries, what is in the middle? With five people, it is the third salary. With six, we average the third and fourth. The newcomer moves our median from $60,000 to $65,000.
Neither summary is dishonest. They answer different questions. To describe what a middle-earning person receives, the median is often more helpful. To work out a total payroll from a head count, use the mean.
A moment to think
The same center can hide different lives
Open Same mean, different spread in the experiment. Both groups average $60,000, but one bunches near that amount and the other stretches from $20,000 to $100,000. Spread describes how far apart values are.
One measure of spread is the standard deviation. It summarizes distance from the mean using squared differences. Its units are the original units: dollars here. You can meet the calculation later; for now, notice how it changes while the mean stays put.
You can now write this
The mean is an equal share
You added the salaries, counted the people, and divided. Here is that whole operation in one line.
x̄ = (x₁ + ⋯ + xₙ) / n
x̄, “x-bar,” names the mean. Each x is one salary; the little number identifies whose. n counts the people. The dots mean “keep adding the remaining salaries.”
An even shorter way to write the addition
x̄ = (1/n) ∑ᵢ₌₁ⁿ xᵢ. The symbol ∑, “sum,” says to add all the values, from person 1 through person n. It is the same equal-sharing calculation.
A mean preserves the total per person. It does not preserve the shape of the group.
Return to the experiment ↑A moment to think
Words you met
| Term | Meaning |
|---|---|
| Observation | One recorded case, such as a person. |
| Variable | An attribute recorded for each case. |
| Mean | The sum of numerical values divided by their count. |
| Median | The middle of the ordered values; average the two middle values for an even count. |
| Spread | How much the values vary. |
Neighbors
- For plots and the standard-deviation calculation, continue into Summarizing Data. Next, we’ll organize a student survey into tables, histograms, and summaries before asking about people we have not measured.
- A mean calculated from data has a mathematical cousin: an average weighted by probabilities. Meet it in the probability textbook’s Expected Value and Variance.
Written by June Kim. The conversational approach was inspired by Danielle Navarro’s Learning Statistics with R (CC BY-SA 4.0). The prose, experiments, and questions here were created for this book. This chapter’s text and illustrations are also shared under CC BY-SA 4.0.