← back to statistics

Populations and Samples

Statistics through experiments · 5 of 16 · CC BY-SA 4.0

You cannot measure everyone. A carefully chosen handful can still tell you something.

A small town wants to know the average height of its 200 adult residents. Measuring everyone would settle it, but suppose you only have time for ten people.

The town is our population: the whole group we want to describe. The ten people we measure form a sample. Choose a sample below, then use its mean as your best guess before revealing the town.

A fictional town of 200 adults

Sample without replacement: nobody is measured twice in a single sample.

Each dot is a resident. Filled dots mark the town-wide sample.

No sample yet. What do you think the town’s average height will be?

One number to find, another to calculate

The town has a definite average height before we measure anyone. That population mean is a parameter. A parameter describes the population, even when we do not know its value.

The mean of the people we actually sampled is a statistic. A statistic is calculated from observed data. When we use it to guess the parameter, it acts as an estimate.

Reveal the town, then draw a few more samples of the same size. The statistic moves. The parameter stays put. That is not a calculation going wrong; different people have different heights.

A moment to think

A sample of ten residents averages 172 cm. The whole town averages 170 cm. Which is the parameter?

What does a larger handful buy us?

Compare several samples of five with several samples of 100. Larger random samples generally give steadier estimates. An individual small sample can still land closer to the truth than an individual large one.

The experiment samples without replacement: once someone is selected, they cannot appear again in that sample. At 200, you have measured the whole town. This is a census, and sampling uncertainty disappears for this population.

Even a census can have measurement errors or missing answers. Here the measurements are exact and everyone responds, so we can concentrate on the uncertainty caused by choosing a sample.

A relationship to remember

The estimate points toward a target

Population (μ) → random sample → estimate (x̄)

μ, “mu,” names the population mean. x̄, “x-bar,” is the sample mean you calculated. Drawing another sample can change x̄ without changing μ.

The arrow matters: an estimate can tell us about the population only when the way we obtained the sample supports that inference.

Return to the experiment ↑

A moment to think

A library samples 50 loans to estimate the mean loan length for all loans last year. What would change if it drew a fresh random sample from that same year?

Words you met

TermMeaning
PopulationThe complete group your question concerns.
SampleThe cases selected from that population.
ParameterA numerical property of the population.
StatisticA quantity calculated from sample data.
CensusMeasuring every member of the population.
Neighbors

Written by June Kim. The conversational approach was inspired by Danielle Navarro’s Learning Statistics with R (CC BY-SA 4.0). The prose, experiments, and questions here were created for this book. This chapter’s text and illustrations are also shared under CC BY-SA 4.0.