Descriptive Statistics and Student Surveys
Statistics through experiments · 4 of 16 · CC BY-SA 4.0
A list of answers becomes useful when we know what each answer measures.
You want to know how students get to school and how long the journey takes. Asking everyone would be ideal, but you have time to ask only ten people. What could those ten answers tell you?
Below is an invented school of sixty students. Draw a random sample: every group of the chosen size has the same chance of being selected. Before drawing again, predict which summaries will stay close and which might move.
A school of 60 fictional students
The full school roster is available. Draw a simple random sample without replacement; each student has the same chance of inclusion. Record usual travel mode and one-way travel time.
No sample yet. Predict whether ten students will describe the whole school exactly.
Read the chart as a table
| Value range | Count |
|---|---|
| 0.0–10.0 | 0 |
| 10.0–20.0 | 0 |
| 20.0–30.0 | 0 |
| 30.0–40.0 | 0 |
| 40.0–50.0 | 0 |
| 50.0–60.0 | 0 |
| 60.0–70.0 | 0 |
| 70.0–80.0 | 0 |
| 80.0–90.0 | 0 |
| 90.0–100.0 | 0 |
| 100.0–110.0 | 0 |
| 110.0–120.0 | 0 |
Each range includes its lower end and excludes its upper end, except the last range, which includes both.
First, name the columns
Each student is an observational unit. Their travel mode and commute time are variables: characteristics we record. A numerical variable describes an amount; a categorical variable places an observation in a group.
A student ID may look numerical, but averaging ID numbers tells us nothing about the students. Likewise, coding bus as 1 and bike as 2 does not make bike twice as much transport. The meaning of a variable determines which summaries make sense.
Switch to travel mode. The frequency is how many sampled students belong to each category; the relative frequency divides that count by the sample size. Three bus riders among ten students gives 3/10 = 30%. That describes the sample. It is not yet a claim that exactly 30% of the school rides the bus.
A moment to think
Look at the shape before the average
Return to commute time. A histogram groups numerical values into intervals and counts how many fall in each. Change the interval width while keeping the same sample. The picture changes, but the measurements have not.
Look for a typical value, spread, gaps, and a long tail. A distribution with a long right tail is right-skewed. One unusually long commute can pull the mean to the right while barely moving the median. Try the 110-minute journey and compare.
The minimum, lower quartile, median, upper quartile, and maximum make the five-number summary. Quartiles mark roughly a quarter and three quarters of the way through the ordered values. Their difference, the interquartile range, describes the spread of the middle half. Software uses several quartile conventions; this activity takes the median of each half.
A common boxplot rule flags observations below Q₁ − 1.5 IQR or above Q₃ + 1.5 IQR. A flag invites a closer look. It is not permission to delete a real student with a difficult commute. Check the measurement, document corrections, and ask how the unusual observation affects the conclusion.
Another way to describe spread
The sample standard deviation measures spread around the mean, in minutes here. It grows when observations sit farther from the mean. Its calculation squares those distances, so unusually large distances matter a lot. The median and IQR are less sensitive to a single extreme value.
An idea to take with you
A summary is a deliberate compression
IQR = Q₃ − Q₁
s = √[Σ(xᵢ − x̄)² / (n − 1)]
The first expression keeps the middle half’s width. In the second, xᵢ is one measurement, x̄ is the sample mean, n is the sample size, and Σ means add. Subtract the mean, square each distance, add them, divide by n − 1, then take the square root.
Estimating the mean from the same data leaves only n − 1 freely varying deviations: they must sum to zero. Dividing by n − 1 makes s² an unbiased estimate of population variance for independent observations from the same distribution with finite variance. The square root s is the usual estimate of standard deviation.
Return to the experiment ↑A moment to think
Our buttons use a complete roster and everyone responds. A real hallway survey may miss absent students, early arrivals, or people who decline. Bigger samples do not automatically repair those omissions. The next chapters separate sample size from sample selection.
Words you met
| Term | Meaning |
|---|---|
| Observational unit | The person or object represented by one observation. |
| Variable | A measured characteristic, such as travel mode or commute time. |
| Relative frequency | A category’s count divided by the total count. |
| Quartiles and IQR | The quarter-way marks and the distance between the outer two. |
| Standard deviation | A measure of spread around the mean, in the measurement’s units. |
| Descriptive statistics | Summaries of the data actually observed. |
Neighbors
- More examples: Summarizing Data.
- The next chapter distinguishes a sample from its population. For the mathematical version of spread, see Expected Value and Variance.
Written by June Kim. The conversational approach was inspired by Danielle Navarro’s Learning Statistics with R (CC BY-SA 4.0). The prose, experiments, and questions here were created for this book. This chapter’s text and illustrations are also shared under CC BY-SA 4.0.