← back to statistics

Random Variables and Distributions

Statistics through experiments · 2 of 16 · CC BY-SA 4.0

One handful of coins gives you a result. Many handfuls give you a shape.

Take ten fair coins and flip them together. Count the heads. You might get four, six, or some other number. Call that one round.

Now imagine doing it again and again. Before you try, which seems more likely: a count somewhere around the middle, or all ten coins landing heads?

Start with Flip one round. The coins show its result, and the chart records its head count. Each round adds one entry to the chart, however many coins you flipped.

Make a distributionOne round, one entry

Each coin is independent. Start with a 50% chance of heads.

Start with one round. Its head count will add one entry to the chart.

0 rounds recorded. Bar height shows the percentage of rounds, so you can compare runs of different lengths.

Current experiment · 10 coins per round · 0 rounds
Current experiment: distribution of head countsEach bar shows the percentage of rounds with that result. Run a round to add the first result.0%5%10%Share of rounds0 heads (0%): 0 rounds1 heads (10%): 0 rounds2 heads (20%): 0 rounds3 heads (30%): 0 rounds4 heads (40%): 0 rounds5 heads (50%): 0 rounds6 heads (60%): 0 rounds7 heads (70%): 0 rounds8 heads (80%): 0 rounds9 heads (90%): 0 rounds10 heads (100%): 0 rounds0510Number of heads in one roundYour first round will go here.
Read current chart as a table
Current experiment: every possible head count
HeadsProportionRounds
00%0
110%0
220%0
330%0
440%0
550%0
660%0
770%0
880%0
990%0
10100%0
Another thing to try: change the coin

So far, heads and tails have had equal chances. What shape would you expect if heads were more likely?

Changing the number of coins or their chance of heads clears the current experiment. A saved comparison stays labeled with its own settings.

A shape made of ordinary surprises

Run a hundred rounds, then a thousand more. Each individual count is still uncertain, but the collection begins to have a recognizable shape. With fair coins, the middle counts tend to turn up more often than the extremes.

Two coins make it easier to see why. There are four equally likely sequences: HH, HT, TH, TT. Two of them give one head. Only one gives two heads, and only one gives zero.

More coins allow many more sequences, and many lead to a middle-sized head count. An all-heads result still has just one sequence.

The bars show an observed distribution: how the recorded rounds are spread across possible head counts. Bar height is the percentage of rounds at each count. The model also has a probability distribution, which assigns a chance to every possible count. More rounds help the observed picture resemble it.

A moment to think

You keep ten coins per round and run another thousand rounds. What have you changed?

More coins, or more rounds?

These are different knobs, and it’s worth spending a moment with each.

Try two coins per round, run 1,000 rounds, and choose Keep for comparison. Then choose twenty coins per round and run 1,000 more. The old chart stays below, with its original settings.

In the Number of heads view, the two experiments occupy different ranges: zero to two heads, and zero to twenty. Switch to Proportion of heads to compare them on the same zero-to-100% scale.

Half the coins being heads is now in the same place on both charts. With more independent fair coins per round, the proportion tends to stay closer to half. All-heads rounds are common enough with two coins; with twenty, they are much rarer.

Pay attention to the units. The number of heads has more room to vary when you flip more coins. The proportion of heads varies less. Both statements can be true.

A name for the number we record

The coins give us a sequence like H, T, H, H, T. We turn that sequence into a number: three heads.

A rule that assigns a number to a random outcome is called a random variable. Here the rule is “count the heads.” We’ll call that count X. It can change from round to round.

The number of coins you chose gets a shorter name too: n. To find the proportion of heads, divide the head count by the coin count. For that five-coin round, it is 3 ÷ 5 = 0.6, or 60%.

A moment to think

You run 100 rounds with five coins in each round. The bar at three heads says 30%. What does it mean?

The whole calculation, in three symbols

You already know the operation: heads divided by coins. We can give its result a name, p̂, pronounced “p-hat.” The hat tells us it is an estimate made from observations.

You can now write this

=

Choose a piece to read it in words.

Read this as “p-hat.” It is the observed proportion of heads. The little hat marks an estimate of the coin’s probability, made from data.

Three heads in five flips gives p-hat = 0.6. The true probability can still be 0.5.

Find it in your experiment ↑

Remember p, the probability? Our fair coin has p = 0.5. A round with three heads out of five gives p̂ = 0.6. One describes the model; the other describes our evidence. There is no contradiction.

A moment to think

You inspect a batch of 20 seeds. Fifteen sprout. If X counts sprouts and n counts seeds, what is p̂?

Words you met

TermMeaning
RoundOne complete experiment: flip the chosen coins and count their heads.
Random variableA rule assigning a number to a random outcome. Here, X counts heads.
DistributionHow values are spread out, together with their frequencies or probabilities.
ProportionA part divided by the whole: three heads out of five coins is 0.6.
EstimateA quantity calculated from data to learn about an unknown quantity.
Neighbors
  • This experiment has a formal name: the binomial model. You can meet its formula in Distributions. Or continue to Center and Spread to start thinking about observations beyond coins.
  • To see why middle counts are so common, explore Combinatorics in the probability textbook. Then Important Distributions derives the binomial model used here.

Written by June Kim. The conversational approach was inspired by Danielle Navarro’s Learning Statistics with R (CC BY-SA 4.0). The prose, experiments, and questions here were created for this book. This chapter’s text and illustrations are also shared under CC BY-SA 4.0.