HomeLearning HubA Level MathsS1 1: Representation of data
S1 1

Representation of data

Probability & Statistics 1 · Paper 5

🎯What you need to be able to do

  • Construct and interpret stem-and-leaf diagrams, including back-to-back ones.
  • Construct and interpret box-and-whisker diagrams, and identify outliers.
  • Draw and read histograms with unequal class widths, using frequency density.
  • Use a cumulative frequency graph to estimate the median, quartiles and percentiles.
  • Calculate the mean and standard deviation, from raw data and from summary totals.
  • Use coded data and adjust the results back.
  • Compare two distributions in context.

📚The mathematics

Stem-and-leaf diagrams

These keep the original data values while showing the shape. Leaves must be ordered, and a key is compulsory — without it the diagram is meaningless, and its absence loses a mark every time.

Back-to-back diagrams compare two sets against a shared stem. Read the left-hand leaves outwards from the stem, so they run right to left.

Box-and-whisker diagrams

A box plot shows five numbers: minimum, lower quartile, median, upper quartile, maximum. The box spans the interquartile range, \( \mathrm{IQR} = Q_3 - Q_1 \).

An outlier is conventionally a value more than \( 1.5 \times \mathrm{IQR} \) beyond the nearer quartile. Check the definition the question gives, since some use a different multiplier. When outliers are present, the whiskers extend only to the most extreme value that is not an outlier, and the outliers are plotted separately as crosses.

Comparing two box plots is a standard question, and the answer needs two elements in context: a comparison of centre (medians) and a comparison of spread (IQR or range). “The medians are 42 and 51, so group B scored higher on average; the IQRs are 15 and 8, so group B was also more consistent” earns the marks. Naming the numbers without interpreting them does not.

Histograms

In a histogram it is the area of each bar that represents frequency, not its height. When class widths differ you must therefore plot

\[ \text{frequency density} = \frac{\text{frequency}}{\text{class width}} \]
Get the class boundaries right, not the class labels. For continuous data recorded to the nearest whole number, a class written “20–29” has boundaries 19.5 and 29.5, so its width is 10, not 9. For a class written \( 20 \le x < 30 \) the width really is 10. Reading widths off the labels rather than the boundaries distorts every bar and is the most common histogram error.

There are no gaps between bars in a histogram, because the data is continuous. To read a frequency back off a histogram, multiply the height by the width.

Cumulative frequency

Plot cumulative frequency against the upper class boundary — not the mid-point — and join the points with a smooth curve. Then read across from the appropriate fraction of the total and down to the axis:

median at \( \tfrac{1}{2}n \)
\( Q_1 \) at \( \tfrac{1}{4}n \)
\( Q_3 \) at \( \tfrac{3}{4}n \)

These are estimates, because the individual values within each class are no longer known.

Mean and standard deviation

\( \bar{x} = \dfrac{\sum x}{n} \)
\( \bar{x} = \dfrac{\sum fx}{\sum f} \)
\( \sigma^{2} = \dfrac{\sum x^{2}}{n} - \bar{x}^{2} \)
\( \sigma^{2} = \dfrac{\sum fx^{2}}{\sum f} - \bar{x}^{2} \)

The variance formula — “the mean of the squares minus the square of the mean” — is what makes questions solvable from summary totals \( \sum x \) and \( \sum x^{2} \) without the original data. The standard deviation is \( \sigma = \sqrt{\sigma^{2}} \), and it carries the same units as the data.

For grouped data, use the mid-interval value of each class as \(x\). The result is an estimate, because you have thrown away the individual values.

Adding data to an existing set, or removing it, is handled through the totals: recover \( \sum x \) and \( \sum x^{2} \) from the given mean and standard deviation, adjust them, then recompute. That is far cleaner than trying to work with the means directly.

Coded data

To simplify arithmetic, questions often give totals for a coded variable such as \( y = x - 200 \) or \( y = \dfrac{x - 50}{10} \). The effects are:

\( \bar{x} = a\bar{y} + b \) for \( y = \dfrac{x-b}{a} \)
\( \sigma_x = a\,\sigma_y \)

In words: a shift changes the mean but leaves the standard deviation unchanged, while a scaling changes both. Sliding a data set along does not spread it out. Forgetting to decode at the end — or decoding the standard deviation as though it were a mean, and adding the shift — are the two errors examiners flag most.

✏️Worked example

(a) The times, in minutes, taken by 60 students to complete a task are summarised by \( \sum x = 1530 \) and \( \sum x^{2} = 41\,220 \). Find the mean and standard deviation. (b) A 61st student took 40 minutes. Find the new mean and standard deviation. (c) The lengths of 40 items are summarised using the coding \( y = x - 150 \), giving \( \sum y = 96 \) and \( \sum y^{2} = 1150 \). Find the mean and standard deviation of \(x\).

(a) The mean is \( \bar{x} = \dfrac{1530}{60} = 25.5 \) minutes. For the variance:

\[ \sigma^{2} = \frac{41\,220}{60} - 25.5^{2} = 687 - 650.25 = 36.75 \]

so \( \sigma = \sqrt{36.75} = 6.06 \) minutes.

(b) Update the totals rather than the statistics: \( \sum x = 1530 + 40 = 1570 \) and \( \sum x^{2} = 41\,220 + 1600 = 42\,820 \), over \( n = 61 \).

\( \bar{x} = \dfrac{1570}{61} = 25.74 \)
\( \sigma^{2} = \dfrac{42\,820}{61} - 25.74^{2} \)

That is \( 701.97 - 662.43 = 39.54 \), so \( \sigma = 6.29 \) minutes.

(c) The coding is a pure shift, so \( \bar{x} = \bar{y} + 150 \) and \( \sigma_x = \sigma_y \).

\[ \bar{y} = \frac{96}{40} = 2.4 \;\Longrightarrow\; \bar{x} = 152.4 \]

and \( \sigma_y^{2} = \dfrac{1150}{40} - 2.4^{2} = 28.75 - 5.76 = 22.99 \), so \( \sigma_y = 4.795 \) and therefore \( \sigma_x = 4.80 \).

Check it. In (b), the new value of 40 is well above the old mean of 25.5, so both the mean and the spread should increase — and they do, from 25.5 to 25.74 and from 6.06 to 6.29. A single added value moving the mean only slightly is right for \( n = 60 \). In (c), the standard deviation of \(x\) equals that of \(y\) exactly; if you found yourself adding 150 to it, that is the error the coding rules exist to prevent.
Work with \( \sum x \) and \( \sum x^{2} \), not with the mean and standard deviation. In (b) it is tempting to average 25.5 with 40 somehow, but there is no valid shortcut — means do not combine by averaging unless the group sizes are equal. Recovering the totals, adjusting them, and recomputing is the only reliable route, and it is quick once the habit is there.

📝Practise

Work through these, then reveal the answer.

1. A data set has \( n = 10 \), \( \sum x = 84 \) and \( \sum x^{2} = 802 \). Find the mean and standard deviation.
\( \bar{x} = \dfrac{84}{10} = 8.4 \). Variance \( = \dfrac{802}{10} - 8.4^{2} = 80.2 - 70.56 = 9.64 \), so \( \sigma = \sqrt{9.64} = 3.10 \).
2. A histogram has a bar covering \( 10 \le x < 25 \) with height (frequency density) 3.2. How many values lie in that class?
Frequency \( = \) frequency density \( \times \) class width \( = 3.2 \times 15 = 48 \). Reading the height as the frequency, and answering 3.2, is the standard error — in a histogram it is area, not height, that represents frequency.
3. A box plot has \( Q_1 = 32 \), \( Q_3 = 50 \). Determine whether the values 5 and 78 are outliers, using the \( 1.5 \times \mathrm{IQR} \) rule.
\( \mathrm{IQR} = 18 \), so \( 1.5 \times \mathrm{IQR} = 27 \). Lower boundary: \( 32 - 27 = 5 \). Upper boundary: \( 50 + 27 = 77 \). The value 5 is exactly on the boundary, so it is not beyond it and is not an outlier. The value 78 exceeds 77, so it is an outlier. Borderline cases like the 5 need the inequality stated carefully.
4. For 25 values coded by \( y = \dfrac{x - 40}{5} \), it is found that \( \sum y = 30 \) and \( \sum y^{2} = 120 \). Find the mean and standard deviation of \(x\).
\( \bar{y} = \dfrac{30}{25} = 1.2 \), so \( \bar{x} = 5(1.2) + 40 = 46 \). Variance of \(y\): \( \dfrac{120}{25} - 1.44 = 4.8 - 1.44 = 3.36 \), so \( \sigma_y = 1.833 \). Since the coding scales by 5, \( \sigma_x = 5 \times 1.833 = 9.17 \). Here the scaling does affect the standard deviation, unlike a pure shift.
5. The mean of 8 numbers is 15. One number, 22, is removed. Find the new mean.
The total was \( 8 \times 15 = 120 \). Removing 22 leaves a total of 98 over 7 values, so the new mean is \( \dfrac{98}{7} = 14 \). Removing a value above the mean pulls the mean down, as it should.
6. Two classes sat the same test. Class A: median 58, IQR 12. Class B: median 61, IQR 25. Compare the two distributions.
Class B has the higher median (61 against 58), so on average Class B performed slightly better. However Class B has a much larger interquartile range (25 against 12), so its results were far more spread out — Class A was more consistent. Both a centre comparison and a spread comparison are needed, and both must be phrased in the context of test performance rather than as bare numbers.

🔗Go deeper — other people’s work

These are external resources, not mine. If one stops working, tell me and everything above it on this page still stands.

  • Seeing Theory (Brown University) — interactive visualisations of spread and centre
  • Khan Academy — box plots, histograms and standard deviation
  • Cambridge examiner reports — histogram class boundaries appear in almost every one