The problem with raw numbers
A teacher records the test scores of 30 students. The numbers sit in a list, unordered and unprocessed: 45, 72, 68, 51, 90, 33, 78, 62, 55, 81, 47, 69, 74, 58, 66, 43, 77, 85, 60, 52, 71, 64, 49, 83, 56, 70, 61, 75, 67, 59. What can you actually say about this class? Not much, until you organise, summarise, and represent the data. That process of turning raw figures into meaningful insight is exactly what the IGCSE Mathematics statistics section tests.
The logical approach breaks into three stages: classify and organise the data, calculate summary measures (averages and spread), then represent the data visually so patterns become clear. Each stage has specific methods and specific pitfalls. Working through them systematically is the fastest route to full marks.
Classifying and organising data
Before any calculation, you need to know what kind of data you're handling. The classification determines which averages, which diagrams, and which techniques apply.
| Classification | Definition | Example |
|---|---|---|
| Qualitative | Non-numerical, describes qualities or categories | Favourite colour, type of transport |
| Quantitative | Numerical, measures or counts something | Height in cm, number of siblings |
| Discrete | Quantitative data that takes specific values (usually whole numbers from counting) | Number of pets, shoe size |
| Continuous | Quantitative data that can take any value within a range (usually from measuring) | Time in seconds, mass in kg |
| Primary | Data you collect yourself for a specific purpose | Survey results from your class |
| Secondary | Data collected by someone else, used for a different purpose | Government census figures |
Organising data into a frequency table is the first practical step. For discrete data, list each value and count how many times it appears. For continuous data, group the values into class intervals. The choice of intervals matters: too few and you lose detail, too many and the table becomes as unwieldy as the raw list.
Averages and range: choosing the right measure
Three averages appear on the IGCSE syllabus. Each answers a slightly different question, and knowing which to use (and which the examiner is asking for) prevents lost marks.
| Measure | How to calculate | Best used when | Weakness |
|---|---|---|---|
| Mean | Sum of all values divided by the number of values | Data is fairly evenly spread without extreme outliers | Distorted by very high or very low values |
| Median | Middle value when data is arranged in order | Data contains outliers or is skewed | Ignores the actual size of most values |
| Mode | Most frequently occurring value | Qualitative data, or when the most common value matters (e.g. most popular shoe size for stock ordering) | May not exist, or there may be several |
The range (highest value minus lowest value) measures spread. A small range means values are clustered; a large range means they're dispersed. The range is simple but sensitive to outliers: one extreme value inflates it dramatically.
Worked Example 1: Mean, median, mode, and range from a list
Problem: Find the mean, median, mode, and range of: 4, 7, 2, 7, 9, 3, 7, 5, 8
Step 1: Arrange in order: 2, 3, 4, 5, 7, 7, 7, 8, 9
Step 2 (Mean): Sum = 2 + 3 + 4 + 5 + 7 + 7 + 7 + 8 + 9 = 52. There are 9 values. Mean = 52 / 9 = 5.78 (to 3 s.f.)
Step 3 (Median): 9 values, so the median is the (9 + 1) / 2 = 5th value. Counting along the ordered list: the 5th value is 7. Median = 7.
Step 4 (Mode): 7 appears three times, more than any other value. Mode = 7.
Step 5 (Range): 9 - 2 = 7.
Worked Example 2: Estimated mean from a grouped frequency table
Problem: The table shows the times (in minutes) taken by 40 students to complete a puzzle.
| Time (t minutes) | Frequency |
|---|---|
| 0 < t ≤ 5 | 6 |
| 5 < t ≤ 10 | 14 |
| 10 < t ≤ 15 | 12 |
| 15 < t ≤ 20 | 8 |
Step 1: Find the midpoint of each class interval. For 0 < t ≤ 5, the midpoint is (0 + 5) / 2 = 2.5. For 5 < t ≤ 10, it is 7.5. For 10 < t ≤ 15, it is 12.5. For 15 < t ≤ 20, it is 17.5.
Step 2: Multiply each midpoint by its frequency: 2.5 x 6 = 15, 7.5 x 14 = 105, 12.5 x 12 = 150, 17.5 x 8 = 140.
Step 3: Sum of (midpoint x frequency) = 15 + 105 + 150 + 140 = 410.
Step 4: Estimated mean = 410 / 40 = 10.25 minutes.
For grouped data, the modal class replaces the mode: it is the class interval with the highest frequency. Here, the modal class is 5 < t ≤ 10 (frequency 14). The class containing the median is found by locating the (n/2)th value: the 20th value falls in the 5 < t ≤ 10 class (cumulative frequency reaches 20 at the end of that class).
Statistical charts and diagrams
Choosing the correct diagram for the data type is a frequent exam question. The decision follows a logical sequence.
- Bar chart: categorical or discrete data. Bars are separate (gaps between them). Heights represent frequency.
- Dual/compound bar chart: comparing two or more datasets for the same categories. Bars are grouped side by side (dual) or stacked (compound).
- Pie chart: showing proportions of a whole. Each sector's angle = (frequency / total) x 360.
- Stem-and-leaf diagram: preserves individual data values while showing shape. The stem is the leading digit(s); the leaf is the trailing digit. Always include a key.
- Frequency polygon: plotted at the midpoint of each class interval, points joined by straight lines. Useful for comparing two distributions on the same axes.
- Histogram (Extended): for continuous data with equal or unequal class widths. With equal widths, the y-axis is frequency. With unequal widths, the y-axis is frequency density = frequency / class width. The area of each bar represents the frequency.
Worked Example 3: Pie chart angles
Problem: 60 students chose their favourite sport: Football 24, Tennis 15, Swimming 12, Athletics 9. Draw a pie chart.
Step 1: Calculate each angle. Football: (24/60) x 360 = 144 degrees. Tennis: (15/60) x 360 = 90 degrees. Swimming: (12/60) x 360 = 72 degrees. Athletics: (9/60) x 360 = 54 degrees.
Step 2: Check: 144 + 90 + 72 + 54 = 360 degrees. The angles sum correctly.
Step 3: Draw the circle, measure each sector with a protractor, and label each sector with the category name.
Worked Example 4: Frequency density for a histogram (Extended)
Problem: A grouped frequency table has unequal class widths:
| Mass (m kg) | Frequency | Class width | Frequency density |
|---|---|---|---|
| 0 < m ≤ 10 | 8 | 10 | 0.8 |
| 10 < m ≤ 20 | 15 | 10 | 1.5 |
| 20 < m ≤ 40 | 24 | 20 | 1.2 |
| 40 < m ≤ 70 | 18 | 30 | 0.6 |
Frequency density = frequency / class width. For the third class: 24 / 20 = 1.2. The y-axis of the histogram shows frequency density, and each bar spans the full width of its class interval with no gaps. The area of each bar (width x height) equals the frequency for that class.
Scatter diagrams and correlation
Scatter diagrams plot paired data as points on a coordinate grid. Each point represents one individual or observation. The pattern of points reveals whether a relationship (correlation) exists between the two variables.
| Pattern | Name | Meaning |
|---|---|---|
| Points slope upward from left to right | Positive correlation | As one variable increases, the other tends to increase |
| Points slope downward from left to right | Negative correlation | As one variable increases, the other tends to decrease |
| Points show no clear pattern | No correlation | No consistent relationship between the variables |
The line of best fit is a straight line drawn through the data that best represents the trend. It should pass through or near the mean point (the point whose coordinates are the mean of all x-values and the mean of all y-values). Roughly equal numbers of points should fall above and below the line.
Interpolation means reading a value within the range of the data. This is generally reliable. Extrapolation means reading a value beyond the range of the data. This is unreliable because you're assuming the pattern continues, and it may not. Examiners regularly ask candidates to distinguish between the two and explain why extrapolation is less trustworthy.
Worked Example 5: Using a line of best fit
Problem: A scatter diagram shows the relationship between hours of revision (x-axis, range 2 to 12) and test score (y-axis). The line of best fit passes through (4, 35) and (10, 65). Estimate the test score for a student who revised for 7 hours. Would you trust a prediction for a student who revised for 20 hours?
Step 1: Find the gradient: (65 - 35) / (10 - 4) = 30 / 6 = 5. The equation of the line: y - 35 = 5(x - 4), so y = 5x + 15.
Step 2: For x = 7: y = 5(7) + 15 = 50. Estimated score: 50 marks. This is interpolation (7 is within the data range of 2 to 12), so the estimate is reasonably reliable.
Step 3: For x = 20: this is extrapolation (20 is well beyond the data range). The prediction would be unreliable because we have no evidence the linear pattern continues that far. Scores may plateau, or the relationship may change at higher revision hours.
Cumulative frequency and box-and-whisker plots (Extended)
Cumulative frequency is a running total of frequencies. It answers the question: "How many data values are less than or equal to this boundary?" Plotting cumulative frequency against the upper class boundary produces a cumulative frequency curve (an S-shaped, or ogive, curve).
Worked Example 6: Drawing and reading a cumulative frequency curve
Problem: Using the puzzle data from Worked Example 2:
| Time (t minutes) | Frequency | Cumulative frequency |
|---|---|---|
| 0 < t ≤ 5 | 6 | 6 |
| 5 < t ≤ 10 | 14 | 20 |
| 10 < t ≤ 15 | 12 | 32 |
| 15 < t ≤ 20 | 8 | 40 |
Step 1: Plot the points at the upper class boundaries: (5, 6), (10, 20), (15, 32), (20, 40). Also plot (0, 0) as the starting point.
Step 2: Join the points with a smooth curve (not straight lines between points).
Step 3: Read key values from the curve:
- Median: the value at the n/2 = 20th position. Read across from 20 on the y-axis to the curve, then down to the x-axis. Approximately 10 minutes.
- Lower quartile (Q1): the value at the n/4 = 10th position. Approximately 7 minutes.
- Upper quartile (Q3): the value at the 3n/4 = 30th position. Approximately 14 minutes.
- Interquartile range (IQR): Q3 - Q1 = 14 - 7 = 7 minutes.
The IQR is a more robust measure of spread than the range because it ignores the extreme values and focuses on the middle 50% of the data.
Box-and-whisker plots
A box-and-whisker plot (box plot) displays five key values on a single diagram: minimum, lower quartile, median, upper quartile, and maximum. The "box" spans from Q1 to Q3, with a line inside at the median. The "whiskers" extend from the box to the minimum and maximum values.
For the puzzle data: minimum = 0, Q1 = 7, median = 10, Q3 = 14, maximum = 20. The box runs from 7 to 14, the median line sits at 10, and whiskers reach to 0 and 20.
Box plots are particularly useful for comparing two distributions side by side. If Class A's box sits entirely above Class B's box, you can say Class A generally scored higher. If Class A's box is narrower, the scores in Class A were more consistent.
Common mistakes and how to avoid them
| Mistake | Why it costs marks | Correct approach |
|---|---|---|
| Using raw frequencies on a histogram with unequal class widths | Wider classes appear disproportionately large; the diagram misrepresents the data | Calculate frequency density (frequency / class width) for each class |
| Forgetting to use midpoints for estimated mean from grouped data | You cannot add up the individual values because they aren't given; the calculation requires midpoints | Find the midpoint of each class, multiply by frequency, sum, then divide by total frequency |
| Writing "mean" instead of "estimated mean" for grouped data | Examiners expect the word "estimated" because you don't know exact values within each group | Always say "estimated mean" when data is grouped |
| Drawing bar charts for continuous data (or histograms for discrete data) | Bar charts have gaps (for separate categories); histograms have no gaps (for continuous intervals) | Check whether the data is discrete or continuous, then select the appropriate diagram |
| Reading cumulative frequency at the midpoint instead of the upper boundary | Cumulative frequency is plotted at the upper boundary of each class, not the midpoint | Always plot cumulative frequency at the upper class boundary |
| Confusing interpolation with extrapolation | Extrapolation is unreliable; claiming it's reliable loses evaluation marks | State that predictions within the data range (interpolation) are more reliable than those outside it (extrapolation) |
| Finding the median position incorrectly | For n values, the median is the (n+1)/2 th value in a list, but the n/2 th value on a cumulative frequency curve | Use (n+1)/2 for raw data lists; use n/2 for cumulative frequency diagrams |
Self-check questions
- Classify each of the following as qualitative or quantitative, and if quantitative, as discrete or continuous: (a) number of goals scored, (b) temperature in Celsius, (c) brand of phone, (d) mass of a parcel.
- The ages of 7 people are: 23, 19, 31, 27, 19, 35, 24. Find the mean, median, mode, and range.
- A grouped frequency table shows heights of 50 plants. The class intervals are 0-10 cm (frequency 8), 10-20 cm (frequency 17), 20-30 cm (frequency 15), 30-40 cm (frequency 10). Calculate the estimated mean height.
- In the same table, identify the modal class and the class containing the median.
- A pie chart shows how 90 students travel to school. If 30 students walk, what angle should the "walk" sector be?
- A histogram has class intervals 0-5 (frequency 10), 5-10 (frequency 20), 10-20 (frequency 30), 20-40 (frequency 16). Calculate the frequency density for each class and identify which bar will be the tallest.
- Describe the type of correlation you would expect between: (a) hours of sunshine and ice cream sales, (b) age of a car and its resale value, (c) shoe size and IQ.
- A line of best fit on a scatter diagram covers data from x = 5 to x = 25. Explain why a prediction at x = 15 is more reliable than a prediction at x = 40.
- From a cumulative frequency curve of 80 data values, at what cumulative frequency would you read (a) the median, (b) the lower quartile, (c) the upper quartile?
- Two box plots show test results for Class A (minimum 30, Q1 45, median 58, Q3 70, maximum 85) and Class B (minimum 40, Q1 52, median 60, Q3 65, maximum 78). Which class had more consistent results? Justify your answer using a specific measure of spread.
Working through these questions with full written solutions, rather than mental estimates, builds the precision that IGCSE mark schemes reward. Each calculation step earns method marks even if the final answer contains an arithmetic slip, so showing clear working is not optional: it is the difference between partial credit and zero.
A methodical guide to the Statistics section of Cambridge IGCSE Mathematics (0580), covering data classification, averages and range, charts and diagrams, scatter diagrams, and cumulative frequency, with step-by-step worked examples and common mistakes to avoid.
Sharhi(ko)