1.Truncated Axes: Making Small Differences Look Big
The heights of the bars in a bar chart are there so you can compare the sizes of the values by eye. But if the vertical axis starts somewhere other than 0, the ratio of the bar heights no longer matches the ratio of the actual values. A small difference ends up looking like a difference of several times.
Truncating the vertical axis is not always wrong. Line graphs that show small changes in values that never come near 0, such as body temperature or stock prices, often cut the axis. The problem is cutting the axis while inviting you to compare sizes by bar length. When you look at a graph, check the starting value and the tick marks on the vertical axis first.
- Step 1: Bar heights as shown on the chart: A is 100 − 98 = 2 units, and B is 104 − 98 = 6 units.
- Step 2: Apparent ratio: 6 ÷ 2 = 3 times.
- Step 3: Actual ratio: 104 ÷ 100 = 1.04, so B is only 4% higher than A.
- Check: If you draw the vertical axis from 0, the bar heights are 100 and 104, a difference that is hard to tell apart by eye.
2.Correlation Is Not Causation
When two values rise and fall together, they are correlated. In summer, ice cream sales go up and so do swimming accidents. The two are clearly correlated, but ice cream does not cause accidents. A 3rd factor, temperature, moves them both. A hidden factor like this is called a confounding variable.
Before reading a correlation as causation, ask three questions. Is there a hidden common cause? Could the direction of cause and effect be reversed (are people who exercise healthy, or do healthy people exercise more)? Could it be a coincidence? The surest way to establish causation is an experiment comparing two randomly assigned groups, but such an experiment isn't always possible.
- Common cause: temperature → ice cream sales, temperature → swimming accidents
- Reversed direction: "People who go to the hospital often are sicker"; they go to the hospital because they are sick
- Coincidence: two unrelated statistics can move in step if you choose the time period carefully
3.Sampling Bias: Whom Did You Ask?
Since you can't survey everyone, you draw a part (a sample) and use it to estimate the whole. If the sample doesn't resemble the whole, the results are skewed no matter how much data you collect. A survey run only inside a particular app reflects the opinions of people who use that app, and a survey of only people who volunteered to answer is closer to the opinions of people with a strong interest in the topic.
Looking only at the survivors is also sampling bias. The claim "successful people all got up at dawn" leaves out the people who got up at dawn and did not succeed. When you look at data, ask "who was chosen, how, and who was left out?" before you look at the numbers. How a sample is drawn often matters more than how big it is.
AI has the same problem. If the training data is skewed toward a particular group, the AI's judgments are skewed in that direction too. The data bias covered in the AI Basics course is exactly this sampling bias.
4.Rates vs. Counts, Relative vs. Absolute Change
If you directly compare the number of accidents in a city of 1,000,000 people and a city of 100,000 people, the big city will always look worse. For a fair comparison, you need rates converted to the same base, such as per 100,000 people. Conversely, if you look only at rates and miss the counts, you misjudge the actual scale.
"The risk has doubled" carries the same problem. Relative change (doubling, a 100% increase) sounds big, but if the original risk was very small, the absolute change may be tiny. You need to look at both to judge.
- Step 1: City A: 500 ÷ 1,000,000 × 100,000 = 50 accidents.
- Step 2: City B: 100 ÷ 100,000 × 100,000 = 100 accidents.
- Step 3: In counts, A has 5 times as many, but relative to population, B's rate is 2 times as high.
- Check: A has 10 groups of 100,000 people, and 500 ÷ 10 = 50; B has 1 group, so 100. Same result.
- Step 1: Relative change: (4 − 2) ÷ 2 = 1 → a 100% increase (doubling).
- Step 2: Absolute change: from 2 ÷ 10,000 = 0.02% to 4 ÷ 10,000 = 0.04%, an increase of 0.02 percentage points.
- Check: 0.04 − 0.02 = 0.02 percentage points, which is the same as 2 more people out of 10,000.
5.The Traps of Averages and Grouping
Can you wade across a river whose average depth is 1 m? If it is shallow at the edges and 3 m deep in the middle, you can't, even though the average is 1 m. An average tells you nothing about spread or the maximum. As you saw in Lesson 8, the median and the spread belong next to the mean.
Sometimes the results for grouped data and for the separate parts point in opposite directions. This is called Simpson's paradox. In the example below, Hospital X has a higher success rate than Y for both easy and difficult surgeries, but combined, Y looks far better. That's because X took on far more difficult surgeries.
| Easy surgery | Difficult surgery | Overall | |
|---|---|---|---|
| Hospital X | 9/10 (90%) | 30/90 (about 33%) | 39/100 (39%) |
| Hospital Y | 85/100 (85%) | 3/10 (30%) | 88/110 (80%) |
- Step 1: X overall: (9 + 30) ÷ (10 + 90) = 39 ÷ 100 = 39%.
- Step 2: Y overall: (85 + 3) ÷ (100 + 10) = 88 ÷ 110 = 80%.
- Step 3: For X, 90 of 100 patients had difficult surgery; for Y, only 10 of 110 did. The overall success rate depends heavily on the mix of patients each hospital took on.
- Check: Easy surgery 90% > 85%, difficult surgery about 33.3% > 30%, so X is higher in both groups. Yet overall it is 39% < 80%.
📌 Key points
- If a bar chart's vertical axis doesn't start at 0, differences look exaggerated
- Correlation is not causation; first suspect a common cause, the direction, or coincidence
- How a sample is drawn matters more than its size; ask who was left out
- Look at counts and rates, and at relative and absolute change, together
- A single average can't show spread, and grouped results can be the opposite of the broken-down ones
🤖 Try asking AI like this
Copy a prompt and replace the [ ] parts with your own situation. Don't take the answer on trust — check it against this lesson.
When you want to read a graph in a news story or report critically
Look at the description of the graph (or figures) below and go through a checklist of which traps might apply: truncated axes, the base (denominator), sampling bias, correlation vs. causation, rates vs. counts, and grouped statistics. For each item, also write what additional information would be needed to check it. Content: [paste here]
When you want to practice spotting graph traps
Make 5 misleading graph descriptions using made-up numbers. Keep the trap in each one hidden, and show me the answer and the correct calculation after I've found it. Be sure to label them "Example" so they don't look like real statistics.
When you receive a data summary made by an AI
For every figure you used in the summary you just wrote, state its source and its base (what you used as the denominator). Mark separately any figures whose source can't be verified.
- Standard high school math textbook content (statistics, sample surveys)
- Standard explanations in introductory statistics textbooks (correlation and causation, Simpson's paradox)
Reached every goal above? Mark the lesson complete.
Storage is unavailable in this browser, so this lasts only for this page.📐 Math Basics
- 1Numbers and Operations: Fractions and Decimals Revisited
- 2Ratios and Rates: Reading Percentages Correctly
- 3Equations: A Balance for Finding Unknown Numbers
- 4Functions and Graphs: An Eye for Change
- 5Exponents and Logarithms: A World That Grows by Multiplying
- 6Geometry Basics: Area and Pythagoras
- 7Probability: Putting Numbers on Uncertainty
- 8Statistics: Mean, Median, and Variance
- 9Reading Data: The Traps in Graphs
- 10The Math for Understanding AI