B Binance · The world's largest crypto exchangeBinance Sign up → AD OKX OKX · A leading global crypto exchangeOKX Sign up → AD
📐 Math Basics · Lesson 9 / 10

Reading Data: The Traps in Graphs

Numbers don't lie, but the way they are presented easily creates misunderstandings. The habit of checking the axes, the base, the sample, and what is being compared is what lets you read data correctly.

⏱ About 16 min ✍️ 4 practice questions Updated 2026-10-08
🎯 By the end of this lesson you can
  • Recalculate the real difference shown in a graph with a truncated vertical axis
  • Tell correlation from causation and think of hidden factors
  • Ask the questions that reveal whether a sample is biased
  • Convert between counts and rates, and between relative and absolute change

1.Truncated Axes: Making Small Differences Look Big

The heights of the bars in a bar chart are there so you can compare the sizes of the values by eye. But if the vertical axis starts somewhere other than 0, the ratio of the bar heights no longer matches the ratio of the actual values. A small difference ends up looking like a difference of several times.

Truncating the vertical axis is not always wrong. Line graphs that show small changes in values that never come near 0, such as body temperature or stock prices, often cut the axis. The problem is cutting the axis while inviting you to compare sizes by bar length. When you look at a graph, check the starting value and the tick marks on the vertical axis first.

ExampleExample: Product A has a satisfaction score of 100, and Product B has 104. In a bar chart whose vertical axis starts at 98, how many times as tall does B's bar look compared with A's? What is the real difference in percent?
  1. Step 1: Bar heights as shown on the chart: A is 100 − 98 = 2 units, and B is 104 − 98 = 6 units.
  2. Step 2: Apparent ratio: 6 ÷ 2 = 3 times.
  3. Step 3: Actual ratio: 104 ÷ 100 = 1.04, so B is only 4% higher than A.
  4. Check: If you draw the vertical axis from 0, the bar heights are 100 and 104, a difference that is hard to tell apart by eye.
AnswerIt looks like 3 times on the chart, but the real difference is 4%

2.Correlation Is Not Causation

When two values rise and fall together, they are correlated. In summer, ice cream sales go up and so do swimming accidents. The two are clearly correlated, but ice cream does not cause accidents. A 3rd factor, temperature, moves them both. A hidden factor like this is called a confounding variable.

Before reading a correlation as causation, ask three questions. Is there a hidden common cause? Could the direction of cause and effect be reversed (are people who exercise healthy, or do healthy people exercise more)? Could it be a coincidence? The surest way to establish causation is an experiment comparing two randomly assigned groups, but such an experiment isn't always possible.

  • Common cause: temperature → ice cream sales, temperature → swimming accidents
  • Reversed direction: "People who go to the hospital often are sicker"; they go to the hospital because they are sick
  • Coincidence: two unrelated statistics can move in step if you choose the time period carefully

3.Sampling Bias: Whom Did You Ask?

Since you can't survey everyone, you draw a part (a sample) and use it to estimate the whole. If the sample doesn't resemble the whole, the results are skewed no matter how much data you collect. A survey run only inside a particular app reflects the opinions of people who use that app, and a survey of only people who volunteered to answer is closer to the opinions of people with a strong interest in the topic.

Looking only at the survivors is also sampling bias. The claim "successful people all got up at dawn" leaves out the people who got up at dawn and did not succeed. When you look at data, ask "who was chosen, how, and who was left out?" before you look at the numbers. How a sample is drawn often matters more than how big it is.

AI has the same problem. If the training data is skewed toward a particular group, the AI's judgments are skewed in that direction too. The data bias covered in the AI Basics course is exactly this sampling bias.

4.Rates vs. Counts, Relative vs. Absolute Change

If you directly compare the number of accidents in a city of 1,000,000 people and a city of 100,000 people, the big city will always look worse. For a fair comparison, you need rates converted to the same base, such as per 100,000 people. Conversely, if you look only at rates and miss the counts, you misjudge the actual scale.

"The risk has doubled" carries the same problem. Relative change (doubling, a 100% increase) sounds big, but if the original risk was very small, the absolute change may be tiny. You need to look at both to judge.

ExampleExample: City A has a population of 1,000,000 and 500 accidents a year, and City B has a population of 100,000 and 100 accidents. Compare them by accidents per 100,000 people.
  1. Step 1: City A: 500 ÷ 1,000,000 × 100,000 = 50 accidents.
  2. Step 2: City B: 100 ÷ 100,000 × 100,000 = 100 accidents.
  3. Step 3: In counts, A has 5 times as many, but relative to population, B's rate is 2 times as high.
  4. Check: A has 10 groups of 100,000 people, and 500 ÷ 10 = 50; B has 1 group, so 100. Same result.
AnswerPer 100,000 people: A 50, B 100. By rate, B is 2 times as high
ExampleExample: A side effect went from 2 people in 10,000 to 4 people. Find the relative change and the absolute change.
  1. Step 1: Relative change: (4 − 2) ÷ 2 = 1 → a 100% increase (doubling).
  2. Step 2: Absolute change: from 2 ÷ 10,000 = 0.02% to 4 ÷ 10,000 = 0.04%, an increase of 0.02 percentage points.
  3. Check: 0.04 − 0.02 = 0.02 percentage points, which is the same as 2 more people out of 10,000.
AnswerRelatively, it doubled (a 100% increase); absolutely, it rose by 0.02 percentage points (2 people out of 10,000)

5.The Traps of Averages and Grouping

Can you wade across a river whose average depth is 1 m? If it is shallow at the edges and 3 m deep in the middle, you can't, even though the average is 1 m. An average tells you nothing about spread or the maximum. As you saw in Lesson 8, the median and the spread belong next to the mean.

Sometimes the results for grouped data and for the separate parts point in opposite directions. This is called Simpson's paradox. In the example below, Hospital X has a higher success rate than Y for both easy and difficult surgeries, but combined, Y looks far better. That's because X took on far more difficult surgeries.

Example: surgery success rates at two hypothetical hospitals
Easy surgeryDifficult surgeryOverall
Hospital X9/10 (90%)30/90 (about 33%)39/100 (39%)
Hospital Y85/100 (85%)3/10 (30%)88/110 (80%)
ExampleCalculate the overall success rate of each hospital in the table above yourself, and explain why the conclusion flips.
  1. Step 1: X overall: (9 + 30) ÷ (10 + 90) = 39 ÷ 100 = 39%.
  2. Step 2: Y overall: (85 + 3) ÷ (100 + 10) = 88 ÷ 110 = 80%.
  3. Step 3: For X, 90 of 100 patients had difficult surgery; for Y, only 10 of 110 did. The overall success rate depends heavily on the mix of patients each hospital took on.
  4. Check: Easy surgery 90% > 85%, difficult surgery about 33.3% > 30%, so X is higher in both groups. Yet overall it is 39% < 80%.
AnswerX 39%, Y 80%. Because the patient mix (the share of difficult surgeries) differs, combining the groups flips the conclusion
A data-reading checklist: Where does the axis start? What is the base (the denominator)? Who was surveyed? Is it a rate or a count? Is it grouped or broken down? Is there another explanation (a hidden factor)?

📌 Key points

  • If a bar chart's vertical axis doesn't start at 0, differences look exaggerated
  • Correlation is not causation; first suspect a common cause, the direction, or coincidence
  • How a sample is drawn matters more than its size; ask who was left out
  • Look at counts and rates, and at relative and absolute change, together
  • A single average can't show spread, and grouped results can be the opposite of the broken-down ones

✍️ Practice questions

Answer first, then open "Answer and explanation".

Q1. Example: In a bar chart whose vertical axis starts at 90, A is 95 and B is 100. How many times as long does B's bar look compared with A's?

⭕ Correct

❌ Not quite — see the explanation

Answer and explanation
Answer ② 2 times

The visible heights are 95 − 90 = 5 for A and 100 − 90 = 10 for B, so it looks 2 times as long. The actual ratio is 100 ÷ 95 ≈ 1.05 times.

Q2. A correlation has been observed: the more coffee people drink, the less they sleep. Which statement is the most appropriate based on this alone?

⭕ Correct

❌ Not quite — see the explanation

Answer and explanation
Answer ② It may be that people who lack sleep drink more coffee

Correlation alone can't tell you the direction of causation. The reverse direction, drinking more coffee because of lack of sleep, or a common cause such as working late, is also possible.

Q3. Example: A city of 200,000 people had 60 accidents in a year. How many accidents is that per 100,000 people?

Answer and explanation
Answer 30 accidents

60 ÷ 200,000 × 100,000 = 30 accidents. Since there are 2 groups of 100,000 people, you can also get it as 60 ÷ 2 = 30.

Q4. Team A has a higher success rate than Team B for both easy and difficult tasks, yet Team B's overall success rate came out higher. What is the most plausible reason?

⭕ Correct

❌ Not quite — see the explanation

Answer and explanation
Answer ② Team A took on far more difficult tasks

This is Simpson's paradox. When the share of each group differs, the overall success rate depends more on the mix of tasks taken on than on skill.

🤖 Try asking AI like this

Copy a prompt and replace the [ ] parts with your own situation. Don't take the answer on trust — check it against this lesson.

When you want to read a graph in a news story or report critically

Look at the description of the graph (or figures) below and go through a checklist of which traps might apply: truncated axes, the base (denominator), sampling bias, correlation vs. causation, rates vs. counts, and grouped statistics. For each item, also write what additional information would be needed to check it. Content: [paste here]

When you want to practice spotting graph traps

Make 5 misleading graph descriptions using made-up numbers. Keep the trap in each one hidden, and show me the answer and the correct calculation after I've found it. Be sure to label them "Example" so they don't look like real statistics.

When you receive a data summary made by an AI

For every figure you used in the summary you just wrote, state its source and its base (what you used as the denominator). Mark separately any figures whose source can't be verified.
References
  • Standard high school math textbook content (statistics, sample surveys)
  • Standard explanations in introductory statistics textbooks (correlation and causation, Simpson's paradox)

Reached every goal above? Mark the lesson complete.

📐 Math Basics

  1. 1Numbers and Operations: Fractions and Decimals Revisited
  2. 2Ratios and Rates: Reading Percentages Correctly
  3. 3Equations: A Balance for Finding Unknown Numbers
  4. 4Functions and Graphs: An Eye for Change
  5. 5Exponents and Logarithms: A World That Grows by Multiplying
  6. 6Geometry Basics: Area and Pythagoras
  7. 7Probability: Putting Numbers on Uncertainty
  8. 8Statistics: Mean, Median, and Variance
  9. 9Reading Data: The Traps in Graphs
  10. 10The Math for Understanding AI
📚 Worth reading
🧠What Generative AI Does and Where It Fails→ ✍️How to Write a Good Prompt→ 🔍Checking AI Answers→ 📚Using AI for Study and Work Without Plagiarism→
← Foundations for the AI Era