Topic 4.13 · teacher page · AA Higher Level

The base rate is the whole story

The most consequential probability error there is, and the theorem that fixes it.

The one thing to do with the animation

Ask for the answer before you run it: a 99% accurate test comes back positive, what is the chance you have it?

Nearly everyone says 99%. The answer is about 9%, and the thousand dots show why: there are 999 healthy people to get wrong, and 1% of 999 beats 99% of 1.

Then raise the prevalence slowly and stop at 1 in 100, where it is exactly a coin flip. That crossover is the number worth leaving the lesson with, because it makes the dependence on the base rate concrete.

Why this is worth more than one lesson

This is the error behind misreported medical statistics, the prosecutor's fallacy in court, and most misunderstandings of screening programmes. It is the single most useful thing in Topic 4 outside an examination hall.

It also explains why doctors retest. A second independent positive starts from a prior of 0.09 rather than 0.001, and the answer climbs steeply. That is worth saying, because it turns a trick question into a piece of working knowledge.

The answers

1. P(A₁ | B)0.5 × 0.02 = 0.01, over 0.045, so 0.222.
2. The 50/50 prevalence0.01, one in a hundred, where true and false positives balance.
3. The positive resultB. About 9%, and a second independent test is worth doing.

Where the marks go

1 markA denominator built from total probability, with every branch present.

1 markEvents defined in words before any numbers go down.

1 markInterpreting the answer for the person in the question, not just reporting a decimal.

What each wrong answer tells you

They giveWhat it means
0.99 or 99%The base rate fallacy itself. Do not just correct it; make them count the false positives in the figure, or it comes straight back.
0.01 (Q1)The numerator alone, with no division by P(B).
0.02 (Q1)Gave P(B | A₁), the conditional they started with. Bayes reverses it, and recognising that is the point.
0.5 (Q1)Gave the prior. Worth asking what observing B was for, if the answer is unchanged.
"The test is useless" (Q3)Over-correction, and worth taking seriously. It moved them from 0.1% to 9%, a ninety-fold increase; that is a great deal of information, just not certainty.

Other things they will say

"So the test is broken?" No. It is working exactly as specified. The surprise comes from the rarity of the disease, not from the quality of the test, and separating those two is the lesson.

"Why do they test twice?" Because the second test starts from a much higher prior. Work it through if there is time; it is the most satisfying five minutes in the sub-topic.

"Is this on the AI course?" No, Bayes is AA Higher Level only. Worth saying if you teach both.

A possible order

 What is happening
1Take the prediction first, in writing, before anything is revealed. Then run the figure.
2Count the dots. Build the table of true and false positives by hand.
3The formula, derived from the table rather than quoted at them.
4The three-event version, with the posteriors summing to 1 as the check.
5The retesting question, and where this appears outside mathematics.

Two things not to say

Do not open with the formula. The formula is the easy part and leading with it wastes the one sub-topic students will still remember in ten years.

Do not leave them thinking the test is bad. It is a good test; the conclusion people draw from it is the bad part.

Questions to set

Three tiers, ramping the way practice should: the method on its own, then the method inside something real, then a challenge. Set the tier the class in front of you needs rather than one undifferentiated sheet. Answers are given so these can go straight onto a board.

1Fluency

The method on its own, with friendly numbers. Set these first and move on quickly once they are secure.

  1. The two ways to be positiveA condition affects 1 in 1000 and a test is right 99% of the time both ways. In 1000 people, find the true positives and the false positives.
    About 0.99 true positives against 9.99 false positives, so roughly 1 against 10.
  2. P(positive)Find the probability of testing positive.
    0.001(0.99) + 0.999(0.01) = 0.011

2In context

The same skill inside a real situation, where the first job is working out what is being asked.

  1. Apply BayesFind the probability of having the condition given a positive test.
    0.00099 / 0.011 = 0.0902, so about 9%.
  2. Change the prevalenceRepeat with a prevalence of 10% instead, test unchanged.
    0.099 / 0.108 = 0.917, so about 92%. The test never changed; only how many people actually have the condition.

3Challenge

Reasoning, working backwards, or spotting an error. These are where the top grades are decided.

  1. The base rate fallacyA patient is told a 99% accurate test means a 99% chance they have the condition. Explain the error and the number that is missing.
    The accuracy of the test is not the probability of the condition. What is missing is the prevalence. At 1 in 1000 the answer is 9%, because the far larger healthy group generates ten times more false positives than there are true cases.
  2. Why screening is targetedUse the two answers above to explain why screening programmes are offered to high-risk groups rather than everyone.
    The same test gives 9% in the general population and 92% in a group with 10% prevalence. Screening a low-prevalence population produces mostly false positives, with the cost and alarm that follow, which is an argument about arithmetic and not about the quality of the test.

Practicalities

Works on a phone. Nothing is loaded from any other site, so it runs behind a school firewall, and nothing a student does is saved or sent anywhere.