Topic 4.15 · teacher page · Higher Level

The data stays skewed

The single most misapplied theorem in the course, and the two-graph figure that fixes it.

The one thing to do with the animation

Point at the top graph and say "this one never changes" before you press anything.

Students arrive believing the central limit theorem makes data normal. The population graph stays stubbornly skewed at every value of n while the lower graph becomes a bell, and the two being on screen at once is what makes the distinction stick.

At n = 1 the lower graph is the population, which is worth pausing on. Everything after that is the theorem doing work.

What the theorem is actually about

It is a statement about the distribution of X̅, the sample mean, and about nothing else. Incomes stay skewed however many people you survey.

Every confidence interval and every test in this topic rests on it, which is why it is worth getting exactly right rather than approximately. If they think it applies to the data, they will use σ where they need σ/√n for the rest of the course.

The answers

1. se with σ = 6, n = 366 over √36 = 6/6 = 1.
2. Factor to halve the se4. The root of 4 is 2, so the standard error is divided by 2.
3. Which is normalB, the distribution of the sample mean. The 100 incomes keep the shape of the population they came from.

Where the marks go

1 markWriting X̅ ~ N(μ, σ²/n) with the numbers substituted.

1 markUsing σ/√n and not σ. This single slip invalidates every probability after it.

1 markJustifying normality: either the population is normal, or n is large enough.

What each wrong answer tells you

They giveWhat it means
6 (Q1)Gave σ itself. They have not distinguished the spread of individual values from the spread of the mean, which is the whole sub-topic.
0.1667 (Q1)Divided by n rather than by √n.
36 (Q1)Gave n.
2 (Q2)Doubling n divides the standard error by √2, about 1.41. A reasonable guess that the square root punishes.
16 (Q2)That quarters it. They have the right idea one step too far.
"Both" or "the 100 incomes" (Q3)The core misconception. If sampling made data normal, skewed data would not exist; that sentence usually ends it.

Other things they will say

"How large is large?" More than 30 for examination purposes, more if the parent is very skewed, and exactly normal for any n at all if the parent is normal. Say all three, because questions use all three.

"Why the square root?" Because variances add and standard deviations do not. It is also why precision is expensive: ten times as precise costs a hundred times the data.

"Does it work for proportions?" Yes, and that is where they will meet it again. Worth flagging without developing it here.

A possible order

 What is happening
1The two graphs, with the top one named as the one that never changes.
2The statement of the theorem, written out with X̅ in it.
3Standard error, the square root, and the three questions.
4A probability question about a sample mean, done wrong with σ first, deliberately, then right.
5Say out loud that confidence intervals next lesson are this theorem with brackets round it.

Two things not to say

Do not say “everything becomes normal”. It is the sentence that creates the misconception this entire page exists to remove.

Do not introduce the standard error without naming it. Questions use the term and students who have only seen “sigma over root n” freeze at it.

Questions to set

Three tiers, ramping the way practice should: the method on its own, then the method inside something real, then a challenge. Set the tier the class in front of you needs rather than one undifferentiated sheet. Answers are given so these can go straight onto a board.

1Fluency

The method on its own, with friendly numbers. Set these first and move on quickly once they are secure.

  1. Standard errorA population has σ = 15. Find the standard error of the mean for n = 36.
    15 / √36 = 2.5
  2. Halve the errorWhat sample size would halve that standard error?
    n = 144. The error goes with √n, so four times the data for half the error.

2In context

The same skill inside a real situation, where the first job is working out what is being asked.

  1. Probability for a meanThat population has μ = 100. For a sample of 36, find P(x̅ > 105).
    z = 5 / 2.5 = 2, so the probability is 0.0228.
  2. Which one is normalThai household incomes are strongly skewed. A sample of 100 is taken. State what is approximately normal and what is not.
    The distribution of the sample MEAN is approximately normal. The 100 incomes themselves stay as skewed as the population, however large the sample.

3Challenge

Reasoning, working backwards, or spotting an error. These are where the top grades are decided.

  1. The cost of the error lawA survey of 400 is criticised for being too small, and the budget allows 25% more. By what factor does the standard error fall, and is it worth it?
    n = 500 gives a factor of √(400/500) = 0.894, so the error falls by about 11% for 25% more cost. The square root is why large surveys stop getting much better.
  2. When it does not helpThe theorem says the sample mean tends to normal. Give a situation where a bigger sample does not rescue the study.
    When the sample is biased. The theorem is about random error around the sampling distribution's centre; if that centre is in the wrong place, more data gives a tighter estimate of the wrong number.

Practicalities

Works on a phone. Nothing is loaded from any other site, so it runs behind a school firewall, and nothing a student does is saved or sent anywhere.