Why there are two regression lines, which predicts what, and the one place they agree.
Before you slide anything, ask how many regression lines a scatter has.
Almost everyone says one. There are two, and the figure draws both: they cross at the mean point and open out into an X as you move away from it.
Then slide y to the top of the data and read the two predictions: 8.70 the right way and 10.1 the wrong way. That gap is a whole mark, and it is invisible if you only ever predict near the middle of the data, which is exactly where textbook examples sit.
Multiply the two gradients and you get r². Here 1.405 × 0.4917 = 0.691 = r². So at r = ±1 the product is 1, the lines are exact inverses, and they coincide.
The angle between the two lines is a picture of how much r is missing. That one sentence ties 4.10 back to correlation and stops it being an arbitrary second formula to memorise.
| 1. Predict x at y = 10.5 | 5.5, the mean of x. Every least squares line passes through the mean point. |
| 2. r² from the gradients | 1.405 × 0.4917 = 0.691. |
| 3. Arm span from height | A, the line of arm span on height. What you are predicting goes first. |
1 markChoosing the line by what is being predicted, not by what the question names first.
1 markSubstituting correctly into the chosen line, without rearranging the other one.
1 markCommenting on reliability: weak r, or a value well outside the data.
| They give | What it means |
|---|---|
| 10.1 instead of 8.70 | They rearranged y on x to get x. This is the error the sub-topic exists to catch, and it produces a plausible number, which is what makes it dangerous. |
| 10.5 (Q1) | Gave y-bar when x-bar was asked for. A reading slip rather than a method one, but worth separating. |
| 0.831 (Q2) | Gave r, not r². Worth pointing out that the gradients multiply to r², never to r. |
| 1.897 (Q2) | Divided the gradients instead of multiplying. |
| "Either, they're the same" (Q3) | They have not accepted that there are two different lines. Send them back to the figure and slide away from the mean. |
"Which one is the real line?" Both. They answer different questions, and least squares only ever minimises in one direction at a time. Neither is a better fit than the other; they fit different things.
"Why do they cross at the mean?" Because both are constructed to pass through (x-bar, y-bar). That is worth stating as a fact they can use: any least squares line through your data hits the mean point.
"Does AI need this?" No. Applications has no x on y line at all, which is one of the cleaner differences between the two courses and worth naming if you teach both.
| What is happening | |
|---|---|
| 1 | How many regression lines are there? Collect answers, then show both. |
| 2 | Slide y to the extremes. Read the two predictions. Name the error. |
| 3 | The gradients-multiply-to-r-squared fact, checked on the calculator. |
| 4 | Calculator practice: get both lines from the same list, which many students have never done. |
| 5 | A question in context where the wrong line gives a believable answer. |
Do not say “the regression line” once this lesson has started. The definite article is the misconception.
Do not demonstrate only near the mean. The two lines agree there, so the demonstration proves the opposite of what you want.
Three tiers, ramping the way practice should: the method on its own, then the method inside something real, then a challenge. Set the tier the class in front of you needs rather than one undifferentiated sheet. Answers are given so these can go straight onto a board.
The method on its own, with friendly numbers. Set these first and move on quickly once they are secure.
The same skill inside a real situation, where the first job is working out what is being asked.
Reasoning, working backwards, or spotting an error. These are where the top grades are decided.
Works on a phone. Nothing is loaded from any other site, so it runs behind a school firewall, and nothing a student does is saved or sent anywhere.