Topic 4.13 · teacher page · HL only

Running non-linear regression

The overfitting demonstration, the numbers behind it, and why this lesson is about judgement rather than curve fitting.

The one thing to do with the widget

Click through linear, quadratic, cubic and have them call out the SSres each time. It falls every time, and they will start to believe the cubic is the answer. Let them.

Then ask what the coffee is doing an hour later. That is the moment the lesson turns, and it is far more effective than warning them about overfitting beforehand.

The numbers

Every figure below was computed from the page's own data, not estimated.

ModelSSresAt 60 minutes, room is 24
Linear211.50−24.4, the coffee has frozen
Quadratic7.41+89.1, it has reheated to near its starting temperature
Cubic0.31−37.9, colder still
Exponential4.6925.7, just above room temperature

The cubic's 0.31 is worth pointing at. With fifteen data points it is very nearly threading every one, which is the curve memorising the measurement noise rather than learning the physics.

The point that makes the lesson. All three polynomials eventually run away, and they do it in different directions. A polynomial has no asymptote to settle onto, so none of them can ever describe something that approaches a limit and stops. No amount of fitting could have told you that. The physics chose the model and SSres only compared the candidates that were already sensible.

Answers

QuestionAnswer
1. Contribution to SSres7.84. 41.0 − 38.2 = 2.8, and 2.8² = 7.84.
2. Cubic beats quadraticB. More parameters always fit at least as well.
3. Predicting at 60 minutesB. The exponential, because cooling approaches room temperature.

The worked example on the student page: observed 51.4, predicted 54.0, so the residual is −2.6 and it contributes 6.76.

What each wrong answer tells you

They enterWhat it means
2.8 (Q1)They gave the residual rather than its square. Very common, and worth one sentence: the second S in SSres is for squared.
−2.8 (Q1)Predicted minus observed. The convention is observed minus predicted, though the square is the same either way. Worth correcting now because the sign matters when they plot residuals.
"Too close to compare" (Q2)They think it is a precision problem. It is not: even a tiny improvement is expected purely from the extra parameter.
"SSres cannot compare models" (Q2)Over-corrected. It compares models with the same number of parameters perfectly well. Praise the caution, fix the scope.
"The cubic" (Q3)The one that matters. They have followed the number. Send them back to the widget and ask where the cubic is at 60 minutes, which is minus 37.9 degrees.
"None, it is extrapolation" (Q3)Good instinct from 4.4, slightly too strong here. A model justified by the physics can be used a little beyond the data if you declare it. Worth a short discussion of the difference between extrapolating a fitted line and extrapolating a mechanism.

Where the marks go

Computing SSres from a table is routine and marked on method: observed minus predicted, square, add. They keep marks through an arithmetic slip if the three steps are visible.

Justifying a model needs the context. "The exponential, because cooling tends towards room temperature" earns the mark. "The exponential, because SSres was smallest" often does not, because on this data it would have picked the cubic.

If they fit with technology, they should state which family they asked for and why. The calculator cannot know that a cooling curve needs an asymptote.

A possible order

 What is happening
1Linear first. Let them see a line visibly failing on obviously curved data, which is also a callback to r only measuring linear association in 4.4.
2Residuals and SSres, with the worked example done by hand once.
3Click up through the polynomials, calling out SSres. Build the belief that smaller is better before breaking it.
4The 60 minute question. This is the lesson. Give it room.
5Questions 1 to 3, then the general principle: choose the family from the context, then use SSres to pick within it.

Two things not to say

Do not warn them about overfitting before they have watched SSres fall three times. The warning costs nothing to give and nothing is learned from it. The surprise is the teaching.

Do not say "the exponential is the best model" without saying why. On SSres alone it is not the best, it is third. If a student notices that and you have no answer ready, the lesson inverts.

Questions to set

Three tiers, ramping the way practice should: the method on its own, then the method inside something real, then a challenge. Set the tier the class in front of you needs rather than one undifferentiated sheet. Answers are given so these can go straight onto a board.

1Fluency

The method on its own, with friendly numbers. Set these first and move on quickly once they are secure.

  1. LineariseTo fit T = aebt, what do you plot against t to get a straight line?
    ln T. The gradient is b and the vertical intercept is ln a.
  2. Read the fitA fit gives gradient −0.173 and a = 80.0. Write the model.
    T = 80.0e−0.173t

2In context

The same skill inside a real situation, where the first job is working out what is being asked.

  1. Predict and judgeCoffee temperatures above room temperature at t = 0, 2, 4, 6, 8 minutes are 80, 56.6, 40, 28.3, 20. Use the model to predict t = 12.
    80.0e−0.173(12) = 10.0 degrees above room temperature.
  2. Half-lifeFind the time for the excess temperature to halve, and check it against the data.
    ln2 / 0.173 = 4.0 minutes. The data agrees: 80 to 40 takes 4 minutes, and 40 to 20 takes another 4.

3Challenge

Reasoning, working backwards, or spotting an error. These are where the top grades are decided.

  1. Choose between modelsA straight line also fits this coffee data with r = −0.97. Give two reasons to prefer the exponential.
    The line must cross zero and then predict a drink colder than the room, which is physically impossible. And cooling has a mechanism, proportional loss, that the exponential encodes and the line does not. A high r does not make a model right.
  2. Spot the invalid stepA student fits ln T against t, gets a gradient of −0.173, and reports that T falls by 0.173 degrees a minute. Correct them.
    The gradient belongs to ln T, not T. It means T falls by a constant 17.3% of itself per minute roughly, not a constant 0.173 degrees. Units from a transformed axis are the easiest marks to lose.

Practicalities

Works on a phone, though the curves past the data are easier to see on something larger. The library is served from this site rather than a public CDN, so it works behind a school firewall. Nothing a student types is saved or sent anywhere.