IB Mathematics Internal Assessment
A calculator or a spreadsheet will fit almost any function family to almost any shape of data, and it will hand you an R² value that looks reassuring. That is exactly the trap. A model earns marks when you can say why this function and not another, and when it survives a test you did not have to run. Here is where the two get confused, and how to stay on the right side of the line.
Curve fitting has a familiar shape. A student collects some points, looks at the scatter, picks a function family that roughly matches the outline, runs a regression, and reports the R² value as if it settles the matter. The working is real mathematics and the graph looks tidy, but nothing in the process asked whether that function made sense for that situation. Swap the data for a different but similarly shaped dataset and the same method would produce the same conclusion, because the method was never actually about the situation at all.
Modelling asks a different question first. Not "what shape fits", but "what process would produce a shape like this, and does that process match a function I can justify". The arithmetic can look almost identical on the page. The reasoning around it is what a reader is actually marking.
Every standard function family carries an assumption about how a quantity behaves. A straight line assumes a constant rate of change. An exponential assumes growth or decay proportional to the current value. A sine curve assumes a repeating cycle with a fixed period. A logistic curve assumes growth that slows as it approaches a limit. If you can name the assumption and explain why your situation should behave that way, you are modelling. If you cannot, you are pattern matching.
Four candidate functions get tried in the calculator's regression menu, and the cubic wins on R², so the cubic goes in the report. Nothing is said about what a cubic would mean for the quantity being modelled, and a fifth candidate might well have won on a slightly different data set. The choice is arithmetic, not reasoning.
The quantity is a population, a temperature difference, or an amount of a radioactive substance, and the situation is one where the rate of change genuinely depends on the current amount. The exponential is chosen before the regression is run, because the mechanism predicts it, and the regression is then used to estimate parameters, not to go hunting for a shape.
A high R² tells you a function passes close to the points you already have. It does not tell you the function means anything, and it becomes actively misleading once you are choosing between functions with different numbers of parameters. A high degree polynomial can be made to pass almost exactly through a small set of points, giving an R² near 1, while doing nothing but memorising your data and producing nonsense the moment you step outside it. If your fitted model is a polynomial with more terms than you have a mechanism to justify, that is usually a sign the model was chosen to match the shape rather than the situation.
Look at what your model predicts just past the edge of your data, and ask whether that prediction is remotely sensible. A model of somebody's running time that predicts they finish in negative minutes if the race gets a little longer has failed a test that R² will never catch, because R² only ever looks backward at the points you already had.
The strongest thing you can do with a model is ask it to predict something you did not use to build it, and then check. If you collected data, hold a portion back before you fit anything, fit the model on the rest, and compare its prediction against the values you set aside. If you cannot hold data back, look instead at the residuals, the gaps between what the model predicts and what actually happened, and ask whether they are small and patternless or whether they show a shape of their own. A residual pattern is the model quietly telling you it is missing something about the situation.
This is also where a comparison between two reasonable candidate models becomes genuinely valuable, rather than decorative. Fitting two functions and picking whichever has the marginally higher R² is still curve fitting. Fitting two functions, each justified on different grounds, and then showing with residuals or held-back data which one actually describes the behaviour better, is modelling with an argument behind it.
Students often treat a poor fit as something to quietly improve by trying another function until the number looks better. Treated honestly, a poor fit is one of the most useful things that can happen in an exploration. It means your assumption about the mechanism was wrong, or only partly right, and working out which part is exactly the kind of thinking that earns marks in evaluation and reflection. Say specifically what the model got wrong, where, and what that implies about the real situation, rather than simply swapping in a better looking curve and moving on.
Both routes are marked against the same five criteria, and modelling belongs in either. AI leans naturally toward modelling because its syllabus content is built around real situations and data, so a modelling exploration often sits close to the centre of what the course expects. AA can model just as well, but the justification for the function often has to come from more structural or analytic reasoning rather than from a described real-world mechanism. What changes most is not AA against AI but SL against HL: at HL a single unjustified regression is not going to reach the top band, and the expectation is that you handle the choice, the testing and the limitations of your model with real depth.
Not sure whether your model is earning its place?
Send me the function you have chosen and why. I read real submitted explorations every year, and I will tell you honestly whether it is modelling or curve fitting, while there is still time to change it.
Get the IA sorted with meMore like this: all Maths IA guides. Related: using your own data in the Maths IA
A new IA guide goes up most nights. If you would rather not keep checking, leave an email and I will send the one that matters that week. No selling, and leave whenever you like.
If you are under 18, use a parent's email. I do not correspond privately with students.
These pages are free and stay free, but they are general and your IA is not. Send me your research question, or whatever exists so far, and I will tell you in writing whether the topic has a ceiling on it, where the marks are going, and what to change first. That costs nothing and it comes back within 24 hours.
Written by a serving IB Diploma and Career-related Programme Coordinator and Head of Mathematics, who reads internal assessments across every subject group every year. If you then want the whole draft reviewed properly against all five criteria, that is the paid one, and it is refunded if it does not name at least three specific things to fix.
Get a free verdict Full written review, $99
I never write any part of it. Not a sentence, not a calculation, not your data. Under 18: a parent buys this and the thread is with them. I do not work with students at my own school.