HomeLearning HubIB DP PhysicsI.3 Concluding and evaluating
I.3

Concluding and evaluating

Tools and inquiry · skills assessed across the whole course

The shortest stage and the most revealing. Anybody can collect data; a conclusion shows whether you understood what you collected, and an evaluation shows whether you know which part of your own experiment was the weakest.

🎯What you need to be able to do

  • State a conclusion that answers the research question, supported by the data.
  • Compare a result with an accepted value using the uncertainty, not the percentage difference alone.
  • Say what the gradient and intercept of your graph physically mean.
  • Distinguish the effect of random error from that of systematic error on your result.
  • Identify methodological weaknesses and rank them by how much they mattered.
  • Propose realistic improvements, each addressing a named weakness.
  • Suggest a sensible extension to the investigation.

Writing a conclusion

A conclusion answers the research question — in a sentence, using the data. Everything else in it is support for that sentence. Three things belong there:

the answerthe relationship you found, stated in words
the evidencethe gradient, intercept or value, with its uncertainty
the comparisonhow that sits against theory or an accepted value

Say what the gradient and intercept actually mean. A gradient of 4.02 s² m−1 is not a conclusion; “the gradient is \( 4\pi^{2}/g \), giving \( g = 9.8 \pm 0.6 \) m s−2” is. And an intercept that theory says should be zero is worth a sentence whether it was zero or not.

⚖️Comparing with an accepted value

Two results for the acceleration of free fall plotted as horizontal uncertainty bars against a scale from 9.0 to 10.2 metres per second squared, with the accepted value of 9.81 marked by a dashed vertical line. The first result, 9.4 plus or minus 0.6, has a bar that reaches past the accepted value, and is labelled as agreeing with it. The second, 9.4 plus or minus 0.2, has a bar that stops short, and is labelled as not agreeing, which is itself a finding. Two panels explain what to write in each case: if the ranges overlap, say plainly that the accepted value lies within the uncertainty range so the experiment is consistent with theory, noting that this does not prove the theory but fails to contradict it; if they do not overlap, something is wrong and finding out what is the interesting part, whether an unidentified systematic error, an underestimated uncertainty or an inapplicable theory. A closing note observes that both results are 4 per cent from the accepted value, so a percentage difference on its own answers nothing.
Both results are the same 4% from 9.81. One agrees with theory and one does not, and the only difference is the uncertainty — which is why the comparison has to be made that way round.

So the question is never “how close was I?” but “does the accepted value lie inside my uncertainty range?” If it does, the experiment is consistent with theory. If it does not, that is a real result and worth chasing: it means an unidentified systematic error, an underestimated uncertainty, or a theory that does not apply in your conditions.

The trap: a result that disagrees is not a failed experiment. The temptation is to inflate the uncertainty until the accepted value creeps inside, or to describe a 12% discrepancy as “fairly close”. Neither is physics. A disagreement that you investigate — and can attribute to a specific systematic effect — is worth considerably more than an agreement you arrived at by rounding generously.

🔍What each kind of error did

The evaluation has to distinguish two quite different things, because they have different fixes and they show up differently in the data.

random errorshows up as SCATTER of points about the line, and as the size of the error bars
systematic errorshows up as an unexpected INTERCEPT, or as a gradient consistently off

That second one is the useful diagnostic. If your graph should pass through the origin and does not, you have measured a systematic error rather than merely suspected one — and its size is the intercept. A zero error on an instrument, a background count, a lead resistance, all announce themselves this way.

✏️Worked example 1 — a conclusion, written out

An experiment plots \( T^{2} \) against \(l\) for a pendulum. The best-fit line has gradient \( 4.02 \pm 0.26 \) s² m−1 and an intercept of \( 0.05 \pm 0.08 \) s². Write the conclusion.

The relationship. The graph of \( T^{2} \) against \(l\) is a straight line, so \( T^{2} \propto l \) and therefore \( T \propto \sqrt{l} \), as theory predicts.

What the gradient gives. Theory says the gradient is \( 4\pi^{2}/g \):

\[ g = \frac{4\pi^{2}}{4.02} = 9.82\ \text{m s}^{-2} \]
\[ \frac{\Delta g}{g} = \frac{0.26}{4.02} = 6.5\%, \qquad \Delta g = 0.64\ \text{m s}^{-2} \]
\[ g = 9.8 \pm 0.6\ \text{m s}^{-2} \]

The comparison. The accepted value of 9.81 m s−2 lies well inside that range, so the result is consistent with theory.

The intercept. Theory predicts zero, and the measured intercept is \( 0.05 \pm 0.08 \) s² — which includes zero. So there is no evidence of a systematic error in the length or timing measurements, and saying so is part of the conclusion.

What makes this a conclusion rather than a summary. It answers the question (\( T \propto \sqrt{l} \)), gives a value with its uncertainty, compares that against the accepted figure using the uncertainty, and interprets the intercept as well as the gradient. Note the last point especially: an intercept consistent with zero is a positive finding, not a non-event, and most reports forget to mention it at all.

🔧Weaknesses and improvements

A table pairing six methodological weaknesses with what each did to the result and a specific improvement. Timing by hand introduced a random reaction-time error of about 0.2 seconds, fixed by light gates and a data logger. The wire warming up made its resistance rise during readings, a systematic effect, fixed by using a lower current and switching off between readings. Only five values over a narrow range left the gradient barely determined, fixed by widening the range and taking eight to ten values. No repeats meant no way to spot an anomalous reading, fixed by repeating three times and taking a mean. A thermometer reading only to one degree dominated the uncertainty budget, fixed by a thermocouple or temperature probe. Ignoring air resistance made the model fail at high speeds, addressed by stating it as a limit or modelling it iteratively. Two panels contrast improvements that are not improvements, such as be more careful or use better equipment, with the test a real improvement passes: it names a specific change and says which weakness it addresses.
Each row is a triple: the weakness, what it did to the result, and a change that would fix it. A weakness with no consequence attached, or an improvement with no weakness attached, is half an answer.

The single most useful thing an evaluation can do is rank the weaknesses. Go back to the uncertainty budget: whichever measurement contributed the largest percentage is where the experiment was really limited, and that is the one to lead with. A list of ten equal-sounding problems says you have not worked out which mattered.

The trap: “human error” is not a weakness. Neither is “be more careful”, “use better equipment” or “repeat more times” when the dominant error was systematic. Each names a feeling rather than a mechanism. A usable weakness says what went wrong, which direction it pushed the result, and roughly how much — and its improvement names a specific change to the apparatus or the method.

✏️Worked example 2 — diagnosing a disagreement

A student measures the resistivity of a wire and gets \( (5.8 \pm 0.3) \times 10^{-7}\ \Omega\ \text{m} \). The accepted value for the alloy is \( 4.9 \times 10^{-7}\ \Omega\ \text{m} \). Their graph of \(R\) against \(l\) is straight, with an intercept of \( +0.42\ \Omega \) where theory predicts zero. Evaluate.

Does it agree? The range is \( 5.5 \) to \( 6.1 \times 10^{-7} \), and 4.9 is outside it. So no — and the disagreement is substantial, roughly three times the quoted uncertainty.

Which way is it wrong? The measured resistivity is too high. Since \( \rho = RA/l \), that means either \(R\) came out too high, \(A\) too high, or \(l\) too low.

The intercept is the clue. A graph of \(R\) against \(l\) should pass through the origin: zero length, zero resistance. An intercept of \( +0.42\ \Omega \) says every resistance reading was about \( 0.42\ \Omega \) too high — almost certainly the resistance of the leads and the crocodile clips, in series with the wire in every measurement.

Does that explain the size of the discrepancy? If the wire resistances were of order a few ohms, an extra \( 0.42\ \Omega \) on each is a substantial systematic addition — in the right direction, and plausibly of the right size.

And the fix. Use the gradient rather than individual \( R/l \) values: a constant offset shifts the intercept but leaves the gradient untouched, so the resistivity from the gradient would have been free of this error all along. Alternatively, measure the lead resistance directly by touching the clips together and subtract it from every reading. Note how much better this is than “there may have been systematic errors” — the intercept identified the error, its sign, roughly its size, and the fix.

➡️Extending the investigation

A good extension asks the next question rather than repeating this one more carefully. It usually comes from one of three places: a variable you held constant that could now be varied; a range you could not reach with the apparatus available; or an assumption the analysis relied on that could be tested directly.

from a control variablehaving done \(R\) against \(l\), now do \(R\) against cross-sectional area
from a range limitdoes the pendulum result still hold at amplitudes above 15°?
from an assumptionthe analysis assumed air resistance was negligible — measure at what speed it stops being

📝Practise

Work through these, then reveal the answer. Each question targets a different objective from the list above.

1. A student concludes: “My value for \(g\) was 9.4 m s−2, which is 4% from the accepted value, so the experiment was quite successful.” Explain what is missing.
The uncertainty. A percentage difference on its own cannot say whether a result agrees with theory, because agreement is a question about whether the accepted value falls inside the uncertainty range.
If the result were \( 9.4 \pm 0.6 \) m s−2, then 9.81 lies inside the range and the experiment is consistent with theory. If it were \( 9.4 \pm 0.2 \), then 9.81 lies outside and something is wrong — despite being the same 4% away.
“Quite successful” is also doing no work: the conclusion should state whether the result is consistent with theory and, if not, what that suggests.
2. A result is \( 2.6 \pm 0.2 \) and the accepted value is \( 3.1 \). State whether it agrees, and describe the two possibilities.
The uncertainty range is 2.4 to 2.8, and 3.1 lies outside it, so the result does not agree with the accepted value.
Two possibilities, and they need different responses:
An unidentified systematic error. Something shifted every reading in the same direction — a zero error, a miscalibrated instrument, a consistently flawed technique, or a quantity that was not held constant. Look for an unexpected intercept on the graph.
An underestimated uncertainty. The true uncertainty may be larger than ± 0.2 — perhaps a source of error was not included in the budget at all. That is a genuine possibility, but it must be argued from evidence, not assumed because it would be convenient.
3. Explain how random error and systematic error each show up differently on a graph.
Random error shows up as scatter: the points do not lie exactly on the best-fit line, but they fall on both sides of it with no pattern. It also determines the size of the error bars, and hence the uncertainty in the gradient found from the steepest and shallowest lines.
Systematic error does not scatter the points at all — it shifts them. It typically shows up as an unexpected intercept: a graph that theory says should pass through the origin but does not. It can also appear as a gradient that is consistently wrong, if the error scales with the measurement rather than being a fixed offset.
The practical consequence: scattered points around a line through the origin means poor precision but no obvious systematic problem, while tightly grouped points on a line with an unexplained intercept means the opposite.
4. A graph of extension against load should pass through the origin but has an intercept of −0.4 cm. Suggest a cause and a correction.
A negative intercept means that at zero load the graph predicts a negative extension — that is, the measured extensions are all about 0.4 cm too small. The likely cause is that the unstretched length was measured incorrectly: perhaps the ruler's zero was not aligned with the top of the spring, or the initial length was recorded while the spring was already under a small load such as the mass of the hanger.
That is a systematic error: every extension carries the same offset.
Correction: add 0.4 cm to every extension, or better, use the gradient to find the spring constant. A constant offset shifts the intercept but not the gradient, so \(k\) obtained from the gradient is unaffected by this error — which is one of the main reasons for plotting a graph rather than calculating from single readings.
5. Explain why “human error” and “be more careful” are not acceptable in an evaluation.
Neither identifies a mechanism. “Human error” could mean reaction time, parallax, misreading a scale, or transcribing a number wrongly — each with a different effect and a different fix, so the phrase communicates nothing about what actually went wrong.
“Be more careful” is not a change to the method: the same apparatus, operated by the same person with the same technique, will produce the same uncertainty however much resolve is applied.
A usable weakness names what went wrong, which direction it pushed the result and roughly how much; a usable improvement names a specific change — “replace hand timing with light gates, removing the ± 0.2 s reaction-time uncertainty”.
6. Why is “take more repeats” a poor improvement when the dominant error is systematic?
Repeating reduces random error, because the scatter falls on both sides of the true value and averaging converges on it. A systematic error shifts every reading in the same direction by the same amount, so every repeat contains it equally and so does their mean.
Taking a hundred readings with a micrometer that has a − 0.03 mm zero error gives a beautifully precise answer that is still 0.03 mm wrong. The extra effort improves the precision, does nothing for the accuracy, and may even be counter-productive if it creates false confidence.
The fix for a systematic error is to find its cause and correct for it — measure the zero error and subtract it, measure the lead resistance and subtract it, subtract the background count.
7. Given an uncertainty budget of: length 0.2%, diameter 5.2%, current 4%, potential difference 0.7% — state which weakness to lead with and what to propose.
Lead with the diameter, at 5.2%. It is the largest single contribution, and it dominates because the cross-sectional area depends on \(d^{2}\), so the 2.6% uncertainty in the diameter measurement doubles.
Proposals, in order of usefulness: measure the diameter at more points along the wire and average, since the wire may not be perfectly uniform; use a thicker wire, so the same ± 0.01 mm micrometer uncertainty is a smaller percentage; or use a micrometer with a finer resolution if one is available.
Note what not to lead with: improving the length measurement from 0.2% would be invisible in the final answer, and proposing it first suggests the budget was never examined.
8. Distinguish a weakness from a limitation, and give an example of each.
A weakness is a flaw in how the investigation was carried out — something that could have been done better with the same apparatus and the same afternoon. Example: taking only one reading at each value instead of three, so an anomalous result could not be identified.
A limitation is a constraint on what the investigation could establish at all, given the apparatus, the time or the physics. Example: the conclusion applies only to amplitudes below about 15°, because the small-angle approximation underpins the theory being tested, and larger amplitudes were not investigated.
Both belong in an evaluation, and they lead to different things: a weakness leads to an improvement, a limitation leads to an extension.
9. An investigation into how the resistance of a wire depends on its length is complete. Suggest two genuine extensions and say where each came from.
From a control variable: having varied length with the diameter held constant, now vary the cross-sectional area with the length held constant, to test whether \( R \propto 1/A \). Together with the first result this would give the full \( R = \rho l / A \) relationship, and a second, independent value for the resistivity.
From an assumption: the analysis assumed the wire's temperature stayed constant, which is why the current was kept low. Test it directly — investigate how the resistance depends on temperature, using a water bath over a controlled range.
Both ask a new question rather than repeating the old one more carefully, which is what makes them extensions rather than improvements.
10. Write a two-sentence conclusion for an experiment whose \( T^{2} \) against \(l\) graph gave a gradient of \( 4.15 \pm 0.30 \) s² m−1 and an intercept of \( 0.02 \pm 0.06 \) s².
First work out what the gradient gives: \[ g = \frac{4\pi^{2}}{4.15} = 9.51\ \text{m s}^{-2}, \qquad \frac{0.30}{4.15} = 7.2\%, \qquad \Delta g = 0.69\ \text{m s}^{-2} \] so \( g = 9.5 \pm 0.7 \) m s−2.

“The graph of \( T^{2} \) against \(l\) is a straight line, confirming that \( T \propto \sqrt{l} \); its gradient of \( 4.15 \pm 0.30 \) s² m−1 corresponds to \( g = 9.5 \pm 0.7 \) m s−2, and the accepted value of 9.81 m s−2 lies within that range, so the result is consistent with theory. The intercept of \( 0.02 \pm 0.06 \) s² includes zero, as theory requires, giving no evidence of a systematic error in the timing or length measurements.”

Both sentences do work: the first answers the research question and compares with theory, the second interprets the intercept — which most reports omit entirely.

🔗Go deeper — other people’s work

These are external resources, not mine. If one stops working, tell me and everything above it on this page still stands.

  • The IB Physics guide internal assessment criteria — conclusion and evaluation are a third of the marks
  • NIST reference values — for the accepted figures your result should be compared against