HomeLearning HubIB DP PhysicsI.2 Collecting and processing data
I.2

Collecting and processing data

Tools and inquiry · skills assessed across the whole course

Between taking a reading and drawing a conclusion sits the part nobody photographs: writing the numbers down properly, turning them into something plottable, and working out what shape they make. It is where most marks are won and lost.

🎯What you need to be able to do

  • Record raw data in a table with correct headings, units, uncertainties and decimal places.
  • Record qualitative observations alongside the quantitative ones.
  • Process raw data into the quantity you actually want to plot, showing one sample calculation.
  • Identify an outlier, and describe what to do about it — including what to write down.
  • Distinguish accuracy, precision, reliability and validity.
  • Recognise the common graph shapes and say what relationship each one indicates.
  • Choose what to plot so that a curve becomes a straight line.

📋Recording the data

A results table is a piece of communication, and it is marked as one. Raw data — the numbers you actually read off instruments — goes in first, exactly as read, before anything is done to it.

A results table for a wire-resistance experiment with five columns: length in metres to plus or minus 0.001, potential difference in volts to plus or minus 0.01, current in amperes to plus or minus 0.01, resistance in ohms marked as calculated, and mean resistance from three repeats. Five rows of data run from 0.100 metres to 0.500 metres. Every value in a column carries the same number of decimal places, because the instrument does not change between rows. Below, the five things a marker looks for in the heading row: a quantity and its symbol, a unit after a solidus, the uncertainty given once in the heading, consistent decimal places down each column, and raw data kept apart from processed data. A closing note urges recording qualitative observations too, since they usually explain an anomalous point later.
The unit belongs in the heading, once — the column is headed \( R\,/\,\Omega \) and the cells hold bare numbers. So does the uncertainty, unless it genuinely differs from row to row.

Qualitative observations count as data. The wire glowing faintly, the pendulum bob starting to swing in an ellipse rather than a plane, a meter reading that would not settle — these are the things that explain an anomalous point three hours later, and you will not remember them if you do not write them down at the time.

The trap: decimal places must be consistent down a column. A balance that reads to 0.01 g gives 5.30 g, not 5.3 g — the trailing zero is a real piece of information, saying the instrument resolved that digit and found it to be zero. Writing 5.3 in one row and 5.30 in the next claims the instrument changed halfway through. Raw data is transcribed exactly as the instrument presented it.

🔢Processing it

Processing is anything you do to the raw numbers: averaging repeats, subtracting a background, converting units, calculating a derived quantity, squaring a value so the graph comes out straight. Processed data goes in its own columns, clearly separated from the raw.

Show one sample calculation in full, for one row, so a reader can check what you did — then present the rest as a column of results. Repeating the arithmetic ten times proves nothing and wastes space.

✏️Worked example 1 — processing one row properly

Three timings for 20 oscillations of a pendulum of length 0.600 m are 31.2 s, 31.5 s and 31.0 s. Process this row for a graph of \( T^{2} \) against \(l\), and state the uncertainty in \( T^{2} \).

Mean of the repeats.

\[ \bar{t} = \frac{31.2 + 31.5 + 31.0}{3} = 31.23\ \text{s} \]

Uncertainty from the spread. Largest − mean = 0.27; mean − smallest = 0.23. Take the larger and round to one significant figure: ± 0.3 s.

Divide by the number of oscillations. Value and absolute uncertainty both divide:

\[ T = \frac{31.23}{20} = 1.562\ \text{s}, \qquad \Delta T = \frac{0.3}{20} = 0.015\ \text{s} \]

Square it. A power multiplies the percentage uncertainty:

\[ T^{2} = 1.562^{2} = 2.440\ \text{s}^{2} \]
\[ \frac{\Delta T}{T} = \frac{0.015}{1.562} = 0.96\%, \qquad \frac{\Delta(T^{2})}{T^{2}} = 2 \times 0.96\% = 1.9\% \]
\[ \Delta(T^{2}) = 0.019 \times 2.440 = 0.046 \approx 0.05\ \text{s}^{2} \]

So this row contributes the point \( (0.600, 2.44 \pm 0.05) \), and that ± 0.05 s² is the error bar.

Sanity check. Squaring doubled the percentage uncertainty, as it must. Note the chain: the raw timings carried ± 0.3 s, dividing by 20 shrank that to ± 0.015 s without changing the percentage, and squaring then doubled the percentage. Every step is one of the rules from T.3, applied in order.

⚠️Outliers

A graph of seven points, six of which lie close to a straight line while one sits well above it, circled and labelled. Alongside, what to do about it in order: first go back to your notes and ask whether anything happened during that reading; second repeat that measurement, which is the cheapest and most convincing test; third, if it repeats then it is real and the physics may be more interesting than expected; fourth, if it does not repeat then say so and record that it was repeated and rejected. Two warnings follow: that deleting an inconvenient point without saying so is falsifying data rather than a rounding error, and that sometimes the outlier is the discovery, as when the elastic limit of a spring announces itself through points that stop obeying a law that held perfectly until then.
The instinct is to drop the awkward point. The correct move is to repeat it — which either removes it honestly or turns it into the most interesting thing in the data set.

🎯Four words that are judged separately

Four terms defined in rows. Accurate means close to the true value, is spoiled by systematic error, and is checked by comparing with an accepted value. Precise means readings close to each other, is spoiled by random error, and is checked by looking at the spread of repeats. Reliable means getting the same result if the experiment is done again, is spoiled by an uncontrolled variable, and is checked by repeating on another day or by someone else. Valid means the experiment actually measures what it claims to, is spoiled by a design flaw, and is checked by asking whether the method answers the question. A closing warning notes that an experiment can be precise, reliable and completely invalid, giving the example of a wire warmed by its own current, which yields beautifully repeatable resistances that answer a question about temperature rather than the one about length that was asked.
The last row is the one that catches people. Repeatable, tightly grouped results feel authoritative — and tell you nothing if the experiment was measuring the wrong thing.

📈Looking for the trend

Plot the data and the shape tells you what kind of relationship you have. Eight shapes cover almost everything at this level, and being able to name them on sight is worth real marks.

Eight small graphs labelled A to H. A is directly proportional, a straight line through the origin, y equals k x. B is linear with a positive intercept, y equals m x plus c. C is linear with a negative gradient, y equals c minus m x. D is inversely proportional, a falling hyperbola, y equals k over x. E is a square relationship, y proportional to x squared, curving upward. F is a square root, y proportional to root x, curving and flattening. G is exponential decay, y equals y nought e to the minus k t. H is exponential rise to a limit, y equals y nought times one minus e to the minus k t. A closing note explains that only A is proportional, that the difference from B is the whole point, and that an unexpected intercept where theory predicted the origin is a systematic error you have just measured.
Shapes D through H all look like “a curve” at a glance and mean quite different things. Telling them apart is exactly what the linearising techniques of T.3 are for.
The trap: “proportional” means through the origin. A straight line showing \(y\) rising steadily with \(x\) is linear. It is only proportional if it passes through the origin, because only then does doubling \(x\) double \(y\). If theory predicted proportionality and your line has an intercept, do not describe the result as “roughly proportional” — identify what that intercept is. It is usually a systematic error you have just successfully measured.

✏️Worked example 2 — deciding what to plot

A student measures the intensity \(I\) of light from a lamp at distances \(d\) from 0.20 to 1.00 m and gets a falling curve. They believe \( I \propto 1/d^{2} \). What should they plot, and what would confirm the belief?

Rearrange into \( y = mx + c \) form. If \( I = k/d^{2} \), then

\[ I = k\left(\frac{1}{d^{2}}\right) \]

which is \( y = mx \) with \( y = I \), \( x = 1/d^{2} \) and \( m = k \).

So build a new column. For each distance, calculate \( 1/d^{2} \):

\( d = 0.20 \) m\( 1/d^{2} = 25.0\ \text{m}^{-2} \)
\( d = 0.50 \) m\( 1/d^{2} = 4.0\ \text{m}^{-2} \)
\( d = 1.00 \) m\( 1/d^{2} = 1.0\ \text{m}^{-2} \)

What confirms it. A plot of \(I\) against \( 1/d^{2} \) that is a straight line through the origin. Both parts matter: straightness confirms the inverse-square form, and passing through the origin confirms there is no constant background light being picked up.

Sanity check. Notice how uneven the new x-values are — 25.0, 4.0, 1.0. The readings taken close to the lamp are spread far apart on the new axis, and the distant ones are bunched near the origin. That is worth knowing before collecting data: to get points evenly spread along the \( 1/d^{2} \) axis, choose distances that are evenly spaced in \( 1/d^{2} \), not in \(d\). Deciding what to plot is a design decision, not just an analysis one.

📝Practise

Work through these, then reveal the answer. Each question targets a different objective from the list above.

1. A student writes a column heading as “Length (cm)” and records 12.4, 15, 18.60, 21.2. Identify three faults.
Inconsistent decimal places. 12.4, 15, 18.60 and 21.2 imply the instrument's resolution changed between rows, which it did not. All should be to the same precision — 12.4, 15.0, 18.6, 21.2 if the rule is 0.1 cm.
No uncertainty. The heading should carry it once, for example “length, \( l\,/\,\text{cm}\ (\pm 0.1) \)”.
The unit convention. The IB convention is a solidus, so \( l\,/\,\text{cm} \) rather than “Length (cm)”. A symbol for the quantity should also be given, since the graph axes and any calculations will use it.
2. Explain the difference between raw and processed data, and state what a report must show for each.
Raw data is what you read directly off an instrument, transcribed exactly — three separate timings, a potential difference, a scale reading. It must be presented in full, unaltered, with its units and uncertainties.
Processed data is anything calculated from it: a mean of repeats, a background-corrected count, a unit conversion, a derived quantity like resistance, or a squared value plotted to straighten a graph.
The report must show all the raw data, at least one sample calculation written out in full for each kind of processing, and then the processed results as a column. Showing every calculation is unnecessary; showing none makes the processing unverifiable.
3. Three timings for 20 oscillations give 24.6, 24.9 and 24.4 s. Find the period with its uncertainty.
Mean: \( (24.6 + 24.9 + 24.4)/3 = 24.63 \) s.
Spread: largest − mean = 0.27; mean − smallest = 0.23. Take the larger, round to one significant figure: ± 0.3 s.
Divide by 20 — both the value and the absolute uncertainty divide: \[ T = \frac{24.63}{20} = 1.232\ \text{s}, \qquad \Delta T = \frac{0.3}{20} = 0.015\ \text{s} \] So \( T = 1.23 \pm 0.02 \) s, rounding the uncertainty to one significant figure and matching the value to the same decimal place.
4. A count rate is measured as 418 counts per minute. The background is 32 counts per minute. State the corrected rate and explain why the correction must come first.
Corrected rate = \( 418 - 32 = 386 \) counts per minute.
It must come first because the background is present in every reading, so it is a systematic addition to all of them. Halving a raw count rate to find a half-life, or plotting \( \ln R \) against \(t\), gives the wrong answer if the background is still in there — the numbers being halved or logged are not the source's count rate at all.
The effect is worst at low count rates: when the source has decayed to 40 counts per minute, an unsubtracted background of 32 makes the reading nearly twice what it should be, and the tail of a decay curve is exactly where that distortion shows.
5. One point on an otherwise clean straight-line graph lies well off the trend. Describe what you should do, in order, and what the report must say.
1. Check your notes for anything recorded at the time — a disturbance, a misread scale, a flickering meter.
2. Repeat that measurement, if the apparatus is still set up. This is the cheapest and most convincing test available.
3. If it repeats, it is a real feature of the system and not an error — and the physics may be more interesting than expected.
4. If it does not repeat, the original was a one-off mistake and may be excluded.

The report must state which point was excluded, why, and what was done to check. A documented exclusion is fine and often earns credit; a point that silently disappears is falsification of data, which is an academic honesty matter rather than a technical one.
6. Distinguish reliability from validity, and give an example of an experiment that is highly reliable and completely invalid.
Reliability is about repeatability: would the same method, repeated on another day or by another person, give the same result? It is undermined by uncontrolled variables.
Validity is about whether the method measures what it claims to. It is undermined by a flaw in the design, and no amount of repetition can reveal it.

Example: measuring the resistance of a wire while a large current flows through it. The wire warms up, its resistance rises, and the readings settle to a repeatable value — the same tomorrow, the same for anyone else. Highly reliable. But the experiment claims to measure how resistance depends on length while it is really measuring how resistance depends on temperature, so it is invalid. Reliable and wrong.
7. A graph of \(y\) against \(x\) is a straight line that does not pass through the origin. State the relationship, and explain why it is wrong to call it proportional.
It is a linear relationship, \( y = mx + c \), with \(c\) the intercept.
It is not proportional, because proportionality means \( y = kx \) with no constant term — and only then does doubling \(x\) double \(y\). Here, doubling \(x\) does not double \(y\), because the constant \(c\) is carried along unchanged.
If theory predicted proportionality, that intercept is evidence of something: most often a systematic error such as a zero error or a background reading. Identifying what it is turns a disappointing graph into a successful measurement of the error itself.
8. For each, state what happens to \(y\) when \(x\) doubles: (a) \( y = kx \); (b) \( y = kx^{2} \); (c) \( y = k/x \); (d) \( y = k\sqrt{x} \).
(a) \(y\) doubles — direct proportion.
(b) \(y\) becomes four times as large, since \( 2^{2} = 4 \).
(c) \(y\) halves — inverse proportion.
(d) \(y\) increases by a factor of \( \sqrt{2} = 1.41 \).
This is a quick way to identify a relationship from a table without plotting anything: find two rows where the independent variable doubles and see what the dependent variable did. It is also a good check on a graph you have already drawn.
9. A student expects \( I \propto 1/d^{2} \). State what they should plot, and what two features of the resulting graph would confirm the relationship.
Plot \(I\) on the y-axis against \( 1/d^{2} \) on the x-axis. Rearranging \( I = k/d^{2} \) into \( I = k(1/d^{2}) \) gives \( y = mx \), so this should be a straight line.
Two features confirm it:
Straightness — confirms the inverse-square form specifically, as opposed to some other falling curve.
Passing through the origin — confirms there is no constant offset, such as background light reaching the sensor. An intercept would mean the sensor reads something even when \( 1/d^{2} \to 0 \), that is, at infinite distance.
10. Explain why plotting \(I\) against \( 1/d^{2} \) should influence which distances you choose to measure at.
Evenly spaced distances do not give evenly spaced values on a \( 1/d^{2} \) axis. Distances of 0.20, 0.50 and 1.00 m give \( 1/d^{2} \) values of 25.0, 4.0 and 1.0 m−2 — the near readings are spread right across the axis while the far ones bunch up near the origin.
That leaves most of the graph determined by a handful of close-range points, which is exactly the narrow-range problem from I.1 in disguise.
The fix is to choose distances that are evenly spaced in \( 1/d^{2} \) rather than in \(d\) — for \( 1/d^{2} \) values of 4, 8, 12, 16, 20, the distances are 0.50, 0.35, 0.29, 0.25 and 0.22 m. Deciding what to plot is therefore part of designing the experiment, not something to leave until the data is already collected.

🔗Go deeper — other people’s work

These are external resources, not mine. If one stops working, tell me and everything above it on this page still stands.

  • The IB Physics guide internal assessment criteria — how data collection and processing are marked
  • Vernier and PASCO teacher notes — worked examples of well-laid-out results tables