Planning, Analysis and Evaluation
🎯What you need to be able to do
- State a prediction linked to an underlying hypothesis, and identify the independent, dependent and standardised variables.
- Describe how to vary and measure the variables, prepare concentrations by serial or proportional dilution, and design appropriate controls.
- Describe a procedure in a logical sequence and explain how the quality and validity of results would be assessed.
- Prepare a simple risk assessment, taking account of the severity of hazards and the probability of a problem.
- Process data: means, percentages, rates, percentage change, standard deviation, standard error and 95% confidence intervals.
- Choose and justify an appropriate statistical test, state a null hypothesis, and interpret the result.
- Recognise categoric, ordinal and continuous data and select the correct form of display.
- Draw conclusions and evaluate an investigation: anomalies, replication, range and intervals, control of variables and confidence in the conclusion.
📚How the paper works
Paper 5 is a written paper of 1 hour 15 minutes worth 30 marks, taken at A Level only. It is 11.5% of the A Level and, like Paper 3, is assessed entirely on AO3. It requires no laboratory — but it assumes you have spent many hours in one, because it asks you to design and criticise experiments you have never seen.
There are two or more questions, and the marks divide almost equally:
- Planning — 14–16 marks: defining the problem, and methods.
- Analysis, conclusions and evaluation — 14–16 marks: dealing with data, conclusions, and evaluation.
Contexts may come from anywhere in the AS or A Level syllabus and may be entirely unfamiliar; where they involve theory or apparatus you would not know, the information is given in the question. Knowledge of the mathematical requirements in section 6 of the syllabus is assumed.
📝Planning
Defining the problem
Start with a prediction linked to a hypothesis. A prediction says what will happen; the hypothesis says why. “As temperature increases from 10 to 40 °C the rate will increase, because molecules have more kinetic energy so there are more successful collisions” is a prediction with its hypothesis attached. A prediction may also be given as a sketch graph of the expected result, and if the question offers that option it is usually the quicker answer.
Then identify the independent and dependent variables, and the key variables to be standardised. Again, variables with a minimal effect need not be listed.
Method
Write it so that another student could follow it without asking you anything. That means quantities and numbers, not adjectives:
- how to vary the independent variable — state the actual values, at least five, evenly spaced, and how each is produced. Concentrations are made by serial dilution (each step diluting the previous one by a fixed factor) or proportional dilution (mixing stock and water in stated proportions). Give the volumes.
- how to measure the dependent variable, with what instrument, and to what precision;
- how to standardise each key variable — not “keep temperature the same” but “use a thermostatically controlled water bath at 25 °C, and equilibrate the solutions in it for 5 minutes before mixing”;
- volumes and concentrations of reagents, with concentrations in % (w/v) or mol dm−3;
- appropriate control experiments;
- the steps in a logical sequence, including how the apparatus is used to collect the results;
- the number of replicates.
Quality, validity and risk
The syllabus separates two ideas that students routinely merge:
- Quality of results — assessed by looking for anomalous results and by examining the spread of repeats, using standard deviation, standard error or 95% confidence intervals.
- Validity — assessed by considering both the accuracy of the measurements and the repeatability of the results, and whether the investigation actually tests the hypothesis stated.
The risk assessment should take account of both the severity of each hazard and the probability that a problem occurs, and then state the precautions that reduce the risk.
📈Dealing with data
Types of data
- Qualitative, categoric → nominal data: observations sorted into categories with no order, such as flower colour. Display as a bar chart.
- Qualitative, ordered → ordinal data: values that can be ranked but where the intervals between them are not necessarily equal — the order in which tubes decolourise, or an abundance scale.
- Quantitative → continuous data: any value within a range, such as body mass or leaf length. Display as a line graph, or a histogram for frequency data.
Spread and reliability
All three formulae are provided. What is examined is what they mean:
- Standard deviation measures the spread of the individual data about the mean. A large value means variable data.
- Standard error measures how reliable the mean itself is as an estimate of the true population mean. Note that increasing \( n \) reduces SE even if the spread stays the same — more data gives a better estimate of the mean.
- 95% confidence intervals plotted as error bars give a visual significance test: if the error bars of two means do not overlap, the difference between them is likely to be significant; if they overlap substantially, it is unlikely to be, and a statistical test is needed to decide.
Choosing the test
This is the single most reliably examined skill on the paper. Four tests, and the choice is made from the type of data and the question being asked:
- Chi-squared — testing whether observed frequencies of nominal data differ significantly from expected ones. Genetics crosses, ecological distributions. Uses counts. \( \nu = c - 1 \).
- t-test — comparing the means of two samples of continuous data from normally distributed populations with approximately equal standard deviations. Works with fewer than 30 values. \( \nu = n_1 + n_2 - 2 \).
- Pearson’s linear correlation — testing for correlation between two sets of continuous, normally distributed data where a scatter diagram suggests a linear relationship. At least five paired observations, ideally ten or more.
- Spearman’s rank correlation — testing for correlation where the data are not normally distributed or are ordinal or can be ranked, and the relationship is increasing or decreasing but not necessarily linear. More than five pairs, ideally 10–30.
The two degrees-of-freedom formulae are not provided. Both correlation coefficients run from −1 through 0 to +1.
Always state a null hypothesis before testing: that there is no significant difference (or no significant correlation), and that any observed difference is due to chance. Then compare the calculated value with the critical value at p = 0.05: calculated value greater than or equal to the critical value means the result is significant and the null hypothesis is rejected.
🔍Conclusions and evaluation
A conclusion should summarise the main findings, quote specific figures from the data, state whether the hypothesis is supported and how far, give a scientific explanation, and where appropriate make further predictions. Discuss the strengths and weaknesses of the evidence, not just the result.
The evaluation checklist — work down it and something will always apply:
- are there anomalous values, and what might explain them?
- were there enough replicates for a reliable mean?
- was the range of the independent variable wide enough, and were the intervals small enough to locate the feature of interest?
- was the method of measuring the dependent variable appropriate and precise enough?
- were the other variables effectively controlled?
- how much confidence can be placed in the conclusion, and to what extent can these data test the hypothesis at all?
✏️Worked example
(a) Using \( \mathrm{SE} = s / \sqrt{n} \):
Then 95% CI \( = \bar{x} \pm (2 \times \mathrm{SE}) \):
(b) The two intervals do not overlap — A ends at 26.0 and B begins at 26.1. This suggests the difference between the means is likely to be significant, and is unlikely to be due to chance alone. Because the margin is narrow, this is a suggestion rather than a demonstration, and a formal test is needed.
The appropriate test is the t-test, because we are comparing the means of two samples of continuous data. It is valid here provided the populations are approximately normally distributed and the two standard deviations are approximately equal — 4.0 and 6.0 are of the same order, so this condition is reasonably met.
(c) Null hypothesis: there is no significant difference between the mean leaf lengths of the plants at site A and site B; any difference observed is due to chance.
Degrees of freedom: \( \nu = n_1 + n_2 - 2 = 16 + 25 - 2 = \mathbf{39} \). This formula is not provided in the exam.
Calculate t from the formula given, then compare it with the critical value at p = 0.05 for 39 degrees of freedom. If the calculated t is greater than or equal to the critical value, the probability that a difference this large arose by chance is less than 5%, so the difference is significant and the null hypothesis is rejected. If t is less than the critical value, the difference is not significant and the null hypothesis is accepted.
📝Practise
Work through these, then reveal the answer. Each question targets a different objective from the list above.
1. Write a prediction, with its underlying hypothesis, for an investigation into the effect of sucrose concentration on the rate of respiration in yeast.
2. Describe how to prepare 10 cm³ each of 0.10, 0.08, 0.06, 0.04 and 0.02 mol dm−3 solutions from a 0.10 mol dm−3 stock.
3. An investigation compares the number of stomata per field of view on the upper and lower surfaces of leaves from 20 plants. State the appropriate statistical test, the null hypothesis, and the degrees of freedom.
4. Explain the difference between standard deviation and standard error, and state which you would plot as error bars when comparing two means.
5. A student concludes: “The t-test showed no significant difference, so the two fertilisers have the same effect on growth.” Criticise this conclusion.
6. An investigation into the effect of temperature on an enzyme uses values of 20, 30, 40, 50 and 60 °C, with one reading at each. Evaluate this design and suggest three improvements.
🔗Go deeper — other people’s work
These are external resources, not mine. If one stops working, tell me and everything above it on this page still stands.
- Cambridge International specimen Paper 5 and its mark scheme — planning marks are awarded for very specific things, and the mark scheme is the only reliable guide to what they are
- Any set of statistical tables for chi-squared, t and the correlation coefficients — practise locating critical values quickly, since the tables are provided but the search costs time
- The Field Studies Council and Nuffield Foundation statistics guides — written at exactly this level, with worked examples of each of the four tests