HomeLearning HubA Level BiologyPaper 5: Planning, Analysis and Evaluation
Paper 5

Planning, Analysis and Evaluation

A Level · Practical assessment · 1 hour 15 minutes · 30 marks

🎯What you need to be able to do

  • State a prediction linked to an underlying hypothesis, and identify the independent, dependent and standardised variables.
  • Describe how to vary and measure the variables, prepare concentrations by serial or proportional dilution, and design appropriate controls.
  • Describe a procedure in a logical sequence and explain how the quality and validity of results would be assessed.
  • Prepare a simple risk assessment, taking account of the severity of hazards and the probability of a problem.
  • Process data: means, percentages, rates, percentage change, standard deviation, standard error and 95% confidence intervals.
  • Choose and justify an appropriate statistical test, state a null hypothesis, and interpret the result.
  • Recognise categoric, ordinal and continuous data and select the correct form of display.
  • Draw conclusions and evaluate an investigation: anomalies, replication, range and intervals, control of variables and confidence in the conclusion.

📚How the paper works

Paper 5 is a written paper of 1 hour 15 minutes worth 30 marks, taken at A Level only. It is 11.5% of the A Level and, like Paper 3, is assessed entirely on AO3. It requires no laboratory — but it assumes you have spent many hours in one, because it asks you to design and criticise experiments you have never seen.

There are two or more questions, and the marks divide almost equally:

  • Planning — 14–16 marks: defining the problem, and methods.
  • Analysis, conclusions and evaluation — 14–16 marks: dealing with data, conclusions, and evaluation.

Contexts may come from anywhere in the AS or A Level syllabus and may be entirely unfamiliar; where they involve theory or apparatus you would not know, the information is given in the question. Knowledge of the mathematical requirements in section 6 of the syllabus is assumed.

📝Planning

Defining the problem

Start with a prediction linked to a hypothesis. A prediction says what will happen; the hypothesis says why. “As temperature increases from 10 to 40 °C the rate will increase, because molecules have more kinetic energy so there are more successful collisions” is a prediction with its hypothesis attached. A prediction may also be given as a sketch graph of the expected result, and if the question offers that option it is usually the quicker answer.

Then identify the independent and dependent variables, and the key variables to be standardised. Again, variables with a minimal effect need not be listed.

Method

Write it so that another student could follow it without asking you anything. That means quantities and numbers, not adjectives:

  • how to vary the independent variable — state the actual values, at least five, evenly spaced, and how each is produced. Concentrations are made by serial dilution (each step diluting the previous one by a fixed factor) or proportional dilution (mixing stock and water in stated proportions). Give the volumes.
  • how to measure the dependent variable, with what instrument, and to what precision;
  • how to standardise each key variable — not “keep temperature the same” but “use a thermostatically controlled water bath at 25 °C, and equilibrate the solutions in it for 5 minutes before mixing”;
  • volumes and concentrations of reagents, with concentrations in % (w/v) or mol dm−3;
  • appropriate control experiments;
  • the steps in a logical sequence, including how the apparatus is used to collect the results;
  • the number of replicates.

Quality, validity and risk

The syllabus separates two ideas that students routinely merge:

  • Quality of results — assessed by looking for anomalous results and by examining the spread of repeats, using standard deviation, standard error or 95% confidence intervals.
  • Validity — assessed by considering both the accuracy of the measurements and the repeatability of the results, and whether the investigation actually tests the hypothesis stated.

The risk assessment should take account of both the severity of each hazard and the probability that a problem occurs, and then state the precautions that reduce the risk.

📈Dealing with data

Types of data

  • Qualitative, categoricnominal data: observations sorted into categories with no order, such as flower colour. Display as a bar chart.
  • Qualitative, orderedordinal data: values that can be ranked but where the intervals between them are not necessarily equal — the order in which tubes decolourise, or an abundance scale.
  • Quantitativecontinuous data: any value within a range, such as body mass or leaf length. Display as a line graph, or a histogram for frequency data.

Spread and reliability

\( s = \sqrt{\dfrac{\sum (x - \bar{x})^{2}}{n-1}} \)
\( \mathrm{SE} = \dfrac{s}{\sqrt{n}} \)
95% CI \( = \bar{x} \pm (2 \times \mathrm{SE}) \)

All three formulae are provided. What is examined is what they mean:

  • Standard deviation measures the spread of the individual data about the mean. A large value means variable data.
  • Standard error measures how reliable the mean itself is as an estimate of the true population mean. Note that increasing \( n \) reduces SE even if the spread stays the same — more data gives a better estimate of the mean.
  • 95% confidence intervals plotted as error bars give a visual significance test: if the error bars of two means do not overlap, the difference between them is likely to be significant; if they overlap substantially, it is unlikely to be, and a statistical test is needed to decide.

Choosing the test

This is the single most reliably examined skill on the paper. Four tests, and the choice is made from the type of data and the question being asked:

  • Chi-squared — testing whether observed frequencies of nominal data differ significantly from expected ones. Genetics crosses, ecological distributions. Uses counts. \( \nu = c - 1 \).
  • t-test — comparing the means of two samples of continuous data from normally distributed populations with approximately equal standard deviations. Works with fewer than 30 values. \( \nu = n_1 + n_2 - 2 \).
  • Pearson’s linear correlation — testing for correlation between two sets of continuous, normally distributed data where a scatter diagram suggests a linear relationship. At least five paired observations, ideally ten or more.
  • Spearman’s rank correlation — testing for correlation where the data are not normally distributed or are ordinal or can be ranked, and the relationship is increasing or decreasing but not necessarily linear. More than five pairs, ideally 10–30.

The two degrees-of-freedom formulae are not provided. Both correlation coefficients run from −1 through 0 to +1.

Always state a null hypothesis before testing: that there is no significant difference (or no significant correlation), and that any observed difference is due to chance. Then compare the calculated value with the critical value at p = 0.05: calculated value greater than or equal to the critical value means the result is significant and the null hypothesis is rejected.

🔍Conclusions and evaluation

A conclusion should summarise the main findings, quote specific figures from the data, state whether the hypothesis is supported and how far, give a scientific explanation, and where appropriate make further predictions. Discuss the strengths and weaknesses of the evidence, not just the result.

The evaluation checklist — work down it and something will always apply:

  • are there anomalous values, and what might explain them?
  • were there enough replicates for a reliable mean?
  • was the range of the independent variable wide enough, and were the intervals small enough to locate the feature of interest?
  • was the method of measuring the dependent variable appropriate and precise enough?
  • were the other variables effectively controlled?
  • how much confidence can be placed in the conclusion, and to what extent can these data test the hypothesis at all?
“Not significant” does not mean “no difference”. It means the data provide insufficient evidence that the difference is anything other than chance — which may be because there is genuinely no effect, or because the sample was too small or too variable to detect one. Saying “the test proves the two are the same” misstates what a significance test can do, and it is a mark lost every session.

✏️Worked example

Two populations of a plant were sampled and leaf length measured. Site A: mean 24.0 mm, standard deviation 4.0 mm, n = 16. Site B: mean 28.5 mm, standard deviation 6.0 mm, n = 25. (a) Calculate the standard error and the 95% confidence interval for each mean. (b) State what the intervals suggest, and which statistical test should be used to decide. (c) State the null hypothesis and the degrees of freedom, and explain how you would reach a conclusion.

(a) Using \( \mathrm{SE} = s / \sqrt{n} \):

Site A: \( \mathrm{SE} = \dfrac{4.0}{\sqrt{16}} = \dfrac{4.0}{4} = 1.0 \) mm
Site B: \( \mathrm{SE} = \dfrac{6.0}{\sqrt{25}} = \dfrac{6.0}{5} = 1.2 \) mm

Then 95% CI \( = \bar{x} \pm (2 \times \mathrm{SE}) \):

Site A: \( 24.0 \pm 2.0 \), so 22.0 to 26.0 mm
Site B: \( 28.5 \pm 2.4 \), so 26.1 to 30.9 mm

(b) The two intervals do not overlap — A ends at 26.0 and B begins at 26.1. This suggests the difference between the means is likely to be significant, and is unlikely to be due to chance alone. Because the margin is narrow, this is a suggestion rather than a demonstration, and a formal test is needed.

The appropriate test is the t-test, because we are comparing the means of two samples of continuous data. It is valid here provided the populations are approximately normally distributed and the two standard deviations are approximately equal — 4.0 and 6.0 are of the same order, so this condition is reasonably met.

(c) Null hypothesis: there is no significant difference between the mean leaf lengths of the plants at site A and site B; any difference observed is due to chance.

Degrees of freedom: \( \nu = n_1 + n_2 - 2 = 16 + 25 - 2 = \mathbf{39} \). This formula is not provided in the exam.

Calculate t from the formula given, then compare it with the critical value at p = 0.05 for 39 degrees of freedom. If the calculated t is greater than or equal to the critical value, the probability that a difference this large arose by chance is less than 5%, so the difference is significant and the null hypothesis is rejected. If t is less than the critical value, the difference is not significant and the null hypothesis is accepted.

Check it. Check the standard errors against intuition: site B has the larger spread (s = 6.0 against 4.0) but the larger sample (25 against 16), and the two effects nearly cancel, giving similar standard errors of 1.2 and 1.0. That is the standard error doing its job — it measures the reliability of the mean, which improves with sample size. If your SE came out larger than the standard deviation, you have multiplied by \( \sqrt{n} \) instead of dividing.
Confusing standard deviation with standard error, and reading overlapping bars as proof. Standard deviation describes the data; standard error describes the mean, and error bars drawn from one are not interchangeable with the other — always state which you have plotted. And the overlap rule is a guide, not a test: non-overlapping intervals suggest significance and overlapping ones suggest its absence, but only the statistical test decides, which is exactly why part (b) asks for one. Note also the direction of the comparison: unlike some tests, for t and for chi-squared a calculated value above the critical value means reject the null hypothesis.

📝Practise

Work through these, then reveal the answer. Each question targets a different objective from the list above.

1. Write a prediction, with its underlying hypothesis, for an investigation into the effect of sucrose concentration on the rate of respiration in yeast.
Prediction: as the sucrose concentration is increased from 0 to a moderate value, the rate of respiration of the yeast — measured as the volume of carbon dioxide produced per minute — will increase; above a certain concentration the rate will level off. Hypothesis: sucrose is the respiratory substrate, so at low concentrations the rate is limited by the availability of substrate — increasing it means more substrate molecules bind to the active sites of the respiratory enzymes per unit time, so more enzyme–substrate complexes form. Above a certain concentration all the active sites are occupied and the enzymes are saturated, so the rate is limited by the number of enzyme molecules and adding more sucrose has no further effect. (A further, testable extension: at very high concentrations the rate may fall, because the very low water potential of the solution causes water to leave the yeast cells by osmosis.) Note that a prediction on its own does not earn the mark — it must be linked to the biological reason.
2. Describe how to prepare 10 cm³ each of 0.10, 0.08, 0.06, 0.04 and 0.02 mol dm−3 solutions from a 0.10 mol dm−3 stock.
This is a proportional dilution: mix stock solution and distilled water in the proportion required, keeping the total volume constant at 10 cm³ so that only the concentration varies. 0.10: 10.0 cm³ stock + 0 water. 0.08: 8.0 cm³ stock + 2.0 cm³ water. 0.06: 6.0 + 4.0. 0.04: 4.0 + 6.0. 0.02: 2.0 + 8.0. Add a control of 10.0 cm³ distilled water (0.00 mol dm−3). Measure both liquids with a syringe or graduated pipette, which is more precise than a measuring cylinder, using a clean one for each solution to avoid contamination, and mix thoroughly before use. Label each tube. The alternative, serial dilution, dilutes each solution from the previous one by a constant factor — useful for spanning several orders of magnitude, but it compounds any error at each step and here would not give the evenly spaced values required.
3. An investigation compares the number of stomata per field of view on the upper and lower surfaces of leaves from 20 plants. State the appropriate statistical test, the null hypothesis, and the degrees of freedom.
The data are counts of stomata, which are continuous enough to treat as quantitative and are being compared as two sample means — upper surface and lower surface. The appropriate test is the t-test, which compares the means of two samples. It is valid provided the data are drawn from approximately normally distributed populations and the two standard deviations are similar; if the counts are highly skewed, a rank-based approach would be preferable. Null hypothesis: there is no significant difference between the mean number of stomata per field of view on the upper and lower surfaces; any difference observed is due to chance. Degrees of freedom: \( \nu = n_1 + n_2 - 2 = 20 + 20 - 2 = 38 \). Compare the calculated t with the critical value at p = 0.05 for 38 degrees of freedom: if t is greater, reject the null hypothesis and conclude the difference is significant.
4. Explain the difference between standard deviation and standard error, and state which you would plot as error bars when comparing two means.
Standard deviation (s) measures the spread of the individual data values about their mean: a large standard deviation means the measurements are widely scattered, whatever the sample size. Standard error (SE = s/√n) measures the reliability of the mean itself as an estimate of the true population mean. The two are related but answer different questions, and crucially SE decreases as sample size increases even when the spread of the data does not, because a larger sample gives a better estimate of the mean. When comparing two means, plot error bars based on standard error, or better on 95% confidence intervals (mean ± 2 × SE), because the question being asked is about the means, not about the spread of individuals. If the 95% confidence intervals of the two means do not overlap, the difference is likely to be significant. Always state on the graph which quantity the bars represent — bars of unstated origin cannot be interpreted.
5. A student concludes: “The t-test showed no significant difference, so the two fertilisers have the same effect on growth.” Criticise this conclusion.
The conclusion overstates what the test shows. A non-significant result means only that the data provide insufficient evidence that the difference between the means is anything other than chance — it does not prove that the two are the same. There are at least three reasons a real difference could go undetected: the sample size may be too small, so the standard error is large and the test lacks the power to detect a genuine effect; the data may be too variable, because other factors such as light, water or soil were not adequately standardised, and that variation masks the treatment effect; and the range or duration may have been inadequate — a difference in growth may take longer to appear than the investigation allowed. A properly worded conclusion is: there is no significant difference between the mean growth of plants given the two fertilisers at the p = 0.05 level, so the null hypothesis is accepted; these data provide no evidence that the fertilisers differ in their effect. The student should then suggest increasing the sample size, controlling variables more tightly and repeating over a longer period.
6. An investigation into the effect of temperature on an enzyme uses values of 20, 30, 40, 50 and 60 °C, with one reading at each. Evaluate this design and suggest three improvements.
Strengths: five values are used, meeting the minimum, and they are evenly spaced across a range broad enough to include both the rising part of the curve and the region of denaturation. Weaknesses and improvements: (i) there is only one reading at each temperature, so anomalies cannot be identified, no mean can be calculated, and the spread of the data is unknown — take at least three replicates at each temperature and calculate a mean, with standard deviation or standard error to show the spread. (ii) The interval of 10 °C is too large to locate the optimum, which lies somewhere between 40 and 50 °C but cannot be pinned down — take additional readings at smaller intervals, such as 42, 44, 46 and 48 °C, once the approximate position is known. (iii) The design says nothing about how temperature is held constant during each run — use a thermostatically controlled water bath and equilibrate the enzyme and substrate separately in it before mixing, since otherwise the mixture is not at the stated temperature for the early part of the reaction. A fourth: include a control with boiled (denatured) enzyme to confirm that any change observed is caused by enzyme activity.

🔗Go deeper — other people’s work

These are external resources, not mine. If one stops working, tell me and everything above it on this page still stands.

  • Cambridge International specimen Paper 5 and its mark scheme — planning marks are awarded for very specific things, and the mark scheme is the only reliable guide to what they are
  • Any set of statistical tables for chi-squared, t and the correlation coefficients — practise locating critical values quickly, since the tables are provided but the search costs time
  • The Field Studies Council and Nuffield Foundation statistics guides — written at exactly this level, with worked examples of each of the four tests