Classification and cladistics
🎯What you need to be able to do
- Explain why organisms need to be classified.
- Outline the traditional hierarchy of taxa and explain why it does not always match evolutionary history.
- Explain the advantages of a classification based on evolutionary relationships.
- Define a clade, and explain why sequence data is the most objective evidence for clades.
- Explain the molecular clock and why it gives only estimates.
- Construct a cladogram from base or amino acid sequences using parsimony.
- Analyse cladograms using the terms root, node and terminal branch.
- Explain how cladistics has shown some traditional classifications to be false, using the figwort family.
- Outline the classification of all organisms into three domains based on rRNA sequences.
📚The biology
Why classify?
Around two million species have been named and many millions more are thought to exist. With diversity on that scale, a system for organizing organisms is essential. Once an organism is classified, a great deal of further study becomes easier: a new species can be compared with its relatives, information about one species can be found and shared under a single name, and features of the group can be predicted.
The traditional hierarchy
Linnaean classification places each species in a nested series of ranks, or taxa:
For example, the lion is in kingdom Animalia, phylum Chordata, class Mammalia, order Carnivora, family Felidae, genus Panthera, species leo.
The problem is that evolution does not produce ranks. It produces a continuously branching tree, with divergences at every possible depth. Deciding that one branch is an “order” and another a “family” is arbitrary: there is no biological test for what makes a group a class rather than an order. And traditional groups were often based on similar appearance, which does not always reflect common ancestry. Cladistics replaces fixed ranks with unranked clades, which only have to reflect the branching pattern. The switch is an example of a paradigm shift — a change in the whole framework within which a science works.
Classification by evolutionary relationships
The ideal classification follows evolutionary history, so that all members of a group are descended from one common ancestor. This has two advantages:
- Predictive power. Members of such a group share characteristics inherited from their common ancestor, so if one member has a feature — a biochemical pathway, a gene, a useful chemical — its relatives probably do too. This is used, for example, to search for medicinal compounds in plants related to known sources.
- It reflects reality. The groups correspond to real events in the history of life, not to human judgement about which features matter.
Clades
A clade is a group of organisms consisting of a common ancestor and all of its descendants. Members of a clade share characteristics inherited from that ancestor.
The most objective evidence for placing organisms in a clade comes from base sequences of genes or amino acid sequences of proteins. Sequences can be compared position by position and differences counted, which avoids subjective judgement about how similar two structures look. Morphological traits can also be used, but they are more open to misinterpretation, because similar structures can evolve independently in unrelated groups (convergent evolution, A4.1).
The molecular clock
Once two lineages split, mutations accumulate independently in each, so their sequences gradually diverge. If differences accumulate at a roughly constant rate, the number of differences between two species indicates how long ago they shared a common ancestor. Calibrated against fossils of known age, this gives a molecular clock.
The clock only gives estimates, because the rate at which differences accumulate is not truly constant. It is affected by:
- generation time — organisms with short generations copy their DNA more often per year, so accumulate mutations faster;
- population size, which affects how quickly new mutations are lost or spread;
- the intensity of selection — genes under strong selection change more slowly, because most changes are harmful and removed;
- other factors, such as differences in mutation rate between species and between genes.
Building a cladogram from sequences
A cladogram is a branching diagram showing a hypothesis of how organisms are related. To build one from sequence data:
- Align the sequences so that equivalent positions line up.
- Count differences between each pair of organisms.
- Group the two most similar organisms together as sister groups; then add the next most similar, and so on.
- Check possible alternative trees and choose the one requiring the fewest changes.
Step 4 is parsimony analysis: the most probable cladogram is the one that accounts for all the observed sequence differences with the smallest number of changes. It is a criterion for choosing between hypotheses, and it is not the only possible criterion — other methods, which weight different kinds of change differently, can favour different trees. Different criteria for judgement can lead to different hypotheses, and that is normal in science.
Reading a cladogram
The base of the cladogram, representing the common ancestor of every organism on it.
A branching point. Each node represents a hypothetical common ancestor of the groups that branch from it.
A branch ending in a present-day (or extinct) group, shown at the tips.
Everything branching from one node: that ancestor and all its descendants. Cutting any single branch removes one clade.
Two rules make reading cladograms reliable. Relatedness depends on the most recent common ancestor: two groups are more closely related if the node joining them is more recent (further from the root). And branches can rotate around a node without changing the meaning, so the left-to-right order of the tips tells you nothing. Species next to each other at the tips are not necessarily closest relatives.
Cladograms usually do not show time to scale, although some are drawn with branch lengths proportional to the number of sequence differences or to estimated time.
Cladistics and reclassification
When cladistic analysis is applied to groups that were classified by appearance, it sometimes shows that the traditional group does not correspond to a clade. The figwort family (Scrophulariaceae) is a well-known case. It had long been a large family of flowering plants, grouped mainly because of similarities in their flowers. When the DNA sequences of several genes were compared, the family turned out to contain plants from several separate lineages. Many genera were moved to other families, some to families that had previously been considered quite distinct, and the figwort family shrank to a fraction of its former size.
The similar flowers had arisen by convergent evolution — similar adaptations, for example to similar pollinators — not by common ancestry. This is an example of a scientific knowledge claim being falsified by new evidence and replaced, which is how science is supposed to work.
Three domains
Until the 1970s all living things were divided into prokaryotes and eukaryotes, and life into kingdoms. In 1977, Carl Woese and colleagues compared the base sequences of ribosomal RNA (rRNA), which is found in every organism and changes slowly. They found that prokaryotes fell into two groups as different from each other as either is from eukaryotes. They proposed a new taxonomic level above kingdoms: the domain.
Prokaryotes; the familiar bacteria, including pathogens and cyanobacteria.
Prokaryotes that look like bacteria but differ in rRNA, membrane lipids and much of their biochemistry; many live in extreme environments.
All organisms with eukaryotic cells: protists, fungi, plants and animals.
This reclassification was revolutionary: it showed that “prokaryote” is not a single natural group, and that archaea are in some respects more closely related to eukaryotes than to bacteria.
✏️Worked example
A C G T A C G T A CA C G T A C G A A CA T G T T C G A A CG T G T T C C A A T(b) Construct the most parsimonious cladogram, and use parsimony to justify it.
(c) For the full-length gene, W and Z differ at 2% of positions. If this gene diverges between two lineages by 1% every 5 million years, estimate when W and Z last shared a common ancestor, and give one reason the estimate may be inaccurate.
(a) Compare position by position:
(b) W and X are the most similar (1 difference), so they are sister species sharing the most recent common ancestor. Y is next closest to that pair (2 differences from X). Z differs most from all the others and branches off first, nearest the root. The cladogram is:
To justify it by parsimony, count the minimum number of base changes each possible grouping requires. The grouping that pairs W with X and Y with Z (as this cladogram does) needs 6 changes; the two alternatives, pairing W with Y or W with Z, each need 8 changes. The chosen cladogram explains all the differences with the fewest changes, so it is the most parsimonious.
(c) 1% divergence per 5 million years, so 2% corresponds to:
The estimate may be inaccurate because the rate of the molecular clock is not constant — it varies with generation time, population size and the strength of selection on the gene, and it depends on calibration against fossils that may themselves be imprecisely dated.
📝Practise
Work through these on paper, then reveal the answer. All are HL only.
1. Define the term clade and explain why a node on a cladogram is described as a hypothetical common ancestor.
2. Explain why base sequences provide more objective evidence for clades than morphology.
3. Outline why the molecular clock can only give estimates of divergence times.
4. State what is meant by parsimony analysis and why it is used to choose between cladograms.
5. Using the figwort family as an example, explain how cladistics can show a traditional classification to be false.
6. Explain the evidence that led to the classification of organisms into three domains, and name the domains.
🔗Go deeper — other people’s work
These are external resources, not mine. If one stops working, tell me and everything above it on this page still stands.
- Understanding Evolution (UC Museum of Paleontology), Reading trees — the clearest short guide to nodes, clades and why tip order does not matter.
- OneZoom Tree of Life — an interactive, zoomable tree of all described species built from phylogenetic data.
- HHMI BioInteractive — activities on building cladograms from DNA and protein sequences.