HomeLearning HubIB DP BiologyD1.3 Mutation and gene editing
D1.3

Mutation and gene editing

Theme D · Continuity and change · Molecules · SL and HL · plus additional higher level

A mutation is a change to the base sequence of DNA. Most are harmless, some are harmful, and a very few are useful — but without them there would be no genetic variation and no evolution at all. This topic explains the types of gene mutation, their causes and consequences, and why they are random. At HL it turns to the tools that now let scientists change genes deliberately.

🎯What you need to be able to do

  • Distinguish base substitutions, insertions and deletions as types of gene mutation.
  • Explain the consequences of base substitutions, including SNPs and the role of degeneracy.
  • Explain the consequences of insertions and deletions, including frameshifts.
  • Outline the causes of gene mutation: chemical mutagens, radiation, and errors in replication or repair.
  • Explain why mutation is random.
  • Compare the consequences of mutation in germ cells and somatic cells, including cancer.
  • Explain mutation as the original source of all genetic variation.
  • AHL Explain gene knockout, and the use of CRISPR sequences and Cas9 in gene editing.
  • AHL Evaluate hypotheses to explain conserved and highly conserved sequences.

📚The biology

Types of gene mutation

A gene mutation is a structural change to a gene at the molecular level: a change in its base sequence. There are three types:

Substitution
One base is replaced by a different base.
GAA → GUA
Insertion
One or more bases are added to the sequence.
GAA → GAAA
Deletion
One or more bases are removed from the sequence.
GAA → GA

Consequences of base substitutions

A base substitution changes only one codon. Because the genetic code is degenerate — most amino acids have several codons — the consequence depends on the new codon:

  • Silent (same-sense): the new codon codes for the same amino acid, so the polypeptide is unchanged. Example: GAA → GAG, both glutamic acid.
  • Missense: the new codon codes for a different amino acid, changing one amino acid in the polypeptide. The effect may be negligible, or it may change the protein’s shape and function, as in sickle cell anaemia (D1.2).
  • Nonsense: the new codon is a stop codon, so translation ends early and the polypeptide is truncated, usually non-functional.

Single-nucleotide polymorphisms (SNPs) — positions where individuals of a species differ by a single base — are the result of base substitution mutations that have become established in a population. Because of degeneracy, a SNP may or may not change an amino acid.

Consequences of insertions and deletions

mRNA is read in consecutive triplets from the start codon. Inserting or deleting a single base (or any number not a multiple of three) shifts the reading frame for every codon after it: a frameshift. Every amino acid downstream is likely to change, and a stop codon often appears early. The polypeptide is very likely to cease to function.

Insertions or deletions of three bases (or a multiple of three) add or remove whole codons without a frameshift, but major insertions or deletions of many codons are still likely to destroy the polypeptide’s function.

Causes of gene mutation

Mutations arise in two main ways.

Mutagens
Agents that increase the rate of mutation.
Chemical mutagens: compounds in tobacco smoke (such as benzo[a]pyrene), mustard gas, aflatoxin from moulds on stored grain and nuts, nitrosamines.
Mutagenic radiation: ultraviolet light (causing adjacent thymines to bond together), X-rays, and gamma rays and other ionizing radiation from radioactive materials, which break DNA strands.
Errors in replication or repair
DNA polymerase occasionally inserts the wrong nucleotide and proofreading misses it, or bases are added or skipped at repetitive sequences. Repair enzymes that fix damaged DNA can also introduce errors.

Randomness in mutation

Mutations can occur anywhere in the base sequence of a genome. They are random in the sense that their position and effect are not directed: some bases do have a higher probability of mutating than others (for example cytosines that are methylated, or sequences that are frequently repeated), but no position is protected. Importantly, no natural mechanism is known that makes a deliberate change to a particular base in order to change a trait. An organism in a cold environment does not produce mutations for cold tolerance; mutations happen regardless of whether they would be useful, and selection then acts on whatever arises.

Germ cells and somatic cells

Mutations in germ cells
Germ cells give rise to gametes. A mutation there can be passed on to offspring, and present in every cell of the child. It is inherited, and so can contribute to genetic diseases and to evolution.
Mutations in somatic cells
Somatic (body) cells do not form gametes, so these mutations are not inherited. They affect only the individual, and only the descendants of that one cell. If mutations occur in genes that control cell division, the result can be uncontrolled division and cancer (D2.1).

Mutation as the source of variation

Gene mutation is the original source of all genetic variation. Every allele of every gene first arose by mutation. Sexual reproduction reshuffles existing alleles into new combinations, but only mutation creates new ones.

For an individual organism, most mutations are either neutral (no effect, for instance silent substitutions or changes in non-coding DNA) or harmful. Only occasionally is a mutation beneficial. But in a species, over the long term, mutation is essential: without a continual supply of new variation, natural selection would have nothing to act on, and populations could not adapt to changing environments (D4.1).

Commercial genetic tests, sold directly to consumers, analyse SNPs to estimate a person’s risk of certain diseases. Without expert interpretation, this information can be problematic: an increased relative risk may sound alarming but still be small in absolute terms; a negative result may give false reassurance because the test examines only some variants; and results can cause anxiety or affect relatives.

Gene knockout AHL

One way to find out what a gene does is to stop it working and see what changes. Gene knockout is a technique in which a gene is altered to make it inoperative in an organism. Comparing the knockout organism with a normal one reveals the gene’s function: if mice lacking a particular gene become obese, for example, that gene is involved in regulating appetite or metabolism. For some model species used in research, such as mice, fruit flies, zebrafish, yeast and the plant Arabidopsis, libraries of knockout organisms are available, with a separate strain for many genes. (You do not need to know the details of how knockouts are made.)

CRISPR–Cas9 gene editing AHL

CRISPR–Cas9 is a technique for editing the genome at a precise location. It uses two components, originally from a bacterial defence system:

  • Guide RNA, designed with a sequence of about 20 bases complementary to the target DNA sequence. It directs the enzyme to exactly that place in the genome.
  • Cas9, an enzyme (a nuclease) that cuts both strands of the DNA at the site where the guide RNA binds.

After the cut, the cell repairs the break. Repair is often imperfect, introducing small insertions or deletions that knock out the gene. Alternatively, if a piece of DNA with the desired sequence is supplied, the cell can use it as a template during repair, inserting or correcting a sequence. CRISPR is faster, cheaper and far more precise than earlier methods.

An example of successful use: in sickle cell disease and beta thalassaemia, a patient’s own blood stem cells are removed and edited with CRISPR–Cas9 to switch back on the production of foetal haemoglobin, which does not sickle. The edited cells are returned to the patient. In clinical trials most patients no longer needed blood transfusions or suffered painful crises, and the first such therapy was approved for medical use in 2023.

Some potential uses raise serious ethical issues that must be addressed before implementation — above all, editing human embryos or germ cells, which would change the genes of future generations who cannot consent, and which in 2018 was carried out without approval. Scientists in different countries are subject to different regulatory systems, so there is an international effort to harmonize regulation of genome-editing technologies.

Conserved sequences AHL

Conserved sequences are base sequences that are identical or very similar across a species or a group of species. Highly conserved sequences are identical or similar over long periods of evolution, across distantly related organisms — for example the genes for histones, ribosomal RNA and cytochrome c. Two hypotheses account for them:

Functional requirement
The gene product has an essential function that depends on a precise sequence. Almost any mutation would disrupt it and be harmful, so individuals carrying such mutations fail to survive or reproduce: natural selection removes the changes. The sequence stays the same because change is selected against.
Slower mutation rate
Some regions of the genome may genuinely mutate less often, for example because they are better protected or repaired more efficiently, so fewer changes arise in the first place.

The two can be distinguished by evidence. If a sequence is conserved because of selection, then silent changes within it (which do not alter the protein) should still accumulate, while changes to amino acids should not. Comparisons of this kind generally support the functional-requirement hypothesis for most conserved coding sequences.

✏️Worked example

The mRNA of a short gene is 5′–AUG CCU GAA UGC AAA GGC UAA–3′, coding for Met–Pro–Glu–Cys–Lys–Gly. (Codons needed: ACC = Thr; UGA = stop; CUG = Leu; AAU = Asn; GCA = Ala; AAG = Lys; GCU = Ala.)
Deduce the effect on the polypeptide of each mutation, and name the type:
(a) an extra A inserted after the first codon;
(b) the C at position 5 (the second base of codon 2) deleted;
(c) UGC changed to UGA.
(d) Explain which of these mutations is least likely to leave a functional protein, and why.

(a) Insertion. The mRNA becomes AUG ACC UGA AUG…, read as AUG | ACC | UGA. The polypeptide is Met–Thr, then a stop codon: a frameshift changes every codon after the insertion, and a stop codon appears immediately.

(b) Deletion. Removing the C gives AUG CUG AAU GCA AAG GCU A…. The polypeptide becomes Met–Leu–Asn–Ala–Lys–Ala…: a frameshift. Every amino acid after the first is different, and the original stop codon is no longer in frame, so translation would continue past it until another stop codon happened to be reached.

(c) Substitution. UGC (Cys) becomes UGA, a stop codon. The polypeptide is Met–Pro–Glu, cut short: a nonsense mutation.

(d) All three are severe here, but the frameshift mutations (a) and (b) are the least likely to leave a functional protein in a real gene hundreds of codons long, because they change every amino acid downstream of the mutation. A single substitution usually alters only one codon; it is only severe here because it happened to create a stop codon. Most substitutions are silent or change a single amino acid.

Check it. After an insertion or deletion, re-divide the entire downstream sequence into triplets from the start codon — do not just edit one codon. A reliable check is to count the bases: the original has 21; after the insertion there are 22 and after the deletion 20, so neither can still be read in the original frame.
“A substitution always changes the amino acid.” Because the code is degenerate, many substitutions — especially in the third base of a codon — give a codon for the same amino acid, and have no effect. State the three possibilities (silent, missense, nonsense) rather than assuming one.

📝Practise

Work through these on paper, then reveal the answer. Questions 5 and 6 are AHL.

1. Explain why a base substitution mutation may have no effect on the polypeptide produced.
The genetic code is degenerate: most amino acids are coded for by more than one codon. If a substitution changes a codon to another codon for the same amino acid (for example GAA to GAG, both glutamic acid), the amino acid sequence of the polypeptide is unchanged. This is a silent mutation. (A substitution in a non-coding region, such as an intron, would also have no effect on the polypeptide.)
2. Explain why the insertion of a single base usually has a greater effect than a base substitution.
mRNA is read in consecutive triplets. Inserting one base causes a frameshift: every codon after the insertion is read differently, so many amino acids change, and a premature stop codon is likely to appear. The polypeptide is very likely to be non-functional. A substitution changes only one codon, so at most one amino acid changes (unless it creates a stop codon), and it may be silent because of degeneracy.
3. Distinguish between the consequences of mutations in germ cells and in somatic cells.
A mutation in a germ cell may be passed into a gamete and so be inherited by offspring, appearing in every cell of the offspring’s body; it can cause inherited disease and contributes to the gene pool and evolution. A mutation in a somatic cell is not inherited; it affects only that cell and its descendants in the individual. If it is in a gene that controls the cell cycle, it can lead to uncontrolled cell division and cancer.
4. Explain why mutation is essential for evolution even though most mutations are neutral or harmful.
Mutation is the original source of all new alleles, and so of all genetic variation. Sexual reproduction only reshuffles existing alleles. Natural selection can only act on heritable variation that already exists. Although most mutations are neutral or harmful to an individual, the occasional beneficial mutation increases fitness in a particular environment; selection increases its frequency over generations. Without a continual supply of mutations, variation would be used up and populations could not adapt to changing environments.
5. AHL Outline how CRISPR–Cas9 is used to edit a specific gene.
A guide RNA is designed with a base sequence complementary to the target DNA sequence. It is combined with the enzyme Cas9 and delivered into cells. The guide RNA binds to the matching DNA sequence by base pairing, directing Cas9 to that site, where Cas9 cuts both strands of the DNA. The cell repairs the break: imperfect repair often introduces insertions or deletions that knock out the gene, or a supplied DNA template can be used to insert or correct a sequence. Example: editing blood stem cells to treat sickle cell disease.
6. AHL The gene for histone H4 has almost the same sequence in peas and cows. Suggest two hypotheses to explain this.
This is a highly conserved sequence. Hypothesis 1 — functional requirement: histone H4 has an essential function in packaging DNA into nucleosomes, and its structure must be precise; almost any change to its amino acid sequence would impair this, so mutations are selected against and removed by natural selection over hundreds of millions of years. Hypothesis 2 — slower mutation rate: the region of DNA containing the gene may mutate less frequently than other regions, so few changes arise. (Evidence that silent base changes have still accumulated while amino acid changes have not would support hypothesis 1.)

🔗Go deeper — other people’s work

These are external resources, not mine. If one stops working, tell me and everything above it on this page still stands.

  • HHMI BioInteractive — the animated explanation of CRISPR–Cas9 and its uses.
  • Learn.Genetics (University of Utah) — The Outcome of Mutation, showing silent, missense, nonsense and frameshift mutations.
  • National Human Genome Research Institute — fact sheets on gene editing ethics, knockout mice and genetic testing.