Nucleic acids and protein synthesis
🎯What you need to be able to do
- Describe the structure of nucleotides, including the phosphorylated nucleotide ATP.
- State that adenine and guanine are purines (double ring) and cytosine, thymine and uracil are pyrimidines (single ring).
- Describe the DNA double helix: antiparallel strands, complementary base pairing, the difference in hydrogen bonding between C–G and A–T, and phosphodiester bonds.
- Describe semi-conservative replication, the roles of DNA polymerase and DNA ligase, and the difference between leading and lagging strand replication.
- Describe the structure of mRNA and explain the universal genetic code.
- Describe transcription and translation, including the roles of RNA polymerase, mRNA, codons, tRNA, anticodons and ribosomes.
- Explain the removal of introns and joining of exons, and outline how substitution, deletion and insertion mutations affect the polypeptide.
📚The biology
Nucleotides
A nucleotide has three parts: a pentose sugar (deoxyribose in DNA, ribose in RNA), a phosphate group, and a nitrogenous base.
- Purines — adenine and guanine — have a double ring structure.
- Pyrimidines — cytosine, thymine and uracil — have a single ring.
A purine always pairs with a pyrimidine, which is why the double helix has a constant width. You do not need structural formulae for the bases.
ATP is a phosphorylated nucleotide: adenine, ribose and three phosphate groups. Hydrolysing the terminal phosphate releases energy for cellular work, and ATP is resynthesised in respiration (Topic 12). Recognising ATP as a nucleotide is the link between this topic and every energy question.
The double helix
Nucleotides are joined into a strand by phosphodiester bonds between the phosphate of one and the sugar of the next, forming a sugar–phosphate backbone. Two strands wind into a double helix, held by hydrogen bonds between complementary bases:
The two strands are antiparallel: one runs 5′ to 3′, the other 3′ to 5′. That detail looks like trivia until replication, where it is the whole explanation for the lagging strand.
Because C–G has three hydrogen bonds and A–T only two, DNA rich in C and G is harder to separate and needs a higher temperature to melt — a favourite application question.
Semi-conservative replication
Replication happens in the S phase of interphase (Topic 5). The hydrogen bonds between the bases break and the strands separate. Each acts as a template: free nucleotides pair with the exposed bases according to the base-pairing rules, and DNA polymerase joins them by forming phosphodiester bonds.
“Semi-conservative” means each new molecule contains one original strand and one new strand. That is what conserves the sequence and what the Meselson–Stahl experiment demonstrated.
DNA polymerase can only add nucleotides in the 5′ to 3′ direction. Since the two template strands are antiparallel, this has an awkward consequence:
- the leading strand is synthesised continuously, following the replication fork;
- the lagging strand must be made in short fragments running away from the fork, and those fragments are then joined by DNA ligase.
You are not required to know the other enzymes involved, or the different types of DNA polymerase.
The genetic code
A gene is a sequence of nucleotides forming part of a DNA molecule that codes for a polypeptide. Three bases — a triplet in DNA, a codon in mRNA — specify one amino acid, or a start or stop signal. The code is:
- universal — the same codons mean the same amino acids in nearly all organisms, which is what makes genetic engineering possible (Topic 19);
- degenerate — most amino acids have more than one codon;
- non-overlapping — each base belongs to one codon only.
Transcription
In the nucleus, the DNA of one gene unwinds and the hydrogen bonds break. Only one strand is used: the transcribed (template) strand; the other is the non-transcribed strand. Free RNA nucleotides pair with the exposed bases — with uracil in place of thymine opposite adenine — and RNA polymerase joins them into a single-stranded molecule. This leaves the nucleus through a nuclear pore.
In eukaryotes the molecule made first is a primary transcript containing both coding and non-coding sequences. The non-coding introns are removed and the coding exons are joined together to form the mature mRNA. Prokaryotes do not do this, which matters when you engineer a human gene into a bacterium.
Translation
mRNA binds to a ribosome, which holds two codons at a time. tRNA molecules each carry a specific amino acid and have an anticodon of three bases. A tRNA whose anticodon is complementary to the exposed codon binds; a peptide bond forms between its amino acid and the growing chain; the ribosome moves on one codon; the empty tRNA leaves to collect another amino acid. This repeats until a stop codon is reached and the polypeptide is released.
The polypeptide then folds into its tertiary structure (Topic 2), and if it is destined for secretion it passes through the rough ER and Golgi (Topic 1).
Gene mutations
A gene mutation is a change in the sequence of base pairs in DNA that may result in an altered polypeptide. Three types:
- Substitution — one base is replaced. Only one codon changes, so at most one amino acid changes. Because the code is degenerate the new codon may still code for the same amino acid, in which case there is no effect at all. If the amino acid does change, the effect depends on where it is: a change far from the active site of an enzyme may do nothing; a change in the active site or in a critical R group can be severe — sickle cell anaemia is a single substitution in the HBB gene.
- Deletion — a base is removed.
- Insertion — a base is added.
Deletion and insertion cause a frameshift: every codon from that point onwards is read in the wrong grouping, so nearly all the amino acids downstream are wrong and the polypeptide is usually non-functional. That is why they are typically far more damaging than a substitution.
✏️Worked example
(a) mRNA is complementary to the template strand, with uracil replacing thymine. Taking each base in turn: T→A, A→U, C→G, and so on.
Written the other way it would be wrong, and reading the mRNA off the non-transcribed strand — a tempting shortcut, since it has the same sequence with U for T — must still be written 5′ to 3′.
(b) The third mRNA codon is GAA. The tRNA anticodon is complementary to it, so it is CUU. Note that the anticodon is RNA, so it contains uracil, not thymine — writing CTT is the usual slip.
(c) Removing a base is a deletion, and because a base is lost the reading frame shifts: this is a frameshift mutation. Every codon from that point on is read in the wrong grouping, so all the amino acids downstream are likely to be wrong, and a stop codon may appear early or the original one be missed. The polypeptide will almost certainly be non-functional.
A substitution of the same base changes only that one codon. Because the genetic code is degenerate, the new codon may still specify the same amino acid, giving no change at all; if the amino acid does change, only one residue in the whole chain differs, and the protein may still work — unless that residue lies at the active site or is essential to the fold. Deletion is therefore far more likely to be damaging than substitution.
📝Practise
Work through these, then reveal the answer. Each question targets a different objective from the list above.
1. Explain why DNA containing a high proportion of C and G requires a higher temperature to separate its strands than DNA rich in A and T.
2. Describe semi-conservative replication and explain how the Meselson–Stahl style of experiment supports it.
3. Compare the structure of an mRNA molecule with that of a DNA molecule. Give four differences.
4. Explain why the lagging strand is synthesised in fragments while the leading strand is continuous.
5. A human gene is inserted into a bacterium but no functional human protein is produced. Suggest an explanation involving introns, and how the problem is solved.
6. A substitution mutation occurs in the sixth codon of a gene, and the resulting protein is completely non-functional. A substitution in the ninetieth codon of the same gene has no detectable effect. Suggest two reasons for the difference.
🔗Go deeper — other people’s work
These are external resources, not mine. If one stops working, tell me and everything above it on this page still stands.
- DNA Learning Center (Cold Spring Harbor) — the classic animations of replication, transcription and translation, which make the leading and lagging strands obvious in seconds
- “Learn Genetics” (University of Utah) — a build-a-protein interactive where you read codons off an mRNA yourself
- Any standard genetic code table — you are not required to memorise it, but practising reading one under exam conditions is worth doing once