Protein synthesis
🎯What you need to be able to do
- Explain transcription, including the roles of RNA polymerase, hydrogen bonding and complementary base pairing.
- Explain why DNA templates are stable, and why transcription is a key point for switching genes on and off.
- Explain translation, including the roles of mRNA, ribosomes and tRNA, codons and anticodons.
- Explain the features of the genetic code, and use a codon table to deduce an amino acid sequence.
- Describe elongation of the polypeptide, and explain how a point mutation can change protein structure.
- AHL Explain 5′ to 3′ transcription and translation, initiation at the promoter, and non-coding sequences.
- AHL Explain post-transcriptional modification, alternative splicing and the initiation of translation, including the A, P and E sites.
- AHL Explain modification of polypeptides, using insulin, and recycling of amino acids by proteasomes.
📚The biology
Transcription
Transcription is the synthesis of RNA using a DNA template. In eukaryotes it happens in the nucleus. The enzyme RNA polymerase carries out the whole process:
- RNA polymerase binds to DNA at the start of a gene and separates the two strands, breaking the hydrogen bonds between them over a short stretch.
- One strand, the template strand, is used as the pattern. Free RNA nucleotides pair with its exposed bases by complementary base pairing.
- RNA polymerase links the RNA nucleotides together with covalent bonds, forming an RNA strand.
- As it moves along the gene, the RNA separates from the template and the DNA strands rejoin behind it.
The RNA is released when the end of the gene is reached. When the gene codes for a polypeptide, the RNA is messenger RNA (mRNA).
Base pairing in transcription
RNA nucleotides are matched to the template by hydrogen bonding between complementary bases. The rules are the same as in DNA except that RNA contains uracil instead of thymine: adenine on the DNA template pairs with uracil on the RNA; thymine pairs with adenine; guanine and cytosine pair with each other. The hydrogen bonds hold each nucleotide in position just long enough for it to be joined to the growing strand.
Stability of DNA templates
A single DNA strand can be used as a template without its base sequence changing: transcription breaks only hydrogen bonds between the strands, never the covalent backbone, and the strands rejoin afterwards. This matters most in somatic cells that do not divide, such as neurons and muscle cells, which may live for decades. Their genes are transcribed thousands of times, and the DNA sequences must be conserved throughout the life of the cell, since there is no chance to replace them.
Transcription and gene expression
Gene expression is the process by which the information in a gene is used to make a product. Not all genes are expressed in a cell at any given time: a skin cell does not make haemoglobin, and a cell may make an enzyme only when its substrate is present. Transcription is the first stage of gene expression, so it is a key point at which expression can be switched on or off: if a gene is not transcribed, no protein can be made from it (D2.2).
Translation
Translation is the synthesis of a polypeptide using the base sequence of mRNA. It takes place on ribosomes in the cytoplasm. The base sequence of the mRNA is translated into the amino acid sequence of the polypeptide.
mRNA, ribosomes and tRNA
Carries the base sequence of the gene. It binds to the small subunit of the ribosome.
Made of rRNA and protein, in a small and a large subunit. It holds the mRNA and tRNAs in position and catalyses peptide bond formation. Two tRNAs can bind at the same time to the large subunit.
Transfer RNA: a folded RNA molecule with an anticodon of three bases at one end and a binding site for one specific amino acid at the other. Brings amino acids to the ribosome.
Codons and anticodons
The mRNA is read in groups of three bases called codons. Each tRNA has a matching anticodon, three bases complementary to a codon. The anticodon pairs with the codon by complementary base pairing (hydrogen bonds), which ensures that the correct amino acid is brought for each codon. For example, the codon AUG pairs with the anticodon UAC on a tRNA carrying methionine.
Features of the genetic code
- Triplet code. Each codon is three bases. With 4 bases, pairs would give only 42 = 16 combinations, too few for 20 amino acids; triplets give 43 = 64 combinations, enough for all 20 plus start and stop signals.
- Degeneracy. Most amino acids are coded for by more than one codon; for example GAA and GAG both code for glutamic acid. (Degenerate means redundant here, not faulty.)
- Universality. The same codons code for the same amino acids in almost all organisms, evidence of common ancestry (A1.2).
- Start and stop codons. AUG (methionine) starts translation; UAA, UAG and UGA are stop codons, which code for no amino acid and end translation.
Using the codon table
A genetic code table lists all 64 mRNA codons and the amino acid each codes for. To deduce an amino acid sequence:
- If given a DNA template strand, first write the complementary mRNA (using U, not T). If given the coding (non-template) strand, the mRNA has the same sequence with U in place of T.
- Divide the mRNA into codons from the start codon, three bases at a time.
- Look up each codon: the table is usually arranged by first base (rows), second base (columns) and third base (within each box).
- Stop at a stop codon.
Elongation of the polypeptide
- A tRNA carrying an amino acid binds to the ribosome where its anticodon pairs with the codon on the mRNA.
- A second tRNA with the next amino acid binds to the next codon, alongside the first.
- The ribosome catalyses the formation of a peptide bond between the two amino acids. The growing polypeptide is transferred to the second tRNA.
- The ribosome moves along the mRNA by one codon. The first tRNA, now without an amino acid, is released.
- The next tRNA binds to the newly exposed codon, and the cycle repeats, adding one amino acid at a time.
Mutations that change protein structure
A point mutation changes a single base in a gene. Because of the one-to-one link between codons and amino acids, it can change one amino acid in the polypeptide. The classic example is sickle cell anaemia. A single base substitution in the gene for the beta chain of haemoglobin changes the mRNA codon GAG to GUG, so the sixth amino acid becomes valine instead of glutamic acid. Glutamic acid is charged and hydrophilic; valine is non-polar. At low oxygen concentrations the altered haemoglobin molecules stick together into long fibres, distorting red blood cells into a sickle shape that blocks capillaries and is destroyed quickly (D1.3).
Directionality of transcription and translation AHL
- Transcription is 5′ → 3′. RNA polymerase adds nucleotides to the 3′ end of the growing RNA, so the RNA is built in the 5′ to 3′ direction. Because the template is antiparallel, RNA polymerase moves along the template strand in its 3′ to 5′ direction.
- Translation is 5′ → 3′. The ribosome reads the mRNA starting near its 5′ end and moving towards the 3′ end. The first codon read (the start codon) is nearest the 5′ end.
Initiation at the promoter AHL
A promoter is a base sequence in the DNA just before (upstream of) a gene, where transcription begins. RNA polymerase cannot bind to it on its own in eukaryotes: proteins called transcription factors first bind to specific sequences in the promoter, and RNA polymerase then binds to them and starts transcription. Which transcription factors are present in a cell decides which genes are transcribed. (You do not need to name any transcription factors.)
Non-coding sequences AHL
Much of the DNA in eukaryotes does not code for polypeptides. Non-coding sequences include:
promoters, enhancers and other sequences where proteins bind to control transcription.
sequences within a gene that are transcribed but removed from the RNA before translation.
repetitive sequences at the ends of chromosomes, protecting them from damage and loss of genes during replication.
transcribed into functional RNA molecules, but never translated into polypeptides.
Post-transcriptional modification AHL
In eukaryotes, the initial RNA transcript (pre-mRNA) is modified in the nucleus before it is exported:
- Splicing. Introns are removed, and the coding sequences, exons, are spliced together to form mature mRNA.
- A 5′ cap (a modified guanine nucleotide) is added to the 5′ end. It protects the mRNA and helps the ribosome bind.
- A 3′ poly-A tail of many adenine nucleotides is added to the 3′ end.
The cap and tail stabilize the mRNA, protecting it from being broken down by enzymes in the cytoplasm, so it can be translated many times.
Alternative splicing AHL
The exons of a gene need not always be spliced together in the same combination. Alternative splicing joins different combinations of exons, including some and leaving out others, so that one gene can code for several different polypeptides. This is one reason humans can make far more different proteins than they have genes.
Initiation of translation AHL
- The small ribosomal subunit attaches to the 5′ end of the mRNA (at the cap) and moves along it until it reaches the start codon, AUG.
- An initiator tRNA, with anticodon UAC and carrying methionine, pairs with the start codon.
- The large subunit attaches, forming a complete ribosome.
- A second tRNA, carrying the next amino acid, binds to the next codon, and elongation begins.
The ribosome has three binding sites for tRNA, used in turn during elongation:
where an incoming tRNA carrying an amino acid binds to the codon.
holds the tRNA attached to the growing polypeptide chain. The peptide bond forms between the chain here and the amino acid at the A site.
where the tRNA, now without an amino acid, sits before it leaves the ribosome.
As the ribosome moves one codon along, each tRNA shifts from A to P to E.
Modifying polypeptides into a functional state AHL
Many polypeptides must be modified after translation before they can function — by folding, by removal of sections, by adding carbohydrate or other groups, or by joining with other polypeptides. Insulin is made in two stages of modification:
- It is translated on ribosomes of the rough ER as pre-proinsulin (110 amino acids). A signal sequence at the start directs it into the ER, and is then cut off, giving proinsulin (86 amino acids), which folds and forms disulfide bonds.
- In the Golgi and secretory vesicles, a central section, the C-peptide, is removed from proinsulin, leaving the two chains of active insulin (21 and 30 amino acids) held together by disulfide bonds.
Proteasomes AHL
Proteins do not last forever. Damaged, misfolded or no-longer-needed proteins are tagged and broken down by proteasomes, large barrel-shaped protein complexes that hydrolyse proteins into short peptides and amino acids. The amino acids are recycled into new proteins. Sustaining a functional proteome — the full set of proteins a cell contains — requires this constant balance of breakdown and synthesis, allowing a cell to change its proteins as conditions change.
✏️Worked example
(a) Write the mRNA transcribed from it, with its 5′ and 3′ ends labelled.
(b) Use the genetic code to deduce the amino acid sequence. (AUG = Met; CCU = Pro; GAA = Glu; GUA = Val; GAG = Glu; UGC = Cys; UAA = stop.)
(c) Give the anticodon of the tRNA that brings the second amino acid.
(d) A mutation changes the template triplet CTT to CAT. Deduce the effect on the polypeptide. A different mutation changes CTT to CTC. Explain why this has no effect.
(a) Pair each template base with its RNA complement (A→U, T→A, G→C, C→G). The mRNA is antiparallel to the template:
(b) Read the codons from the 5′ end: AUG → Met; CCU → Pro; GAA → Glu; UGC → Cys; UAA → stop. Polypeptide: Met–Pro–Glu–Cys.
(c) The second codon is CCU, so the anticodon is its complement: GGA (written 3′–GGA–5′ to pair with 5′–CCU–3′).
(d) CAT on the template gives the codon GUA, which codes for valine instead of glutamic acid. The polypeptide becomes Met–Pro–Val–Cys. Glutamic acid is charged and hydrophilic while valine is non-polar, so the protein’s folding and function may change — the same kind of change as in sickle cell haemoglobin. CTC on the template gives the codon GAG, which also codes for glutamic acid: because the code is degenerate, this substitution is silent and the polypeptide is unchanged.
📝Practise
Work through these on paper, then reveal the answer. Questions 4–6 are AHL.
1. Outline the role of RNA polymerase in transcription.
2. Explain why the genetic code must be a triplet code.
3. Describe how the ribosome, mRNA and tRNA interact to add an amino acid to a growing polypeptide.
4. AHL Outline the modifications made to pre-mRNA in eukaryotic cells before translation.
5. AHL Explain the roles of the A, P and E sites of the ribosome during elongation.
6. AHL Describe how pre-proinsulin is modified to produce functional insulin.
🔗Go deeper — other people’s work
These are external resources, not mine. If one stops working, tell me and everything above it on this page still stands.
- DNA Learning Center (Cold Spring Harbor) — the Transcription and Translation 3D animations.
- Learn.Genetics (University of Utah) — the interactive Transcribe and Translate a Gene activity.
- RCSB Protein Data Bank, Molecule of the Month — the ribosome, RNA polymerase and the proteasome.