HomeLearning HubIB DP BiologyA1.2 Nucleic acids
A1.2

Nucleic acids

Theme A · Unity and diversity · Molecules · SL and HL · plus additional higher level

DNA stores the instructions for building an organism, and RNA helps carry them out. Both are chains of just four kinds of nucleotide, and almost everything interesting about them follows from two simple facts: the backbone is held together by strong covalent bonds, and the bases pair with each other using weak hydrogen bonds in only one way. Strong enough to keep information safe, weak enough to open up and read it.

🎯What you need to be able to do

  • State that DNA is the genetic material of all living organisms (and that some viruses use RNA).
  • Draw a nucleotide using a circle, a pentagon and a rectangle, and draw a short RNA strand showing how nucleotides are linked by condensation.
  • Name the four bases in DNA and in RNA, and explain how the sugar–phosphate backbone forms.
  • Draw DNA as two antiparallel strands joined by hydrogen bonds between A–T and G–C pairs.
  • Compare DNA and RNA: number of strands, bases and sugar — including a sketch of ribose and deoxyribose.
  • Explain how complementary base pairing lets genetic information be copied and expressed, why DNA can store almost unlimited information, and why a shared genetic code points to common ancestry.
  • AHL Explain 5′ to 3′ directionality, why purine–pyrimidine pairing keeps the helix a constant width, and the structure of a nucleosome.
  • AHL Explain how the Hershey–Chase experiment and Chargaff’s data provided evidence about DNA.

📚The biology

DNA as the genetic material

Every living organism — bacterium, fungus, plant, animal — stores its genetic information as DNA (deoxyribonucleic acid). Some viruses use RNA (ribonucleic acid) as their genetic material instead, but viruses are not considered living (A2.3), so the statement “all living things use DNA” still holds.

Nucleotides

DNA and RNA are polymers whose monomers are nucleotides. Every nucleotide has three parts:

Phosphate group
Drawn as a circle. Carries a negative charge, which is why nucleic acids are acidic.
Pentose sugar
Drawn as a pentagon. A five-carbon sugar: deoxyribose in DNA, ribose in RNA.
Nitrogenous base
Drawn as a rectangle. One of four bases, attached to the sugar.

In a diagram the phosphate and the base are both attached to the sugar, on opposite sides of it: the phosphate joins carbon 5 of the pentose and the base joins carbon 1.

The sugar–phosphate backbone

Nucleotides are joined by condensation reactions: the phosphate of one nucleotide bonds covalently to the sugar of the next, and a molecule of water is released. Repeating this produces a continuous chain of alternating sugar and phosphate — the sugar–phosphate backbone — with a base sticking out from every sugar.

Because every link in the backbone is a covalent bond, the strand is strong and stable. The information is carried by the sequence of bases along it, and that sequence is not disturbed when the two strands of DNA separate, because separation breaks only the weak hydrogen bonds between bases, not the backbone.

The bases and the code

There are four bases in each nucleic acid:

DNA
adenine (A), thymine (T), guanine (G), cytosine (C)
RNA
adenine (A), uracil (U), guanine (G), cytosine (C)

The order of bases along a strand is a code. Genes are sequences of bases, read three at a time during protein synthesis (D1.2), and the sequence determines the amino acid sequence of the polypeptide made.

RNA as a polymer

RNA is normally a single strand of nucleotides linked by condensation reactions into a sugar–phosphate backbone. To draw it, draw several nucleotides in a column, each phosphate (circle) linked to the sugar (pentagon) of the nucleotide above or below, with the bases (rectangles) all projecting to one side.

The DNA double helix

DNA is made of two strands of nucleotides wound into a double helix. Three features define it:

  • The strands are antiparallel — they run in opposite directions. In a diagram, draw one strand with its sugars the right way up and the other upside down.
  • Bases pair across the middle, joined by hydrogen bonds, with the backbones on the outside.
  • Pairing is complementary: adenine always pairs with thymine, and guanine always pairs with cytosine.

You do not need to draw the helical twist, memorize which bases are larger, or learn how many hydrogen bonds hold each pair together.

DNA and RNA compared

Strands
DNA: two, forming a double helix
RNA: usually one
Sugar
DNA: deoxyribose
RNA: ribose
Bases
DNA: A, T, G, C
RNA: A, U, G, C

The two sugars differ at a single carbon. Ribose has an –OH group on carbon 2; deoxyribose (“de-oxy”, one oxygen fewer) has only –H there. To sketch them, draw the same five-membered ring (four carbons and an oxygen) with carbon 5 sticking up outside the ring, and change only the group on carbon 2.

Examples of nucleic acids worth knowing by name: the DNA of chromosomes; messenger RNA (mRNA), which carries a copy of a gene to ribosomes; transfer RNA (tRNA), which brings amino acids; and ribosomal RNA (rRNA), part of the structure of ribosomes.

Why complementary pairing matters

Because A pairs only with T (or U) and G pairs only with C, the base sequence of one strand determines the sequence of the other. That one rule does two jobs.

  • Replication. The two strands separate and each acts as a template: new nucleotides are added by complementary pairing, so the two new molecules have the same sequence as the original (D1.1).
  • Expression. In transcription, an RNA copy of a gene is made by pairing RNA nucleotides with one DNA strand; in translation, tRNA anticodons pair with mRNA codons (D1.2).

In each case, the specificity comes from hydrogen bonding: only the correct base pair has hydrogen-bonding groups in the right positions to fit.

The capacity of DNA to store information

Any base can follow any other base, and a DNA molecule can be any length. So the number of possible sequences is enormous: with four bases, a sequence \( n \) bases long can be arranged in \( 4^{n} \) ways. A sequence just 10 bases long already has over a million possibilities; the human genome has about 3 billion base pairs.

DNA is also extraordinarily compact: the entire genome of a human cell, about two metres of DNA, fits inside a nucleus a few micrometres across. No human-made storage medium holds so much information in so little space.

A universal code and common ancestry

With very few, minor exceptions, every organism uses the same genetic code: the same three-base codon specifies the same amino acid in bacteria, plants and humans. There is no functional reason why one particular set of codon meanings should be the only one that works. The simplest explanation for all life sharing it is that all life inherited it from a single common ancestor. This is why a human gene inserted into a bacterium can be translated into the correct human protein.

Directionality AHL

The carbons of the pentose sugar are numbered 1′ to 5′. In the backbone, each phosphate links the 5′ carbon of one sugar to the 3′ carbon of the next. As a result every strand has two different ends: a 5′ end with a free phosphate on carbon 5′, and a 3′ end with a free –OH on carbon 3′. A strand therefore has a direction, written 5′ → 3′.

In double-stranded DNA, “antiparallel” means exactly this: one strand runs 5′ to 3′ and its partner runs 3′ to 5′ alongside it. Directionality matters because the enzymes that work on nucleic acids can only work in one direction:

  • Replication: DNA polymerase only adds nucleotides to the 3′ end of a growing strand, which is why the two strands are copied differently (D1.1).
  • Transcription: RNA is also built 5′ → 3′.
  • Translation: ribosomes read mRNA from its 5′ end towards its 3′ end.

Purine–pyrimidine pairing AHL

Adenine and guanine are purines, with a double-ring structure. Thymine, cytosine and uracil are pyrimidines, with a single ring. Every base pair in DNA is one purine with one pyrimidine (A–T, G–C), so every base pair is the same length. The two backbones are therefore always the same distance apart, and the double helix has the same regular three-dimensional shape whatever the base sequence. That regularity is part of what makes the helix stable, and it means enzymes can bind DNA in the same way anywhere along its length.

Nucleosomes AHL

In eukaryotes, DNA is wound around proteins called histones. A nucleosome is a length of DNA wrapped around a core of eight histone proteins, with a further histone protein attached to the linker DNA that runs between nucleosomes, holding the arrangement together. A string of nucleosomes looks like beads on a thread, and it can be coiled further to condense chromosomes for division (D2.1).

DNA is negatively charged (its phosphates), and histones are rich in positively charged amino acids, so the two are held together by attraction between opposite charges. Molecular visualization software (for example the RCSB Protein Data Bank viewer) lets you rotate a nucleosome and see the DNA wrapped around its histone core.

Hershey and Chase: DNA, not protein AHL

By the 1950s it was known that chromosomes contain both protein and DNA, and many scientists expected protein, which is far more varied, to be the genetic material. Alfred Hershey and Martha Chase settled the question using bacteriophages, viruses that consist only of a protein coat and DNA and that inject their genetic material into bacteria.

  • One batch of viruses was grown with radioactive sulfur-35. Sulfur is in protein but not DNA, so this labelled the protein coats.
  • Another batch was grown with radioactive phosphorus-32. Phosphorus is in DNA but hardly at all in protein, so this labelled the DNA.
  • Each batch infected bacteria. The mixture was agitated in a blender to knock the virus coats off the bacteria, then centrifuged: the heavier bacteria formed a pellet and the lighter virus coats stayed in the liquid.

Result: the radioactive sulfur was mostly in the liquid, with the empty coats; the radioactive phosphorus was mostly in the pellet, inside the bacteria, which went on to produce new viruses. So the material entering the cell and directing the production of new viruses was DNA.

The experiment was only possible because radioisotopes had become available as research tools — a good example of new technology opening up a new kind of experiment.

Chargaff’s data AHL

An early idea, the tetranucleotide hypothesis, proposed that DNA was a simple repeating sequence of the four bases in equal amounts — too dull a molecule to carry genetic information. Erwin Chargaff measured the proportions of the bases in DNA from many different organisms and found two patterns:

  • the proportions of bases differ between species, so the four bases are not present in equal amounts;
  • but within any species, the amount of A equals T and the amount of G equals C, so purines = pyrimidines.

The first pattern falsified the tetranucleotide hypothesis. This is a point about how science works: however many observations agree with a hypothesis, they cannot prove it true (the “problem of induction”), but one reliable set of contradicting observations can prove it false. The second pattern was a clue that Watson and Crick later explained with complementary base pairing.

✏️Worked example

A sample of double-stranded DNA from a bacterium contains 21% adenine.
(a) Deduce the percentages of thymine, guanine and cytosine.
(b) The base sequence of part of one strand is A T G C C G A A. Write the sequence of the complementary DNA strand, and of an RNA strand complementary to the first strand.
(c) Calculate how many different base sequences are possible for a DNA strand 10 bases long.
(d) AHL The strand in (b) is written 5′ to 3′. Write the complementary DNA strand so that it is also read 5′ to 3′.

(a) In double-stranded DNA every A is paired with a T, so T = 21%. Together A and T make up 42%, which leaves 58% for G and C. G pairs with C, so they are equal: G = 29% and C = 29%.

(b) Pair each base in turn. DNA uses T opposite A; RNA uses U opposite A.

Original strand
A T G C C G A A
Complementary DNA
T A C G G C T T
Complementary RNA
U A C G G C U U

(c) Each of the 10 positions can be any of 4 bases, independently:

\[ 4^{10} = 1\,048\,576 \]

So over a million different sequences are possible from a strand only 10 bases long.

(d) The complementary strand runs antiparallel. Written underneath the original, base by base, it reads 3′–T A C G G C T T–5′. To write it 5′ to 3′, reverse the whole sequence: 5′–T T C G G C A T–3′. The complement alone is not enough — a strand written in the wrong direction describes a different molecule.

Check it. In (a) the four percentages must add to 100: 21 + 21 + 29 + 29 = 100. In (b), every column should be a valid pair: A–T, T–A, G–C, C–G; the RNA column contains U and no T. In (d), reverse-complement the answer again and you should get back the original strand.
Assuming all four bases are 25%. Chargaff’s rule is A = T and G = C, not A = G. Given one base, you can find its partner directly, and the other two only by subtracting from 100 and halving. And the rule only applies to double-stranded DNA — a single strand, or RNA, need not obey it at all.

📝Practise

Work through these on paper, then reveal the answer. Questions 5 and 6 are AHL.

1. State three differences between the structure of DNA and RNA.
(1) DNA has two strands (a double helix); RNA is usually a single strand. (2) DNA contains the sugar deoxyribose; RNA contains ribose, which has an –OH instead of –H on carbon 2. (3) DNA contains the base thymine; RNA contains uracil instead. (The other three bases, A, G and C, are found in both.)
2. Explain how the structure of DNA allows it to be copied accurately.
The two strands are held together by hydrogen bonds between bases, which are weak enough to break so the strands can separate, while the sugar–phosphate backbones are held by strong covalent bonds, so the base sequence of each strand is not changed. Each strand acts as a template. Bases pair complementarily — A only with T, G only with C — so the sequence of the new strand is determined exactly by the template. Each new molecule therefore has the same base sequence as the original.
3. A DNA sample contains 18% cytosine. Calculate the percentage of adenine.
C pairs with G, so G = 18%. C + G = 36%. The remaining 100 − 36 = 64% is A + T, and A = T, so A = 32%.
4. Explain why the universality of the genetic code is regarded as evidence for common ancestry.
Almost all organisms use the same codons for the same amino acids. The particular assignment of codons to amino acids appears to be largely arbitrary — a different code would work just as well chemically — so it is very unlikely that unrelated lineages would arrive at the same code independently. The simplest explanation is that the code evolved once, in a common ancestor, and has been inherited by all its descendants. Changing it would be strongly selected against, because changing a codon’s meaning would alter many proteins at once, so it has been conserved.
5. AHL In the Hershey–Chase experiment, explain why radioactive sulfur and radioactive phosphorus were used, and what result showed that DNA is the genetic material.
The two isotopes labelled the two components of the virus separately: sulfur is present in proteins (in some amino acids) but not in DNA, while phosphorus is present in DNA (in its phosphate groups) but almost absent from proteins. After infection, blending and centrifuging, the 32P label was found mainly in the pellet of bacteria, while the 35S label remained mainly in the liquid with the empty virus coats. So DNA, not protein, entered the bacteria, and since the infected bacteria went on to make new viruses, DNA must carry the genetic information.
6. AHL Explain why the DNA double helix has a constant diameter regardless of its base sequence.
Adenine and guanine are purines (two rings) and thymine and cytosine are pyrimidines (one ring). Complementary pairing always joins a purine to a pyrimidine (A–T, G–C). A purine–purine pair would be too wide and a pyrimidine–pyrimidine pair too narrow; but every purine–pyrimidine pair is the same length. So the distance between the two sugar–phosphate backbones is the same at every base pair, giving the helix a uniform width and a stable, regular shape whatever the sequence.

🔗Go deeper — other people’s work

These are external resources, not mine. If one stops working, tell me and everything above it on this page still stands.

  • RCSB Protein Data Bank, Molecule of the Month — short illustrated articles on DNA and on the nucleosome, with a 3D viewer for the AHL visualization skill.
  • DNA Learning Center (Cold Spring Harbor) — animations of the double helix and a narrated account of the Hershey–Chase experiment.
  • Khan Academy — Nucleic acids and DNA structure videos, useful for drawing nucleotides and antiparallel strands correctly.