DNA base pairing
Statement
Hydrogen bonding and the double helix.
Why it matters
The chemistry of life depends on molecules that can both store information reliably and copy that information with extremely high fidelity, and DNA base pairing is the specific chemical mechanism that makes both possible at once. A small set of hydrogen-bonding rules between just four nucleobases lets one strand of DNA specify the exact sequence of its partner strand, which is what makes DNA replication, transcription, and the entire flow of genetic information mechanistically possible in the first place — and it sets up michaelis-menten and enzyme-inhibition's later treatment of the polymerases and other enzymes that read and act on this base-pairing chemistry.
Base pairing is also a clean, concrete example of the same principle protein-folding-thermodynamics addresses in a more general form: that biological specificity and structure very often arise not from covalent bonds at all, but from the cumulative effect of many individually weak, individually reversible non-covalent interactions acting together.
Hypotheses
Proof
Result
Reading. Watson–Crick base pairing restricts each of the four standard nucleobases to exactly one specific, geometrically and electronically complementary partner, and this strict complementarity is what allows one DNA strand's sequence to fully determine its partner strand's sequence.
Scope. Applies to standard Watson–Crick base pairing in double-stranded B-form DNA; alternative, non-standard pairing geometries (wobble pairing in RNA, Hoogsteen pairing in certain triplex or quadruplex structures) exist under specific structural circumstances and are treated as extensions beyond this core result.
Corollaries & converses
- Chargaff's empirical observation that the amount of adenine in any DNA sample equals the amount of thymine, and the amount of guanine equals the amount of cytosine, follows directly and necessarily from Step 3's strict one-to-one pairing rule, rather than being an independent chemical fact.
- protein-folding-thermodynamics addresses the more general principle, illustrated concretely here, that a large number of individually weak, individually reversible non-covalent interactions (hydrogen bonds in this case) can together produce a highly specific and reasonably stable overall structure.
- Converse: measuring a DNA sample's melting temperature \(T_m\) experimentally provides an indirect but reliable estimate of its overall G···C content, via the direct proportionality of Step 4.
Fails without
- Allow purine–purine pairing (violating the purine/pyrimidine complementarity hypothesis): two fused-ring purines together are too wide to fit within the uniform backbone geometry the double helix requires (Step 2), distorting or breaking the regular helical structure at that position.
- Treat the two DNA strands as running parallel rather than antiparallel (Hypotheses): this mismatches both the helix's true geometry and the direction in which DNA polymerase actually reads a template strand during replication, making the resulting model mechanistically inconsistent with how replication is observed to proceed.
Common errors
- Believing DNA's two strands are held together by covalent bonds; the base-pairing interactions linking the strands are hydrogen bonds (Hypotheses), while covalent bonds exist only within each individual strand's sugar-phosphate backbone.
- Assuming purines can pair with purines, or pyrimidines with pyrimidines, rather than recognising the strict purine–pyrimidine requirement of Step 2 that maintains the helix's uniform width.
- Forgetting that G···C pairs are held together more strongly than A···T pairs (three hydrogen bonds versus two, Step 1), and consequently mispredicting the relative thermal stability of DNA regions with different base composition.
- Treating the two DNA strands as running in the same (parallel) direction rather than antiparallel (Hypotheses), a detail essential to both the helix's actual geometry and to how DNA polymerase reads a template strand during replication.
Discussion
James Watson and Francis Crick published the double-helix structure of DNA in 1953, proposing base pairing as the specific structural feature that immediately suggested, in their own words, a possible copying mechanism for the genetic material. Their model drew directly on Rosalind Franklin and Maurice Wilkins's X-ray diffraction data, which constrained the helix's overall geometry, and on Erwin Chargaff's earlier compositional observation that adenine and thymine, and separately guanine and cytosine, occur in equal amounts within any given DNA sample — an empirical regularity that Watson and Crick's specific base-pairing scheme (Step 1) explained mechanistically for the first time.
Beyond the two standard Watson–Crick pairs, other hydrogen-bonding arrangements between the same four bases are chemically possible and biologically important in specific structural contexts: Hoogsteen base pairing, using a different face of the purine ring, stabilises DNA triplex structures and certain protein–DNA recognition motifs, while wobble base pairing (permitting some non-standard pairings at the third codon position) is essential to how a comparatively small set of transfer RNA molecules can still recognise the full genetic code's redundancy.
Common misconception: that hydrogen bonds, being individually much weaker than covalent bonds, therefore make DNA base pairing a "weak" or unreliable interaction overall. In fact it is precisely the combination of many hydrogen bonds acting cooperatively along an extended sequence (Corollaries) that gives a full DNA duplex very high overall stability and sequence specificity, while still allowing the individual, localised strand separation that replication and transcription both require.
Worked examples
Reading. Base-pairing complementarity, applied position by position with correct attention to strand direction, converts any given DNA sequence directly into its unique partner sequence.
Scope. The identical procedure applies to a sequence of any length, and is the direct chemical basis for how DNA polymerase synthesises a new complementary strand during replication.
Problems
- A double-stranded DNA sample is found to be \(42\%\) adenine by base composition. State the percentage of each of the other three bases, using Chargaff's rule (Corollaries).
Solution
Since A pairs exclusively with T (Step 3), thymine must also be \(42\%\). The remaining \(100-42-42=16\%\) is split equally between guanine and cytosine (which likewise pair exclusively with each other), giving \(8\%\) guanine and \(8\%\) cytosine. - Two DNA duplexes of equal length have measurably different melting temperatures, with duplex X melting at a higher temperature than duplex Y. Predict which duplex has the higher G···C content, and justify your answer using Step 4.
Solution
Duplex X has the higher G···C content. Since G···C pairs are held together by three hydrogen bonds versus only two for A···T pairs (Step 1), a duplex with more G···C pairs requires more thermal energy to fully denature, and therefore has a higher melting temperature; the directly proportional relationship of Step 4 lets the higher-melting duplex X be identified as the one richer in G···C content without needing to sequence either sample directly. - Explain, using the Hypotheses, why an adenine base cannot form a stable, correctly geometrically aligned pair with a guanine base, even though both are purines capable of hydrogen bonding.
Solution
Two structural problems arise together. First, adenine and guanine's specific hydrogen-bond donor/acceptor patterns (Hypotheses) are not mutually complementary in the way A's pattern matches T's, or G's matches C's, so any hydrogen bonds that did form would be poorly aligned and weak. Second, and independently, both adenine and guanine are purines (two fused rings), and a purine–purine pair is too wide to fit within the uniform backbone geometry the double helix requires (Step 2) — a geometric mismatch that would distort the regular helical structure even if the hydrogen-bonding pattern were otherwise compatible.