biology2u
Tier
⌕ Search ⌘K
Concept

DNA sequencing

T-080Home BU-303Threads information
Statement

Reading the order of bases in a genome.

Why it matters

polymerase-chain-reaction established how to amplify a specific DNA region into a usable quantity; DNA sequencing is what reads out its exact base order once amplified, converting a physical DNA sample into the actual information content operon-model, eukaryotic-gene-regulation, and epigenetic-inheritance all discuss in the abstract. It is also the direct enabling technology behind crispr-cas9's practical use, since confirming that an intended edit actually occurred, and only the intended edit, depends on sequencing the edited locus afterward.

Hypotheses
DNA polymerase can be made to terminate synthesis at a specific, base-identifiable point, rather than only ever synthesising a complete, full-length strand.Sanger sequencing's entire method depends on this controlled termination: without some means of stopping synthesis at a known, base-specific position, the reaction would only ever produce full-length product, carrying no positional information about individual base identities along the way. Fragments differing in length by as little as a single nucleotide can be physically resolved and ordered from shortest to longest.Reading the sequence out of a set of terminated fragments requires being able to tell, unambiguously, which fragment is one base longer than which other; gel or capillary electrophoresis, separating fragments by size, provides this resolution for Sanger sequencing, while modern methods achieve the equivalent resolution electronically or optically, base by base, in real time. Because a single sequencing read is typically limited in length (hundreds to a few thousand bases, depending on the platform), sequencing a genome or a large fragment requires breaking it into many overlapping smaller pieces first and computationally reassembling the reads afterward, rather than reading the whole molecule in one continuous pass.
Proof
1
\text{Sanger sequencing: a DNA synthesis reaction includes a small proportion of chain-terminating dideoxynucleotides (ddNTPs), each labelled by base identity, alongside normal nucleotides.}
A ddNTP lacks the 3' hydroxyl group DNA polymerase needs to add the next nucleotide, so whenever one is incorporated by chance instead of the normal nucleotide, synthesis of that particular strand copy stops permanently at that exact position, while other copies of the template continue further before their own random termination. A
2
\text{Because incorporation of a ddNTP at any given position is random across many copies of the template, the reaction produces a full range of fragment lengths, each terminating at a different, base-identified position.}
Given enough template copies, essentially every possible termination position along the sequenced region is represented at least once in the resulting mixture of fragments, collectively spanning the entire length being sequenced. A
3
\text{Separating these fragments by length (Hypotheses) and reading off which base-specific label terminates each length, from shortest to longest, reconstructs the sequence directly.}
Because each fragment's terminal base is known from its label, and the fragments are ordered strictly by length (each one nucleotide longer than the last), reading the terminal-base labels in order of increasing fragment length reconstructs the template's complementary sequence one base at a time. A
4
\text{Next-generation sequencing platforms parallelise this same base-by-base readout across millions of DNA fragments simultaneously, rather than one fragment at a time.}
Rather than a single sequencing reaction generating one long read via electrophoretic separation (Steps 1–3), massively parallel platforms fragment the DNA sample first and sequence very large numbers of short fragments simultaneously, detecting base incorporation optically or electronically in real time; the resulting short reads are computationally reassembled afterward using their overlaps (Hypotheses, t3). B
Result
\text{Base-specific, length-resolved chain termination (or its massively parallel equivalent) converts DNA structure into a readable base sequence.}

Reading. Every DNA sequencing method, however different in throughput and chemistry, works by the same underlying logic: generating a signal that identifies a specific base at a specific, resolvable position, then reading many such signals out in positional order.

Scope. Sanger sequencing remains the standard for short, high-accuracy single-fragment reads (e.g. confirming a cloning construct or a CRISPR edit); next-generation platforms (Step 4) are the standard for large-scale, whole-genome-level sequencing, trading some per-base accuracy and read length for enormously higher throughput.

Corollaries & converses
  • crispr-cas9's outcome (whether an intended edit, an unintended indel, or no change occurred at the target locus) is routinely confirmed by directly sequencing the edited region, applying this result's readout capability to verify that result's edit.
  • Reference genome assembly, the reconstruction of a full genome sequence from many overlapping short reads (Step 4, t3), depends on sequence redundancy (each genomic position being covered by multiple independent reads) to resolve ambiguity and correct occasional individual read errors.
  • Converse: if a genomic region's exact base sequence is known with confidence, any DNA sample can, in principle, be tested for whether it matches that reference sequence at a given position, the logical basis of sequence-based diagnostic and forensic applications.
Fails without
  • Drop controlled, base-identifiable chain termination (Hypotheses): without a means of stopping synthesis at a known, base-specific position, a synthesis reaction produces only full-length double-stranded product, carrying no information about individual base positions along the way; Sanger sequencing's readout mechanism has no signal to read at all.
  • Drop single-nucleotide length resolution (Hypotheses): if fragments differing by one base could not be reliably distinguished and ordered, adjacent positions in the sequence would become confused or skipped, directly limiting the achievable read accuracy and length; this resolution requirement is exactly why fragment separation (electrophoresis, historically) had to be developed to very high precision for sequencing to be practical at all.
Common errors
  • Confusing PCR (polymerase-chain-reaction), which amplifies a DNA region into many copies without reading its sequence, with sequencing itself, which determines the actual base order; the two are frequently used together (amplify first, then sequence) but accomplish entirely different things.
  • Assuming a single sequencing read can cover an entire chromosome or genome directly; individual reads are length-limited (Hypotheses, t3), and large-scale sequencing always requires fragmenting and computationally reassembling many overlapping shorter reads.
  • Assuming ddNTPs (Step 1) behave identically to normal nucleotides except for their fluorescent or radioactive label; their defining functional property is the missing 3' hydroxyl group, which is what actually terminates synthesis, with the label serving only to identify which base was incorporated at that position.
  • Treating all sequencing platforms as interchangeable in accuracy and read length; Sanger sequencing (Steps 1–3) and next-generation methods (Step 4) have distinct trade-offs, and the appropriate choice depends on the specific application (Result, Scope).
Discussion

Frederick Sanger developed the chain-termination sequencing method in 1977, work that, together with his earlier development of protein sequencing methods, earned him a share of two separate Nobel Prizes in Chemistry. Sanger sequencing remained the dominant sequencing technology for roughly three decades, including underpinning the Human Genome Project's initial reference sequence, before next-generation, massively parallel platforms (Step 4) began displacing it for large-scale applications from the mid-2000s onward, owing to their vastly higher throughput and correspondingly lower per-base cost.

More recent long-read sequencing technologies extend read length substantially beyond both classical Sanger and early next-generation short-read methods, directly by reading a single, much longer DNA (or, in some methods, RNA) molecule continuously rather than via chain termination or short-fragment sequencing at all; longer reads simplify genome assembly (Step 4, t3) by spanning repetitive regions that short reads alone cannot resolve unambiguously.

Common misconception: that "sequencing a genome" means reading every base in one continuous pass from one end to the other. In practice, essentially all sequencing (Sanger and next-generation methods alike) works on comparatively short DNA fragments, with the full genome sequence reconstructed afterward by computationally identifying and merging the overlaps between many separate reads (Hypotheses, t3), not by any single, unbroken read spanning an entire chromosome.

Worked examples
1
\text{Electrophoresis of a Sanger sequencing reaction resolves fragments (shortest to longest) with terminal bases: G, A, T, C, C, G, A, T.}
Reading the terminal-base labels strictly in order of increasing fragment length, per Step 3, directly reconstructs the sequence of the strand synthesised (complementary to the original template strand). A
5'\text{-GATCCGAT-}3'

Reading. Ordering the terminal-base identifications by fragment length, from shortest (nearest the primer) to longest, reads the sequence directly in the \(5'\) to \(3'\) direction of the newly synthesised strand.

Scope. The same length-ordered readout procedure applies regardless of how many bases are being read, up to the practical read-length limit of the specific sequencing chemistry used.

Problems
  1. A researcher runs a Sanger sequencing reaction but forgets to include any ddNTPs, using only normal dNTPs. Predict the outcome and explain using Step 1.
    SolutionWithout any chain-terminating ddNTPs, DNA polymerase has nothing to force premature termination at any position; every reaction proceeds to produce only full-length, complete product (Step 1). No range of differently terminated fragment lengths is generated, so there is no positional, base-specific signal to read, and the sequencing reaction yields no sequence information.
  2. A genome is sequenced using short reads of 150 bases each, with each genomic position covered by an average of 30 independent reads (30x coverage). Explain why this redundancy is useful, referencing the Discussion's account of assembly and error correction.
    SolutionBecause reads are generated independently, an occasional sequencing error in any one read at a given position is very unlikely to be repeated identically across most of the other reads covering that same position; comparing the 30 independent reads spanning each position allows the consensus (most common) base call to be trusted with much higher confidence than any single read alone, and also helps correctly resolve the overlaps needed to assemble the full sequence.
  3. After using CRISPR-Cas9 to attempt a targeted gene knockout, a researcher needs to confirm whether the intended edit actually occurred at the target locus, and whether any unintended additional changes are present. Explain which technique from this result is appropriate, and why PCR alone (without sequencing) would be insufficient.
    SolutionSequencing the target locus directly (Steps 1–4) is the appropriate technique, since only base-level sequence readout can confirm the exact nature of the edit, including small insertions or deletions from NHEJ repair, distinguishing an intended precise change from an unintended one. PCR alone only amplifies the region and, at best via simple size comparison, might reveal a large insertion or deletion by a shift in overall fragment size; it cannot resolve small (single- or few-base) changes or confirm the exact resulting sequence, which requires sequencing the amplified product afterward.