DNA sequencing
Statement
Reading the order of bases in a genome.
Why it matters
polymerase-chain-reaction established how to amplify a specific DNA region into a usable quantity; DNA sequencing is what reads out its exact base order once amplified, converting a physical DNA sample into the actual information content operon-model, eukaryotic-gene-regulation, and epigenetic-inheritance all discuss in the abstract. It is also the direct enabling technology behind crispr-cas9's practical use, since confirming that an intended edit actually occurred, and only the intended edit, depends on sequencing the edited locus afterward.
Hypotheses
Proof
Result
Reading. Every DNA sequencing method, however different in throughput and chemistry, works by the same underlying logic: generating a signal that identifies a specific base at a specific, resolvable position, then reading many such signals out in positional order.
Scope. Sanger sequencing remains the standard for short, high-accuracy single-fragment reads (e.g. confirming a cloning construct or a CRISPR edit); next-generation platforms (Step 4) are the standard for large-scale, whole-genome-level sequencing, trading some per-base accuracy and read length for enormously higher throughput.
Corollaries & converses
- crispr-cas9's outcome (whether an intended edit, an unintended indel, or no change occurred at the target locus) is routinely confirmed by directly sequencing the edited region, applying this result's readout capability to verify that result's edit.
- Reference genome assembly, the reconstruction of a full genome sequence from many overlapping short reads (Step 4, t3), depends on sequence redundancy (each genomic position being covered by multiple independent reads) to resolve ambiguity and correct occasional individual read errors.
- Converse: if a genomic region's exact base sequence is known with confidence, any DNA sample can, in principle, be tested for whether it matches that reference sequence at a given position, the logical basis of sequence-based diagnostic and forensic applications.
Fails without
- Drop controlled, base-identifiable chain termination (Hypotheses): without a means of stopping synthesis at a known, base-specific position, a synthesis reaction produces only full-length double-stranded product, carrying no information about individual base positions along the way; Sanger sequencing's readout mechanism has no signal to read at all.
- Drop single-nucleotide length resolution (Hypotheses): if fragments differing by one base could not be reliably distinguished and ordered, adjacent positions in the sequence would become confused or skipped, directly limiting the achievable read accuracy and length; this resolution requirement is exactly why fragment separation (electrophoresis, historically) had to be developed to very high precision for sequencing to be practical at all.
Common errors
- Confusing PCR (polymerase-chain-reaction), which amplifies a DNA region into many copies without reading its sequence, with sequencing itself, which determines the actual base order; the two are frequently used together (amplify first, then sequence) but accomplish entirely different things.
- Assuming a single sequencing read can cover an entire chromosome or genome directly; individual reads are length-limited (Hypotheses, t3), and large-scale sequencing always requires fragmenting and computationally reassembling many overlapping shorter reads.
- Assuming ddNTPs (Step 1) behave identically to normal nucleotides except for their fluorescent or radioactive label; their defining functional property is the missing 3' hydroxyl group, which is what actually terminates synthesis, with the label serving only to identify which base was incorporated at that position.
- Treating all sequencing platforms as interchangeable in accuracy and read length; Sanger sequencing (Steps 1–3) and next-generation methods (Step 4) have distinct trade-offs, and the appropriate choice depends on the specific application (Result, Scope).
Discussion
Frederick Sanger developed the chain-termination sequencing method in 1977, work that, together with his earlier development of protein sequencing methods, earned him a share of two separate Nobel Prizes in Chemistry. Sanger sequencing remained the dominant sequencing technology for roughly three decades, including underpinning the Human Genome Project's initial reference sequence, before next-generation, massively parallel platforms (Step 4) began displacing it for large-scale applications from the mid-2000s onward, owing to their vastly higher throughput and correspondingly lower per-base cost.
More recent long-read sequencing technologies extend read length substantially beyond both classical Sanger and early next-generation short-read methods, directly by reading a single, much longer DNA (or, in some methods, RNA) molecule continuously rather than via chain termination or short-fragment sequencing at all; longer reads simplify genome assembly (Step 4, t3) by spanning repetitive regions that short reads alone cannot resolve unambiguously.
Common misconception: that "sequencing a genome" means reading every base in one continuous pass from one end to the other. In practice, essentially all sequencing (Sanger and next-generation methods alike) works on comparatively short DNA fragments, with the full genome sequence reconstructed afterward by computationally identifying and merging the overlaps between many separate reads (Hypotheses, t3), not by any single, unbroken read spanning an entire chromosome.
Worked examples
Reading. Ordering the terminal-base identifications by fragment length, from shortest (nearest the primer) to longest, reads the sequence directly in the \(5'\) to \(3'\) direction of the newly synthesised strand.
Scope. The same length-ordered readout procedure applies regardless of how many bases are being read, up to the practical read-length limit of the specific sequencing chemistry used.
Problems
- A researcher runs a Sanger sequencing reaction but forgets to include any ddNTPs, using only normal dNTPs. Predict the outcome and explain using Step 1.
Solution
Without any chain-terminating ddNTPs, DNA polymerase has nothing to force premature termination at any position; every reaction proceeds to produce only full-length, complete product (Step 1). No range of differently terminated fragment lengths is generated, so there is no positional, base-specific signal to read, and the sequencing reaction yields no sequence information. - A genome is sequenced using short reads of 150 bases each, with each genomic position covered by an average of 30 independent reads (30x coverage). Explain why this redundancy is useful, referencing the Discussion's account of assembly and error correction.
Solution
Because reads are generated independently, an occasional sequencing error in any one read at a given position is very unlikely to be repeated identically across most of the other reads covering that same position; comparing the 30 independent reads spanning each position allows the consensus (most common) base call to be trusted with much higher confidence than any single read alone, and also helps correctly resolve the overlaps needed to assemble the full sequence. - After using CRISPR-Cas9 to attempt a targeted gene knockout, a researcher needs to confirm whether the intended edit actually occurred at the target locus, and whether any unintended additional changes are present. Explain which technique from this result is appropriate, and why PCR alone (without sequencing) would be insufficient.
Solution
Sequencing the target locus directly (Steps 1–4) is the appropriate technique, since only base-level sequence readout can confirm the exact nature of the edit, including small insertions or deletions from NHEJ repair, distinguishing an intended precise change from an unintended one. PCR alone only amplifies the region and, at best via simple size comparison, might reveal a large insertion or deletion by a shift in overall fragment size; it cannot resolve small (single- or few-base) changes or confirm the exact resulting sequence, which requires sequencing the amplified product afterward.