biology2u
Tier
⌕ Search ⌘K
Concept

Mutations

T-035Home BU-201Threads information
Statement

How changes in DNA sequence alter proteins.

Why it matters

central-dogma establishes that genetic information flows from DNA through RNA to protein, and translation-genetic-code shows exactly how a sequence of codons is read into a sequence of amino acids; mutations are what happens when the DNA sequence feeding that whole pipeline is itself altered, and understanding the different types of mutation is what lets you predict how, and how severely, a given DNA change propagates through transcription and translation to affect the resulting protein. semiconservative-replication is also directly relevant here, since DNA replication errors are a major source of the mutations this result classifies, alongside external mutagens.

The classification matters because not all DNA sequence changes are equal in consequence: some are functionally silent, others are catastrophic, and the difference is determined entirely by exactly what kind of change occurred and where in the gene it occurred — the subject of this result.

Hypotheses
The genetic code is redundant (degenerate): most amino acids are specified by more than one codon, and codons differing only in the third ("wobble") position frequently specify the same amino acid.Without this redundancy, every single-nucleotide change in a coding sequence would necessarily alter the encoded amino acid; redundancy is exactly why some point mutations have no effect on the protein at all (Step 1), a possibility that would not exist under a strictly one-to-one code. The reading frame of a coding sequence is fixed by its start codon and read in consecutive, non-overlapping triplets.This is what makes insertions or deletions qualitatively different in consequence from substitutions: because the ribosome reads triplets from a fixed starting point, adding or removing nucleotides in a number not divisible by three shifts every subsequent triplet's grouping, an effect substitutions (which change content but not the grouping) never produce. Not every mutation occurs in a protein-coding region; mutations in regulatory sequences (promoters, splice sites, untranslated regions) can alter gene expression level or splicing pattern without changing the encoded protein sequence at all, a category the classification below (organised around effects on coding sequence) does not directly address.
Proof
1
\text{Point substitution: silent (synonymous)}, \text{missense (non-synonymous)}, \text{or nonsense (premature stop)}
A single nucleotide substitution changes one codon; by the redundancy of Step Hypotheses, this may leave the encoded amino acid unchanged (silent), change it to a different amino acid (missense), or convert it to a stop codon, truncating the protein (nonsense) — the same type of underlying DNA change (one base for another) can therefore have three qualitatively different consequences depending purely on which codon and which position within it is affected. A
2
\text{Insertion or deletion of } n \text{ nucleotides: frameshift if } n \not\equiv 0 \pmod 3\text{, in-frame if } n \equiv 0 \pmod 3
Because translation reads fixed, non-overlapping triplets (Hypotheses), inserting or deleting a number of bases not divisible by three shifts every downstream codon's grouping, typically scrambling the entire downstream amino acid sequence and often introducing a premature stop soon after; insertions or deletions in multiples of three instead simply add or remove whole codons, leaving the reading frame downstream intact. B
3
\text{Chromosomal-scale mutations: deletion, duplication, inversion, translocation of large DNA segments.}
Beyond single- or few-nucleotide changes, mutation can also occur at the scale of entire chromosome segments — losing a segment (deletion), gaining a duplicate copy (duplication), reversing a segment's orientation (inversion), or moving a segment to a different chromosome (translocation) — each capable of disrupting multiple genes simultaneously or altering gene dosage without changing any individual gene's own sequence. A
4
\text{Somatic mutations affect only the individual bearing them; germline mutations are heritable.}
A mutation's consequences depend critically on which cell lineage it occurs in: a mutation in a somatic (non-reproductive) cell affects only that individual and its descendant cells (relevant to cancer, oncogenes-tumour-suppressors), while a mutation in a germline cell can be transmitted to offspring and becomes subject to natural-selection acting across generations, not just within one organism's lifetime. A
5
\text{Mutagens (chemical, radiation) raise the baseline mutation rate; proofreading and repair systems reduce it.}
Baseline mutation rate reflects a balance between processes that damage DNA or introduce replication errors (semiconservative-replication is not error-free) and the cell's proofreading and repair machinery that corrects most such errors before they become permanent; mutagens shift this balance by increasing the rate of DNA damage beyond what repair can fully correct. A
Result
\text{Point: silent / missense / nonsense}\quad\big|\quad\text{Indel: in-frame / frameshift}\quad\big|\quad\text{Chromosomal: deletion / duplication / inversion / translocation}

Reading. A mutation's functional consequence is determined jointly by its scale (single nucleotide versus segment-sized), its effect on reading frame, and its location relative to coding and regulatory sequence — not by mutation "severity" as some single, undifferentiated property.

Scope. This classification covers DNA-sequence-level effects; whether a given mutation is ultimately harmful, neutral, or beneficial is a separate question addressed by natural-selection and neutral-theory, not determined by mutation type alone.

Corollaries & converses
  • neutral-theory relies directly on this classification: silent substitutions (Step 1) are the paradigm case of a mutation with no effect on protein sequence, and hence a strong candidate for evolving at the neutral, mutation-rate-set pace that result describes.
  • Frameshift mutations (Step 2) are, on average, far more disruptive than in-frame indels or missense substitutions, because they typically corrupt every downstream codon rather than altering or removing only a localised stretch of the protein — a direct, general consequence of the fixed-reading-frame logic in Hypotheses.
  • Converse: observing that a protein's function is completely lost is not, by itself, enough to identify which type of mutation caused it; nonsense, frameshift, and even certain missense mutations at critical residues can all independently produce a complete loss-of-function phenotype, so sequencing is required to distinguish them.
Fails without
  • Drop codon redundancy (Hypotheses): if the genetic code were strictly one-to-one (no synonymous codons at all), every single-nucleotide substitution in a coding region would necessarily change the encoded amino acid; the silent-mutation category of Step 1 would not exist, and neutral-theory's reliance on synonymous substitutions as a baseline neutral rate would have no natural source of such mutations to draw on.
  • Ignore the fixed reading frame (Hypotheses): without a fixed triplet grouping read from a defined start point, there would be no meaningful distinction between in-frame and frameshift indels (Step 2) — the entire, sharp difference in typical severity between an indel of 3 nucleotides and one of 4 nucleotides, otherwise nearly identical in physical scale, would disappear.
Common errors
  • Assuming all point mutations change the encoded protein; silent (synonymous) substitutions (Step 1) do not, by definition, given codon redundancy.
  • Assuming any insertion or deletion causes a frameshift; only those not divisible by three do (Step 2) — a common source of confusion is treating "indel" and "frameshift" as synonyms rather than as a general category and one specific consequence within it.
  • Confusing mutation (a change to DNA sequence itself) with epigenetic change (a heritable change in gene expression without altered DNA sequence, covered separately) — the two are mechanistically distinct even though both can alter phenotype.
  • Assuming somatic mutations are heritable across generations; only germline mutations (Step 4) are transmitted to offspring, regardless of how severe a somatic mutation's effect on the individual bearing it may be.
Discussion

Hermann Muller's 1927 demonstration that X-ray irradiation dramatically increased mutation rate in Drosophila was the first direct experimental evidence that mutation could be artificially induced rather than only arising spontaneously, work that also first drew broad public and scientific attention to the mutagenic hazards of ionising radiation.

Sickle-cell disease is a well-known single, precisely characterised missense mutation (a single amino acid substitution in haemoglobin, glutamic acid to valine), illustrating that "missense" does not automatically mean "mild" — a change at a single, critical residue can be as consequential as a much larger-scale disruption, depending entirely on that residue's structural and functional role, not on the physical size of the underlying DNA change.

Common misconception: that larger-scale mutations (chromosomal, Step 3) are always more harmful than small-scale ones (point mutations, Step 1). Severity depends on which specific gene or regulatory region is affected and how, not on physical scale alone: a single nonsense mutation early in an essential gene can be more damaging than a large duplication in a region with no critical function.

Worked examples
1
\text{Wild-type codon UUU (Phe)} \to \text{UUC (Phe): silent}; \quad \to \text{UCU (Ser): missense}; \quad \to \text{UAA (stop): nonsense}
Three different single-nucleotide substitutions at different positions within the identical starting codon (UUU) illustrate all three point-mutation outcomes of Step 1 side by side: a third-position change preserving the amino acid (silent, exploiting codon redundancy), a second-position change altering it (missense), and a change generating a premature stop (nonsense, truncating the protein). A
\text{Same starting codon, three distinct substitution outcomes, purely from which base and which position changed}

Reading. Mutation consequence cannot be predicted from "how many bases changed" alone; the specific position and identity of the change, read against the redundancy of the genetic code, determines the outcome.

Scope. The same logic (Step 1) applies to any codon; degeneracy is concentrated at the third position, which is why third-position changes are disproportionately likely to be silent.

Problems
  1. A coding sequence has 2 nucleotides deleted partway through the gene. Using Step 2, predict whether this is an in-frame or frameshift mutation, and describe the expected effect on the downstream protein sequence.
    Solution2 is not divisible by 3 (\(2\not\equiv0\pmod3\)), so this is a frameshift mutation (Step 2). Every codon downstream of the deletion is shifted by two positions, changing essentially all subsequent triplet groupings and typically producing a scrambled amino acid sequence downstream, often terminating early at an out-of-frame stop codon.
  2. A single-nucleotide substitution changes a codon from CGA (Arg) to CGG (still Arg). Classify this mutation using Step 1, and explain which Hypothesis makes this outcome possible.
    SolutionSince the encoded amino acid (arginine) is unchanged, this is a silent (synonymous) substitution (Step 1). It is possible specifically because the genetic code is redundant (Hypotheses): multiple codons, differing here only at the third, "wobble" position, specify the same amino acid.
  3. Compare the likely severity of a 3-nucleotide (single-codon) in-frame deletion versus a 4-nucleotide deletion at the same position in a gene, using Step 2, and explain why the size difference of only one nucleotide produces such different expected outcomes.
    SolutionThe 3-nucleotide deletion removes exactly one codon (\(3\equiv0\pmod3\)), leaving the reading frame downstream intact and simply removing one amino acid from the protein — often tolerated, especially outside a critical functional region. The 4-nucleotide deletion (\(4\not\equiv0\pmod3\)) shifts the reading frame for every codon downstream (Step 2), typically corrupting the entire remainder of the protein sequence; a difference of a single nucleotide in deletion size produces a qualitatively different outcome purely because of the fixed-triplet reading frame (Hypotheses), not because of any difference in the physical scale of DNA lost.