Translation and the genetic code
Statement
Reading codons to build a polypeptide.
Why it matters
transcription produces a messenger RNA transcript, but an RNA sequence is not itself a protein; translation is the process that actually reads that sequence and builds a polypeptide, and the genetic code is the fixed dictionary that specifies exactly which amino acid each three-nucleotide codon calls for. Together they complete the arrow from RNA to protein that central-dogma states only in outline, and they are the step at which mutation-types occurring in the coding sequence are finally converted into an altered, or unaltered, protein product.
The genetic code's near-universality across all known life is also one of the strongest single pieces of evidence cited under evidence-common-descent: an arbitrary, 64-codon assignment shared, with only minor exceptions, across every domain of life is far more consistent with common ancestry than with independent origin.
Hypotheses
Proof
Result
Reading. Every possible sequence of mRNA, read three bases at a time from a fixed starting point, is deterministically translated into a specific sequence of amino acids (or terminated) by the fixed, near-universal genetic-code dictionary, physically implemented through codon–anticodon base-pairing at the ribosome.
Scope. The standard code applies with only minor, well-catalogued exceptions (certain codon reassignments in mitochondria and a small number of other organisms); the core triplet, non-overlapping, fixed-reading-frame logic of Steps 1–3 is universal across all known translation systems.
Corollaries & converses
- mutation-types acting within a coding sequence produce qualitatively different outcomes depending on exactly how they interact with this Result: a substitution within a codon can be silent (degenerate code, Hypotheses), missense (changes the amino acid), or nonsense (creates a premature stop codon, Step 3), while an insertion or deletion not divisible by three triggers the far more disruptive frameshift of Step 5.
- central-dogma's RNA-to-protein arrow is precisely this process; transcription supplies the input (mature mRNA) that translation, governed by the fixed genetic-code dictionary, converts into the final protein product.
- Converse: given a known protein's amino acid sequence, the set of possible mRNA (and hence DNA) sequences that could have encoded it can be inferred directly from the genetic code's degeneracy (Hypotheses, Tier 3) — though because most amino acids have multiple synonymous codons, this inference is generally many-to-one rather than fully determining a single unique sequence.
Fails without
- Drop the fixed, non-overlapping reading frame (Hypotheses): if codons could instead overlap or be read starting from any arbitrary position, the same mRNA sequence would correspond to many different possible amino acid sequences simultaneously, with no way for the ribosome (or the cell) to reliably determine which one was intended — protein synthesis would be effectively undefined rather than deterministic.
- Drop accurate tRNA charging (Hypotheses' adaptor mechanism): if aminoacyl-tRNA synthetases attached the wrong amino acid to a given tRNA, that tRNA's anticodon would still correctly pair with its matching mRNA codon (Step 2), but the ribosome would incorporate the wrong amino acid regardless — the fidelity of the entire genetic code depends on this charging step being accurate, not merely on correct codon–anticodon base-pairing at the ribosome.
Common errors
- Assuming DNA is translated directly; translation acts on the mature mRNA transcript produced by transcription, not on DNA itself, which never leaves the nucleus in eukaryotic cells and is not physically read by the ribosome at all.
- Treating all base substitutions within a coding sequence as equally consequential; Step 3's stop-codon mechanism and the code's degeneracy (Hypotheses) mean substitution effects range from entirely silent to catastrophic (premature stop), depending on exactly which codon position changes and to what.
- Confusing an insertion/deletion mutation's severity with its physical size; a single-base insertion (a frameshift, Step 5) is typically far more disruptive to the downstream protein than a three-base insertion or deletion (which adds or removes one whole codon but preserves the reading frame for everything after it).
- Assuming the genetic code is arbitrary and therefore uninformative about evolutionary history; its near-universal sharing across all domains of life (Step 4) is in fact one of the strongest arguments for common ancestry precisely because the specific codon assignments have no obvious chemical necessity forcing them to be identical across independently originating systems.
Discussion
The genetic code was deciphered experimentally through the 1960s, chiefly by Marshall Nirenberg and Har Gobind Khorana, whose synthetic-RNA experiments systematically determined which amino acid each codon specifies — work recognised with a Nobel Prize. Francis Crick had earlier reasoned, on largely theoretical grounds, that the code must be a non-overlapping triplet code, a prediction the subsequent decipherment confirmed directly.
The genetic code is not perfectly universal: mitochondrial genomes in many organisms use a handful of reassigned codons that differ from the standard nuclear code (for instance, some mitochondrial codes read UGA as an amino acid rather than as a stop signal), and a very small number of other organisms show additional, independently documented exceptions. These deviations are themselves informative, generally consistent with a shared ancestral standard code that subsequently diverged slightly in specific, well-characterised lineages, rather than evidence against the code's overall near-universality.
Common misconception: that each amino acid corresponds to exactly one codon. The code's degeneracy (Hypotheses, Tier 3) means most amino acids are specified by several synonymous codons (leucine and serine, for instance, are each specified by six), and this many-to-one mapping is precisely what allows some point mutations to be phenotypically silent.
Worked examples
Reading. The identical starting sequence produces either a short, correctly terminated peptide or a frameshifted, almost certainly non-functional downstream sequence, depending entirely on whether the total number of bases inserted or deleted is a multiple of three.
Scope. This frame-sensitivity applies to any coding sequence of any length; it is the direct mechanistic reason why insertion/deletion mutations are, on average, far more disruptive to protein function than single-base substitutions.
Problems
- Translate the mRNA sequence \(5'\text{-AUG UUU GGA UGA-}3'\) into its amino acid sequence, using the standard genetic code (AUG = Met/start, UUU = Phe, GGA = Gly, UGA = stop).
Solution
Reading in fixed triplets from the start codon: AUG (Met), UUU (Phe), GGA (Gly), UGA (stop, terminates translation). Resulting polypeptide: Met-Phe-Gly. - A point mutation changes the third position of a codon from GCA to GCG. Both GCA and GCG specify alanine. Predict the effect on the resulting protein, and name this type of mutation.
Solution
Because GCA and GCG are synonymous codons both specifying alanine (the genetic code's degeneracy, Hypotheses), the amino acid incorporated at this position is unchanged; the resulting protein sequence is identical to the unmutated version. This is a silent (synonymous) point mutation. - Explain why a three-base (single-codon) deletion within a coding sequence is generally far less disruptive to the resulting protein than a one-base deletion at the same position, using Step 5.
Solution
A three-base deletion removes exactly one whole codon; because \(3\) is a multiple of \(3\), the reading frame for every codon after the deletion point is unaffected (Step 5's condition for a frameshift is not met), so the result is simply a protein missing one amino acid, otherwise unchanged. A one-base deletion, by contrast, is not a multiple of three, so it shifts the reading frame for every subsequent codon, typically producing a completely different, non-functional amino acid sequence from that point onward — a much more severe disruption despite removing far less genetic material.