biology2u
Tier
⌕ Search ⌘K
Concept

Translation and the genetic code

T-034Home BU-201Threads information
Statement

Reading codons to build a polypeptide.

Why it matters

transcription produces a messenger RNA transcript, but an RNA sequence is not itself a protein; translation is the process that actually reads that sequence and builds a polypeptide, and the genetic code is the fixed dictionary that specifies exactly which amino acid each three-nucleotide codon calls for. Together they complete the arrow from RNA to protein that central-dogma states only in outline, and they are the step at which mutation-types occurring in the coding sequence are finally converted into an altered, or unaltered, protein product.

The genetic code's near-universality across all known life is also one of the strongest single pieces of evidence cited under evidence-common-descent: an arbitrary, 64-codon assignment shared, with only minor exceptions, across every domain of life is far more consistent with common ancestry than with independent origin.

Hypotheses
The genetic code reads mRNA in non-overlapping triplets (codons), starting from a fixed reading frame established at the start codon, with each codon specifying exactly one amino acid or a stop signal.Because there are 4 possible bases and codons are triplets, there are \(4^3=64\) possible codons, comfortably more than the 20 standard amino acids that must be specified; without a fixed reading frame, the same sequence of bases could be read as several entirely different amino acid sequences depending on where reading begins, which is exactly why insertion or deletion mutations that are not a multiple of three bases are so disruptive (mutation-types). Each amino acid is delivered to the ribosome attached to a specific transfer RNA (tRNA) whose anticodon base-pairs with the mRNA codon (the adaptor hypothesis).The ribosome itself has no intrinsic means of distinguishing one amino acid from another directly from the mRNA sequence; tRNA is the physical adaptor that links a codon's nucleotide sequence to the correct amino acid, and the accuracy of this link depends entirely on aminoacyl-tRNA synthetase enzymes correctly charging each tRNA with its matching amino acid in the first place. The genetic code is degenerate (most amino acids are specified by more than one codon) but unambiguous (each codon specifies only one amino acid or stop).Degeneracy means that some point mutations, especially at a codon's third position, can change the codon without changing the amino acid it specifies (a "silent" or synonymous mutation, mutation-types); this partly buffers the coding sequence against the effect of certain mutations, though not against most first- or second-position changes.
Proof
1
\text{Initiation: the small ribosomal subunit, an initiator tRNA carrying methionine, and mRNA assemble at the start codon (AUG), fixing the reading frame.}
Locating the correct AUG (typically aided by surrounding sequence context and the 5′ cap in eukaryotes) sets, once and for all, how every subsequent triplet downstream will be grouped into codons; the large ribosomal subunit then joins to complete the functional ribosome. A
2
\text{Elongation: for each codon in turn, a matching aminoacyl-tRNA (Hypotheses' adaptor) enters the ribosome, its anticodon base-pairs with the codon, and the ribosome catalyses formation of a peptide bond between the new amino acid and the growing chain.}
Correct codon–anticodon pairing (standard Watson–Crick base-pairing, applied here between mRNA and tRNA) is what ensures the amino acid actually added matches the one specified by the genetic code for that codon; the ribosome then translocates exactly one codon along the mRNA, exposing the next codon for the same process to repeat. A
3
\text{Termination: when a stop codon (UAA, UAG, or UGA) enters the ribosome, no matching tRNA exists; a release factor instead binds, releasing the completed polypeptide.}
Because no tRNA carries an anticodon complementary to any stop codon, the ribosome cannot continue elongation past this point; release factors recognise stop codons directly and trigger hydrolysis of the bond linking the completed polypeptide to the final tRNA, freeing the finished chain. A
4
\text{The genetic code's codon-to-amino-acid assignment is fixed and (with rare, well-documented exceptions) identical across essentially all known organisms.}
This was established empirically (Discussion) by systematically testing which amino acid each of the 64 possible codons specifies; the near-universal sharing of this arbitrary assignment across bacteria, archaea, and eukaryotes is strong independent evidence, alongside other homologies (evidence-common-descent), for a single common ancestor from which this coding system was inherited rather than independently re-invented. A
5
\text{A shift in reading frame (insertion or deletion of a number of bases not divisible by three) changes every downstream codon grouping from that point onward.}
Because codons are read as fixed, non-overlapping triplets (Hypotheses), removing or adding bases in a quantity that is not a multiple of three regroups every subsequent triplet differently, almost always producing a completely different, and usually non-functional, amino acid sequence downstream of the mutation site — far more disruptive, on average, than a single-codon substitution. B
Result
\text{mRNA codon (triplet)} \xrightarrow{\text{tRNA anticodon pairing}} \text{amino acid} \ ;\quad 4^3=64\text{ codons} \to 20\text{ amino acids} + \text{stop}

Reading. Every possible sequence of mRNA, read three bases at a time from a fixed starting point, is deterministically translated into a specific sequence of amino acids (or terminated) by the fixed, near-universal genetic-code dictionary, physically implemented through codon–anticodon base-pairing at the ribosome.

Scope. The standard code applies with only minor, well-catalogued exceptions (certain codon reassignments in mitochondria and a small number of other organisms); the core triplet, non-overlapping, fixed-reading-frame logic of Steps 1–3 is universal across all known translation systems.

Corollaries & converses
  • mutation-types acting within a coding sequence produce qualitatively different outcomes depending on exactly how they interact with this Result: a substitution within a codon can be silent (degenerate code, Hypotheses), missense (changes the amino acid), or nonsense (creates a premature stop codon, Step 3), while an insertion or deletion not divisible by three triggers the far more disruptive frameshift of Step 5.
  • central-dogma's RNA-to-protein arrow is precisely this process; transcription supplies the input (mature mRNA) that translation, governed by the fixed genetic-code dictionary, converts into the final protein product.
  • Converse: given a known protein's amino acid sequence, the set of possible mRNA (and hence DNA) sequences that could have encoded it can be inferred directly from the genetic code's degeneracy (Hypotheses, Tier 3) — though because most amino acids have multiple synonymous codons, this inference is generally many-to-one rather than fully determining a single unique sequence.
Fails without
  • Drop the fixed, non-overlapping reading frame (Hypotheses): if codons could instead overlap or be read starting from any arbitrary position, the same mRNA sequence would correspond to many different possible amino acid sequences simultaneously, with no way for the ribosome (or the cell) to reliably determine which one was intended — protein synthesis would be effectively undefined rather than deterministic.
  • Drop accurate tRNA charging (Hypotheses' adaptor mechanism): if aminoacyl-tRNA synthetases attached the wrong amino acid to a given tRNA, that tRNA's anticodon would still correctly pair with its matching mRNA codon (Step 2), but the ribosome would incorporate the wrong amino acid regardless — the fidelity of the entire genetic code depends on this charging step being accurate, not merely on correct codon–anticodon base-pairing at the ribosome.
Common errors
  • Assuming DNA is translated directly; translation acts on the mature mRNA transcript produced by transcription, not on DNA itself, which never leaves the nucleus in eukaryotic cells and is not physically read by the ribosome at all.
  • Treating all base substitutions within a coding sequence as equally consequential; Step 3's stop-codon mechanism and the code's degeneracy (Hypotheses) mean substitution effects range from entirely silent to catastrophic (premature stop), depending on exactly which codon position changes and to what.
  • Confusing an insertion/deletion mutation's severity with its physical size; a single-base insertion (a frameshift, Step 5) is typically far more disruptive to the downstream protein than a three-base insertion or deletion (which adds or removes one whole codon but preserves the reading frame for everything after it).
  • Assuming the genetic code is arbitrary and therefore uninformative about evolutionary history; its near-universal sharing across all domains of life (Step 4) is in fact one of the strongest arguments for common ancestry precisely because the specific codon assignments have no obvious chemical necessity forcing them to be identical across independently originating systems.
Discussion

The genetic code was deciphered experimentally through the 1960s, chiefly by Marshall Nirenberg and Har Gobind Khorana, whose synthetic-RNA experiments systematically determined which amino acid each codon specifies — work recognised with a Nobel Prize. Francis Crick had earlier reasoned, on largely theoretical grounds, that the code must be a non-overlapping triplet code, a prediction the subsequent decipherment confirmed directly.

The genetic code is not perfectly universal: mitochondrial genomes in many organisms use a handful of reassigned codons that differ from the standard nuclear code (for instance, some mitochondrial codes read UGA as an amino acid rather than as a stop signal), and a very small number of other organisms show additional, independently documented exceptions. These deviations are themselves informative, generally consistent with a shared ancestral standard code that subsequently diverged slightly in specific, well-characterised lineages, rather than evidence against the code's overall near-universality.

Common misconception: that each amino acid corresponds to exactly one codon. The code's degeneracy (Hypotheses, Tier 3) means most amino acids are specified by several synonymous codons (leucine and serine, for instance, are each specified by six), and this many-to-one mapping is precisely what allows some point mutations to be phenotypically silent.

Worked examples
1
\text{mRNA: } 5'\text{-AUG GCC UAA-}3'
Reading in fixed, non-overlapping triplets from the start codon (Step 1): AUG specifies methionine (also the start signal); GCC specifies alanine; UAA is a stop codon (Step 3), terminating translation. The resulting dipeptide is Met-Ala. A
2
\text{A single base is deleted from this sequence between the first and second codon: } 5'\text{-AUG GCU AA-}3' \text{ (frameshifted from that point)}
Deleting a single base (not a multiple of three) shifts every subsequent triplet grouping (Step 5): reading resumes as AUG (Met), GCU (still alanine, coincidentally, since GCU is a synonymous codon for alanine here), then AA... with the reading frame now misaligned relative to any further downstream sequence not shown, illustrating how a frameshift's damage compounds with every base after the deletion point, even when an early codon happens to still specify a similar amino acid by chance. B
\text{In-frame sequence: Met-Ala-stop} \qquad \text{After single-base deletion: reading frame shifted for all downstream codons}

Reading. The identical starting sequence produces either a short, correctly terminated peptide or a frameshifted, almost certainly non-functional downstream sequence, depending entirely on whether the total number of bases inserted or deleted is a multiple of three.

Scope. This frame-sensitivity applies to any coding sequence of any length; it is the direct mechanistic reason why insertion/deletion mutations are, on average, far more disruptive to protein function than single-base substitutions.

Problems
  1. Translate the mRNA sequence \(5'\text{-AUG UUU GGA UGA-}3'\) into its amino acid sequence, using the standard genetic code (AUG = Met/start, UUU = Phe, GGA = Gly, UGA = stop).
    SolutionReading in fixed triplets from the start codon: AUG (Met), UUU (Phe), GGA (Gly), UGA (stop, terminates translation). Resulting polypeptide: Met-Phe-Gly.
  2. A point mutation changes the third position of a codon from GCA to GCG. Both GCA and GCG specify alanine. Predict the effect on the resulting protein, and name this type of mutation.
    SolutionBecause GCA and GCG are synonymous codons both specifying alanine (the genetic code's degeneracy, Hypotheses), the amino acid incorporated at this position is unchanged; the resulting protein sequence is identical to the unmutated version. This is a silent (synonymous) point mutation.
  3. Explain why a three-base (single-codon) deletion within a coding sequence is generally far less disruptive to the resulting protein than a one-base deletion at the same position, using Step 5.
    SolutionA three-base deletion removes exactly one whole codon; because \(3\) is a multiple of \(3\), the reading frame for every codon after the deletion point is unaffected (Step 5's condition for a frameshift is not met), so the result is simply a protein missing one amino acid, otherwise unchanged. A one-base deletion, by contrast, is not a multiple of three, so it shifts the reading frame for every subsequent codon, typically producing a completely different, non-functional amino acid sequence from that point onward — a much more severe disruption despite removing far less genetic material.