biology2u
Tier
⌕ Search ⌘K
Concept

The central dogma

T-031Home BU-201Threads information
Statement

Information flows DNA to RNA to protein.

Why it matters

The central dogma is the organising statement for the entire flow of genetic information in the cell, and this unit's remaining results are, in effect, an expansion of its three individual steps: semiconservative-replication describes how DNA is copied to make more DNA, transcription describes the DNA-to-RNA step in detail, and translation-genetic-code describes the RNA-to-protein step in detail. Stating the overall direction of information flow first, before any of the individual mechanisms, gives a map for where each subsequent result fits and why they occur in this particular order.

It also draws a sharp, testable line — information flows from nucleic acid to protein, never the reverse — that turns out to matter well beyond textbook description: mutation-types' account of how DNA sequence changes propagate into altered proteins depends entirely on this directionality, and the line the dogma draws between permitted and forbidden directions of information flow is itself an experimentally supported claim, not simply definitional (Discussion).

Hypotheses
Genetic information is encoded as a linear sequence — of nucleotide bases in DNA and RNA, of amino acids in protein.The dogma's claim about the direction of information "flow" only makes sense for information encoded linearly, where one sequence can be said to specify another by a defined mapping (as the genetic code does between RNA codons and amino acids); without linear encoding there would be no well-defined sense in which one molecule "contains the information for" another. The mapping from nucleic acid sequence to protein sequence (the genetic code) is essentially one-directional in normal cellular chemistry: it can be read from nucleic acid to protein, but there is no cellular machinery that reads a protein's amino acid sequence back into a nucleic acid sequence.This is the specific claim that distinguishes the dogma from a trivial observation about molecular order; it says something about which biochemical processes exist, not merely about which order events happen to occur in. Once assembled, a protein's amino acid sequence is not used as a template to specify further nucleic acid or protein synthesis.This is what actually rules out "reverse translation": there is no known biochemical mechanism that reads an existing protein and produces a corresponding nucleic acid or new protein sequence from it, a stronger and more specific claim than simply noting that DNA happens to be copied before protein is made.
Proof
1
\text{DNA} \xrightarrow{\text{replication}} \text{DNA}
A cell must be able to make more DNA from DNA, since each cell division (cell-cycle-mitosis) requires a complete duplicate genome for the daughter cell; semiconservative-replication describes the specific mechanism, in which each strand of the double helix templates a new complementary strand. A
2
\text{DNA} \xrightarrow{\text{transcription}} \text{RNA}
A gene's DNA sequence is copied into a complementary messenger RNA sequence; transcription describes this step's mechanism and machinery in full. This step, not translation, is the point at which a specific gene is "read" out of the genome for use. A
3
\text{RNA} \xrightarrow{\text{translation}} \text{Protein}
Messenger RNA is read three bases (one codon) at a time and converted into a corresponding sequence of amino acids according to the genetic code; translation-genetic-code describes this mapping and the machinery that carries it out in full. A
4
\text{No general cellular pathway exists for Protein} \to \text{RNA or Protein} \to \text{DNA.}
Because the genetic code maps many possible codons to each amino acid (the code is degenerate, Discussion) and because no biochemical machinery exists to read a protein sequence and reconstruct a corresponding nucleic acid template (Hypotheses, third assumption), information cannot in general flow backward from protein to nucleic acid — even in principle, a given protein sequence would not specify a unique nucleic acid sequence to reverse-translate into. A
5
\text{RNA} \xrightarrow{\text{reverse transcription}} \text{DNA (retroviruses only)}
Certain viruses (retroviruses) carry an enzyme, reverse transcriptase, that synthesises DNA from an RNA template — a genuine, well-documented flow of information from RNA back to DNA. This does not violate Step 4's protein-to-nucleic-acid claim (protein is never the source), and Crick's own original formulation of the dogma explicitly anticipated RNA-to-DNA transfer as a chemically permitted, if biologically unusual, direction. B
Result
\text{DNA} \;\underset{\text{reverse transcription}}{\overset{\text{replication}}{\rightleftarrows}}\; \text{RNA} \xrightarrow{\text{translation}} \text{Protein} \quad (\text{Protein} \to \text{nucleic acid: never observed})

Reading. Sequence information moves freely between DNA and RNA in both directions (replication one way, transcription and — in retroviruses only — reverse transcription the other), and from RNA into protein by translation, but there is no known biochemical route by which sequence information moves from protein back into nucleic acid.

Scope. Describes the direction of information flow at the level of sequence (which linear polymer specifies which). It says nothing about whether a protein can influence gene expression by other means — many proteins (regulatory proteins, discussed later in this unit) do exactly this — only that they do not do so by contributing their own amino-acid sequence back into a nucleic acid template.

Corollaries & converses
  • mutation-types' account of how a single DNA base change alters a protein depends entirely on Steps 2–3: because information flows strictly DNA-to-RNA-to-protein and never back, a change introduced at the DNA level propagates forward into every subsequent protein made from that gene, rather than being correctable by the cell examining the resulting protein and inferring the original sequence.
  • Step 5's retroviral exception is itself the direct mechanistic basis of a major class of antiviral drugs, which specifically inhibit reverse transcriptase to block this otherwise-unusual direction of information flow.
  • Converse: observing that an acquired protein-level change (for example, a change caused by an environmental exposure during an organism's lifetime) is not passed on to that organism's offspring is consistent with, and can be explained by, the dogma's Step 4 claim that protein sequence is never copied back into the germline's nucleic acid.
Fails without
  • Drop the one-directional nucleic-acid-to-protein mapping: this would require a cellular mechanism to reverse-translate protein sequence back into RNA or DNA sequence, which no known biological system possesses; its absence is precisely what makes protein the functional endpoint of information flow, distinguishing this step from the genuinely two-directional DNA↔RNA step permitted by reverse transcription.
  • Drop sequence-based information encoding (imagine genetic information were instead stored in some non-linear, non-sequence-based chemical property): the entire mechanistic basis for templated, base-pairing-driven copying — semiconservative-replication, transcription, and translation-genetic-code, all of which depend on reading a linear sequence one unit at a time — would not exist in anything resembling its actual form.
Common errors
  • Stating the dogma as "DNA makes RNA makes protein" as if this were the complete claim; the genuinely restrictive part is Step 4 — that information does not flow back from protein — not merely the forward order of the first two steps.
  • Treating reverse transcription (Step 5) as disproving the central dogma; Crick's original formulation explicitly allowed nucleic-acid-to-nucleic-acid transfers in either direction, and reverse transcription is RNA to DNA, not protein to nucleic acid.
  • Assuming the dogma claims a gene is transcribed and translated only once, or in only one direction of regulation; it describes the direction sequence information can flow, not the number of times or the rate at which each step occurs.
  • Confusing the central dogma (a claim about the direction of sequence information transfer) with the genetic code (translation-genetic-code's specific codon-to-amino-acid mapping) — the two are related but answer different questions.
Discussion

Francis Crick proposed the central dogma in 1958, several years after he and James Watson had described the structure of DNA in 1953. Crick's original formulation was explicitly about the direction sequence information is permitted to flow between the three classes of biological polymer (nucleic acid to nucleic acid, nucleic acid to protein), and specifically did not claim that reverse transcription (RNA to DNA) was impossible — a point sometimes lost in later, looser popular restatements of the dogma as simply "DNA makes RNA makes protein."

The discovery of reverse transcriptase in retroviruses some years after Crick's original 1958 statement is frequently, but incorrectly, presented as having overturned the central dogma; in fact it confirmed exactly the kind of nucleic-acid-to-nucleic-acid transfer Crick's original, more careful formulation had already allowed for. The claim that has never been overturned, and remains the dogma's genuinely load-bearing content, is that protein sequence is never read back into nucleic acid sequence (Step 4).

Common misconception: that the central dogma has been "disproven" by the existence of reverse transcription or by phenomena such as prions (infectious, misfolded proteins that convert other proteins to the same misfolded shape). Prions transmit a change in protein conformation from one protein molecule to another, not a change in amino-acid sequence — no new sequence information is created or copied from protein back into nucleic acid, so this does not violate Step 4 as originally, precisely stated.

Worked examples
1
\text{A single-gene mutation changes one DNA base within a coding sequence.}
By Step 2, transcription copies the altered DNA sequence into messenger RNA, faithfully including the changed base; by Step 3, translation then reads the altered RNA codon and may incorporate a different amino acid at the corresponding position in the resulting protein, depending on which codon the change produced. A
2
\text{The resulting altered protein is now present in the cell in large amounts.}
By Step 4, this altered protein cannot itself correct, propagate, or otherwise write its changed sequence back into the cell's DNA; the DNA sequence remains altered (or reverts only by a further, independent mutation event) regardless of how much of the altered protein the cell subsequently produces. A
\text{DNA change} \to \text{RNA change} \to \text{protein change (one-way)}

Reading. A change introduced anywhere at the DNA level propagates forward through transcription and translation into every protein subsequently made from that gene, but a change confined to an existing protein molecule has no route back into the cell's genetic material.

Scope. This one-way propagation is the mechanistic reason mutation-types can treat DNA sequence as the definitive record of an organism's heritable genetic information, unaffected by whatever happens to the proteins that sequence goes on to produce.

Problems
  1. A researcher isolates a novel protein and wants to determine the DNA sequence that encodes it. Explain, using Step 4, why this cannot be done by directly "reverse-translating" the protein's amino acid sequence with full certainty.
    SolutionThe genetic code is degenerate — most amino acids are specified by more than one codon (translation-genetic-code) — so a given amino acid sequence corresponds to multiple possible underlying RNA (and hence DNA) sequences, not a unique one. Combined with Step 4's claim that no cellular mechanism performs protein-to-nucleic-acid information transfer at all, this means the researcher cannot recover the exact original DNA sequence from protein sequence alone; in practice the corresponding gene must instead be located and sequenced directly at the DNA or RNA level.
  2. HIV is a retrovirus. Explain how its use of reverse transcriptase (Step 5) fits within the central dogma rather than contradicting it.
    SolutionReverse transcriptase converts the virus's RNA genome into DNA, which is then integrated into the host cell's genome. This is a nucleic-acid-to-nucleic-acid transfer (RNA to DNA), the same general category of transfer as ordinary replication (DNA to DNA) and transcription (DNA to RNA), just running in the RNA-to-DNA direction. Crick's original formulation of the dogma explicitly permitted such nucleic-acid-to-nucleic-acid transfers; what remains prohibited, and is not violated here, is any transfer of sequence information from protein into nucleic acid (Step 4).
  3. A student argues that because regulatory proteins can bind DNA and switch genes on or off, proteins must be able to "send information back" to DNA, contradicting the dogma. Evaluate this argument using the Result's Scope.
    SolutionThe argument conflates two different senses of "information." The central dogma concerns the flow of sequence information — which linear polymer's sequence determines which. A regulatory protein binding DNA can influence whether and how much a gene is transcribed (an important and real form of control), but it does not alter the DNA's own base sequence, nor does its own amino-acid sequence get copied into the DNA. This is precisely the distinction the Result's Scope note draws: proteins can influence gene expression by other means without violating the dogma's specific claim about sequence-information flow.