biology2u
Tier
⌕ Search ⌘K
Concept

Transcription

T-033Home BU-201Threads information
Statement

Copying a gene into messenger RNA.

Why it matters

central-dogma states, at a high level, that genetic information flows from DNA to RNA to protein; transcription is the first of those two steps made mechanistically concrete, and it is the step at which a cell actually decides which of its genes get expressed at any given moment. semiconservative-replication copies the entire genome faithfully once per cell cycle, but transcription is selective and repeated: the same gene can be transcribed many times, or not at all, depending on the cell's needs, making it the primary point of control for differential gene expression across cell types and conditions.

Everything translation-genetic-code does afterward depends entirely on the fidelity and processing of the messenger RNA transcript produced here, and mutation-types occurring within a gene's sequence exert their downstream effect specifically by altering the RNA transcript this process produces.

Hypotheses
RNA polymerase synthesises RNA in the 5′ to 3′ direction, reading the DNA template strand 3′ to 5′, and incorporates ribonucleotides complementary to the template (using uracil in place of thymine).This directionality and base-pairing rule is what makes the RNA transcript an accurate, predictable copy of the coding (non-template) strand's sequence, with U substituted for T; without a fixed, enzyme-enforced directionality and pairing rule, transcript sequence would not reliably correspond to the gene's DNA sequence at all. Transcription initiation requires RNA polymerase (in eukaryotes, aided by general transcription factors) to recognise and bind a specific promoter sequence upstream of the gene.Promoter recognition is what makes transcription gene-specific rather than starting at arbitrary points along the DNA; different genes' promoters can differ substantially in sequence and in which additional regulatory proteins they require, which is the molecular basis of differential, cell-type-specific gene expression. In eukaryotes, the primary RNA transcript (pre-mRNA) is processed — capped, polyadenylated, and spliced — before export from the nucleus as mature mRNA.This processing has no equivalent in bacterial transcription, where mRNA is used directly, often even before transcription of it is complete; the eukaryotic-specific splicing step in particular allows alternative splicing, generating multiple distinct mature mRNAs (and so multiple distinct proteins) from a single gene, developed as its own topic given the depth it requires.
Proof
1
\text{Initiation: RNA polymerase (with transcription factors, in eukaryotes) binds the gene's promoter and unwinds a local region of DNA to expose the template strand.}
Promoter recognition (Hypotheses) positions the polymerase precisely at the correct start site and in the correct orientation, ensuring transcription proceeds along the intended template strand in the intended direction, rather than starting non-specifically or reading the wrong strand. A
2
\text{Elongation: RNA polymerase moves along the template strand, synthesising a complementary RNA strand in the 5}'\text{ to 3}'\text{ direction while the DNA re-forms its double helix behind it.}
The enzyme adds ribonucleotides one at a time, each selected by Watson–Crick base-pairing with the template strand (dna-base-pairing's rules extended to an RNA product, with uracil pairing where thymine would); the transcription bubble (locally unwound DNA) moves with the polymerase, so only a short stretch of DNA is single-stranded at any instant. A
3
\text{Termination: RNA polymerase reaches a specific terminator sequence (or, in eukaryotes, a polyadenylation signal coupled to termination), releasing both the completed transcript and the DNA template.}
Termination sequences are recognised by the transcription machinery itself (directly, or via associated termination factors) and mark the defined end of the gene's transcribed region, ensuring the transcript has a defined 3′ end rather than continuing indefinitely along the DNA. A
4
\text{In eukaryotes, the pre-mRNA transcript is capped at its 5}'\text{ end (7-methylguanosine cap) and polyadenylated at its 3}'\text{ end (poly-A tail) before splicing is completed.}
The 5′ cap protects the transcript from degradation and is later required for ribosome recognition during translation initiation; the poly-A tail similarly stabilises the transcript and assists nuclear export, so both modifications prepare the transcript for its subsequent journey to, and use at, the ribosome. B
5
\text{Splicing removes introns (non-coding intervening sequences) from the pre-mRNA and joins the remaining exons together, producing the mature mRNA.}
Carried out by the spliceosome, a large ribonucleoprotein complex that recognises specific sequences at intron boundaries; only the joined exon sequence is retained in the mature transcript that is ultimately translated (translation-genetic-code), meaning the initial pre-mRNA sequence and the final coding message can differ substantially in length and content. B
Result
\text{DNA template} \xrightarrow{\text{RNA polymerase}} \text{pre-mRNA} \xrightarrow{\text{cap, poly-A, splice}} \text{mature mRNA}

Reading. A gene's DNA sequence is copied into a complementary, single-stranded RNA transcript by an enzyme reading the template strand with fixed directionality and base-pairing rules, then, in eukaryotes, that transcript is chemically capped, tailed, and spliced before it is ready to direct protein synthesis.

Scope. The core initiation–elongation–termination mechanism (Steps 1–3) is broadly conserved across bacteria and eukaryotes; the processing steps (Steps 4–5) are specifically eukaryotic, reflecting the physical separation of transcription (nucleus) from translation (cytoplasm) that bacteria, lacking a nucleus, do not have.

Corollaries & converses
  • central-dogma's DNA-to-RNA arrow is precisely this process; transcription is the specific mechanistic content behind that single, high-level arrow, just as translation-genetic-code is the mechanistic content behind the RNA-to-protein arrow that follows it.
  • mutation-types occurring within a promoter sequence can alter transcription rate (or abolish it) without changing the coding sequence at all, while mutations within the coding sequence itself are transcribed into the RNA transcript unchanged and instead exert their effect at translation — the same DNA change can act at either of two distinct downstream steps depending on exactly where it falls.
  • Converse: observing that two cell types produce different proteins from an identical genome (the basis of cellular differentiation) is itself indirect evidence that transcription (specifically, which promoters are active, Step 1) is being differentially regulated between them, since the underlying DNA sequence available to be transcribed is the same in both.
Fails without
  • Drop promoter-specific initiation (Hypotheses, Step 1): if RNA polymerase bound and initiated transcription at arbitrary locations along the DNA rather than at defined promoters, transcripts would not correspond to intact, functional genes, and the cell would have no mechanism for selectively expressing one gene rather than another — gene-specific regulation would be impossible.
  • Drop splicing (Step 5) in a eukaryotic cell: an unspliced pre-mRNA, still containing its introns, would be translated (if it reached the ribosome at all) with intervening non-coding sequence included, almost always disrupting the reading frame or introducing premature stop codons partway through the intended protein sequence (translation-genetic-code), rendering the resulting protein non-functional or truncated.
Common errors
  • Confusing the template strand (read 3′ to 5′, used to synthesise RNA) with the coding strand (the strand whose sequence the RNA transcript actually matches, apart from U replacing T); the RNA transcript is complementary to the template strand but identical in sequence to the coding strand.
  • Assuming the primary RNA transcript is immediately ready for translation in eukaryotes; Steps 4–5 (capping, polyadenylation, splicing) must all occur first, and skipping this distinction is a frequent source of confusion between bacterial and eukaryotic gene expression.
  • Treating introns as "junk" that plays no functional role simply because they are removed before translation; some introns contain regulatory sequences affecting splicing efficiency or timing, and alternative splicing of the same pre-mRNA (retaining or removing different exon combinations) is itself a major, biologically important source of protein diversity from a single gene.
  • Assuming transcription and translation are temporally separated in bacteria the way they are in eukaryotes; because bacteria lack a nucleus, translation of a bacterial mRNA can begin before its transcription is even complete, a coupling that is structurally impossible in eukaryotic cells.
Discussion

The basic scheme of transcription was worked out through the 1960s, building directly on the central-dogma framework Francis Crick had articulated in 1958, and on the identification of messenger RNA itself as the intermediate carrier of genetic information from DNA to the protein-synthesising machinery. RNA polymerase's core catalytic mechanism and its promoter-recognition subunits were characterised over the following decades in both bacterial and, later, eukaryotic systems.

Eukaryotes possess three distinct RNA polymerases (I, II, and III), each dedicated to transcribing a different class of RNA — polymerase II specifically transcribes protein-coding genes into pre-mRNA, while polymerase I and III transcribe most ribosomal and transfer RNA genes respectively — a division of labour with no equivalent in bacteria, which use a single RNA polymerase for essentially all transcription.

Common misconception: that a gene's transcription rate is fixed and identical across all the cells of an organism. In practice, transcription is the primary point at which gene expression is regulated cell-type-by-cell-type and condition-by-condition (Corollaries' converse); the presence of a gene in a cell's genome guarantees nothing about whether, or how much, it is actually being transcribed in that particular cell at a particular time.

Worked examples
1
\text{Template strand (3}'\to5'\text{): } 3'\text{-TACCGGATT-}5'
Applying Step 2's base-pairing rule (A pairs with U, T pairs with A, C pairs with G, G pairs with C, all read in the 5′ to 3′ direction for the new RNA strand) gives the RNA transcript \(5'\text{-AUGGCCUAA-}3'\), synthesised antiparallel to the template as required by Step 2's directionality rule. A
2
\text{This transcript, } 5'\text{-AUGGCCUAA-}3', \text{ begins with the start codon AUG and ends with the stop codon UAA.}
Once capped, polyadenylated, and (if introns were present) spliced (Steps 4–5), this mature mRNA is ready to be read by the ribosome three nucleotides (one codon) at a time, exactly the process translation-genetic-code develops in full. A
3'\text{-TACCGGATT-}5' \ (\text{template}) \ \Rightarrow\ 5'\text{-AUGGCCUAA-}3' \ (\text{mRNA transcript})

Reading. Applying the fixed base-pairing and directionality rules of Steps 1–2 mechanically to any given template sequence deterministically produces its RNA transcript; there is no ambiguity once the template sequence and its 3′-to-5′ orientation are known.

Scope. The identical rule applies to transcribing any gene of any length, in either bacteria or eukaryotes; only the subsequent processing steps (Steps 4–5) differ between the two.

Problems
  1. Given the coding (non-template) strand \(5'\text{-ATGCGTACG-}3'\), write the corresponding mRNA transcript.
    SolutionThe mRNA transcript matches the coding strand's sequence directly, with U substituted for T (since the coding strand and the mRNA are, by definition, identical apart from this substitution — the template strand is the one actually read and is complementary to both). Transcript: \(5'\text{-AUGCGUACG-}3'\).
  2. A point mutation changes a single nucleotide within an intron, several bases away from any splice-site boundary sequence. Predict the most likely effect on the mature mRNA and resulting protein, using Step 5.
    SolutionSince the mutated nucleotide lies within an intron and away from the splice-site sequences the spliceosome recognises (Step 5), the intron containing it is still recognised and removed normally during splicing; the mutated nucleotide is excised along with the rest of the intron and does not appear in the mature mRNA at all, so the resulting protein sequence is most likely unaffected.
  3. Explain why a drug that inhibits RNA polymerase II specifically would block production of new proteins in a eukaryotic cell, while having no direct effect on ribosomal RNA synthesis, using the Discussion's note on the three eukaryotic RNA polymerases.
    SolutionRNA polymerase II is responsible specifically for transcribing protein-coding genes into pre-mRNA (Discussion); inhibiting it blocks Step 1–3 for all such genes, halting new mRNA production and, downstream, new protein synthesis (translation-genetic-code) once existing mRNA is degraded. Ribosomal RNA, however, is transcribed by RNA polymerase I (and some small RNAs by polymerase III), enzymes unaffected by a polymerase-II-specific inhibitor, so ribosomal RNA synthesis continues even as mRNA production for new proteins stops.