biology2u
Tier
⌕ Search ⌘K
Concept

The molecular clock

T-083Home BU-304Threads evolution
Statement

Dating divergence from accumulated substitutions.

Why it matters

phylogenetics establishes how to reconstruct the branching pattern of evolutionary relationships from shared data, but a tree topology alone says nothing about when any given split actually occurred in real time; the molecular clock is what converts branch length (measured in accumulated sequence differences) into an actual date, turning a purely relative family tree into a dated one. neutral-theory (developed alongside this result and sharing its founder) supplies the theoretical justification for why substitutions might accumulate at a sufficiently steady rate to be useful as a clock at all.

It matters practically for dating events with no fossil record of their own — viral outbreaks, the divergence of lineages that left few or no fossils, or splits far too ancient for any direct physical record to survive — provided the clock can be calibrated against at least some independent, dated reference points.

Hypotheses
Substitutions accumulate in a given gene or genomic region at an approximately constant rate over time, at least within a given lineage or class of sites.Without approximate rate constancy there is no fixed conversion factor between sequence divergence and elapsed time at all; the entire method depends on being able to treat divergence as behaving like a linear clock, even though the underlying process is fundamentally stochastic (Step 2) rather than exactly regular tick-by-tick. The clock must be calibrated against at least one node of independently known age (typically from the fossil record or a well-dated geological event).Sequence divergence alone gives a number of substitutions, not a date; without external calibration, the substitution rate itself (substitutions per site per year) is unknown, and divergence can only be expressed in relative terms, not converted into actual years. The "strict" clock assumption of one single rate across an entire tree is now known to be commonly violated (rate variation between lineages, e.g. from differing generation time or metabolic rate); relaxed-clock models explicitly allow rate to vary across branches, trading some of the original method's simplicity for substantially improved accuracy on real, heterogeneous datasets.
Proof
1
d = \text{(number of substitutions per site between two sequences)}
Comparing homologous sequences from two lineages descended from a common ancestor gives a direct count of accumulated differences per site, the raw observable quantity the clock is built from. A
2
\text{Substitutions accumulate as an approximately Poisson process at rate } r \text{ (substitutions/site/year) per lineage.}
If mutations arise randomly over time at a roughly constant underlying rate, and a large enough fraction of them are neutral or nearly so (neutral-theory) to be fixed by drift rather than filtered out by strong, time-varying selection, the number of accumulated substitutions in a given lineage approximates a Poisson counting process, whose expected count grows linearly with elapsed time. B
3
d = 2rt
Since two lineages have each been independently accumulating substitutions since diverging from their common ancestor at time \(t\) in the past, total observed divergence \(d\) between them is the sum of both lineages' independent accumulation, giving \(d=2rt\) (the factor of 2 for the two independently evolving branches). A
4
r = \frac{d_{cal}}{2t_{cal}}\quad(\text{calibration}); \qquad t_{unknown} = \frac{d_{unknown}}{2r}
Given one pair of lineages with both a known divergence \(d_{cal}\) and an independently known divergence time \(t_{cal}\) (Hypotheses), Step 3 is solved for \(r\); this calibrated rate is then applied to any other pair of lineages sharing the same clock, converting their measured divergence \(d_{unknown}\) directly into an estimated divergence time. B
5
\text{Different sites/genes evolve at different rates: synonymous sites (neutral-theory) generally faster than nonsynonymous, and different genes faster or slower depending on functional constraint.}
A single universal clock rate does not apply across the whole genome; rapidly evolving, weakly constrained regions (e.g. synonymous codon positions) are useful for dating recent divergences, while slowly evolving, strongly constrained genes are needed to resolve much older splits without the signal saturating (Fails without). A
Result
d = 2rt \ \ \Longrightarrow\ \ t = \frac{d}{2r}

Reading. Given a calibrated substitution rate, the number of sequence differences accumulated between two lineages converts directly into an estimated time since their common ancestor.

Scope. Requires approximate rate constancy over the interval considered (Hypotheses) and at least one external calibration point; accuracy degrades for very deep divergences (saturation, Fails without) or lineages with substantially different substitution rates (rate variation, Discussion).

Corollaries & converses
  • neutral-theory's prediction that the substitution rate for strictly neutral sites equals the mutation rate itself, independent of population size, is exactly the theoretical reason a molecular clock can be approximately constant across lineages with very different population sizes and generation times, at least for neutral or nearly neutral sites.
  • coevolution can distort local molecular clock behaviour: genes under strong, fluctuating reciprocal selection (e.g. host-pathogen arms races) evolve at rates that deviate from the neutral baseline, so clock-based dating is generally applied to more neutrally evolving regions of the genome instead.
  • Converse: if a gene's substitution rate is found to differ substantially between two lineages sharing a well-established divergence time from other evidence, this is itself direct evidence that a strict, single-rate clock does not apply to that gene, motivating a relaxed-clock model (Hypotheses' t3 note) rather than the simple calculation of Step 4.
Fails without
  • Drop approximate rate constancy (Hypotheses): lineages with markedly different generation times or metabolic rates (both empirically correlated with substitution rate) violate the assumption that \(r\) in Step 2 is shared across the tree; applying one calibrated \(r\) uniformly to such lineages systematically mis-dates whichever lineage's true rate differs from the calibration lineage's.
  • Apply the linear relationship \(d=2rt\) (Step 3) to very deep divergences without correction: at large \(t\), multiple substitutions can occur at the same site (multiple hits), some of which reverse earlier changes; observed \(d\) then underestimates the true number of substitutions that actually occurred, a saturation effect that makes naive linear dating systematically too young for ancient splits unless a model correcting for multiple hits is used instead.
Common errors
  • Assuming a single universal substitution rate applies to every gene and every lineage, ignoring both site-to-site rate variation (Step 5) and lineage-to-lineage rate variation (Fails without, first bullet).
  • Applying an uncorrected linear divergence-to-time conversion (Step 3) to deeply diverged sequences without accounting for multiple-hit saturation (Fails without, second bullet).
  • Treating a molecular-clock date as having the same certainty as a directly dated fossil; clock estimates carry substantial statistical uncertainty from both the stochastic substitution process (Step 2) and calibration error.
  • Confusing the substitution rate (used for dating, generally dominated by neutral or nearly-neutral sites) with the rate of adaptive evolutionary change, which is a much smaller and less clock-like fraction of total substitutions.
Discussion

Emile Zuckerkandl and Linus Pauling proposed the molecular clock hypothesis in 1962, observing that amino acid differences between species' haemoglobin sequences appeared to accumulate roughly in proportion to time since divergence, as independently estimated from the fossil record — an empirical regularity that predated, and helped motivate, Motoo Kimura's neutral theory (1968) as a mechanistic explanation for why such rate constancy might be expected in the first place.

Modern relaxed-clock methods, widely used in current phylogenetic dating, model substitution rate itself as a random variable drawn from a statistical distribution across branches of the tree, rather than assuming one fixed universal rate; this substantially improves accuracy for datasets spanning lineages with genuinely different rates, at the cost of requiring more sophisticated statistical inference than the simple linear calculation of Step 4.

Common misconception: that the molecular clock "ticks" at a fixed number of years per substitution the way a mechanical clock ticks at a fixed time interval. It is a statistical, average tendency over a fundamentally stochastic mutation-and-fixation process (Step 2); individual lineages and individual time intervals can and do depart from the average rate, which is exactly why calibration and, increasingly, relaxed-clock models are needed rather than a single fixed conversion factor applied uncritically.

Worked examples
1
d_{cal}=0.06\ \text{substitutions/site},\quad t_{cal}=30\times10^6\ \text{yr (fossil-calibrated node)}
Applying Step 4's calibration formula: \(r=\dfrac{0.06}{2\times30\times10^6}=1.0\times10^{-9}\) substitutions/site/year, a rate now fixed for this particular gene and lineage group. A
2
d_{unknown}=0.03\ \text{substitutions/site (a different, undated pair of lineages sharing the same gene and clock)}
Using the calibrated rate from Step 1 in Step 4's rearranged formula: \(t=\dfrac{0.03}{2\times1.0\times10^{-9}}=1.5\times10^{7}\) years, an estimated divergence time obtained with no fossil evidence for this particular split at all. A
t_{unknown}\approx1.5\times10^{7}\ \text{years}

Reading. A single well-dated calibration point is enough to convert any other measured sequence divergence, on the same clock, into an estimated absolute date.

Scope. The estimate's reliability depends entirely on how well the calibration lineage's rate represents the lineage being dated (Fails without).

Problems
  1. Using the calibrated rate \(r=1.0\times10^{-9}\) substitutions/site/year from Worked Example 1, find the estimated divergence time for a lineage pair showing \(d=0.10\) substitutions/site.
    Solution\(t=\dfrac{d}{2r}=\dfrac{0.10}{2\times1.0\times10^{-9}}=5.0\times10^{7}\) years, using Step 4's formula directly with the previously calibrated rate.
  2. A researcher calibrates a molecular clock using a fossil-dated primate divergence, then applies the resulting rate to date a divergence between two bacterial lineages with much shorter generation times. Explain, using Fails without, why this cross-application is likely to give an unreliable date.
    SolutionSubstitution rate is empirically correlated with generation time and other lineage-specific factors, so a rate calibrated on primates need not apply to bacteria with a very different generation time (Fails without, first bullet, and the Hypotheses' rate-constancy assumption). Applying a mismatched rate systematically over- or under-estimates the bacterial divergence time, since the true bacterial \(r\) likely differs substantially from the calibrated primate \(r\).
  3. Two deeply diverged lineages show an observed sequence divergence that, when converted via Step 3's simple linear formula, gives an implausibly young estimated divergence time compared to strong independent fossil evidence. Suggest, using Fails without, a likely cause.
    SolutionAt large true divergence times, multiple substitutions can occur at the same site, including changes that revert a site to its original state; the observed \(d\) then undercounts the true number of substitutions that actually occurred (saturation, Fails without, second bullet), making the naive linear estimate of \(t\) too young. A model that explicitly corrects for multiple hits at each site would be needed to recover a more accurate, older estimate consistent with the fossil evidence.