Mutation-selection balance
Statement
Why harmful alleles persist at low frequency.
Why it matters
hardy-weinberg established the null expectation for allele frequencies in an idealised population with no mutation, selection, drift or migration; mutation-selection balance is the first, and one of the most important, results in this unit that puts two of those forces back in simultaneously and shows they can reach a genuine equilibrium together, rather than one simply overwhelming the other. It directly answers a question hardy-weinberg alone cannot: why do clearly harmful, disease-causing alleles persist in every population at low but predictable and stable frequency, rather than being eliminated entirely by selection or drifting to unpredictable levels.
It is also the standard quantitative framework used in medical and population genetics to interpret the frequency of recessive genetic disease alleles, and inbreeding-heterozygosity depends on knowing the underlying allele frequency mutation-selection balance predicts before its own consequences for homozygosity can be assessed.
Hypotheses
Proof
Result
Reading. A harmful recessive allele reaches a stable, predictable equilibrium frequency set by the balance between how often mutation recreates it and how strongly selection removes it, rather than being eliminated entirely or drifting without bound.
Scope. Assumes recurrent mutation, a fixed selection coefficient, and full recessivity (Hypotheses); heterozygote advantage (Step 5) or partial dominance changes the formula and can maintain much higher frequencies than mutation-selection balance alone.
Corollaries & converses
- genetic-drift can push an allele's frequency away from \(\hat q\) temporarily, especially in small populations, but selection and recurrent mutation together continually pull it back toward \(\hat q\), so mutation-selection balance describes the expected long-run average frequency rather than a value fixed rigidly every generation.
- Because \(\hat q\propto\sqrt{\mu/s}\) (Step 4), a more severe (higher-\(s\)) genetic disease is predicted to be rarer at equilibrium than a milder one with the same mutation rate — a testable, and generally confirmed, prediction linking disease severity to population frequency.
- Converse: given an observed disease allele frequency and an independently estimated mutation rate, Step 4 can be rearranged (\(s=\mu/\hat q^2\)) to estimate the selection coefficient acting against the allele, without measuring fitness differences directly.
Fails without
- Stop mutation (\(\mu=0\), Hypotheses): Step 1's input term vanishes, and selection (Step 2) alone steadily removes the harmful allele generation after generation with nothing replacing it; the allele frequency declines toward zero rather than settling at the stable \(\hat q=\sqrt{\mu/s}\) of the Result.
- Selection coefficient \(s=0\) (a neutral allele, Hypotheses): Step 4's formula gives \(\hat q=0\) formally, but the actual behaviour is qualitatively different, not merely a smaller equilibrium value — with no selection at all opposing it, mutation input is unopposed and the allele's frequency instead behaves according to neutral-theory and genetic-drift dynamics, with no stable equilibrium frequency of the mutation-selection type existing at all.
Common errors
- Applying the recessive-case formula \(\hat q=\sqrt{\mu/s}\) (Step 4) to a dominant harmful allele, for which selection acts on both heterozygotes and homozygotes and the correct equilibrium is instead \(\hat q\approx\mu/s\) (linear in \(\mu\), not a square root) — the two formulas are not interchangeable.
- Assuming mutation-selection balance alone explains every case of a harmful allele persisting at unexpectedly high frequency, without checking for heterozygote advantage (Step 5), which can dominate entirely at frequencies mutation rate alone could never sustain.
- Treating \(\hat q\) as a fixed, permanent constant rather than an equilibrium that the population's actual frequency fluctuates around, particularly noticeable via drift in small populations.
- Confusing selection coefficient \(s\) (a measure of relative fitness reduction, ranging 0 to 1) with mutation rate \(\mu\) (a per-generation probability, typically far smaller, on the order of \(10^{-5}\) to \(10^{-6}\) per locus) — the two parameters have very different typical magnitudes and enter Step 4's formula asymmetrically (square-rooted).
Discussion
The sickle-cell allele is the standard textbook illustration of Step 5's heterozygote advantage overriding ordinary mutation-selection balance: heterozygous carriers are substantially protected against severe malaria while homozygotes suffer sickle-cell disease, and in regions with historically high malaria burden the allele is maintained at frequencies far above what recurrent mutation against a purely deleterious, non-advantaged recessive allele could ever sustain by Step 4 alone.
Because \(\hat q\) scales as the square root of \(\mu/s\) rather than linearly, mutation-selection balance is a relatively weak brake on allele frequency for any but the most severe (highest-\(s\)) genetic diseases; this square-root dependence is exactly why even quite low, realistic per-locus mutation rates are sufficient to maintain measurable frequencies of rare recessive disease alleles across human populations, without requiring implausibly high mutation rates.
Common misconception: that a harmful allele persisting in a population must be secretly beneficial in some way (as in heterozygote advantage). Step 4 shows this need not be true at all: a purely, unconditionally deleterious allele can persist indefinitely at a low but nonzero and entirely predictable equilibrium frequency simply because mutation keeps recreating it as fast as selection removes it — no hidden benefit is required.
Worked examples
Reading. Even a fully lethal recessive condition is not eliminated by selection; it persists at a small but stable frequency because mutation continually recreates carrier alleles as fast as selection removes affected homozygotes.
Scope. Halving \(s\) to 0.5 (partial fitness reduction rather than lethality) would raise \(\hat q\) only by a factor of \(\sqrt2\approx1.41\), illustrating the square-root formula's relative insensitivity to \(s\).
Problems
- A recessive genetic disease has mutation rate \(\mu=4\times10^{-6}\) and selection coefficient \(s=0.25\) (homozygotes have 25% reduced fitness). Find the equilibrium allele frequency \(\hat q\) and the homozygote (affected) frequency \(\hat q^2\).
Solution
\(\hat q=\sqrt{\mu/s}=\sqrt{4\times10^{-6}/0.25}=\sqrt{1.6\times10^{-5}}\approx4.0\times10^{-3}\). Affected homozygote frequency: \(\hat q^2=1.6\times10^{-5}\), or about 1.6 in 100,000 births, using hardy-weinberg's \(q^2\) directly. - Two recessive diseases share the same mutation rate \(\mu\), but disease X has \(s=1\) (lethal) and disease Y has \(s=0.01\) (mild fitness reduction). Using Step 4, determine the ratio \(\hat q_Y/\hat q_X\).
Solution
\(\hat q_Y/\hat q_X=\sqrt{\mu/0.01}\big/\sqrt{\mu/1}=\sqrt{1/0.01}=\sqrt{100}=10\). The milder disease (Y) has an equilibrium allele frequency 10 times higher than the lethal disease (X), despite sharing an identical mutation rate — directly confirming the Corollaries' prediction that milder conditions are more common at equilibrium. - A population geneticist observes a recessive disease allele at frequency \(\hat q=0.02\) and independently estimates the mutation rate at \(\mu=8\times10^{-6}\). Using the Converse in Corollaries, estimate the selection coefficient \(s\) acting against the allele.
Solution
Rearranging Step 4: \(s=\mu/\hat q^2=8\times10^{-6}/(0.02)^2=8\times10^{-6}/4\times10^{-4}=0.02\). The estimated selection coefficient is 0.02 — a relatively mild fitness reduction, consistent with the allele persisting at a comparatively high equilibrium frequency.