Genetic drift
Statement
Random sampling changes allele frequencies.
Why it matters
hardy-weinberg's equilibrium assumes an infinitely large population, in which allele frequencies never change from generation to generation absent other forces; genetic-drift is exactly what happens once that idealisation is dropped and population size is treated as finite. It is the force that operates in every real, finite population to some degree, and its strength — unlike selection or mutation, which depend on specific biological effects — depends purely on population size, making it the single most universal departure from Hardy-Weinberg's idealised baseline.
Drift's practical importance is greatest precisely where conservation biology needs it most: small, isolated, or recently bottlenecked populations, where random sampling alone can eliminate genetic variation surprisingly quickly, independent of whether any of the alleles involved actually affect fitness at all.
Hypotheses
Proof
Result
Reading. Genetic drift's strength scales inversely with effective population size: allele frequencies wander randomly and heterozygosity is lost fastest in small populations, while large populations drift negligibly over any given number of generations.
Scope. Applies to any finite, sexually reproducing population absent other forces; effective population size \(N_e\), not raw census size, should be used for accurate prediction, since the two can differ substantially (Hypotheses, Step 5).
Corollaries & converses
- gene-flow (migration) directly counteracts drift's tendency toward local fixation or loss by re-injecting variation from other populations, providing the main opposing force to drift's homogenising-within/differentiating-between effect.
- inbreeding-heterozygosity's decline in a small, closed mating pool is drift's genotypic-level consequence, viewed through the lens of individual relatedness rather than population-wide allele frequency alone.
- hardy-weinberg's equilibrium is recovered exactly in the idealised \(N\to\infty\) limit, where \(\text{Var}(\Delta p)\to0\); real populations never meet this exactly, which is precisely why some degree of drift is always present in real biology.
Fails without
- Population size is effectively infinite (Hardy-Weinberg's idealisation): \(\text{Var}(\Delta p)=p(1-p)/(2N)\to0\) as \(N\to\infty\), and drift vanishes entirely; real, finite populations never meet this exactly, which is exactly why drift is present to some nonzero degree in every real biological population, however large.
- Effective population size is assumed equal to census size when the true \(N_e\) is much smaller (Step 5): predicted rates of heterozygosity loss and allele fixation would be badly underestimated relative to what is actually observed, since high variance in reproductive success or a historical bottleneck can make \(N_e\) far smaller than a simple head-count of living individuals would suggest.
Common errors
- Treating genetic drift as directional (e.g. "drift pushes frequencies toward 0.5" or toward fixation specifically), when in fact \(E[\Delta p]=0\) — drift is undirected in expectation; only its variance, not its mean, is nonzero (Step 2).
- Confusing genetic drift with natural selection as an explanation for an observed frequency change; small-sample random fluctuation can mimic a pattern that superficially looks like selection acting, without any actual fitness difference being present.
- Assuming \(N_e\) equals the census population size, ignoring that unequal sex ratios, high variance in offspring number among individuals, and historical bottlenecks generally make the true effective population size substantially smaller (Step 5).
- Assuming drift only matters in populations that are "small" in some absolute, fixed sense, rather than recognising its strength is a continuous function of \(1/(2N_e)\) that matters to some measurable degree at any finite population size.
Discussion
Sewall Wright and R.A. Fisher independently developed the foundational mathematical treatment of random genetic drift in the early-to-mid 20th century, and the resulting framework (often called the Wright-Fisher model) remains the standard formal starting point for analysing drift's effect on finite populations. Founder effects (a new population established by a small number of colonising individuals) and population bottlenecks (a temporary, sharp reduction in population size) are both real-world scenarios of especially strong drift, corresponding directly to an episode of unusually small \(N_e\) in Step 5's formula.
Because drift's strength depends on \(N_e\) rather than raw census size, two populations with identical current head counts can nonetheless experience very different rates of drift if their variance in reproductive success, sex ratio, or recent demographic history differ — a point of direct practical importance in conservation genetics, where \(N_e\) rather than census size is the relevant quantity for assessing a small population's genetic risk.
Common misconception: that an allele observed rising in frequency over several generations must be under positive selection. Because drift alone can, by chance, produce a run of frequency increases (or decreases) over a limited number of generations even with \(E[\Delta p]=0\) overall, distinguishing a genuine selective signal from a chance drift trajectory generally requires statistical tests that account explicitly for the expected magnitude of drift at the population's actual effective size.
Worked examples
Reading. Even a population of a modest, not extremely small, size shows readily measurable drift effects over a comparatively short number of generations — drift is not a phenomenon confined only to extremely tiny populations.
Scope. The identical two formulas, evaluated at any \(N\) (or, more accurately, \(N_e\)) and any number of generations, predict the expected magnitude of drift for that specific case.
Problems
- A population has \(N=1000\) and starting allele frequency \(p=0.30\). Compute \(\text{Var}(\Delta p)\) for a single generation, and compare qualitatively to the Worked Example's \(N=25\) case.
Solution
\(\text{Var}(\Delta p)=\dfrac{0.30\times0.70}{2000}=\dfrac{0.21}{2000}=0.000105\), roughly 48 times smaller than the Worked Example's \(N=25\) case (\(0.005\)), consistent with Step 4 of the Proof: variance scales as \(1/(2N)\), so a 40-fold increase in \(N\) (from 25 to 1000) reduces the variance by roughly the same factor. - A conservation biologist estimates a population's census size at 500 individuals, but genetic data suggest its effective population size is closer to 50, due to a small number of males siring most offspring. Explain the practical consequence of using the census size rather than \(N_e\) when predicting this population's rate of heterozygosity loss.
Solution
Using the census size of 500 in the heterozygosity-decay formula (Step 3) would predict a much slower rate of heterozygosity loss than the population is actually experiencing, since the true relevant quantity is \(N_e=50\), tenfold smaller (Step 5). The population is genetically behaving like a much smaller one than its head count suggests, and is losing genetic variation, and approaching fixation at individual loci, considerably faster than a census-size-based estimate would indicate — a serious underestimate of genetic risk if not corrected for. - Explain why an allele with no effect on fitness whatsoever can nonetheless become completely fixed (reach a frequency of 1.0) in a small population within a modest number of generations.
Solution
Fixation requires no fitness effect at all under drift; because \(\text{Var}(\Delta p)=p(1-p)/(2N)\) is nonzero for any finite \(N\) (Step 1), and \(E[\Delta p]=0\) means increases and decreases are equally likely at every step (Step 2), a purely neutral allele's frequency performs an unbiased random walk that will, given enough generations, eventually hit either 0 (loss) or 1 (fixation) with a probability set purely by its current frequency and population size — smaller \(N\) makes this random walk's steps proportionally larger relative to the \([0,1]\) frequency range, so fixation or loss happens faster in smaller populations (Step 4), entirely independent of whether the allele affects fitness at all.