biology2u
Tier
⌕ Search ⌘K
Concept

Genetic drift

T-047Home BU-204Threads information · evolution
Statement

Random sampling changes allele frequencies.

Why it matters

hardy-weinberg's equilibrium assumes an infinitely large population, in which allele frequencies never change from generation to generation absent other forces; genetic-drift is exactly what happens once that idealisation is dropped and population size is treated as finite. It is the force that operates in every real, finite population to some degree, and its strength — unlike selection or mutation, which depend on specific biological effects — depends purely on population size, making it the single most universal departure from Hardy-Weinberg's idealised baseline.

Drift's practical importance is greatest precisely where conservation biology needs it most: small, isolated, or recently bottlenecked populations, where random sampling alone can eliminate genetic variation surprisingly quickly, independent of whether any of the alleles involved actually affect fitness at all.

Hypotheses
Each generation's allele pool is a finite random sample of the previous generation's alleles.Sampling error is unavoidable whenever population size \(N\) is finite; only in the idealised infinite-population limit does this sampling variance vanish entirely, which is precisely the limit hardy-weinberg's equilibrium assumes. Absent other forces, drift is directionless in expectation but not varianceless.The expected frequency next generation equals the current frequency (\(E[p']=p\)), so drift carries no systematic bias toward higher or lower frequency on average; but because the variance around that expectation is nonzero, any single finite population's actual trajectory is a random walk, not a flat line.
Proof
1
\text{Var}(\Delta p) = \frac{p(1-p)}{2N}
Modelling allele transmission each generation as a binomial sample of \(2N\) allele copies (diploid, \(N\) individuals) drawn from a population at frequency \(p\), this is the resulting sampling variance of the next generation's frequency — the direct quantitative consequence of the Hypotheses' finite-sampling requirement. B
2
E[p'] = p, \qquad \text{Var}(\Delta p) \neq 0
Because the expected value of the next generation's frequency equals the current frequency, drift has no directional bias on average (Hypotheses); but the nonzero variance from Step 1 means any single finite population's realised trajectory is a random walk around that expectation, not a constant line. A
3
H_t = H_0\left(1-\frac{1}{2N}\right)^t
Heterozygosity declines geometrically under drift alone, generation by generation, since each generation's finite sampling has some chance of failing to pass on one of the two alleles at a heterozygous locus; smaller \(N\) makes the per-generation decline factor \((1-1/2N)\) smaller (further from 1), so heterozygosity is lost faster in small populations. B
4
\text{Because } \text{Var}(\Delta p) \propto 1/(2N)\text{, drift's effect is inversely proportional to population size.}
Large populations drift negligibly from one generation to the next; small populations can undergo large, essentially random frequency swings, including outright fixation or loss of an allele within relatively few generations, purely from sampling variance rather than from any fitness difference between alleles. A
5
\text{Effective population size } N_e \text{ (generally} \le \text{census size, reduced by unequal sex ratio, variance in reproductive success, or historical bottlenecks) should replace } N \text{ in the formulas above.}
Real populations rarely have a single constant, evenly reproducing census size; \(N_e\) is the value that correctly predicts a real population's actual observed rate of drift, and can be substantially smaller than the simple census head-count would suggest. B
Result
\text{Var}(\Delta p)=\frac{p(1-p)}{2N_e}, \qquad H_t=H_0\left(1-\frac{1}{2N_e}\right)^t

Reading. Genetic drift's strength scales inversely with effective population size: allele frequencies wander randomly and heterozygosity is lost fastest in small populations, while large populations drift negligibly over any given number of generations.

Scope. Applies to any finite, sexually reproducing population absent other forces; effective population size \(N_e\), not raw census size, should be used for accurate prediction, since the two can differ substantially (Hypotheses, Step 5).

Corollaries & converses
  • gene-flow (migration) directly counteracts drift's tendency toward local fixation or loss by re-injecting variation from other populations, providing the main opposing force to drift's homogenising-within/differentiating-between effect.
  • inbreeding-heterozygosity's decline in a small, closed mating pool is drift's genotypic-level consequence, viewed through the lens of individual relatedness rather than population-wide allele frequency alone.
  • hardy-weinberg's equilibrium is recovered exactly in the idealised \(N\to\infty\) limit, where \(\text{Var}(\Delta p)\to0\); real populations never meet this exactly, which is precisely why some degree of drift is always present in real biology.
Fails without
  • Population size is effectively infinite (Hardy-Weinberg's idealisation): \(\text{Var}(\Delta p)=p(1-p)/(2N)\to0\) as \(N\to\infty\), and drift vanishes entirely; real, finite populations never meet this exactly, which is exactly why drift is present to some nonzero degree in every real biological population, however large.
  • Effective population size is assumed equal to census size when the true \(N_e\) is much smaller (Step 5): predicted rates of heterozygosity loss and allele fixation would be badly underestimated relative to what is actually observed, since high variance in reproductive success or a historical bottleneck can make \(N_e\) far smaller than a simple head-count of living individuals would suggest.
Common errors
  • Treating genetic drift as directional (e.g. "drift pushes frequencies toward 0.5" or toward fixation specifically), when in fact \(E[\Delta p]=0\) — drift is undirected in expectation; only its variance, not its mean, is nonzero (Step 2).
  • Confusing genetic drift with natural selection as an explanation for an observed frequency change; small-sample random fluctuation can mimic a pattern that superficially looks like selection acting, without any actual fitness difference being present.
  • Assuming \(N_e\) equals the census population size, ignoring that unequal sex ratios, high variance in offspring number among individuals, and historical bottlenecks generally make the true effective population size substantially smaller (Step 5).
  • Assuming drift only matters in populations that are "small" in some absolute, fixed sense, rather than recognising its strength is a continuous function of \(1/(2N_e)\) that matters to some measurable degree at any finite population size.
Discussion

Sewall Wright and R.A. Fisher independently developed the foundational mathematical treatment of random genetic drift in the early-to-mid 20th century, and the resulting framework (often called the Wright-Fisher model) remains the standard formal starting point for analysing drift's effect on finite populations. Founder effects (a new population established by a small number of colonising individuals) and population bottlenecks (a temporary, sharp reduction in population size) are both real-world scenarios of especially strong drift, corresponding directly to an episode of unusually small \(N_e\) in Step 5's formula.

Because drift's strength depends on \(N_e\) rather than raw census size, two populations with identical current head counts can nonetheless experience very different rates of drift if their variance in reproductive success, sex ratio, or recent demographic history differ — a point of direct practical importance in conservation genetics, where \(N_e\) rather than census size is the relevant quantity for assessing a small population's genetic risk.

Common misconception: that an allele observed rising in frequency over several generations must be under positive selection. Because drift alone can, by chance, produce a run of frequency increases (or decreases) over a limited number of generations even with \(E[\Delta p]=0\) overall, distinguishing a genuine selective signal from a chance drift trajectory generally requires statistical tests that account explicitly for the expected magnitude of drift at the population's actual effective size.

Worked examples
1
p=0.5,\quad N=25\ (\text{so } 2N=50)
Applying Step 1 of the Proof: \(\text{Var}(\Delta p)=\dfrac{0.5\times0.5}{50}=\dfrac{0.25}{50}=0.005\), giving a standard deviation of roughly \(0.071\) — a single generation's random frequency swing in a population this small is comparable in size to several percentage points of allele frequency, purely from sampling. A
2
H_{10} = H_0\left(1-\frac{1}{50}\right)^{10} = H_0 \times (0.98)^{10} \approx 0.817\,H_0
Applying Step 3 of the Proof over ten generations at this same small population size: heterozygosity has already fallen to roughly 82% of its starting value after only ten generations, illustrating concretely how quickly drift erodes variation in a small population, well before any change in allele frequency is even necessarily visible as fixation. A
N=25 \Rightarrow \text{substantial per-generation frequency variance and measurable heterozygosity loss within ten generations}

Reading. Even a population of a modest, not extremely small, size shows readily measurable drift effects over a comparatively short number of generations — drift is not a phenomenon confined only to extremely tiny populations.

Scope. The identical two formulas, evaluated at any \(N\) (or, more accurately, \(N_e\)) and any number of generations, predict the expected magnitude of drift for that specific case.

Problems
  1. A population has \(N=1000\) and starting allele frequency \(p=0.30\). Compute \(\text{Var}(\Delta p)\) for a single generation, and compare qualitatively to the Worked Example's \(N=25\) case.
    Solution\(\text{Var}(\Delta p)=\dfrac{0.30\times0.70}{2000}=\dfrac{0.21}{2000}=0.000105\), roughly 48 times smaller than the Worked Example's \(N=25\) case (\(0.005\)), consistent with Step 4 of the Proof: variance scales as \(1/(2N)\), so a 40-fold increase in \(N\) (from 25 to 1000) reduces the variance by roughly the same factor.
  2. A conservation biologist estimates a population's census size at 500 individuals, but genetic data suggest its effective population size is closer to 50, due to a small number of males siring most offspring. Explain the practical consequence of using the census size rather than \(N_e\) when predicting this population's rate of heterozygosity loss.
    SolutionUsing the census size of 500 in the heterozygosity-decay formula (Step 3) would predict a much slower rate of heterozygosity loss than the population is actually experiencing, since the true relevant quantity is \(N_e=50\), tenfold smaller (Step 5). The population is genetically behaving like a much smaller one than its head count suggests, and is losing genetic variation, and approaching fixation at individual loci, considerably faster than a census-size-based estimate would indicate — a serious underestimate of genetic risk if not corrected for.
  3. Explain why an allele with no effect on fitness whatsoever can nonetheless become completely fixed (reach a frequency of 1.0) in a small population within a modest number of generations.
    SolutionFixation requires no fitness effect at all under drift; because \(\text{Var}(\Delta p)=p(1-p)/(2N)\) is nonzero for any finite \(N\) (Step 1), and \(E[\Delta p]=0\) means increases and decreases are equally likely at every step (Step 2), a purely neutral allele's frequency performs an unbiased random walk that will, given enough generations, eventually hit either 0 (loss) or 1 (fixation) with a probability set purely by its current frequency and population size — smaller \(N\) makes this random walk's steps proportionally larger relative to the \([0,1]\) frequency range, so fixation or loss happens faster in smaller populations (Step 4), entirely independent of whether the allele affects fitness at all.