Equivalence of Ensembles in the Thermodynamic Limit
Statement
In the thermodynamic limit \(N\to\infty\) with intensive variables (energy density \(e=E/N\), number density \(n=N/V\)) held fixed, the microcanonical, canonical and grand-canonical ensembles yield identical intensive thermodynamics. The equality follows because the canonical partition function is a Laplace transform of the microcanonical density of states whose integrand, written as \(e^{N\varphi}\), is dominated by a single saddle at the energy where the microcanonical temperature equals the bath temperature; the Legendre transform \(f=e^\*-Ts^\*\) then connects the free-energy densities, while the Gaussian width of the saddle gives relative fluctuations \(\Delta E/\langle E\rangle\sim N^{-1/2}\to 0\). The same argument in the particle number links canonical and grand-canonical potentials with \(\Delta N/\langle N\rangle\sim N^{-1/2}\).
Why it matters
Equivalence is the licence that lets us compute in whichever ensemble is analytically easiest — almost always the canonical or grand-canonical, where sums factorise and no awkward energy-shell constraint survives — and still trust the answer for a real, energy-exchanging or particle-exchanging system. Without it, every result would be ensemble-dependent and the phrase "the entropy of the gas" would be ambiguous.
It also draws the exact boundary where that licence is revoked. The proof rests on the concavity of the entropy density; systems that violate it — first-order coexistence, long-range gravitational or unscreened Coulomb systems, small clusters — are precisely those where microcanonical negative heat capacity, ensemble inequivalence and Maxwell constructions appear. Knowing why the ensembles usually agree tells you exactly when they must not.
Assumptions
Derivation
Reading. The three ensembles are related by successive Legendre transforms of intensive potentials — \(s(e)\) to \(f(T)\) to the grand potential density — each exact in the limit because the transforming integral (or sum) is dominated by a single sharp saddle. "Sharp" is quantitative: the relative width of the energy and number distributions shrinks as \(1/\sqrt N\), so every intensive quantity computed in one ensemble equals its value in the others. The choice of ensemble is a matter of convenience, not of physics.
Units check. \(\varphi=s/k_B-\beta e\) is dimensionless (\(s/k_B\) dimensionless; \(\beta e=e/k_BT\) dimensionless), so \(N\varphi\) is a legitimate exponent. \(f=e-Ts\): \([\text{J}]-[\text{K}][\text{J K}^{-1}]=\text{J}\). \(\operatorname{Var}(E)=k_BT^2C_V\): \([\text{J K}^{-1}][\text{K}^2][\text{J K}^{-1}]=\text{J}^2\), so \(\Delta E\) is an energy and \(\Delta E/\langle E\rangle\) is dimensionless.
Limiting cases
- \(N\to\infty\), fixed \(e,n\): exact equivalence; all three ensembles collapse to the same intensive equation of state.
- Large but finite \(N\): ensembles agree on intensive quantities to \(O(N^{-1})\); differences appear first in the \(O(\ln N/N)\) corrections to \(f\) and in fluctuation-sensitive observables.
- Ideal, non-interacting system: saddle is exactly Gaussian, \(\operatorname{Var}(E)=\tfrac32 Nk_B^2T^2\) (monatomic), \(\operatorname{Var}(N)=\langle N\rangle\) (Poisson) — the \(N^{-1/2}\) law is realised with no corrections beyond the Gaussian.
- At a critical point (\(C_V\to\infty\), \(\kappa_T\to\infty\)): \(\varphi''(e^\*)\to0\); the saddle flattens, fluctuations become long-ranged and the \(N^{-1/2}\) scaling is replaced by an anomalous exponent, though intensive potentials still coincide.
Breaks when
- Long-range interactions (self-gravitating stars, unscreened plasmas, mean-field spin models with all-to-all coupling). Energy is non-additive, \(S\neq Ns(e)\), and the microcanonical ensemble can display negative heat capacity — impossible canonically. The ensembles genuinely disagree; the astrophysical "gravothermal catastrophe" lives here.
- First-order phase coexistence, where \(s(e)\) has a convex (non-concave) segment. The canonical Legendre transform returns the concave-hull common tangent (Maxwell construction), so canonical energy jumps discontinuously while the microcanonical entropy interpolates smoothly across the forbidden region — the microcanonical ensemble resolves states the canonical one cannot.
- Small or mesoscopic systems (clusters, nanoparticles, single molecules). With \(N\) of order \(10^2\)–\(10^4\), relative fluctuations \(\sim N^{-1/2}\) are percent-level or larger; ensemble differences and surface/boundary terms are experimentally resolvable and the "limit" is not reached.
- Critical points, where \(C_V\) or \(\kappa_T\) diverges. \(\varphi''(e^\*)=0\) invalidates the leading Gaussian; fluctuations no longer self-average as \(N^{-1/2}\) and finite-size scaling, not simple equivalence, governs the approach to the limit.
Failure modes
- Confusing "energy is fixed" with "energy does not fluctuate". In the canonical ensemble \(E\) fluctuates; what vanishes is the relative width \(\Delta E/\langle E\rangle\), not \(\Delta E\) itself (\(\Delta E\propto\sqrt N\) grows).
- Claiming the ensembles give identical fluctuations. They do not — microcanonical energy fluctuation is zero by construction. Equivalence is of intensive averages and thermodynamic potentials, not of every moment.
- Applying equivalence at first-order transitions. Students quote "ensembles are equivalent" to justify a canonical calculation across coexistence, silently discarding the microcanonical negative-\(C_V\) branch. The concavity assumption is exactly what fails there.
- Forgetting the concavity/stability requirement. Treating \(s(e)\) as automatically concave; without \(C_V>0\) the saddle is a minimum, the Gaussian integral diverges, and the whole construction collapses.
- Using the \(N^{-1/2}\) law at a critical point. Assuming fluctuations always self-average; near \(T_c\) the diverging susceptibility voids the estimate.
- Dropping the \(N\,de\) Jacobian or the \(O(\ln N)\) Gaussian prefactor and then worrying it changes \(f\). It does not — it is subextensive — but omitting it silently is still an error of bookkeeping that bites when computing finite-size corrections.
Discussion
The physical content is self-averaging. An extensive additive quantity written as a sum of \(N\) weakly correlated microscopic contributions has, by the central limit theorem, a standard deviation \(\propto\sqrt N\) against a mean \(\propto N\); its density therefore has vanishing relative fluctuation. Ensemble equivalence is simply this statement dressed in the language of Laplace transforms: the "constraint" that distinguishes the microcanonical ensemble (fixed \(E\)) from the canonical (fixed \(T\)) becomes immaterial once \(P(E)\) is a delta-like spike. The mathematical engine is Laplace's method, and the thermodynamic content — that the transform between \(s(e)\) and \(f(T)\) is a Legendre transform — is nothing but the saddle-point condition \(\partial s/\partial e=1/T\).
The construction also explains why concavity of the entropy is not a technical nicety but a physical stability condition. A maximum of \(\varphi\) requires \(s''(e)<0\), and \(s''(e)=-1/(T^2C_V)\cdot(\text{positive})\); a concave \(s\) is exactly the statement \(C_V>0\), i.e. that supplying heat raises temperature. Where this fails the canonical ensemble cannot even represent the state — it Legendre-transforms the convex region away — and the microcanonical description becomes the more fundamental one. This is why long-range and finite systems, far from being pathological curiosities, are the natural home of ensemble inequivalence and of genuinely microcanonical phenomena such as negative specific heat.
There is a rigorous backbone beneath the physicist's saddle-point. The large-deviation principle states \(P(E/N=e)\asymp e^{-N I(e)}\) with rate function \(I(e)=\sup_\beta[\beta e-\beta f(\beta)]-\) (constant), and the Gärtner–Ellis theorem guarantees that when the scaled cumulant generating function \(\lambda(\beta)=\lim_{N\to\infty}N^{-1}\ln\langle e^{-\beta E}\rangle\) is differentiable, the rate function is its Legendre–Fenchel transform and the two are convex-conjugate. Ensemble equivalence at the level of thermodynamic potentials is then precisely the statement that \(I(e)\) and \(\lambda(\beta)\) are Legendre duals; equivalence at the level of macrostates (equal typical configurations) is the stronger condition that \(I\) is strictly convex. The whole theory of ensemble (in)equivalence — including the classification of partial and macrostate equivalence for long-range systems — is the theory of when Legendre–Fenchel duality is or is not involutive.
Common misconceptions. Equivalence does not say the ensembles are the same probability distribution — they are demonstrably different for finite \(N\), and their entropies differ by subextensive terms. It says their intensive predictions coincide in the limit. Nor does it require interactions to be weak; it requires them to be short-ranged and the entropy to be concave. Ideal gases satisfy it trivially, but so do dense liquids and typical solids.
Worked examples
Reading. The canonical energy of a mole of gas is defined to better than a part in \(10^{12}\); on any macroscopic scale it is indistinguishable from the microcanonical fixed energy. This is why one never worries which ensemble was used to compute the internal energy of a gas.
Units check. \(\sqrt{k_BT^2C_V}\): \(\sqrt{[\text{J K}^{-1}][\text{K}^2][\text{J K}^{-1}]}=\text{J}\), divided by \(\langle E\rangle\) in J gives a pure number. Correct.
Reading. Opening the system to particle exchange leaves the density fixed to a part in \(10^{10}\). The grand-canonical ensemble, chosen because it factorises quantum statistics cleanly, gives the same intensive equation of state as the fixed-\(N\) canonical ensemble — the \(N^{-1/2}\) collapse in particle number is the number-space analogue of Example 1's energy-space collapse.
Units check. \(\operatorname{Var}(N)=k_BT\rho^2V\kappa_T\): \([\text{J K}^{-1}][\text{K}][\text{m}^{-6}][\text{m}^3][\text{Pa}^{-1}]=\text{J}\,\text{m}^{-3}\,\text{Pa}^{-1}\); with \(\text{Pa}=\text{J m}^{-3}\) this is dimensionless (a pure count squared). Correct.
Problems
- Starting from \(Z=\int dE\,\Omega(E)e^{-\beta E}\) and \(S=k_B\ln\Omega\), derive the saddle-point condition and show it is equivalent to the thermodynamic definition \(1/T=\partial S/\partial E\).
Solution
Write \(Z=\int dE\,e^{g(E)}\) with \(g(E)=S(E)/k_B-\beta E\). The integrand is maximised where \(g'(E^\*)=0\), i.e. \(\tfrac1{k_B}S'(E^\*)-\beta=0\), so \(S'(E^\*)=k_B\beta=1/T\). Since \(S'(E)=\partial S/\partial E\), the dominant energy \(E^\*\) is the one at which the microcanonical temperature \((\partial S/\partial E)^{-1}\) equals the bath temperature \(T\). Because \(g''(E^\*)=S''(E^\*)/k_B=-1/(k_BT^2C_V)<0\) when \(C_V>0\), this stationary point is a maximum, justifying Laplace's method. - One mole of a solid obeys the Dulong–Petit law, \(C_V=3Nk_B\). Compute \(\Delta E/\langle E\rangle\) at \(T=300\,\text{K}\), taking the internal (vibrational) energy \(\langle E\rangle=3Nk_BT\).
Solution
\(\operatorname{Var}(E)=k_BT^2C_V=3Nk_B^2T^2\). Then \(\dfrac{\Delta E}{\langle E\rangle}=\dfrac{\sqrt{3Nk_B^2T^2}}{3Nk_BT}=\dfrac{1}{\sqrt{3N}}\). With \(N=6.0\times10^{23}\): \(\dfrac{1}{\sqrt{1.8\times10^{24}}}=\dfrac{1}{1.34\times10^{12}}=7.5\times10^{-13}\). The fraction is again \(\sim N^{-1/2}\); the numerical prefactor \(1/\sqrt3\) differs from the gas's \(\sqrt{2/3}\) only through \(c_v\). - An ideal two-level paramagnet has \(N\) spins with energies \(0\) and \(\varepsilon\). Show that the relative energy fluctuation scales as \(N^{-1/2}\) and find the temperature at which \(\Delta E/\langle E\rangle\) is largest per spin.
Solution
Per spin, \(\langle e\rangle=\varepsilon\,p\) with \(p=\dfrac{e^{-\beta\varepsilon}}{1+e^{-\beta\varepsilon}}\), and \(\operatorname{Var}(e)=\varepsilon^2 p(1-p)\) (Bernoulli). For \(N\) independent spins \(\langle E\rangle=N\varepsilon p\), \(\operatorname{Var}(E)=N\varepsilon^2p(1-p)\), so \[\frac{\Delta E}{\langle E\rangle}=\frac{\varepsilon\sqrt{Np(1-p)}}{N\varepsilon p}=\frac{1}{\sqrt N}\sqrt{\frac{1-p}{p}}.\] The explicit \(N^{-1/2}\). The prefactor \(\sqrt{(1-p)/p}\) grows without bound as \(p\to0\) (\(T\to0\)): at low temperature the mean energy per spin vanishes faster than its fluctuation, so the relative spread diverges — the fluctuation is dominated by rare excited spins. (The variance itself peaks at \(p=\tfrac12\), i.e. \(T\to\infty\).) - For water at \(T=300\,\text{K}\), \(\kappa_T=4.5\times10^{-10}\,\text{Pa}^{-1}\) and number density \(\rho=3.3\times10^{28}\,\text{m}^{-3}\). For a \(1\,\mu\text{m}^3\) observation volume, compute \(\langle N\rangle\) and \(\Delta N/\langle N\rangle\) using \(\operatorname{Var}(N)=k_BT\rho^2V\kappa_T\).
Solution
\(V=1\,\mu\text{m}^3=1.0\times10^{-18}\,\text{m}^3\). \(\langle N\rangle=\rho V=3.3\times10^{28}\times10^{-18}=3.3\times10^{10}\). \(\operatorname{Var}(N)=k_BT\rho^2V\kappa_T=(1.38\times10^{-23})(300)(3.3\times10^{28})^2(10^{-18})(4.5\times10^{-10})\). Compute stepwise: \(k_BT=4.14\times10^{-21}\); \(\rho^2=1.09\times10^{57}\); \(\rho^2V=1.09\times10^{39}\); \(\times k_BT=4.5\times10^{18}\); \(\times\kappa_T=2.0\times10^{9}\). So \(\operatorname{Var}(N)\approx2.0\times10^{9}\), \(\Delta N=4.5\times10^{4}\), and \(\Delta N/\langle N\rangle=4.5\times10^4/3.3\times10^{10}=1.4\times10^{-6}\). Note \(\operatorname{Var}(N)/\langle N\rangle=2.0\times10^9/3.3\times10^{10}=0.06\ll1\): water is far less compressible than an ideal gas (for which the ratio is 1), so its number fluctuations are suppressed below Poisson. - A model system has entropy density \(s(e)\) with a convex segment on \(e\in(e_1,e_2)\) (an entropy "dip"). Explain, using the saddle-point picture, why the canonical ensemble cannot access the microcanonical states in this interval, and identify the thermodynamic signature.
Solution
The canonical free energy density is \(f(\beta)=\min_e[e-Ts(e)]\), equivalently \(\beta f=-\max_e[s(e)/k_B-\beta e]=-\max_e\varphi(e)\). On the convex segment \(\varphi(e)\) has a local minimum, not a maximum, so no value of \(\beta\) selects an \(e^\*\) inside \((e_1,e_2)\): as \(\beta\) is tuned, the global maximum jumps discontinuously from \(e<e_1\) to \(e>e_2\), skipping the dip. The canonical ensemble therefore realises the concave hull (common-tangent / Maxwell construction) and the interval is thermodynamically inaccessible — it corresponds to phase coexistence, and the microcanonical curvature there, \(s''(e)>0\), means \(C_V=-1/(T^2s''e\text{-scaling})<0\): a negative microcanonical heat capacity. This negative specific heat is the hallmark of ensemble inequivalence, seen in self-gravitating systems and in the melting of small clusters.