physics2u
Tier
⌕ Search ⌘K
Derivation

Equivalence of Ensembles in the Thermodynamic Limit

Statement

In the thermodynamic limit \(N\to\infty\) with intensive variables (energy density \(e=E/N\), number density \(n=N/V\)) held fixed, the microcanonical, canonical and grand-canonical ensembles yield identical intensive thermodynamics. The equality follows because the canonical partition function is a Laplace transform of the microcanonical density of states whose integrand, written as \(e^{N\varphi}\), is dominated by a single saddle at the energy where the microcanonical temperature equals the bath temperature; the Legendre transform \(f=e^\*-Ts^\*\) then connects the free-energy densities, while the Gaussian width of the saddle gives relative fluctuations \(\Delta E/\langle E\rangle\sim N^{-1/2}\to 0\). The same argument in the particle number links canonical and grand-canonical potentials with \(\Delta N/\langle N\rangle\sim N^{-1/2}\).

Why it matters

Equivalence is the licence that lets us compute in whichever ensemble is analytically easiest — almost always the canonical or grand-canonical, where sums factorise and no awkward energy-shell constraint survives — and still trust the answer for a real, energy-exchanging or particle-exchanging system. Without it, every result would be ensemble-dependent and the phrase "the entropy of the gas" would be ambiguous.

It also draws the exact boundary where that licence is revoked. The proof rests on the concavity of the entropy density; systems that violate it — first-order coexistence, long-range gravitational or unscreened Coulomb systems, small clusters — are precisely those where microcanonical negative heat capacity, ensemble inequivalence and Maxwell constructions appear. Knowing why the ensembles usually agree tells you exactly when they must not.

Assumptions
Short-range interactions, so energy and entropy are extensive:if the potential is long-range (gravity, unscreened Coulomb), \(E\) and \(S\) are not proportional to \(N\), the scaling \(S=Ns(e)\) fails, and the ensembles need not agree.
The entropy density \(s(e)\) is concave (equivalently the microcanonical heat capacity is positive):a convex dip in \(s(e)\) is invisible to the canonical Legendre transform, which replaces it by its concave hull; the two ensembles then disagree across the tie-line.
The saddle-point (Laplace) approximation is asymptotically exact: subleading corrections to \(\ln Z\) are \(O(\ln N)\), i.e. \(o(N)\):if \(\varphi''(e^\*)=0\) (a critical point, where \(C_V\) diverges) the Gaussian width no longer shrinks like \(N^{-1/2}\); fluctuations become anomalous and intensive quantities can still coincide but the simple \(N^{-1/2}\) law breaks.
A unique global maximum of \(\varphi(e)\) exists and is non-degenerate:at a first-order transition two saddles of equal height coexist; the integral is shared between them and the canonical energy sits between the two microcanonical branches, the sharpest form of inequivalence.
Derivation
1
\[ Z(\beta,V,N)=\int_0^\infty dE\,\Omega(E,V,N)\,e^{-\beta E},\qquad \Omega=e^{S(E,V,N)/k_B} \]
The canonical weight is the microcanonical count \(\Omega\) times the Boltzmann factor, summed over energy shells; \(S=k_B\ln\Omega\) is the microcanonical entropy. This is an exact Laplace transform. A
2
\[ Z=\int dE\,\exp\!\left[\frac{S(E)}{k_B}-\beta E\right]=\int N\,de\,\exp\!\Big[N\,\varphi(e)\Big],\quad \varphi(e)\equiv \frac{s(e)}{k_B}-\beta e \]
Insert extensivity \(S=Ns(e)\), \(E=Ne\), and change variable \(E\to e\) (\(dE=N\,de\)). The exponent is now \(N\) times an intensive function \(\varphi\) — the structure a saddle-point demands. A
3
\[ \varphi'(e^\*)=0\ \Longrightarrow\ \frac{1}{k_B}\frac{ds}{de}\bigg|_{e^\*}=\beta \ \Longleftrightarrow\ \frac{\partial S}{\partial E}\bigg|_{E^\*}=\frac{1}{T} \]
For large \(N\) the integrand \(e^{N\varphi}\) is dominated by the maximum of \(\varphi\) (Laplace's method). The stationarity condition identifies the dominant energy \(E^\*=Ne^\*\) as the one whose microcanonical temperature \((\partial S/\partial E)^{-1}\) equals the bath temperature \(T=1/(k_B\beta)\). C
4
\[ \ln Z = N\,\varphi(e^\*) - \tfrac12\ln\!\big(2\pi N|\varphi''(e^\*)|^{-1}\big)+\dots = N\varphi(e^\*)+O(\ln N) \]
Expand \(\varphi(e)=\varphi(e^\*)+\tfrac12\varphi''(e^\*)(e-e^\*)^2+\dots\) with \(\varphi''(e^\*)=s''(e^\*)/k_B<0\) (concavity ensures a maximum) and do the Gaussian integral. The correction is \(O(\ln N)\), negligible against the extensive term. C
5
\[ f\equiv\frac{F}{N}=-\frac{k_BT}{N}\ln Z \xrightarrow{N\to\infty} -k_BT\,\varphi(e^\*)=e^\*-T\,s^\*,\qquad s^\*\equiv s(e^\*) \]
Divide by \(N\), drop the \(O(N^{-1}\ln N)\) term, and substitute \(\varphi(e^\*)=s^\*/k_B-\beta e^\*\). The Helmholtz free-energy density is the Legendre transform of the entropy density — the exact bridge between the two ensembles' potentials. B
6
\[ \operatorname{Var}(E)=\frac{\partial^2\ln Z}{\partial\beta^2}=k_BT^2C_V,\qquad C_V=Nc_v \]
The curvature of the saddle is exactly the canonical energy variance; by the fluctuation–response identity (result energy-fluctuations-and-heat-capacity) it equals \(k_BT^2C_V\), which is extensive because \(C_V=Nc_v\). B
7
\[ \frac{\Delta E}{\langle E\rangle}=\frac{\sqrt{k_BT^2C_V}}{\,Ne\,}=\frac{\sqrt{k_BT^2\,c_v}}{e}\,\frac{1}{\sqrt{N}}\ \propto\ N^{-1/2}\xrightarrow{N\to\infty}0 \]
Numerator \(\propto\sqrt{N}\), denominator \(\propto N\); the ratio vanishes as \(N^{-1/2}\). The canonical energy distribution collapses onto the single value \(E^\*\), so canonical averages equal the microcanonical value at that energy. A
8
\[ \Xi(\beta,\mu,V)=\sum_N e^{\beta\mu N}Z(\beta,V,N)=\sum_N \exp\!\Big[N\big(\beta\mu-\beta f(n)\big)\Big] \]
Repeat the construction one level up: the grand partition function is a discrete Laplace transform of \(Z\) over particle number, with intensive exponent per particle. A
9
\[ \frac{\partial f}{\partial n}\bigg|_{n^\*}=\mu,\qquad \operatorname{Var}(N)=k_BT\,\frac{\partial\langle N\rangle}{\partial\mu}=k_BT\,\rho^2 V\kappa_T,\qquad \frac{\Delta N}{\langle N\rangle}\propto N^{-1/2} \]
The number saddle fixes the chemical potential \(\mu=\partial f/\partial n\) (result grand-canonical-ensemble-and-fugacity); the same fluctuation–response identity gives \(\operatorname{Var}(N)\) in terms of the isothermal compressibility \(\kappa_T\), extensive in \(V\), so the relative number fluctuation again vanishes as \(N^{-1/2}\). C
\[ \boxed{\,s_{\text{micro}}(e)=\frac{f_{\text{can}}(T)-e^\*}{-T}\ \Longleftrightarrow\ f=e^\*-Ts^\*,\qquad \frac{\Delta E}{\langle E\rangle},\ \frac{\Delta N}{\langle N\rangle}\ \sim\ \frac{1}{\sqrt N}\to 0\,} \]

Reading. The three ensembles are related by successive Legendre transforms of intensive potentials — \(s(e)\) to \(f(T)\) to the grand potential density — each exact in the limit because the transforming integral (or sum) is dominated by a single sharp saddle. "Sharp" is quantitative: the relative width of the energy and number distributions shrinks as \(1/\sqrt N\), so every intensive quantity computed in one ensemble equals its value in the others. The choice of ensemble is a matter of convenience, not of physics.

Units check. \(\varphi=s/k_B-\beta e\) is dimensionless (\(s/k_B\) dimensionless; \(\beta e=e/k_BT\) dimensionless), so \(N\varphi\) is a legitimate exponent. \(f=e-Ts\): \([\text{J}]-[\text{K}][\text{J K}^{-1}]=\text{J}\). \(\operatorname{Var}(E)=k_BT^2C_V\): \([\text{J K}^{-1}][\text{K}^2][\text{J K}^{-1}]=\text{J}^2\), so \(\Delta E\) is an energy and \(\Delta E/\langle E\rangle\) is dimensionless.

Limiting cases
  • \(N\to\infty\), fixed \(e,n\): exact equivalence; all three ensembles collapse to the same intensive equation of state.
  • Large but finite \(N\): ensembles agree on intensive quantities to \(O(N^{-1})\); differences appear first in the \(O(\ln N/N)\) corrections to \(f\) and in fluctuation-sensitive observables.
  • Ideal, non-interacting system: saddle is exactly Gaussian, \(\operatorname{Var}(E)=\tfrac32 Nk_B^2T^2\) (monatomic), \(\operatorname{Var}(N)=\langle N\rangle\) (Poisson) — the \(N^{-1/2}\) law is realised with no corrections beyond the Gaussian.
  • At a critical point (\(C_V\to\infty\), \(\kappa_T\to\infty\)): \(\varphi''(e^\*)\to0\); the saddle flattens, fluctuations become long-ranged and the \(N^{-1/2}\) scaling is replaced by an anomalous exponent, though intensive potentials still coincide.
Breaks when
  • Long-range interactions (self-gravitating stars, unscreened plasmas, mean-field spin models with all-to-all coupling). Energy is non-additive, \(S\neq Ns(e)\), and the microcanonical ensemble can display negative heat capacity — impossible canonically. The ensembles genuinely disagree; the astrophysical "gravothermal catastrophe" lives here.
  • First-order phase coexistence, where \(s(e)\) has a convex (non-concave) segment. The canonical Legendre transform returns the concave-hull common tangent (Maxwell construction), so canonical energy jumps discontinuously while the microcanonical entropy interpolates smoothly across the forbidden region — the microcanonical ensemble resolves states the canonical one cannot.
  • Small or mesoscopic systems (clusters, nanoparticles, single molecules). With \(N\) of order \(10^2\)–\(10^4\), relative fluctuations \(\sim N^{-1/2}\) are percent-level or larger; ensemble differences and surface/boundary terms are experimentally resolvable and the "limit" is not reached.
  • Critical points, where \(C_V\) or \(\kappa_T\) diverges. \(\varphi''(e^\*)=0\) invalidates the leading Gaussian; fluctuations no longer self-average as \(N^{-1/2}\) and finite-size scaling, not simple equivalence, governs the approach to the limit.
Failure modes
  • Confusing "energy is fixed" with "energy does not fluctuate". In the canonical ensemble \(E\) fluctuates; what vanishes is the relative width \(\Delta E/\langle E\rangle\), not \(\Delta E\) itself (\(\Delta E\propto\sqrt N\) grows).
  • Claiming the ensembles give identical fluctuations. They do not — microcanonical energy fluctuation is zero by construction. Equivalence is of intensive averages and thermodynamic potentials, not of every moment.
  • Applying equivalence at first-order transitions. Students quote "ensembles are equivalent" to justify a canonical calculation across coexistence, silently discarding the microcanonical negative-\(C_V\) branch. The concavity assumption is exactly what fails there.
  • Forgetting the concavity/stability requirement. Treating \(s(e)\) as automatically concave; without \(C_V>0\) the saddle is a minimum, the Gaussian integral diverges, and the whole construction collapses.
  • Using the \(N^{-1/2}\) law at a critical point. Assuming fluctuations always self-average; near \(T_c\) the diverging susceptibility voids the estimate.
  • Dropping the \(N\,de\) Jacobian or the \(O(\ln N)\) Gaussian prefactor and then worrying it changes \(f\). It does not — it is subextensive — but omitting it silently is still an error of bookkeeping that bites when computing finite-size corrections.
Discussion

The physical content is self-averaging. An extensive additive quantity written as a sum of \(N\) weakly correlated microscopic contributions has, by the central limit theorem, a standard deviation \(\propto\sqrt N\) against a mean \(\propto N\); its density therefore has vanishing relative fluctuation. Ensemble equivalence is simply this statement dressed in the language of Laplace transforms: the "constraint" that distinguishes the microcanonical ensemble (fixed \(E\)) from the canonical (fixed \(T\)) becomes immaterial once \(P(E)\) is a delta-like spike. The mathematical engine is Laplace's method, and the thermodynamic content — that the transform between \(s(e)\) and \(f(T)\) is a Legendre transform — is nothing but the saddle-point condition \(\partial s/\partial e=1/T\).

The construction also explains why concavity of the entropy is not a technical nicety but a physical stability condition. A maximum of \(\varphi\) requires \(s''(e)<0\), and \(s''(e)=-1/(T^2C_V)\cdot(\text{positive})\); a concave \(s\) is exactly the statement \(C_V>0\), i.e. that supplying heat raises temperature. Where this fails the canonical ensemble cannot even represent the state — it Legendre-transforms the convex region away — and the microcanonical description becomes the more fundamental one. This is why long-range and finite systems, far from being pathological curiosities, are the natural home of ensemble inequivalence and of genuinely microcanonical phenomena such as negative specific heat.

There is a rigorous backbone beneath the physicist's saddle-point. The large-deviation principle states \(P(E/N=e)\asymp e^{-N I(e)}\) with rate function \(I(e)=\sup_\beta[\beta e-\beta f(\beta)]-\) (constant), and the Gärtner–Ellis theorem guarantees that when the scaled cumulant generating function \(\lambda(\beta)=\lim_{N\to\infty}N^{-1}\ln\langle e^{-\beta E}\rangle\) is differentiable, the rate function is its Legendre–Fenchel transform and the two are convex-conjugate. Ensemble equivalence at the level of thermodynamic potentials is then precisely the statement that \(I(e)\) and \(\lambda(\beta)\) are Legendre duals; equivalence at the level of macrostates (equal typical configurations) is the stronger condition that \(I\) is strictly convex. The whole theory of ensemble (in)equivalence — including the classification of partial and macrostate equivalence for long-range systems — is the theory of when Legendre–Fenchel duality is or is not involutive.

Common misconceptions. Equivalence does not say the ensembles are the same probability distribution — they are demonstrably different for finite \(N\), and their entropies differ by subextensive terms. It says their intensive predictions coincide in the limit. Nor does it require interactions to be weak; it requires them to be short-ranged and the entropy to be concave. Ideal gases satisfy it trivially, but so do dense liquids and typical solids.

Worked examples
1
Monatomic ideal gas, \(N=6.0\times10^{23}\) atoms at \(T=300\,\text{K}\): how sharp is the canonical energy?
Symbols first. \(E=\tfrac32 Nk_BT\), \(C_V=\tfrac32 Nk_B\), \(\operatorname{Var}(E)=k_BT^2C_V=\tfrac32 Nk_B^2T^2\). Hence \[\frac{\Delta E}{\langle E\rangle}=\frac{\sqrt{\tfrac32 Nk_B^2T^2}}{\tfrac32 Nk_BT}=\sqrt{\frac{2}{3N}}.\] A
2
Numbers: \(\dfrac{\Delta E}{\langle E\rangle}=\sqrt{\dfrac{2}{3\times6.0\times10^{23}}}=\sqrt{1.11\times10^{-24}}=1.05\times10^{-12}.\)
The absolute values, for a units sanity check: \(\langle E\rangle=\tfrac32(6.0\times10^{23})(1.38\times10^{-23})(300)=3.73\times10^{3}\,\text{J}\), \(\Delta E=1.05\times10^{-12}\times3.73\times10^{3}=3.9\times10^{-9}\,\text{J}\). A
\(\Delta E/\langle E\rangle\approx1.1\times10^{-12}\)

Reading. The canonical energy of a mole of gas is defined to better than a part in \(10^{12}\); on any macroscopic scale it is indistinguishable from the microcanonical fixed energy. This is why one never worries which ensemble was used to compute the internal energy of a gas.

Units check. \(\sqrt{k_BT^2C_V}\): \(\sqrt{[\text{J K}^{-1}][\text{K}^2][\text{J K}^{-1}]}=\text{J}\), divided by \(\langle E\rangle\) in J gives a pure number. Correct.

1
Grand-canonical classical ideal gas, \(\langle N\rangle=1.0\times10^{20}\): relative number fluctuation and the implied compressibility.
Symbols first. For the ideal gas the grand partition function gives \(\langle N\rangle=e^{\beta\mu}\,z_1\) with \(\operatorname{Var}(N)=\langle N\rangle\) (Poisson statistics). Using the general identity \(\operatorname{Var}(N)=k_BT\,\rho^2V\kappa_T\) with the ideal-gas \(\kappa_T=1/(\rho k_BT)\) reproduces \(\operatorname{Var}(N)=\rho V=\langle N\rangle\), consistent. Then \[\frac{\Delta N}{\langle N\rangle}=\frac{\sqrt{\langle N\rangle}}{\langle N\rangle}=\frac{1}{\sqrt{\langle N\rangle}}.\] B
2
Numbers: \(\dfrac{\Delta N}{\langle N\rangle}=\dfrac{1}{\sqrt{1.0\times10^{20}}}=\dfrac{1}{1.0\times10^{10}}=1.0\times10^{-10}.\)
Absolute spread: \(\Delta N=\sqrt{10^{20}}=1.0\times10^{10}\) particles — enormous in count, negligible as a fraction. A
\(\Delta N/\langle N\rangle=1.0\times10^{-10}\)

Reading. Opening the system to particle exchange leaves the density fixed to a part in \(10^{10}\). The grand-canonical ensemble, chosen because it factorises quantum statistics cleanly, gives the same intensive equation of state as the fixed-\(N\) canonical ensemble — the \(N^{-1/2}\) collapse in particle number is the number-space analogue of Example 1's energy-space collapse.

Units check. \(\operatorname{Var}(N)=k_BT\rho^2V\kappa_T\): \([\text{J K}^{-1}][\text{K}][\text{m}^{-6}][\text{m}^3][\text{Pa}^{-1}]=\text{J}\,\text{m}^{-3}\,\text{Pa}^{-1}\); with \(\text{Pa}=\text{J m}^{-3}\) this is dimensionless (a pure count squared). Correct.

Problems
  1. Starting from \(Z=\int dE\,\Omega(E)e^{-\beta E}\) and \(S=k_B\ln\Omega\), derive the saddle-point condition and show it is equivalent to the thermodynamic definition \(1/T=\partial S/\partial E\).
    Solution Write \(Z=\int dE\,e^{g(E)}\) with \(g(E)=S(E)/k_B-\beta E\). The integrand is maximised where \(g'(E^\*)=0\), i.e. \(\tfrac1{k_B}S'(E^\*)-\beta=0\), so \(S'(E^\*)=k_B\beta=1/T\). Since \(S'(E)=\partial S/\partial E\), the dominant energy \(E^\*\) is the one at which the microcanonical temperature \((\partial S/\partial E)^{-1}\) equals the bath temperature \(T\). Because \(g''(E^\*)=S''(E^\*)/k_B=-1/(k_BT^2C_V)<0\) when \(C_V>0\), this stationary point is a maximum, justifying Laplace's method.
  2. One mole of a solid obeys the Dulong–Petit law, \(C_V=3Nk_B\). Compute \(\Delta E/\langle E\rangle\) at \(T=300\,\text{K}\), taking the internal (vibrational) energy \(\langle E\rangle=3Nk_BT\).
    Solution \(\operatorname{Var}(E)=k_BT^2C_V=3Nk_B^2T^2\). Then \(\dfrac{\Delta E}{\langle E\rangle}=\dfrac{\sqrt{3Nk_B^2T^2}}{3Nk_BT}=\dfrac{1}{\sqrt{3N}}\). With \(N=6.0\times10^{23}\): \(\dfrac{1}{\sqrt{1.8\times10^{24}}}=\dfrac{1}{1.34\times10^{12}}=7.5\times10^{-13}\). The fraction is again \(\sim N^{-1/2}\); the numerical prefactor \(1/\sqrt3\) differs from the gas's \(\sqrt{2/3}\) only through \(c_v\).
  3. An ideal two-level paramagnet has \(N\) spins with energies \(0\) and \(\varepsilon\). Show that the relative energy fluctuation scales as \(N^{-1/2}\) and find the temperature at which \(\Delta E/\langle E\rangle\) is largest per spin.
    Solution Per spin, \(\langle e\rangle=\varepsilon\,p\) with \(p=\dfrac{e^{-\beta\varepsilon}}{1+e^{-\beta\varepsilon}}\), and \(\operatorname{Var}(e)=\varepsilon^2 p(1-p)\) (Bernoulli). For \(N\) independent spins \(\langle E\rangle=N\varepsilon p\), \(\operatorname{Var}(E)=N\varepsilon^2p(1-p)\), so \[\frac{\Delta E}{\langle E\rangle}=\frac{\varepsilon\sqrt{Np(1-p)}}{N\varepsilon p}=\frac{1}{\sqrt N}\sqrt{\frac{1-p}{p}}.\] The explicit \(N^{-1/2}\). The prefactor \(\sqrt{(1-p)/p}\) grows without bound as \(p\to0\) (\(T\to0\)): at low temperature the mean energy per spin vanishes faster than its fluctuation, so the relative spread diverges — the fluctuation is dominated by rare excited spins. (The variance itself peaks at \(p=\tfrac12\), i.e. \(T\to\infty\).)
  4. For water at \(T=300\,\text{K}\), \(\kappa_T=4.5\times10^{-10}\,\text{Pa}^{-1}\) and number density \(\rho=3.3\times10^{28}\,\text{m}^{-3}\). For a \(1\,\mu\text{m}^3\) observation volume, compute \(\langle N\rangle\) and \(\Delta N/\langle N\rangle\) using \(\operatorname{Var}(N)=k_BT\rho^2V\kappa_T\).
    Solution \(V=1\,\mu\text{m}^3=1.0\times10^{-18}\,\text{m}^3\). \(\langle N\rangle=\rho V=3.3\times10^{28}\times10^{-18}=3.3\times10^{10}\). \(\operatorname{Var}(N)=k_BT\rho^2V\kappa_T=(1.38\times10^{-23})(300)(3.3\times10^{28})^2(10^{-18})(4.5\times10^{-10})\). Compute stepwise: \(k_BT=4.14\times10^{-21}\); \(\rho^2=1.09\times10^{57}\); \(\rho^2V=1.09\times10^{39}\); \(\times k_BT=4.5\times10^{18}\); \(\times\kappa_T=2.0\times10^{9}\). So \(\operatorname{Var}(N)\approx2.0\times10^{9}\), \(\Delta N=4.5\times10^{4}\), and \(\Delta N/\langle N\rangle=4.5\times10^4/3.3\times10^{10}=1.4\times10^{-6}\). Note \(\operatorname{Var}(N)/\langle N\rangle=2.0\times10^9/3.3\times10^{10}=0.06\ll1\): water is far less compressible than an ideal gas (for which the ratio is 1), so its number fluctuations are suppressed below Poisson.
  5. A model system has entropy density \(s(e)\) with a convex segment on \(e\in(e_1,e_2)\) (an entropy "dip"). Explain, using the saddle-point picture, why the canonical ensemble cannot access the microcanonical states in this interval, and identify the thermodynamic signature.
    Solution The canonical free energy density is \(f(\beta)=\min_e[e-Ts(e)]\), equivalently \(\beta f=-\max_e[s(e)/k_B-\beta e]=-\max_e\varphi(e)\). On the convex segment \(\varphi(e)\) has a local minimum, not a maximum, so no value of \(\beta\) selects an \(e^\*\) inside \((e_1,e_2)\): as \(\beta\) is tuned, the global maximum jumps discontinuously from \(e<e_1\) to \(e>e_2\), skipping the dip. The canonical ensemble therefore realises the concave hull (common-tangent / Maxwell construction) and the interval is thermodynamically inaccessible — it corresponds to phase coexistence, and the microcanonical curvature there, \(s''(e)>0\), means \(C_V=-1/(T^2s''e\text{-scaling})<0\): a negative microcanonical heat capacity. This negative specific heat is the hallmark of ensemble inequivalence, seen in self-gravitating systems and in the melting of small clusters.