Gibbs Entropy and the Maximum-Entropy Derivation of the Ensembles
Statement
Among all probability distributions \( \{p_i\} \) over the microstates of a system, the one that maximizes the Gibbs entropy \( S = -k_B \sum_i p_i \ln p_i \) subject to fixed normalization, fixed average energy, and (where particle exchange is allowed) fixed average particle number, is uniquely the microcanonical, canonical, or grand-canonical distribution respectively. Concretely, maximizing under \( \sum_i p_i = 1 \) alone gives \( p_i = 1/\Omega \); adding \( \sum_i p_i E_i = U \) gives \( p_i = e^{-\beta E_i}/Z \); adding also \( \sum_i p_i N_i = \langle N\rangle \) gives \( p_i = e^{-\beta(E_i-\mu N_i)}/\Xi \), with \( \beta = 1/k_B T \) fixed by the temperature.
Why it matters
The maximum-entropy principle turns the ensembles from postulates justified by counting arguments into consequences of a single inference rule: given only the values of a few conserved averages, the least-biased distribution consistent with them is the one of maximal Gibbs entropy. This is Jaynes' bridge between statistical mechanics and information theory, and it explains why the exponential (Boltzmann) form is universal rather than special.
Practically, the derivation delivers all three canonical ensembles and their partition functions in one stroke, identifies the Lagrange multipliers with the physical fields \( T \) and \( \mu \), and supplies the variational characterization of equilibrium that underlies mean-field theory, the maximum-entropy method in spectroscopy and imaging, and modern approaches to non-equilibrium steady states.
Assumptions
Derivation
Result
Reading. Each ensemble is the maximum-entropy (least-biased) distribution given exactly the macroscopic information imposed on it: nothing but normalization gives the flat microcanonical weight; adding the mean energy tilts the weights into the Boltzmann exponential with slope \( \beta \); adding the mean particle number adds the fugacity factor \( e^{\beta\mu N_i} \). The Lagrange multipliers are not free parameters — they are the inverse temperature and chemical potential that tune the averages to their prescribed values.
Units check. The \( p_i \) are pure numbers. In the exponent \( \beta E_i \) has units \( (\mathrm{J}^{-1})(\mathrm{J}) = 1 \) and \( \beta\mu N_i \) likewise, so every exponential is dimensionless. \( S=-k_B\sum p_i\ln p_i \) carries the units of \( k_B \), namely \( \mathrm{J\,K^{-1}} \), as an entropy must.
Limiting cases
- \( T\to\infty \) (\( \beta\to0 \)): all Boltzmann factors \( \to1 \), the canonical distribution flattens to the microcanonical \( p_i=1/\Omega \), and \( S\to k_B\ln\Omega \) (maximal).
- \( T\to0 \) (\( \beta\to\infty \)): weight collapses onto the ground state(s); \( p_i\to 1/g_0 \) on the \( g_0 \)-fold degenerate ground manifold and \( S\to k_B\ln g_0 \) (third law).
- No energy constraint: the canonical result degenerates exactly to the microcanonical, showing the ensembles are nested by how much macroscopic data is imposed.
- \( \mu\to-\infty \): the fugacity \( e^{\beta\mu}\to0 \) suppresses all \( N_i>0 \) states, \( \Xi\to1 \), and the grand-canonical system empties out (\( \langle N\rangle\to0 \)).
- Dilute / classical limit \( e^{\beta(\mu-\epsilon)}\ll1 \): the grand-canonical occupations reduce to Maxwell–Boltzmann statistics, \( \langle n_\epsilon\rangle\approx e^{-\beta(\epsilon-\mu)} \).
Breaks when
- Long-range interactions / non-additive energy. If \( E \) does not scale extensively (gravitating systems, unscreened Coulomb), the constraint \( \sum p_i E_i=U \) no longer factorizes over subsystems, ensembles become inequivalent, and the canonical construction can give negative specific heat or fail to exist.
- Additional conserved quantities. For integrable or weakly-thermalizing systems the true long-time distribution maximizes entropy subject to all conserved charges, yielding a generalized Gibbs ensemble \( p_i\propto e^{-\sum_a\lambda_a Q_i^{(a)}} \); imposing only \( U \) (and \( N \)) is then simply wrong.
- Small systems / large fluctuations. When \( \Omega \) is not large the ensembles are physically inequivalent (fluctuations of \( E \) or \( N \) are not negligible), and which constraint one imposes materially changes predictions — there is no thermodynamic limit to hide the difference.
- Ill-defined constraints (unbounded spectrum). If \( E_i \) is unbounded above and one tries to fix \( U \) above the spectrum's mean, no normalizable maximum-entropy distribution exists (the would-be \( \beta<0 \) diverges the sum) — as for the kinetic energy of a free particle constrained to a too-high mean.
Failure modes
- Forgetting the normalization multiplier. Dropping \( \alpha \) and only imposing the energy constraint yields an unnormalized \( e^{-\beta E_i} \) and a wrong \( S \); \( \alpha \) is what produces the partition-function denominator.
- Sign error on the chemical-potential term. Writing \( e^{-\beta(E_i+\mu N_i)} \) makes larger \( \mu \) suppress particles; the physically correct grand-canonical weight is \( e^{-\beta(E_i-\mu N_i)} \), so \( \mu>\epsilon \) promotes occupation.
- Treating \( \beta \) as a fitting constant, not \( 1/k_BT \). Failing to identify the multiplier via \( dS=dU/T \) leaves the result thermodynamically meaningless.
- Maximizing \( \sum p_i\ln p_i \) (wrong sign) or minimizing \( S \). This finds a delta-function distribution (minimum entropy), the exact opposite of equilibrium.
- Confusing the Gibbs entropy of the ensemble with the fine-grained entropy under unitary evolution. The maximization applies to the coarse-grained equilibrium distribution; the fine-grained Gibbs entropy of an isolated system is constant in time (Liouville).
- Assuming the maximum could be a saddle. Skipping the concavity check invites the error of thinking multiple stationary distributions compete; the Hessian rules that out.
Discussion
The deepest content of this derivation is inferential: the exponential Boltzmann form is not a special property of thermal systems but the generic answer to "what distribution assumes the least beyond the data?" when the data are averages of additive quantities. The logarithm in \( S \) forces an exponential in the constraints; the additivity of the constrained quantities makes the exponent linear; together they give \( p_i\propto e^{-\sum_a\lambda_a Q_i^{(a)}} \) for whatever conserved charges \( Q^{(a)} \) are fixed. Temperature and chemical potential enter only as the multipliers that enforce the prescribed averages — they are properties of the inference, promoted to physical fields once identified through \( dS=(dU-\mu\,d\langle N\rangle)/T \).
This unifies the three ensembles as a single family indexed by which reservoirs the system is coupled to: an isolated system fixes \( (E,N) \) exactly (microcanonical); coupling to a heat bath trades fixed \( E \) for fixed \( T \) (canonical); coupling to a particle reservoir trades fixed \( N \) for fixed \( \mu \) (grand canonical). Each Legendre-style trade replaces a hard constraint by its conjugate multiplier, and the associated free energy (\( F=-k_BT\ln Z \), \( \Phi=-k_BT\ln\Xi \)) is the Legendre transform of the entropy. The equivalence of ensembles in the thermodynamic limit is then the statement that fluctuations of the traded quantity are relatively negligible.
At the sharpest level the construction exposes why equilibrium statistical mechanics is consistent with reversible microscopic dynamics. The fine-grained Gibbs entropy \( -k_B\int\rho\ln\rho\,d\Gamma \) is a constant of Liouville evolution, so it cannot itself increase; the second law lives in the coarse-grained distribution, and the maximum-entropy distribution is precisely the coarse-graining that retains only the imposed averages and forgets all else. The variational principle is thus a statement about our description, made objective by the enormous phase-space volume that concentrates almost all microstates near the maximum-entropy weights — which is why the same exponentials that we derive as "least-biased inference" also emerge from blind microcanonical counting of a system plus its reservoir.
Common misconceptions. The maximum-entropy distribution is not "the most probable microstate" (that would minimize entropy) — it is the distribution over microstates that is maximally noncommittal given the constraints. And \( \beta \) is not defined to be positive: for systems with a bounded spectrum (spins in a field) population inversion gives a genuine negative-temperature equilibrium, still a bona-fide maximum-entropy state.
Worked examples
Reading. At room temperature the gap is about two \( k_BT \), so the upper level is sparsely but non-negligibly occupied; the entropy \( 0.38\,k_B \) sits below the \( \ln2=0.69 \) that a degenerate pair would give, reflecting the partial ordering imposed by the gap.
Units check. \( \beta\varepsilon \) dimensionless; \( S \) inherits \( k_B=1.381\times10^{-23}\,\mathrm{J\,K^{-1}} \), so \( 0.379\times1.381\times10^{-23}=5.24\times10^{-24}\,\mathrm{J\,K^{-1}} \).
Reading. Because the level sits about \( 3.9\,k_BT \) above the chemical potential, it is almost empty; the maximum-entropy grand-canonical construction reproduces Fermi–Dirac statistics with no extra postulate, and \( \langle n\rangle=1/2 \) exactly when \( \varepsilon=\mu \).
Units check. \( \beta(\varepsilon-\mu) \) is dimensionless (\( \mathrm{eV}/\mathrm{eV} \)); \( \langle n\rangle \) is a pure number in \( [0,1] \), as an occupation of a single fermionic level must be.
Problems
- Microcanonical from maximum entropy. A system has \( n \) accessible microstates and no constraint beyond normalization. Show that \( S \) is maximized by the uniform distribution and evaluate \( S \) for a fair die (\( n=6 \)).
Solution
Maximize \( -\sum_i p_i\ln p_i \) under \( \sum p_i=1 \). Stationarity of \( \mathcal{L}=-\sum p_i\ln p_i-(\alpha-1)(\sum p_i-1) \) gives \( -\ln p_i-\alpha=0 \), i.e. \( p_i=e^{-\alpha} \), the same for all \( i \); normalization gives \( p_i=1/n \). Then \( S=-k_B\,n\cdot\tfrac1n\ln\tfrac1n=k_B\ln n \). For \( n=6 \): \( S=k_B\ln6=1.792\,k_B=1.792\times1.381\times10^{-23}=2.47\times10^{-23}\ \mathrm{J\,K^{-1}} \). - Fixing the temperature from an occupation. For the two-level system (\( 0,\varepsilon \)) with \( \varepsilon=0.10\ \mathrm{eV} \), at what temperature is the excited-state probability \( p_1=0.25 \)?
Solution
\( \dfrac{p_1}{p_0}=e^{-\beta\varepsilon}=\dfrac{0.25}{0.75}=\dfrac13 \Rightarrow \beta\varepsilon=\ln3=1.0986 \). Thus \( k_BT=\varepsilon/\ln3=0.10/1.0986=0.09103\ \mathrm{eV} \), and \( T=0.09103/(8.617\times10^{-5})=1057\ \mathrm{K} \). - Average energy from the partition function. Show from the canonical distribution that \( U=\langle E\rangle=-\dfrac{\partial \ln Z}{\partial\beta} \), and apply it to the two-level system to get \( U(T) \).
Solution
\( \langle E\rangle=\sum_i E_i\dfrac{e^{-\beta E_i}}{Z}=-\dfrac1Z\dfrac{\partial}{\partial\beta}\sum_i e^{-\beta E_i}=-\dfrac{\partial\ln Z}{\partial\beta} \). For \( Z=1+e^{-\beta\varepsilon} \), \( \ln Z=\ln(1+e^{-\beta\varepsilon}) \), so \( U=-\dfrac{-\varepsilon e^{-\beta\varepsilon}}{1+e^{-\beta\varepsilon}}=\dfrac{\varepsilon}{e^{\beta\varepsilon}+1} \). At \( \varepsilon=0.050\ \mathrm{eV},\ \beta\varepsilon=1.934 \): \( U=0.050/(6.92+1)=6.31\times10^{-3}\ \mathrm{eV}=1.01\times10^{-21}\ \mathrm{J} \), consistent with \( U=\varepsilon p_1 \) from worked example 1. - Chemical potential for half-filling. For the single grand-canonical level, find the \( \mu \) that gives \( \langle n\rangle=0.5 \), and interpret.
Solution
\( \langle n\rangle=\dfrac{1}{e^{\beta(\varepsilon-\mu)}+1}=0.5 \Rightarrow e^{\beta(\varepsilon-\mu)}=1 \Rightarrow \varepsilon-\mu=0 \), i.e. \( \mu=\varepsilon \). Half-filling occurs exactly when the chemical potential is level with the state; this is the defining property of the Fermi level at \( T\to0 \) and holds at all \( T \) for a single level by particle-hole symmetry. - Energy fluctuations and heat capacity. Show \( \mathrm{Var}(E)=\dfrac{\partial^2\ln Z}{\partial\beta^2}=k_BT^2 C_V \), and evaluate \( \mathrm{Var}(E) \) for the two-level system at \( \varepsilon=0.050\ \mathrm{eV},\ T=300\ \mathrm{K} \).
Solution
\( \langle E\rangle=-\partial_\beta\ln Z \Rightarrow \partial_\beta^2\ln Z=-\partial_\beta\langle E\rangle=\langle E^2\rangle-\langle E\rangle^2=\mathrm{Var}(E) \). Since \( C_V=\partial U/\partial T=-k_B\beta^2\,\partial_\beta\langle E\rangle=k_B\beta^2\mathrm{Var}(E) \), we get \( \mathrm{Var}(E)=k_BT^2C_V \). For two levels, \( \mathrm{Var}(E)=\varepsilon^2 p_0 p_1 \). With \( p_0=0.8737,\ p_1=0.1263 \) and \( \varepsilon=0.050\ \mathrm{eV} \): \( \mathrm{Var}(E)=(0.050)^2(0.8737)(0.1263)=2.76\times10^{-4}\ \mathrm{eV}^2 \), so the RMS energy fluctuation is \( \sqrt{\mathrm{Var}}=0.0166\ \mathrm{eV}\approx0.64\,k_BT \) — comparable to the gap, as expected for a single small system.