physics2u
Tier
⌕ Search ⌘K
Derivation

Boltzmann Distribution from a Heat Reservoir

Statement

For a system \(S\) that exchanges energy with a large reservoir \(R\) at temperature \(T\), with the composite \(S+R\) isolated at fixed total energy, the probability that \(S\) occupies one specified microstate of energy \(E_i\) is \(P(i)=\dfrac{1}{Z}\,e^{-E_i/k_B T}\), where \(Z=\sum_j e^{-E_j/k_B T}\). This is obtained by expanding the reservoir entropy \(S_R(E_{\text{tot}}-E_i)\) to first order in \(E_i\).

Why it matters

The Boltzmann (Gibbs) factor \(e^{-E/k_B T}\) is the single most-used weight in all of statistical physics: it fixes the population of every quantum level, every conformation of a molecule, every spin in a magnet, and every rung of a chemical reaction coordinate. Almost the entire canonical ensemble, and with it chemistry's Arrhenius law, semiconductor carrier statistics, and the barometric formula, is a corollary of this one distribution.

The derivation is conceptually important because it shows that a small system is not at maximum entropy on its own; instead it inherits its statistics from the reservoir's overwhelming multiplicity. The exponential is nothing more than a Taylor expansion of the log of a very large number, evaluated at the temperature that the reservoir defines.

Assumptions
The composite \(S+R\) is isolated with fixed total energy \(E_{\text{tot}}\).Without a hard constraint \(E_S+E_R=E_{\text{tot}}\) the reservoir count \(\Omega_R(E_{\text{tot}}-E_i)\) is not a function of \(E_i\) alone, and the microcanonical equal-a-priori-probability rule that underlies the whole argument does not apply. Every accessible microstate of the isolated composite is equally probable.Drop this and \(P(i)\) is no longer proportional to \(\Omega_R\); the weight would depend on unknown dynamical details rather than on counting, and the clean exponential is lost. The reservoir is much larger than the system: \(E_i \ll E_{\text{tot}}\) and \(R\) has many more degrees of freedom.If the reservoir is comparable in size the first-order truncation fails; higher derivatives of \(S_R\) survive and introduce a finite heat-capacity correction, giving a non-exponential weight. The reservoir entropy \(S_R(E_R)\) is smooth and analytic near \(E_R=E_{\text{tot}}\), with a well-defined \(\partial S_R/\partial E_R\).At a phase transition \(S_R(E_R)\) can have a kink or discontinuous slope; the Taylor step is then illegitimate and \(T\) is ambiguous, so no single Boltzmann factor describes the system. The system energy \(E_i\) labels a single microstate, not an energy level.If \(E_i\) is a degenerate level of degeneracy \(g_i\), the microstate probabilities still carry \(e^{-E_i/k_BT}\), but the level probability acquires a factor \(g_i\); conflating the two double-counts or omits the degeneracy.
Derivation
1
\[ E_S + E_R = E_{\text{tot}} \quad\Rightarrow\quad E_R = E_{\text{tot}} - E_i \]
Energy conservation for the isolated composite: fixing the system in microstate \(i\) of energy \(E_i\) forces the reservoir energy. A
2
\[ P(i) \;\propto\; \Omega_R\!\left(E_{\text{tot}}-E_i\right) \]
Equal a-priori probabilities: the probability of a composite configuration is proportional to the number of reservoir microstates compatible with system microstate \(i\). Since \(S\) is pinned to one microstate, its own count is \(1\). A
3
\[ \Omega_R(E_R) = \exp\!\left[\frac{S_R(E_R)}{k_B}\right] \]
Invert the Boltzmann entropy relation \(S_R=k_B\ln\Omega_R\) (prior result: Boltzmann entropy from microstate counting) so the enormous count is written as an exponential of an extensive quantity. A
4
\[ S_R(E_{\text{tot}}-E_i) = S_R(E_{\text{tot}}) - E_i\left.\frac{\partial S_R}{\partial E_R}\right|_{E_{\text{tot}}} + \frac{1}{2}E_i^2\left.\frac{\partial^2 S_R}{\partial E_R^2}\right|_{E_{\text{tot}}} - \cdots \]
Taylor-expand the reservoir entropy in the small quantity \(E_i\) about the full energy \(E_{\text{tot}}\). Legal because \(E_i\ll E_{\text{tot}}\) and \(S_R\) is assumed smooth. Expanding the slowly varying \(\ln\Omega_R\) rather than \(\Omega_R\) itself is what makes truncation accurate. B
5
\[ \frac{1}{2}E_i^2\,\frac{\partial^2 S_R}{\partial E_R^2} = -\frac{E_i^2}{2C_R T^2}\;\longrightarrow\;0 \quad(C_R\to\infty) \]
The second-order term is negligible: using \(\partial S_R/\partial E_R = 1/T\) and \(\partial(1/T)/\partial E_R = -1/(T^2 C_R)\), the curvature scales as \(1/C_R\) and vanishes for a macroscopic reservoir heat capacity. This is the formal meaning of "large reservoir." C
6
\[ \left.\frac{\partial S_R}{\partial E_R}\right|_{E_{\text{tot}}} \equiv \frac{1}{T} \]
Definition of thermodynamic temperature from the entropy (prior result: temperature from entropy maximization). \(T\) is the reservoir's temperature evaluated at \(E_{\text{tot}}\), essentially unchanged by the tiny \(E_i\). A
7
\[ S_R(E_{\text{tot}}-E_i) \approx S_R(E_{\text{tot}}) - \frac{E_i}{T} \]
Substitute step 6 and drop the second-order term (step 5). The result is linear in \(E_i\). A
8
\[ P(i) \propto \exp\!\left[\frac{S_R(E_{\text{tot}})}{k_B}\right]\exp\!\left[-\frac{E_i}{k_B T}\right] \]
Insert step 7 into step 3 and factor the exponential. The first factor is a constant independent of \(i\); all system dependence is in the second. B
9
\[ P(i) = \frac{e^{-E_i/k_B T}}{\displaystyle\sum_j e^{-E_j/k_B T}} \equiv \frac{e^{-E_i/k_B T}}{Z} \]
Absorb the \(i\)-independent constant into a normalisation. Requiring \(\sum_i P(i)=1\) fixes it and defines the partition function \(Z\). A
Result
\[ P(i)=\frac{1}{Z}\,e^{-E_i/k_B T},\qquad Z=\sum_j e^{-E_j/k_B T} \]

Reading. The probability of finding the system in a single microstate falls off exponentially with that state's energy, on the scale \(k_B T\). States more than a few \(k_B T\) above the ground state are exponentially rare; states within \(k_B T\) are appreciably populated. The reservoir never appears explicitly — its entire influence is compressed into the single number \(T\). Note this is the probability of one microstate; a level of degeneracy \(g_i\) carries \(g_i e^{-E_i/k_B T}/Z\).

Units check. \([E_i]=\mathrm{J}\), \([k_B]=\mathrm{J\,K^{-1}}\), \([T]=\mathrm{K}\), so \(E_i/(k_B T)\) is dimensionless and the exponential is well defined. \(Z\) is dimensionless, and \(P(i)\) is a pure number in \([0,1]\), as a probability must be.

Limiting cases
  • \(T\to 0\): \(k_B T\to 0\), the exponential collapses onto the lowest-energy microstate(s); \(P\to 1\) for the ground state, \(0\) for all excited states (the system freezes into its ground state).
  • \(T\to\infty\): every \(E_i/k_B T\to 0\), all Boltzmann factors \(\to 1\), and \(P(i)\to 1/\mathcal{N}\) — the uniform (microcanonical-like) distribution over all \(\mathcal{N}\) accessible microstates.
  • Two states, gap \(\Delta E\): the population ratio is \(P_2/P_1=e^{-\Delta E/k_B T}\); at \(k_B T=\Delta E\) the upper state holds a fraction \(1/(1+e)\approx 0.27\).
  • Uniform energy shift \(E_i\to E_i+E_0\): the constant \(E_0\) cancels between numerator and \(Z\); only energy differences matter, so the zero of energy is arbitrary.
Breaks when
  • The "reservoir" is small or comparable to the system. The neglected second-order term \(-E_i^2/(2C_R T^2)\) is no longer negligible; the weight becomes \(e^{-E_i/k_B T}\,e^{-E_i^2/2C_R k_B T^2}\), a Gaussian-corrected non-canonical distribution.
  • The reservoir sits at a first-order phase transition. \(S_R(E_R)\) has a straight segment / kink, \(\partial S_R/\partial E_R\) is not single-valued, and no unique \(T\) can be assigned; the Taylor step in the derivation is illegal.
  • System and reservoir exchange more than energy without accounting for it. If particles or volume are also exchanged, \(\Omega_R\) depends on \(N_i,V_i\) too; expanding only in \(E_i\) omits the \(-\mu N_i/T\) and \(pV_i/T\) terms and gives the wrong (non-grand-canonical) weight.
  • Strong system–reservoir coupling. The additivity \(E_S+E_R=E_{\text{tot}}\) assumed a negligible interaction energy; if the coupling energy is comparable to \(E_i\), \(\Omega_R\) is not a function of \(E_{\text{tot}}-E_i\) alone and the Boltzmann form fails.
Failure modes
  • Maximising the system's own entropy. Students try to maximise \(S_S\) subject to \(\langle E\rangle\) and reproduce the answer, but forget that the distribution comes from the reservoir's multiplicity; the two routes agree only because of the constraint, not by coincidence to be assumed.
  • Expanding \(\Omega_R\) instead of \(\ln\Omega_R\). Taylor-expanding the raw count \(\Omega_R\) is hopeless — it varies over dozens of orders of magnitude. The linearisation is legitimate only for \(S_R=k_B\ln\Omega_R\).
  • Confusing microstate and energy-level probability. Writing \(P(E_i)\propto e^{-E_i/k_B T}\) for a degenerate level while forgetting the degeneracy factor \(g_i\); this misplaces populations whenever levels are degenerate (e.g. atomic \(J\) multiplets).
  • Treating \(Z\) as physically dimensionful or state-specific. Believing \(Z\) depends on \(i\); it is a single sum over all states and cancels the reservoir constant \(e^{S_R(E_{\text{tot}})/k_B}\).
  • Using \(T\) of the system. Inserting the system's "temperature" — a small system may not even have one; \(T\) here is unambiguously the reservoir's.
  • Forgetting energy differences only. Panicking that \(E_i\) has no absolute zero; the offset cancels in \(P(i)\), so any convenient reference is fine.
Discussion

The physical heart of the derivation is a competition the system cannot see directly. When the system takes up energy \(E_i\), it removes exactly that energy from the reservoir, and the reservoir responds by losing multiplicity at the rate \(\partial S_R/\partial E_R = 1/T\). A high-energy system microstate is disfavoured not because it is intrinsically unlikely but because it starves the reservoir of the states it would otherwise have. The exponential is the reservoir's density of states, viewed through the logarithm.

Because only the first derivative of \(S_R\) survives, the reservoir is fully summarised by one intensive parameter, \(T\). This is why temperature is such a powerful idea: an object with \(10^{23}\) degrees of freedom communicates with a small system through a single number. The same logic, expanded in additional conserved quantities, generates the grand-canonical weight \(e^{-(E_i-\mu N_i)/k_B T}\) and the isothermal–isobaric weight — chemical potential and pressure are simply the first derivatives of \(S_R\) with respect to particle number and volume.

At the level of large-deviation theory the derivation is exact rather than approximate: the canonical distribution is the \(N\to\infty\) contraction of the microcanonical one, and the neglected curvature term \(-E_i^2/(2C_R k_B T^2)\) is precisely the leading finite-size correction, controlling the \(O(1/N)\) inequivalence of ensembles and the width of energy fluctuations, \(\langle\Delta E^2\rangle = k_B T^2 C\). The Legendre structure connecting the microcanonical entropy \(S(E)\) and the canonical free energy \(F(T)=-k_BT\ln Z\) is the thermodynamic shadow of this contraction, and it breaks — ensembles become inequivalent — exactly where \(S(E)\) is non-concave, i.e. at first-order transitions.

Common misconceptions. The Boltzmann factor does not say high energies are impossible, only exponentially costly; with enough available states (large \(g_i\) or a rising density of states) the most probable energy can lie far above the ground state, since observed populations weigh \(g_i e^{-E_i/k_BT}\). It is also not a statement about the system reaching maximum entropy — the system's entropy is generally sub-maximal; it is the total entropy of system-plus-reservoir that is maximised.

Worked examples
1
\[ \text{Two-level system: }E_0=0,\;E_1=\Delta,\quad \frac{P_1}{P_0}=e^{-\Delta/k_B T} \]
Set up the population ratio for a spin-\(\tfrac12\) magnetic moment in a field, with level splitting \(\Delta\). Symbols first, numbers next. A
2
\[ \Delta = g\mu_B B,\qquad g=2,\;\mu_B=9.274\times10^{-24}\ \mathrm{J\,T^{-1}},\;B=1.00\ \mathrm{T} \]
Choose a concrete electron-spin splitting in a \(1\ \mathrm{T}\) field. A
3
\[ \Delta = 2\,(9.274\times10^{-24})(1.00)=1.855\times10^{-23}\ \mathrm{J} \]
Evaluate the gap. A
4
\[ k_B T = (1.381\times10^{-23})(300)=4.14\times10^{-21}\ \mathrm{J},\quad \frac{\Delta}{k_B T}=4.48\times10^{-3} \]
Compare with the thermal scale at \(T=300\ \mathrm{K}\). The ratio is tiny — the splitting is far below room-temperature thermal energy. A
5
\[ \frac{P_1}{P_0}=e^{-4.48\times10^{-3}}=0.99553 \]
Exponentiate. The two spin states are almost equally populated, with a fractional excess in the lower state of about \(0.45\%\). A
\[ \frac{P_1}{P_0}=e^{-\Delta/k_BT}\approx 0.9955,\quad \frac{P_0-P_1}{P_0+P_1}\approx 2.2\times10^{-3} \]

Reading. At room temperature an electron spin in \(1\ \mathrm{T}\) is only \(\sim0.2\%\) polarised — precisely why magnetic-resonance signals are weak and why low temperatures or high fields boost them.

Units check. \(g\mu_B B\) has units \(\mathrm{(J\,T^{-1})(T)=J}\); divided by \(k_BT\) in \(\mathrm{J}\) it is dimensionless, and the ratio \(P_1/P_0\) is a pure number.

1
\[ \frac{N(v)}{N(0)}=\exp\!\left[-\frac{E(v)-E(0)}{k_B T}\right],\quad E(v)=\hbar\omega\left(v+\tfrac12\right) \]
Vibrational population of the first excited state \(v=1\) of a diatomic (harmonic) molecule relative to the ground state \(v=0\). Symbols before numbers. A
2
\[ E(1)-E(0)=\hbar\omega,\qquad \frac{N_1}{N_0}=e^{-\hbar\omega/k_B T} \]
The zero-point term \(\tfrac12\hbar\omega\) cancels; only the level spacing \(\hbar\omega\) survives. A
3
\[ \tilde{\nu}=2143\ \mathrm{cm^{-1}}\ (\mathrm{CO}),\quad \hbar\omega=hc\tilde{\nu}=(6.626\times10^{-34})(3.00\times10^{10})(2143) \]
Use the CO stretch wavenumber; convert with \(c=3.00\times10^{10}\ \mathrm{cm\,s^{-1}}\). A
4
\[ \hbar\omega=4.26\times10^{-20}\ \mathrm{J},\qquad k_B T=(1.381\times10^{-23})(300)=4.14\times10^{-21}\ \mathrm{J} \]
Evaluate both energies at \(T=300\ \mathrm{K}\). The vibrational quantum is about ten times the thermal energy. A
5
\[ \frac{\hbar\omega}{k_B T}=\frac{4.26\times10^{-20}}{4.14\times10^{-21}}=10.3,\qquad \frac{N_1}{N_0}=e^{-10.3} \]
Form the dimensionless ratio and exponentiate. A
\[ \frac{N_1}{N_0}=e^{-10.3}\approx 3.4\times10^{-5} \]

Reading. Only about three CO molecules in \(10^5\) are vibrationally excited at room temperature — vibrational modes are effectively "frozen out," which is why they contribute almost nothing to the molar heat capacity at \(300\ \mathrm{K}\).

Units check. \(hc\tilde\nu\) has units \(\mathrm{(J\,s)(cm\,s^{-1})(cm^{-1})=J}\); divided by \(k_BT\) it is dimensionless. \(N_1/N_0\) is a pure number.

Problems
  1. A three-level system has energies \(0,\ \varepsilon,\ 2\varepsilon\), all non-degenerate. Write \(Z\) and find \(P(2\varepsilon)\) at \(k_B T=\varepsilon\).
    Solution With \(x=e^{-\varepsilon/k_BT}\), \(Z=1+x+x^2\). At \(k_BT=\varepsilon\), \(x=e^{-1}=0.3679\), so \(x^2=0.1353\) and \(Z=1+0.3679+0.1353=1.5032\). Then \(P(2\varepsilon)=x^2/Z=0.1353/1.5032=\boxed{0.0900}\), about \(9\%\).
  2. An atomic level lies \(\Delta=2.0\ \mathrm{eV}\) above the ground state and is \(3\)-fold degenerate (the ground state is non-degenerate). Find the population ratio \(N_{\text{exc}}/N_0\) at \(T=6000\ \mathrm{K}\) (photosphere of the Sun).
    Solution \(N_{\text{exc}}/N_0=g\,e^{-\Delta/k_BT}\). \(k_BT=(8.617\times10^{-5}\ \mathrm{eV\,K^{-1}})(6000)=0.5170\ \mathrm{eV}\). \(\Delta/k_BT=2.0/0.5170=3.868\). \(e^{-3.868}=0.02090\). Multiply by \(g=3\): \(N_{\text{exc}}/N_0=3(0.02090)=\boxed{0.0627}\), about \(6\%\).
  3. Show that for a reservoir of finite heat capacity \(C\) the next correction to the Boltzmann factor is \(\exp[-E^2/(2Ck_BT^2)]\), and estimate its size for \(E=k_BT\) when \(C=10^{3}k_B\).
    Solution Expanding \(S_R\) to second order, the extra term in \(S_R/k_B\) is \(\tfrac12 E^2\,\partial^2 S_R/\partial E_R^2/k_B\). Using \(\partial S_R/\partial E_R=1/T\) and \(\partial^2 S_R/\partial E_R^2=\partial(1/T)/\partial E_R=-(1/T^2)(\partial T/\partial E_R)=-1/(T^2 C)\), the correction to \(\ln P\) is \(-E^2/(2Ck_BT^2)\), i.e. a multiplicative factor \(\exp[-E^2/(2Ck_BT^2)]\). For \(E=k_BT\): exponent \(=-(k_BT)^2/(2Ck_BT^2)=-1/(2C/k_B)=-1/(2\times10^3)=-5\times10^{-4}\). Factor \(\approx e^{-0.0005}=0.9995\) — a \(0.05\%\) deviation, negligible, confirming the large-reservoir approximation.
  4. The barometric formula follows from the Boltzmann distribution with \(E=mgh\). For \(\mathrm{N_2}\) (\(m=4.65\times10^{-26}\ \mathrm{kg}\)) at \(T=288\ \mathrm{K}\), find the height at which the number density falls to \(1/e\) of its ground value.
    Solution \(n(h)/n(0)=e^{-mgh/k_BT}\); this equals \(1/e\) when \(mgh=k_BT\), so \(h=k_BT/(mg)\). \(k_BT=(1.381\times10^{-23})(288)=3.977\times10^{-21}\ \mathrm{J}\). \(mg=(4.65\times10^{-26})(9.81)=4.562\times10^{-25}\ \mathrm{N}\). \(h=3.977\times10^{-21}/4.562\times10^{-25}=\boxed{8.72\times10^{3}\ \mathrm{m}}\approx 8.7\ \mathrm{km}\), the scale height of the atmosphere.
  5. A protein has a folded state (energy \(0\), non-degenerate) and \(1000\) unfolded conformations each of energy \(\varepsilon=0.15\ \mathrm{eV}\). At what temperature are folded and unfolded states equally populated?
    Solution Equal total populations require \(1\cdot e^{0}=g\,e^{-\varepsilon/k_BT}\), i.e. \(g\,e^{-\varepsilon/k_BT}=1\Rightarrow e^{\varepsilon/k_BT}=g\Rightarrow k_BT=\varepsilon/\ln g\). With \(g=1000\), \(\ln g=6.908\). \(k_BT=0.15/6.908=0.02172\ \mathrm{eV}\). \(T=0.02172/(8.617\times10^{-5})=\boxed{252\ \mathrm{K}}\). Above this temperature entropy (the \(1000\) unfolded states) wins and the protein denatures; below it the folded state dominates — a minimal model of an unfolding transition where the balance is entropy vs. energy.