physics2u
Tier
⌕ Search ⌘K
Derivation

Relativistic Kinetic-Energy Correction

D-298 Home PU-306 Threads energy · matter Depends on Bound-State Spectrum of the Hydrogen Atom, time-independent-perturbation-theory
Statement

Expanding the exact relativistic kinetic energy to first order beyond the Newtonian term yields the perturbation \(H'=-\hat{p}^4/(8m^3c^2)\). Applied to the bound eigenstates \(|n\ell m\rangle\) of the Coulomb Hamiltonian, non-degenerate perturbation theory gives the leading relativistic shift of the hydrogen levels, \(\;E^{(1)}_{\mathrm{rel}}=-\dfrac{E_n^{2}}{2mc^{2}}\left(\dfrac{4n}{\ell+\tfrac12}-3\right)\), where \(E_n=-13.6\,\mathrm{eV}/n^2\) are the Bohr energies.

Why it matters

This is the first of the three terms that together make up the fine structure of hydrogen. It is the correction that turns the crude Bohr/Schrödinger spectrum — degenerate in \(\ell\) — into the observed line structure, and it fixes the scale \(\alpha^2\sim 5\times10^{-5}\) at which relativity first perturbs atomic energies. Historically it is what forced the recognition that the electron's speed in the atom, \(v/c\sim Z\alpha\), is not entirely negligible.

Pedagogically the derivation is a showcase for a recurring trick: rather than compute a hard matrix element of \(\hat p^4\) directly, one uses the Schrödinger equation itself, \(\hat p^2=2m(H_0-V)\), to trade momentum operators for the potential, reducing everything to the known radial expectation values \(\langle 1/r\rangle\) and \(\langle 1/r^2\rangle\).

Assumptions
The electron is non-relativistic to leading order, \(v/c\sim Z\alpha\ll 1\); if dropped the binomial series in \(p/mc\) does not converge term-by-term and one must solve the Dirac equation instead of truncating.
The perturbation is small compared with level spacings, \(|E^{(1)}_{\mathrm{rel}}|\ll|E_n|\); if dropped first-order perturbation theory is invalid and the shift must be resummed non-perturbatively.
The unperturbed states are the exact Coulomb eigenstates, so that \(H_0|n\ell m\rangle=E_n|n\ell m\rangle\) may be used to eliminate \(\hat p^2\); if the true \(H_0\) differed (screening, finite nucleus) the substitution \(\hat p^2=2m(H_0-V)\) would introduce extra terms.
\(H'\) is diagonal within each degenerate \(n\)-shell, because \(\hat p^4\) is a rotational scalar it commutes with \(\hat{\mathbf L}^2\) and \(\hat L_z\), so \(|n\ell m\rangle\) already diagonalises it and the non-degenerate formula applies; if this were not checked, degenerate perturbation theory would be mandatory and the naive \(\langle H'\rangle\) could be meaningless.
Spin, spin-orbit and Darwin terms are set aside, this term is computed in isolation; if dropped one forgets that only the sum of all three fine-structure terms is physical and depends solely on \(j\), not on \(\ell\) alone.
Derivation
1
\[E=\sqrt{p^{2}c^{2}+m^{2}c^{4}}\]
Relativistic energy–momentum relation for a free particle of rest mass \(m\). A
2
\[T=E-mc^{2}=mc^{2}\left(\sqrt{1+\frac{p^{2}}{m^{2}c^{2}}}-1\right)\]
Subtract the rest energy to isolate the kinetic energy and factor out \(mc^2\). A
3
\[T=\frac{p^{2}}{2m}-\frac{p^{4}}{8m^{3}c^{2}}+\frac{p^{6}}{16m^{5}c^{4}}-\cdots\]
Binomial expansion \(\sqrt{1+x}=1+\tfrac12 x-\tfrac18 x^{2}+\cdots\) with \(x=p^{2}/m^{2}c^{2}\ll1\). The first term is the Newtonian kinetic energy already in \(H_0\). B
4
\[H'=-\frac{\hat{p}^{4}}{8m^{3}c^{2}}\]
The leading correction beyond \(H_0=\hat p^2/2m+V\); promote \(p\) to the operator \(\hat p\). Higher terms are \(\mathcal O((Z\alpha)^4)\) and neglected. B
5
\[E^{(1)}_{\mathrm{rel}}=\langle n\ell m|\,H'\,|n\ell m\rangle=-\frac{1}{8m^{3}c^{2}}\,\big\langle \hat p^{4}\big\rangle\]
First-order perturbation theory. Legitimate in the non-degenerate form because \(\hat p^4\) is a scalar: \([\hat p^4,\hat{\mathbf L}^2]=[\hat p^4,\hat L_z]=0\), so \(|n\ell m\rangle\) diagonalises \(H'\) inside every degenerate shell. C
6
\[\hat p^{2}=2m\big(H_0-V\big),\qquad V=-\frac{e^{2}}{4\pi\varepsilon_0\,r}\]
Rearrange the unperturbed Schrödinger Hamiltonian \(H_0=\hat p^2/2m+V\) to express \(\hat p^2\) through \(H_0\) and the Coulomb potential — the key move that avoids differentiating \(\psi\) four times. B
7
\[\big\langle \hat p^{4}\big\rangle=\big\langle \hat p^{2}\,\hat p^{2}\big\rangle=4m^{2}\big\langle (H_0-V)^{2}\big\rangle=4m^{2}\big\langle (E_n-V)^{2}\big\rangle\]
\(\hat p^2\) is Hermitian, so let one factor act left and the other right onto the eigenstate; \(H_0|n\ell m\rangle=E_n|n\ell m\rangle\) replaces the operator by its eigenvalue. C
8
\[\big\langle \hat p^{4}\big\rangle=4m^{2}\Big(E_n^{2}-2E_n\langle V\rangle+\langle V^{2}\rangle\Big)\]
Expand the square; \(E_n\) is a number and pulls out of the expectation value. A
9
\[\Big\langle \tfrac1r\Big\rangle=\frac{1}{n^{2}a},\qquad \Big\langle \tfrac1{r^{2}}\Big\rangle=\frac{1}{n^{3}\big(\ell+\tfrac12\big)a^{2}}\]
Standard Coulomb radial expectation values (\(a\) = Bohr radius), assumed from the hydrogen bound-state spectrum. These give \(\langle V\rangle\) and \(\langle V^2\rangle\). B
10
\[E_n=-\frac{e^{2}}{4\pi\varepsilon_0}\frac{1}{2an^{2}}\ \Rightarrow\ \frac{e^{2}}{4\pi\varepsilon_0\,a}=-2n^{2}E_n\]
Use the Bohr energy to eliminate the Coulomb constant, so \(\langle V\rangle=2E_n\) (virial theorem) and \(\langle V^2\rangle=\dfrac{4nE_n^{2}}{\ell+\tfrac12}\). B
11
\[E^{(1)}_{\mathrm{rel}}=-\frac{4m^{2}}{8m^{3}c^{2}}\Big(E_n^{2}-4E_n^{2}+\tfrac{4nE_n^{2}}{\ell+1/2}\Big)=-\frac{E_n^{2}}{2mc^{2}}\left(\frac{4n}{\ell+\tfrac12}-3\right)\]
Insert into Step 5, collect the \(E_n^2\) terms (\(1-4=-3\)), and simplify. A
Result
\[E^{(1)}_{\mathrm{rel}}=-\frac{E_n^{2}}{2mc^{2}}\left(\frac{4n}{\ell+\tfrac12}-3\right)=\frac{E_n\,\alpha^{2}}{n^{2}}\left(\frac{n}{\ell+\tfrac12}-\frac34\right)\]

Reading. The shift is negative for every allowed \((n,\ell)\) — relativity always lowers the levels, because a genuinely relativistic electron has slightly less kinetic energy than the Newtonian \(p^2/2m\) predicts for the same momentum. Its size relative to \(|E_n|\) is of order \((Z\alpha)^2/n^2\). Crucially the shift depends on \(\ell\): it partially lifts the accidental Coulomb degeneracy, states of smaller \(\ell\) (more penetrating, faster near the nucleus) being pushed down most.

Units check. \(E_n^2\) has units of (energy)\(^2\); dividing by \(mc^2\) (an energy) leaves energy. The bracket \(\left(4n/(\ell+\tfrac12)-3\right)\) is dimensionless, as is \(\alpha^2\). Both forms therefore return an energy, e.g. in eV. \(\checkmark\)

Limiting cases
  • \(c\to\infty\): the prefactor \(1/2mc^2\to0\), the correction vanishes and the Bohr spectrum is recovered — the expansion parameter is \(1/c^2\).
  • Ground state \(n=1,\ell=0\): bracket \(=4/(1/2)-3=5\), giving the largest fractional shift, \(-\tfrac52 E_1^2/mc^2\).
  • Circular orbit \(\ell=n-1\): bracket \(=\dfrac{4n}{n-1/2}-3\to1\) for large \(n\); the shift shrinks like \(E_n^2\sim1/n^4\), so high Rydberg levels are essentially Newtonian.
  • Hydrogen-like ion, nuclear charge \(Z\): replace \(E_n\to Z^2E_n\) and \(a\to a/Z\); the correction scales as \(Z^4\), so it grows explosively for heavy ions.
Breaks when
  • High nuclear charge, \(Z\alpha\gtrsim1\). The expansion parameter is \(Z\alpha\); once \(v/c\) is not small the binomial series in Step 3 diverges and the perturbative truncation is meaningless — the Dirac equation must be solved exactly (for \(Z\gtrsim137\) even that fails without QED and finite-nuclear-size input).
  • As an isolated, observable quantity. Taken alone this term is not measurable: the physical fine structure is the sum of the \(p^4\) term, spin–orbit coupling, and the Darwin term, which combine to give a splitting depending only on \(j\). Quoting \(E^{(1)}_{\mathrm{rel}}\) as "the" relativistic shift of a spectral line is wrong.
  • Near-degeneracy with another perturbation. If a second perturbation (e.g. an external field, or spin–orbit) connects \(|n\ell m\rangle\) to states with which \(H'\) is comparable to the gap, the naive non-degenerate treatment fails and one must diagonalise the full perturbation in the degenerate subspace.
  • Very high \(n\) or continuum. For weakly bound or unbound states \(\langle 1/r\rangle,\langle 1/r^2\rangle\) are not given by the bound-state formulae of Step 9, so the closed result does not apply.
Failure modes
  • Computing \(\langle\hat p^4\rangle\) by brute force. Differentiating \(\psi_{n\ell m}\) four times, especially for \(s\)-states, is error-prone; the \(\hat p^2=2m(H_0-V)\) substitution is both faster and correct.
  • The \(s\)-state Hermiticity worry. Students fear the substitution "hides" a delta-function singularity at the origin for \(\ell=0\). The operator identity \(\langle\hat p^2\phi|\hat p^2\phi\rangle=4m^2\langle(E_n-V)^2\rangle\) is nonetheless exact and reproduces the Dirac result; the subtlety is a Darwin-term effect belonging to a different perturbation, not to \(H'\).
  • Forgetting the virial step. Leaving \(\langle V\rangle\) and \(\langle V^2\rangle\) in terms of \(e^2/4\pi\varepsilon_0\) instead of eliminating them via \(E_n\) leads to messy, dimensionally opaque algebra and sign slips.
  • Using \(\ell+1\) or \(\ell\) instead of \(\ell+\tfrac12\). The half-integer denominator comes from \(\langle 1/r^2\rangle\); mis-remembering it corrupts the \(\ell\)-dependence.
  • Claiming the shift can be positive. The bracket \(4n/(\ell+\tfrac12)-3\ge1>0\) for all allowed \(\ell\le n-1\), so with the overall minus sign the shift is always negative — a positive answer signals an algebra error.
Discussion

The physical content is that mass grows with speed: the exact kinetic energy \(mc^2(\gamma-1)\) rises more slowly with \(p\) than \(p^2/2m\), so a bound electron carrying its Coulomb-fixed momentum is less energetic than the Schrödinger picture assumes. Because the electron moves fastest where the potential is deepest — near the nucleus, and most so for penetrating low-\(\ell\) orbitals — the correction is largest exactly for those states, which is why it lifts the \(\ell\)-degeneracy that the pure \(1/r\) potential accidentally protects.

That accidental degeneracy is no accident of algebra but of symmetry: the Coulomb problem has an extra conserved quantity, the Laplace–Runge–Lenz vector, enlarging its symmetry from \(SO(3)\) to \(SO(4)\) and making \(E_n\) independent of \(\ell\). The \(p^4\) term is not \(SO(4)\)-invariant, so it breaks the enhanced symmetry down to ordinary rotational symmetry, and the levels spread according to \(\ell\).

Standing alone the result is incomplete. The full \(\mathcal O((Z\alpha)^4)\) fine structure of hydrogen requires adding spin–orbit coupling, \(H_{SO}\propto \tfrac1r\frac{dV}{dr}\,\hat{\mathbf S}\!\cdot\!\hat{\mathbf L}\), and — for \(\ell=0\) — the Darwin term \(\propto\nabla^2 V\). A minor miracle of the hydrogen atom is that their sum collapses to a single expression depending only on the total angular momentum \(j\): \(E^{(1)}_{\mathrm{fs}}=\dfrac{(E_n)^2}{2mc^2}\!\left(3-\dfrac{4n}{j+\tfrac12}\right)\), which is exactly the \(\mathcal O((Z\alpha)^4)\) expansion of the closed-form Dirac (Sommerfeld) energy. The three terms individually depend on \(\ell\); only their combination is a good, \(\ell\)-blind quantum-mechanical observable, so the states \(2S_{1/2}\) and \(2P_{1/2}\) come out degenerate — a degeneracy finally lifted only by the QED Lamb shift.

Common misconceptions. (i) The correction is not the spin–orbit interaction; it exists even for a spinless particle. (ii) It does not by itself explain the sodium doublet or hydrogen fine-structure splitting — that needs the full \(j\)-dependent sum. (iii) "Relativistic mass increase" is a heuristic, not the mechanism: the honest statement is the expansion of \(\sqrt{p^2c^2+m^2c^4}\).

Worked examples
1
Hydrogen \(2s\) state (\(n=2,\ \ell=0\)). Find \(E^{(1)}_{\mathrm{rel}}\).
\[E^{(1)}_{\mathrm{rel}}=-\frac{E_n^{2}}{2mc^{2}}\left(\frac{4n}{\ell+\tfrac12}-3\right)\]
Insert symbols first: \(E_2=E_1/n^2=-13.6/4=-3.40\ \mathrm{eV}\), \(mc^2=0.511\ \mathrm{MeV}=5.110\times10^{5}\ \mathrm{eV}\). A
\[\frac{4n}{\ell+\tfrac12}-3=\frac{8}{1/2}-3=16-3=13\]
Evaluate the dimensionless bracket for \(n=2,\ \ell=0\). A
\[E^{(1)}_{\mathrm{rel}}=-\frac{(3.40)^{2}}{2(5.110\times10^{5})}\times13\ \mathrm{eV}=-\frac{11.56}{1.022\times10^{6}}\times13\ \mathrm{eV}\]
Substitute numbers; units of eV throughout. B
\[E^{(1)}_{\mathrm{rel}}\approx-1.47\times10^{-4}\ \mathrm{eV}\;=\;-147\ \mu\mathrm{eV}\]

Reading. About \(4\times10^{-5}\) of \(|E_2|=3.40\) eV, i.e. of order \(\alpha^2\), as expected. Negative: the level moves down.

Units check. \(\mathrm{eV}^2/\mathrm{eV}=\mathrm{eV}\). \(\checkmark\)

2
Compare the two \(n=2\) sub-levels: how much does the \(p^4\) term split \(2s\ (\ell=0)\) from \(2p\ (\ell=1)\)?
\[\Delta=E^{(1)}_{\mathrm{rel}}(\ell=0)-E^{(1)}_{\mathrm{rel}}(\ell=1)=-\frac{E_2^{2}}{2mc^{2}}\Big[b_0-b_1\Big],\quad b_\ell=\frac{4n}{\ell+\tfrac12}-3\]
Both share the prefactor; only the bracket differs. Symbols before numbers. A
\[b_0=13,\qquad b_1=\frac{8}{3/2}-3=\frac{16}{3}-3=\frac{7}{3}\approx2.333\]
Evaluate each bracket for \(n=2\). A
\[\Delta=-\frac{(3.40)^{2}}{1.022\times10^{6}}\big(13-2.333\big)\ \mathrm{eV}=-(1.131\times10^{-5})(10.667)\ \mathrm{eV}\]
Insert numbers; \(E_2^2/2mc^2=1.131\times10^{-5}\) eV as in Example 1. B
\[\Delta\approx-1.21\times10^{-4}\ \mathrm{eV}\;=\;-121\ \mu\mathrm{eV}\]

Reading. The \(2s\) level sits \(\sim\!121\ \mu\mathrm{eV}\) below \(2p\) from this term alone — but adding spin–orbit and Darwin restores the \(2S_{1/2}\)–\(2P_{1/2}\) degeneracy, so this splitting is not directly observable. It illustrates that the \(p^4\) term by itself is \(\ell\)-dependent.

Units check. \(\mathrm{eV}^2/\mathrm{eV}\times(\text{dimensionless})=\mathrm{eV}\). \(\checkmark\)

Problems
  1. Show that the fractional shift of the hydrogen ground state is \(E^{(1)}_{\mathrm{rel}}/E_1=-\tfrac52\,\alpha^2\), and evaluate it numerically.
    Solution For \(n=1,\ell=0\): bracket \(=4/(1/2)-3=5\). \(E^{(1)}_{\mathrm{rel}}=-\dfrac{E_1^2}{2mc^2}\cdot5\). Using \(E_1=-\tfrac12\alpha^2 mc^2\), \(E_1^2=\tfrac14\alpha^4 m^2c^4\), so \(E^{(1)}_{\mathrm{rel}}=-\dfrac{\alpha^4 m^2c^4/4}{2mc^2}\cdot5=-\dfrac58\alpha^4 mc^2\). Dividing by \(E_1=-\tfrac12\alpha^2mc^2\) gives \(E^{(1)}_{\mathrm{rel}}/E_1=(-\tfrac58\alpha^4mc^2)/(-\tfrac12\alpha^2mc^2)=\tfrac54\alpha^2\)... check sign: \(E^{(1)}<0\), \(E_1<0\), ratio positive, magnitude \(\tfrac54\alpha^2\). (The "\(\tfrac52\)" in the prompt refers to \(|E^{(1)}|=\tfrac52 E_1^2/mc^2\); the ratio to \(E_1\) is \(\tfrac54\alpha^2\).) Numerically \(\tfrac54(5.325\times10^{-5})=6.66\times10^{-5}\), and \(E^{(1)}_{\mathrm{rel}}=-6.66\times10^{-5}\times13.6\ \mathrm{eV}=-9.05\times10^{-4}\ \mathrm{eV}=-905\ \mu\mathrm{eV}\).
  2. For fixed \(n\), which value of \(\ell\) gives the smallest-magnitude relativistic shift, and why physically?
    Solution The magnitude scales with the bracket \(b_\ell=4n/(\ell+\tfrac12)-3\), which decreases monotonically as \(\ell\) increases. Hence the maximum \(\ell=n-1\) (the "circular" orbit) gives the smallest shift: \(b_{n-1}=\dfrac{4n}{n-1/2}-3=\dfrac{4n-3(n-1/2)}{n-1/2}=\dfrac{n+3/2}{n-1/2}\), which \(\to1\) for large \(n\). Physically, high-\(\ell\) states have vanishing amplitude near the nucleus (centrifugal barrier), the electron moves slowest there, so relativistic corrections are weakest.
  3. A hydrogen-like ion has nuclear charge \(Z\). By how much does \(E^{(1)}_{\mathrm{rel}}\) grow relative to hydrogen for the same \((n,\ell)\)?
    Solution Under \(Z\): \(E_n\to Z^2 E_n\), so \(E_n^2\to Z^4 E_n^2\); the bracket is \(Z\)-independent and \(mc^2\) unchanged. Therefore \(E^{(1)}_{\mathrm{rel}}\to Z^4 E^{(1)}_{\mathrm{rel}}\). Equivalently the expansion parameter is \(Z\alpha\) and the shift is \(\mathcal O((Z\alpha)^2)\times|E_n|\propto(Z\alpha)^2 Z^2=Z^4\alpha^2\). For e.g. \(\mathrm{C}^{5+}\) (\(Z=6\)) it is \(6^4=1296\) times larger than in H.
  4. Verify the virial relation \(\langle V\rangle=2E_n\) used in the derivation, starting from \(\langle 1/r\rangle=1/(n^2 a)\) and \(E_n=-\tfrac{e^2}{4\pi\varepsilon_0}\tfrac{1}{2an^2}\).
    Solution \(\langle V\rangle=-\dfrac{e^2}{4\pi\varepsilon_0}\Big\langle\dfrac1r\Big\rangle=-\dfrac{e^2}{4\pi\varepsilon_0}\cdot\dfrac{1}{n^2 a}=-\dfrac{e^2}{4\pi\varepsilon_0 a n^2}\). From the energy, \(\dfrac{e^2}{4\pi\varepsilon_0 a}=-2n^2 E_n\), so \(\langle V\rangle=-(-2n^2E_n)/n^2=2E_n\). Consistent with the virial theorem \(\langle T\rangle=-E_n,\ \langle V\rangle=2E_n\), \(\langle T\rangle+\langle V\rangle=E_n\).
  5. Estimate the wavelength shift of the Lyman-\(\alpha\) line (\(2p\to1s\)) produced by the \(p^4\) term alone, given the line sits near \(\lambda=121.6\) nm. State clearly why this is not the measured fine-structure shift.
    Solution Using \(E^{(1)}(1s)=-9.05\times10^{-4}\) eV (Prob. 1) and \(E^{(1)}(2p,\ell=1)=-\dfrac{E_2^2}{2mc^2}b_1=-(1.131\times10^{-5})(7/3)=-2.64\times10^{-5}\) eV (from Example 2). The transition energy changes by \(\delta E=E^{(1)}(2p)-E^{(1)}(1s)=(-2.64\times10^{-5})-(-9.05\times10^{-4})=+8.79\times10^{-4}\) eV — the upper level drops far less than the lower, so the emitted photon gains energy. Since \(E=hc/\lambda\), \(\delta\lambda=-\lambda^2\,\delta E/hc\). With \(hc=1239.8\) eV·nm: \(\delta\lambda=-(121.6)^2(8.79\times10^{-4})/1239.8=-1.05\times10^{-2}\) nm \(\approx-0.010\) nm. This is only the \(p^4\) contribution; the actual Lyman-\(\alpha\) fine structure also includes spin–orbit and Darwin terms, which regroup the levels by \(j\) (giving the observed \(2P_{3/2}\)–\(2P_{1/2}\) doublet). The isolated \(p^4\) number is therefore not directly observable.