Degenerate Perturbation Theory
Statement
For a Hamiltonian \(\hat{H} = \hat{H}_0 + \lambda \hat{H}'\) whose unperturbed level \(E_n^{(0)}\) is \(g\)-fold degenerate, the first-order energy shifts are the eigenvalues of the \(g \times g\) matrix \(W_{ab} = \langle n,a | \hat{H}' | n,b \rangle\) formed by projecting the perturbation onto the degenerate eigenspace; the eigenvectors of \(W\) select the "good" zeroth-order states \(|n,\alpha\rangle^{(0)}\) that connect smoothly to the exact states as \(\lambda \to 0\).
Why it matters
Nondegenerate perturbation theory divides by \(E_n^{(0)} - E_m^{(0)}\) and therefore diverges the instant two unperturbed levels coincide. Yet degeneracy is the rule, not the exception, whenever a symmetry is present: the hydrogen \(n^2\)-fold shell, the \((2\ell+1)\) magnetic sublevels of a central potential, the two spin states of a free electron. Degenerate perturbation theory is the tool that decides how such a multiplet splits.
The central lesson is that an arbitrary basis of the degenerate subspace is generically the wrong basis. A weak perturbation "chooses" a preferred set of linear combinations, and only those combinations have well-defined energies to first order. Diagonalizing \(W\) is precisely the act of finding that preferred basis, which is why the method is inseparable from the symmetry thread: the good states are usually eigenstates of whatever operator commutes with \(\hat{H}'\).
Assumptions
Derivation
Result
Reading. To split a degenerate level, build the matrix of the perturbation in any orthonormal basis of the degenerate shell and diagonalize it. Its \(g\) eigenvalues are the first-order energy corrections \(E_{n,\alpha}^{(1)}\); its eigenvectors are the unique "good" combinations of the degenerate states that acquire definite energies. If the eigenvalues are distinct, the perturbation has fully lifted the degeneracy at first order.
Units check. \(W_{ab}\) is a matrix element of \(\hat{H}'\), hence an energy; its eigenvalues \(E^{(1)}\) are energies (joules, or eV). The eigenvectors \(\mathbf{c}\) are dimensionless amplitudes with \(\sum_b |c_b|^2 = 1\), consistent with normalized states. Both sides of \(W\mathbf{c}=E^{(1)}\mathbf{c}\) carry units of energy.
Limiting cases
- \(g=1\) (no degeneracy). \(W\) is the single number \(\langle n|\hat{H}'|n\rangle\), and the formula reduces exactly to the nondegenerate first-order shift \(E^{(1)}=\langle n|\hat{H}'|n\rangle\).
- \(W\) already diagonal. If the chosen basis happens to be the good basis (e.g. it diagonalizes a symmetry that \(\hat{H}'\) respects), then \(E_{n,a}^{(1)} = W_{aa}\) directly and no diagonalization is needed.
- \(W \propto \mathbf{1}\). All shifts equal; the degeneracy is unbroken at first order and the good basis is undetermined until higher order.
- Full splitting. If all \(g\) eigenvalues are distinct, the shell fans out into \(g\) nondegenerate levels and ordinary perturbation theory resumes at second order.
Breaks when
- The degeneracy is only partially lifted. If \(W\) has a repeated eigenvalue, a sub-degeneracy survives at first order. The good basis within that surviving subspace is not fixed by \(W\), and one must diagonalize the second-order (or effective) matrix on the remaining sub-block to proceed.
- Nearby levels intrude. When another unperturbed level lies within \(O(\lambda\|\hat{H}'\|)\) of \(E_n^{(0)}\), restricting \(W\) to the single shell is illegitimate: the small denominators \(E_n^{(0)}-E_m^{(0)}\) are no longer large, and the correct procedure is to diagonalize the full near-degenerate block without a \(\lambda\)-expansion.
- The perturbation is comparable to the splitting it induces. If \(\lambda\hat{H}'\) is not small relative to the level spacing, the first-order truncation is meaningless and one needs exact or nonperturbative diagonalization.
Failure modes
- Using a random basis and reading diagonal elements. Students compute \(W_{aa}\) in whatever basis is handy and quote those as the shifts. That is only valid if \(W\) is already diagonal; otherwise the off-diagonal \(W_{ab}\) matter and the true shifts are the eigenvalues, not the diagonal entries.
- Dividing by zero. Applying the nondegenerate formula \(\sum_{m\neq n}|\langle m|\hat{H}'|n\rangle|^2/(E_n^{(0)}-E_m^{(0)})\) to states within the same degenerate shell, producing a \(0/0\) divergence. The degenerate machinery exists precisely to avoid this.
- Forgetting Hermiticity of \(W\). Miscomputing \(W_{ba}\) as unrelated to \(W_{ab}\); in fact \(W_{ba}=W_{ab}^{*}\), and a non-Hermitian \(W\) signals an algebra error.
- Choosing eigenvectors that are not orthonormal. For a degenerate eigenvalue of \(W\), any orthonormal pair in that subspace works, but students sometimes pick non-orthogonal combinations and lose the clean first-order corrections to the wavefunction.
- Ignoring symmetry. Grinding through a full diagonalization when a conserved quantity commuting with \(\hat{H}'\) already block-diagonalizes \(W\), turning a hard problem into an easy one.
Discussion
The conceptual heart of the method is that the labels of a degenerate multiplet are not physical until a perturbation distinguishes them. In the unperturbed problem, any orthonormal basis of the shell is as good as any other, so the states carry no individual identity. A weak \(\hat{H}'\) breaks this democracy: the exact eigenstates, followed continuously back to \(\lambda=0\), land on one specific basis, the eigenvectors of \(W\). This is why the naive perturbation series blows up in a bad basis, the correction \(|\psi^{(1)}\rangle\) would need infinite coefficients to rotate a bad zeroth-order state into a good one, and why the divergence disappears once the good basis is used.
The symmetry connection is the practical engine of the whole subject. If an operator \(\hat{A}\) commutes with both \(\hat{H}_0\) and \(\hat{H}'\), then \(W\) is block-diagonal in the eigenbasis of \(\hat{A}\): matrix elements of \(\hat{H}'\) between different \(\hat{A}\)-eigenvalues vanish. Choosing the degenerate basis to be simultaneous eigenstates of such conserved quantities often makes \(W\) diagonal outright, so the "good basis" is handed to you by the symmetry and no determinant need be solved. The Stark and Zeeman effects, spin-orbit and fine structure all exploit exactly this: parity, \(L_z\), \(J^2\), and \(J_z\) do the block-diagonalization for free.
More formally, degenerate perturbation theory is the leading term of an effective Hamiltonian obtained by projection. Let \(P\) project onto the degenerate shell and \(Q=\mathbf{1}-P\) onto its complement. The Bloch/Löwdin construction produces an effective operator \(H_{\text{eff}} = P\hat{H}'P + P\hat{H}'Q\,(E_n^{(0)}-Q\hat{H}_0 Q)^{-1}Q\hat{H}'P + \cdots\) acting only within the shell; its eigenvalues are the exact shifts to the corresponding order. The first term is exactly \(W\). The second term is the second-order correction that lifts any degeneracy left unbroken by \(W\), the same object that appears when \(W \propto \mathbf{1}\). This viewpoint unifies degenerate and nondegenerate theory: both are eigenvalue problems for \(H_{\text{eff}}\), differing only in the dimension of the projected subspace.
Common misconceptions. Degenerate perturbation theory is not a different theory from the nondegenerate one; it is the same first-order equation, honestly solved when the projection onto the unperturbed level is multidimensional. The determinant condition is not an approximation on top of an approximation, it is the exact statement that \(E^{(1)}\) must be an eigenvalue of \(W\). And "good states" are not chosen by the physicist for convenience; they are dictated by the perturbation, and any other choice simply does not have a well-defined first-order energy.
Worked examples
Reading. The doublet splits by \(2b = 0.10~\text{eV}\) into symmetric and antisymmetric states; the common shift \(a\) just moves the center of the doublet. Units check. \(a,b\) are matrix elements of \(\hat{H}'\), i.e. energies (eV); the shifts and the splitting are in eV.
Reading. The field splits \(n=2\) into three levels: an upper and lower state (mixtures of \(2s\) and \(2p_0\)) shifted by \(\pm 3eEa_0\), and an unshifted doublet (\(m=\pm1\)). The shift is linear in \(E\), the hallmark of the degenerate case, because the \(s\)-\(p\) degeneracy allows a permanent dipole. Units check. \(eEa_0\) has units \((\text{C})(\text{V/m})(\text{m}) = \text{C·V} = \text{J}\), an energy.
Problems
- A two-fold degenerate level at \(E_0=1.00~\text{eV}\) has perturbation matrix \(W=\begin{pmatrix}0.30 & 0.40\\ 0.40 & 0.30\end{pmatrix}\) eV. Find the first-order energies and the good states.
Solution
The secular equation is \((0.30-E^{(1)})^2 = 0.40^2\), so \(E^{(1)} = 0.30 \pm 0.40 = \{0.70,\,-0.10\}~\text{eV}\). Total energies \(E = E_0 + E^{(1)} = \{1.70,\,0.90\}~\text{eV}\). Good states are \(\tfrac{1}{\sqrt2}(|1\rangle+|2\rangle)\) for \(+0.70\) and \(\tfrac{1}{\sqrt2}(|1\rangle-|2\rangle)\) for \(-0.10\). - For \(W=\begin{pmatrix}a & 0\\ 0 & a\end{pmatrix}\) with \(a>0\), explain what happens to the degeneracy at first order and what one must do next.
Solution
Both eigenvalues equal \(a\), so \(W \propto \mathbf{1}\): the degeneracy is not lifted at first order (both states shift by the same \(a\)). The good basis is undetermined at this order. One must proceed to second order, forming the effective matrix \(W^{(2)}_{ab} = \sum_{m\notin \text{shell}} \frac{\langle a|\hat{H}'|m\rangle\langle m|\hat{H}'|b\rangle}{E_n^{(0)}-E_m^{(0)}}\) and diagonalizing it to determine both the second-order splitting and the good states. - A three-fold degenerate level has \(W=\begin{pmatrix}2 & 1 & 0\\ 1 & 2 & 0\\ 0 & 0 & 5\end{pmatrix}\) (units of \(\varepsilon\)). Find all first-order shifts.
Solution
\(W\) is block-diagonal: a \(2\times2\) block \(\begin{pmatrix}2&1\\1&2\end{pmatrix}\) and a \(1\times1\) block \((5)\). The \(2\times2\) block gives \(E^{(1)} = 2\pm1 = \{3,\,1\}\varepsilon\); the isolated block gives \(5\varepsilon\). So the shifts are \(\{3,\,1,\,5\}\varepsilon\), and the degeneracy is fully lifted. Good states: \(\tfrac{1}{\sqrt2}(|1\rangle\pm|2\rangle)\) and \(|3\rangle\). - Consider a spin-1/2 particle with \(\hat{H}_0 = 0\) (both spin states degenerate at \(E=0\)) and \(\hat{H}' = \gamma\,\mathbf{B}\cdot\boldsymbol{\sigma}\) with \(\mathbf{B}=B\hat{x}\). Find the first-order energies and good states, with \(B=0.10~\text{T}\) and \(\gamma = 5.79\times10^{-5}~\text{eV/T}\).
Solution
With \(\mathbf{B}=B\hat{x}\), \(\hat{H}' = \gamma B\,\sigma_x = \gamma B\begin{pmatrix}0&1\\1&0\end{pmatrix}\), which is already \(W\) in the \(\{|\!\uparrow\rangle,|\!\downarrow\rangle\}\) basis. Eigenvalues \(E^{(1)} = \pm\gamma B\); good states are the \(\sigma_x\) eigenstates \(\tfrac{1}{\sqrt2}(|\!\uparrow\rangle\pm|\!\downarrow\rangle)\). Numerically \(\gamma B = (5.79\times10^{-5})(0.10) = 5.79\times10^{-6}~\text{eV}\), so \(E^{(1)} = \pm 5.79\times10^{-6}~\text{eV}\), a splitting of \(1.16\times10^{-5}~\text{eV}\). - A two-fold degenerate level has a complex-coupled perturbation \(W=\begin{pmatrix}0 & -i\Delta\\ i\Delta & 0\end{pmatrix}\) with \(\Delta>0\) real. Show \(W\) is Hermitian, find the shifts, and give the good states.
Solution
Hermiticity: \(W_{21} = i\Delta = (W_{12})^{*} = (-i\Delta)^{*}\) — yes, \(W=W^\dagger\). Secular equation: \(\det\begin{pmatrix}-E^{(1)} & -i\Delta\\ i\Delta & -E^{(1)}\end{pmatrix} = (E^{(1)})^2 - (-i\Delta)(i\Delta) = (E^{(1)})^2 - \Delta^2 = 0\), so \(E^{(1)} = \pm\Delta\). Eigenvectors: for \(+\Delta\), \(-i\Delta c_2 = \Delta c_1 \Rightarrow \mathbf{c}=\tfrac{1}{\sqrt2}(1,\,i)^T\); for \(-\Delta\), \(\mathbf{c}=\tfrac{1}{\sqrt2}(1,\,-i)^T\). These are the good states \(\tfrac{1}{\sqrt2}(|1\rangle\pm i|2\rangle)\), circularly-polarized combinations, and the shifts \(\pm\Delta\) are real as required.