Non-Degenerate Perturbation Theory
Statement
For a Hamiltonian \( \hat{H} = \hat{H}_0 + \lambda \hat{V} \) whose unperturbed part \( \hat{H}_0 \) has a discrete, non-degenerate eigenvalue \( E_n^{(0)} \) with normalized eigenstate \( \lvert n^{(0)} \rangle \), the perturbed eigenvalue and eigenstate admit power-series expansions in the bookkeeping parameter \( \lambda \). This page derives the leading corrections: the energy through second order, \( E_n = E_n^{(0)} + \lambda \langle n^{(0)} \lvert \hat{V} \rvert n^{(0)} \rangle + \lambda^2 \sum_{m \neq n} \frac{\lvert \langle m^{(0)} \lvert \hat{V} \rvert n^{(0)} \rangle \rvert^2}{E_n^{(0)} - E_m^{(0)}} + \cdots \), and the state through first order, \( \lvert n \rangle = \lvert n^{(0)} \rangle + \lambda \sum_{m \neq n} \frac{\langle m^{(0)} \lvert \hat{V} \rvert n^{(0)} \rangle}{E_n^{(0)} - E_m^{(0)}} \lvert m^{(0)} \rangle + \cdots \).
Why it matters
Almost no realistic Hamiltonian is exactly solvable, yet enormous swathes of physics are small deformations of ones that are: the hydrogen atom perturbed by relativistic and spin-orbit terms, an atom in a weak electric or magnetic field, a crystal electron feeling a weak lattice potential, an anharmonic vibration. Non-degenerate perturbation theory turns "the exact answer is unknown" into a controlled, order-by-order approximation built entirely from matrix elements of the known basis.
The structure of the result is as important as its numbers. The first-order energy is simply the perturbation averaged in the unperturbed state; the second-order energy always pushes the ground state down and encodes level repulsion, virtual transitions, and — through the same denominators — the divergences that signal when the expansion breaks. These ideas propagate directly into time-dependent perturbation theory, scattering, and quantum field theory.
Assumptions
Derivation
Result
Reading. The first-order energy shift is the perturbation simply averaged over the unperturbed state — its "expectation value" in \( \lvert n^{(0)} \rangle \). The second-order shift is a sum over virtual admixtures: each other level \( m \) contributes an amount proportional to how strongly \( \hat{V} \) connects it to \( n \), weighted inversely by how far away it sits in energy. The state correction mixes in a little of every other level, again controlled by coupling strength over energy gap.
Units check. Every term carries units of energy. \( \langle n^{(0)} \lvert \hat{V} \rvert n^{(0)} \rangle \) is an energy because \( \hat{V} \) is an energy operator and the states are dimensionless (normalized). In the second-order term the numerator is (energy)\(^2\) and the denominator is energy, giving energy. The state coefficients \( \langle m^{(0)} \lvert \hat{V} \rvert n^{(0)} \rangle / (E_n^{(0)} - E_m^{(0)}) \) are (energy)/(energy) = dimensionless, as required for expansion coefficients.
Limiting cases
- Ground state, \( n = 0 \): every denominator \( E_0^{(0)} - E_m^{(0)} < 0 \), so \( E_0^{(2)} \le 0 \) — the second-order shift always lowers the ground-state energy.
- Diagonal perturbation (\( \hat{V} \) commutes with \( \hat{H}_0 \)): all off-diagonal \( \langle m^{(0)} \lvert \hat{V} \rvert n^{(0)} \rangle = 0 \), so \( E_n^{(2)} = 0 \) and the state is unchanged; the first-order shift is exact.
- Large gaps (\( \lvert E_n^{(0)} - E_m^{(0)} \rvert \to \infty \)): second-order and state corrections \( \to 0 \); the level is "isolated" and the perturbation acts diagonally to leading order.
- Two nearly degenerate levels: the pair with the smallest gap dominates \( E_n^{(2)} \); as the gap shrinks the expansion loses validity and one crosses over to degenerate theory.
Breaks when
- Degeneracy or near-degeneracy: if \( E_m^{(0)} = E_n^{(0)} \) for some \( m \neq n \) with \( \langle m^{(0)} \lvert \hat{V} \rvert n^{(0)} \rangle \neq 0 \), the coefficient \( 1/(E_n^{(0)} - E_m^{(0)}) \) diverges. The remedy is to diagonalize \( \hat{V} \) in the degenerate subspace first.
- Perturbation not small: when \( \lvert \langle m^{(0)} \lvert \hat{V} \rvert n^{(0)} \rangle \rvert \gtrsim \lvert E_n^{(0)} - E_m^{(0)} \rvert \), the state coefficient is order unity or larger and the series ceases to be ordered by size; truncation is meaningless.
- Non-analytic dependence on \( \lambda \): perturbations that change the qualitative nature of the spectrum — e.g. an \( x^4 \) term that makes an unbounded-below or barrier-tunneling situation, or a field that ionizes a bound state into the continuum — give series that are only asymptotic (zero radius of convergence) or miss non-perturbative effects (tunneling \( \sim e^{-c/\lambda} \)) entirely.
- Continuous spectrum overlapping \( E_n^{(0)} \): if \( E_n^{(0)} \) is degenerate with continuum states, the discrete sum must become an integral and a bound state can acquire a finite lifetime (a resonance); ordinary bound-state perturbation theory does not capture this decay.
Failure modes
- Forgetting to exclude \( m = n \): including the \( m = n \) term produces a \( 0/0 \) in the state sum. The exclusion follows from intermediate normalization \( \langle n^{(0)} \lvert n \rangle = 1 \).
- Sign of the denominator: writing \( E_m^{(0)} - E_n^{(0)} \) instead of \( E_n^{(0)} - E_m^{(0)} \) flips the sign of every correction and wrongly predicts the ground-state shift going up.
- Using \( \langle m \lvert \hat{V} \rvert n \rangle^2 \) instead of \( \lvert \langle m \lvert \hat{V} \rvert n \rangle \rvert^2 \): for complex matrix elements the naive square is complex; only the modulus-squared is correct and guarantees a real energy.
- Applying it across a degeneracy: blindly using the non-degenerate formulas when two unperturbed levels coincide — the single most common conceptual error; the answer is not merely large, it is undefined.
- Confusing \( \lambda \) with the physical coupling: leaving \( \lambda \) in final answers, or forgetting that \( \hat{V} \) may already contain the physical small parameter, leading to double counting of orders.
- Assuming the first-order state is normalized: \( \lvert n^{(0)} \rangle + \lvert n^{(1)} \rangle \) has norm \( 1 + O(\lambda^2) \); one must renormalize before computing expectation values to \( O(\lambda^2) \).
Discussion
The second-order energy has a compelling physical reading as level repulsion. Any two levels \( n \) and \( m \) connected by \( \hat{V} \) push apart: the higher one is raised and the lower one is depressed, by equal-and-opposite amounts \( \pm \lvert V_{mn} \rvert^2 / \lvert E_n^{(0)} - E_m^{(0)} \rvert \). This is why energy levels in quantum systems "avoid crossing" as a parameter is varied — a phenomenon (the von Neumann–Wigner theorem) that underlies band gaps in solids and the structure of molecular potential surfaces.
The state correction \( \lvert n^{(1)} \rangle \) says the true eigenstate is the old one with a little of every other level mixed in — "virtual" excitations to intermediate states \( m \) and back. This is the germ of the sum-over-intermediate-states picture that becomes the sum over Feynman diagrams in field theory. The very same denominators \( E_n^{(0)} - E_m^{(0)} \) reappear there as energy denominators in old-fashioned (time-ordered) perturbation theory.
There is a deep link to the ground state through the variational principle. Because \( E_0^{(2)} \le 0 \), second-order theory can be viewed as the leading correction that a trial state \( \lvert 0^{(0)} \rangle + \lvert 0^{(1)} \rangle \) buys you over the bare estimate — perturbation theory and the Rayleigh–Ritz method share the same first-order energy, and the sign of the second-order term reflects that the true ground state does at least as well as any trial state.
Convergence is subtler than the tidy series suggests. Kato and Rellich established that if \( \hat{V} \) is relatively bounded with respect to \( \hat{H}_0 \) (roughly, not too singular), the perturbed eigenvalue is genuinely analytic in \( \lambda \) in a disk and the series converges. But many physically central cases fail this: the quartic anharmonic oscillator \( \hat{H}_0 + \lambda \hat{x}^4 \) has a series whose coefficients grow like \( n! \), giving zero radius of convergence — the series is asymptotic. It is still useful (optimal truncation gives exponentially small error for small \( \lambda \)) and can be resummed (Borel summation), but this is a warning that "small perturbation" does not guarantee "convergent series". Non-perturbative effects such as tunneling, scaling as \( e^{-c/\lambda} \), are invisible to any finite order.
Common misconceptions. A larger perturbation does not always mean a larger shift — a perturbation that is purely off-diagonal produces zero first-order shift no matter how strong. And "second order" does not mean "smaller than first order term for this particular state": if \( V_{nn} = 0 \) the second-order term is the leading effect. Finally, the sum in \( E_n^{(2)} \) runs over all other states, including any continuum — omitting the continuum is a frequent quantitative error in atomic problems (e.g. polarizability).
Worked examples
Reading. The lower level is pushed down by \( 5\ \text{meV} \); the perturbative value \( 0.9950\ \text{eV} \) matches the exact \( 0.99501\ \text{eV} \) to the quoted precision, since the next correction is \( O(v^4/\Delta^3) \sim 10^{-6}\ \text{eV} \).
Units check. \( v^2/\Delta \) is \( (\text{eV})^2/\text{eV} = \text{eV} \); the shift is a genuine energy.
Reading. The \( x^4 \) term raises every level; for the ground state the shift is \( \approx 0.16\ \text{eV} \). Compare with \( \tfrac{1}{2}\hbar\omega = \tfrac{1}{2}(1.055\times10^{-34})(2.0\times10^{15}) = 1.05\times10^{-19}\,\text{J} \approx 0.66\ \text{eV} \): the perturbation is about \( 24\% \) of the level spacing, near the edge of where first order alone is trustworthy.
Units check. \( \gamma \) has \( \text{J/m}^4 \); \( (\hbar/m\omega)^2 \) has \( \text{m}^4 \); the product is \( \text{J} \). Correct.
Problems
- (Foundational) State the first-order energy correction for a level \( n \) in words and symbols, and explain in one sentence why an off-diagonal perturbation gives \( E_n^{(1)} = 0 \).
Solution
\( E_n^{(1)} = \langle n^{(0)} \lvert \hat{V} \rvert n^{(0)} \rangle \), the expectation value (diagonal matrix element) of the perturbation in the unperturbed state. If \( \hat{V} \) is purely off-diagonal in the \( \hat{H}_0 \) basis then \( \langle n^{(0)} \lvert \hat{V} \rvert n^{(0)} \rangle = 0 \) by definition, so there is no first-order shift; the leading effect is then second order. - (Two-level, general) For \( \hat{H}_0 = \operatorname{diag}(E_a, E_b) \) with \( E_a < E_b \) and \( \hat{V} = \begin{pmatrix} d & v \\ v^* & -d \end{pmatrix} \), find \( E_a \) to second order.
Solution
First order: \( E_a^{(1)} = d \). Second order: only \( m=b \) contributes, \( E_a^{(2)} = \lvert v \rvert^2/(E_a - E_b) \). Thus \( E_a \approx E_a + d + \dfrac{\lvert v \rvert^2}{E_a - E_b} \). The gap in the denominator should strictly use the unperturbed energies \( E_a - E_b \); the diagonal \( d \) shifts do not enter the second-order denominator at this order. - (Numeric second order) Three levels: \( E_1^{(0)} = 0,\ E_2^{(0)} = 2\,\text{eV},\ E_3^{(0)} = 5\,\text{eV} \). The perturbation connects level 1 to 2 and to 3 with \( \langle 2 \lvert \hat{V} \rvert 1 \rangle = 0.3\,\text{eV} \), \( \langle 3 \lvert \hat{V} \rvert 1 \rangle = 0.4\,\text{eV} \), and \( \langle 1 \lvert \hat{V} \rvert 1 \rangle = 0 \). Find \( E_1 \) to second order.
Solution
\( E_1^{(1)} = 0 \). \( E_1^{(2)} = \dfrac{(0.3)^2}{0-2} + \dfrac{(0.4)^2}{0-5} = \dfrac{0.09}{-2} + \dfrac{0.16}{-5} = -0.045 - 0.032 = -0.077\,\text{eV} \). So \( E_1 \approx -0.077\,\text{eV} \), pushed down as expected for the lowest level. - (State correction and normalization) For problem 3, write the first-order corrected state \( \lvert 1 \rangle \) and compute the norm-squared through \( O(V^2) \).
Solution
\( \lvert 1 \rangle = \lvert 1^{(0)} \rangle + \dfrac{0.3}{0-2}\lvert 2^{(0)} \rangle + \dfrac{0.4}{0-5}\lvert 3^{(0)} \rangle = \lvert 1^{(0)} \rangle - 0.15\lvert 2^{(0)} \rangle - 0.08\lvert 3^{(0)} \rangle \) (coefficients in units where \( \hat{V} \) is in eV). Norm-squared \( = 1 + (0.15)^2 + (0.08)^2 = 1 + 0.0225 + 0.0064 = 1.0289 \). To use the state for \( O(V^2) \) expectation values one divides by \( \sqrt{1.0289} \approx 1.0143 \); the \( O(V^2) \) excess norm is exactly \( -2E_1^{(2)}\cdot(\text{per-term structure}) \), reflecting intermediate normalization. - (Stark effect, ground-state polarizability) A hydrogen-like ground state in a uniform field \( \mathcal{E} \) along \( z \) has \( \hat{V} = e\mathcal{E}\hat{z} \). Explain why \( E^{(1)} = 0 \), and write the structure of \( E^{(2)} \); identify the polarizability \( \alpha \) via \( E^{(2)} = -\tfrac{1}{2}\alpha \mathcal{E}^2 \).
Solution
The ground state \( \lvert 1s \rangle \) has definite (even) parity, while \( \hat{z} \) is odd, so \( \langle 1s \lvert \hat{z} \rvert 1s \rangle = 0 \) and \( E^{(1)} = 0 \) — hydrogen has no linear (permanent-dipole) Stark shift in its ground state. Second order: \( E^{(2)} = (e\mathcal{E})^2 \sum_{m \neq 1s} \dfrac{\lvert \langle m \lvert \hat{z} \rvert 1s \rangle \rvert^2}{E_{1s}^{(0)} - E_m^{(0)}} \), a sum over \( p \)-states (and the continuum, which is essential). Since every denominator is negative, \( E^{(2)} < 0 \). Comparing with \( E^{(2)} = -\tfrac{1}{2}\alpha\mathcal{E}^2 \) gives \( \alpha = 2e^2 \sum_{m} \dfrac{\lvert \langle m \lvert \hat{z} \rvert 1s \rangle \rvert^2}{E_m^{(0)} - E_{1s}^{(0)}} > 0 \). The exact result (including continuum) is \( \alpha = \tfrac{9}{2}a_0^3 \) in Gaussian units; truncating to bound states alone underestimates it, illustrating the discussion's warning about the continuum.