Basis sets
Statement
Representing orbitals for practical computation.
Why it matters
hartree-fock formulates the many-electron problem as a set of coupled one-electron equations for molecular orbitals, but leaves open the practical question of how to actually represent those orbitals as concrete mathematical functions a computer can manipulate. Basis sets are the answer: a chosen, finite set of functions in terms of which every molecular orbital is expanded, and the specific choice of basis set is one of the two central decisions (along with the choice of electronic-structure method, such as Hartree-Fock or density-functional-theory) governing the accuracy, and the computational cost, of essentially any quantum-chemical calculation.
Because a genuinely complete, infinite basis is never computationally affordable, understanding what a finite basis captures well and what it captures poorly is essential background for potential-energy-surfaces and for interpreting any computed energy, geometry, or property sensibly, including recognising when a result is basis-set-limited rather than a genuine physical prediction.
Hypotheses
Proof
Result
Reading. A larger, more flexible basis set gives a systematically more accurate (lower, closer to exact) computed energy, at systematically higher computational cost; basis-set choice is fundamentally a controllable trade-off between accuracy and affordability, not an all-or-nothing decision.
Scope. Applies within a fixed electronic-structure method (e.g. Hartree-Fock, or a given density functional); comparing energies computed with different basis sets, or different methods, directly is not meaningful without accounting for this dependence.
Corollaries & converses
- hartree-fock's own accuracy ceiling (the "Hartree-Fock limit," the energy obtained as \(M\to\infty\) within the Hartree-Fock method specifically) is a direct consequence of Step 5 applied to that one method; the remaining error relative to experiment beyond that limit is attributed to electron correlation, addressed instead by improving the method itself (e.g. via density-functional-theory or post-Hartree-Fock approaches) rather than by enlarging the basis further.
- Because basis-set incompleteness and method incompleteness are separate sources of error, a genuinely reliable calculation must be converged with respect to both, typically demonstrated by repeating a calculation with successively larger basis sets and confirming the computed property has stabilised.
- Converse: if enlarging a basis set produces no further meaningful change in a computed property, the remaining discrepancy from experiment is attributable to the chosen method's own limitations (e.g. missing electron correlation), not to basis-set incompleteness, a useful diagnostic when troubleshooting a calculation.
Fails without
- Compare raw computed energies from two different, unequal-quality basis sets: since energy depends systematically on basis size (Step 5), any difference between the two numbers conflates genuine chemistry with basis-set incompleteness, making the comparison meaningless unless corrected for.
- Omit diffuse functions when modelling an anion or excited state: compact, atom-centred basis functions cannot represent the more spatially extended electron density these systems require (Step 4), giving qualitatively poor results regardless of how large the rest of the basis is.
Common errors
- Comparing raw computed energies from calculations run with different basis sets (or different methods) as though the difference were chemically meaningful, rather than partly or wholly a basis-set artefact.
- Omitting polarisation functions when studying bond angles or bonding distortions specifically, then attributing the resulting inaccuracy to the electronic-structure method rather than to basis-set inadequacy (Step 4).
- Neglecting diffuse functions when modelling an anion or an excited state, where the extra spatial extent of the electron density is not optional but essential to a qualitatively reasonable result.
- Ignoring basis set superposition error when computing a binding or interaction energy from separately optimised fragments, artificially inflating the apparent binding strength.
Discussion
The shift from Slater-type functions (which more faithfully reproduce the true exponential decay and the sharp cusp of an atomic orbital at the nucleus) to computationally far cheaper Gaussian-type functions was central to making routine molecular quantum chemistry computationally practical; John Pople's development of widely used, systematically constructed Gaussian basis sets (and associated methods) was recognised with a share of the 1998 Nobel Prize in Chemistry.
Because no single finite basis set is Gaussian-shaped exactly like an atomic orbital near the nucleus, several Gaussian primitives are typically combined ("contracted") into one effective basis function to approximate the correct cusp behaviour better than any single Gaussian could alone — a practical compromise between the mathematical convenience of Gaussians and the physically more accurate shape of a true atomic orbital.
Common misconception: that a "bigger" basis set is simply, unconditionally better with no downside beyond cost. In practice a basis set must also be balanced — for instance, comparably flexible for every atom type and bonding situation present in a molecule — since an unbalanced combination (very large on one atom, minimal on another) can introduce its own artefacts rather than uniformly improving accuracy.
Worked examples
Reading. Basis-set quality affects more than the raw computed energy; structural predictions such as bond angles and bond lengths are similarly sensitive to whether the basis is flexible enough to describe the true, distorted orbital shapes present in a bonded molecule.
Scope. The same qualitative lesson — that polarisation and, where relevant, diffuse functions materially change predicted structure and energetics, not only convergence speed — holds generally across organic and inorganic molecules alike.
Problems
- Explain why the product of two Gaussian functions centred at different points is itself a single Gaussian, and why this matters computationally.
Solution
The Gaussian product theorem states that the product of two Gaussians centred at different points is a third Gaussian centred somewhere on the line between them (with a modified exponent and a scaling prefactor); because a molecular integral over two, three, or four basis functions ultimately reduces to integrals over such products, this theorem lets a computer evaluate them essentially in closed form, extremely quickly, in contrast to the far more computationally demanding integrals arising from Slater-type functions, which lack this convenient algebraic property. - A calculation of a small molecule's energy is repeated with basis sets of increasing size, giving energies of \(-76.02\), \(-76.05\), and \(-76.06\) hartree respectively. Is this trend consistent with the variational principle, and what does the shrinking gap between successive values suggest?
Solution
Yes: the energies decrease monotonically as the basis grows, exactly as Step 5 requires for a variational method. The shrinking gap between successive values (\(0.03\) then \(0.01\) hartree) indicates the calculation is approaching convergence toward the complete-basis-set limit for this method; a further increase in basis size would be expected to change the energy by even less. - Why is basis set superposition error (BSSE) a particular concern for weakly bound complexes (e.g. hydrogen-bonded dimers) but comparatively less important for strongly, covalently bonded molecules?
Solution
BSSE artificially lowers a complex's computed energy because each fragment can use the other fragment's nearby basis functions to improve its own description. For a weakly bound complex, the true interaction energy being measured is itself small, so even a modest artificial lowering from BSSE can be a large fraction of, or even exceed, the genuine binding energy, badly distorting the result. For a strongly covalent bond, the true bonding energy is far larger than typical BSSE magnitudes, so the same absolute error is comparatively insignificant.