Gaussian Beams from the Paraxial Wave Equation
Statement
Starting from the paraxial Helmholtz equation for the slowly varying envelope of a monochromatic scalar field, we derive the fundamental (TEM\(_{00}\)) Gaussian mode: its transverse profile, the complex beam parameter \(q(z)\), the waist \(w_0\), the Rayleigh range \(z_R\), the spot size \(w(z)\), the wavefront radius \(R(z)\), and the Gouy phase \(\psi(z)\). We then show that \(q\) transforms under free-space and thin-lens propagation exactly as a ray coordinate does through an ABCD matrix, giving the ABCD law \(q' = (Aq+B)/(Cq+D)\).
Why it matters
Every real laser beam is a finite-aperture bundle, not a plane wave and not a geometric ray. The fundamental Gaussian mode is the exact self-consistent solution of the paraxial wave equation that a stable spherical-mirror cavity selects, so it is the working description of laser light for beam delivery, focusing, coupling into fibres and cavities, and Gaussian optics design.
The complex beam parameter compresses the entire transverse structure — spot size and wavefront curvature — into a single complex number that propagates by the same ABCD matrices used for rays. This unifies ray optics (PU thread ray-transfer-abcd-matrices) and wave optics into one algebra, so a system already characterised by its ray matrix immediately tells you what it does to a Gaussian beam.
Assumptions
Derivation
Result
Reading. A single complex number \(q\) carries everything: its imaginary part fixes the beam width through \(w\), its real part fixes the wavefront curvature through \(R\). At the waist \(q=iz_R\) is purely imaginary (flat wavefront, minimum spot \(w_0\)). Moving a distance \(z\) just adds \(z\) to \(q\); a lens subtracts \(1/f\) from \(1/q\); any paraxial system acts on \(q\) through its ray ABCD matrix by the same fractional-linear map that would move a ray. The beam expands hyperbolically, has far-field half-angle \(\theta=\lambda/(\pi w_0)\), and accumulates an extra \(\pi\) of Gouy phase as it passes through focus.
Units check. \(z_R=\pi w_0^2/\lambda\): \([\text{m}^2]/[\text{m}]=[\text{m}]\), a length, as required for \(q=z+iz_R\). \(w(z)=w_0\sqrt{1+(z/z_R)^2}\): argument dimensionless, so \([w]=[\text{m}]\). \(1/q\) has units \([\text{m}^{-1}]\); \(1/R\) is \([\text{m}^{-1}]\) and \(\lambda/(\pi w^2)=[\text{m}]/[\text{m}^2]=[\text{m}^{-1}]\) match. In \(q_2=(Aq_1+B)/(Cq_1+D)\), \(A,D\) are dimensionless, \(B\) is a length, \(C\) is inverse length, so \(Aq+B\) is a length, \(Cq+D\) is dimensionless, and \(q_2\) is a length. Consistent.
Limiting cases
- At the waist (\(z=0\)): \(w=w_0\), \(R\to\infty\) (flat wavefront), \(\psi=0\); the field is a pure real Gaussian times \(e^{-ikz}\).
- Near field (\(|z|\ll z_R\)): \(w\approx w_0\), the beam is essentially collimated; \(R\approx z_R^2/z\to\infty\).
- Far field (\(|z|\gg z_R\)): \(w\approx w_0 z/z_R=\theta z\) with divergence \(\theta=\lambda/(\pi w_0)\); \(R\approx z\) (spherical wave from the waist); \(\psi\to\pm\pi/2\).
- Plane-wave limit (\(w_0\to\infty\), \(z_R\to\infty\)): \(q\to iz_R\to i\infty\), \(1/q\to 0\), envelope flat — recovers the uniform plane wave.
- Ray limit: identify \(q=x/x'\) (position over slope); a real \(q\) is a geometric ray and the ABCD law reduces to the ray-transfer law from ray-transfer-abcd-matrices.
- Total Gouy shift: from \(z=-\infty\) to \(+\infty\), \(\Delta\psi=\pi\) (a half-integer wave of extra phase through focus).
Breaks when
- Tight focusing / high NA: when \(w_0\lesssim\lambda\) the divergence \(\theta=\lambda/(\pi w_0)\) is no longer small, the paraxial assumption (Step 3) fails, \(\partial_z^2\psi\) cannot be dropped, and the true focal field requires vector (Richards–Wolf) diffraction. The scalar Gaussian underestimates the spot and ignores longitudinal field components.
- Nonlinear or inhomogeneous media: intensity-dependent index (Kerr self-focusing, filamentation) or a graded/aberrated index makes the medium non-uniform, so the separable single-\(q\) ansatz (Step 4) is invalid and \(q\) no longer obeys a linear ABCD law.
- Hard apertures / truncation: clipping the beam on a finite aperture injects higher-order and diffraction-ring content; the field is no longer a single Gaussian mode and \(w\), \(R\) lose their simple meaning (Fresnel diffraction of the truncated field is needed).
- Gain, loss, or strong dispersion: a complex or frequency-dependent \(k\) (amplifier, absorber, or short pulse) breaks the real-\(k\) monochromatic assumption; the beam parameter becomes complex-\(k\) or frequency dependent and pulses acquire spatio-temporal coupling.
Failure modes
- Confusing \(1/e\) field radius with \(1/e^2\) intensity radius: \(w\) is where the field falls to \(1/e\) and the intensity to \(1/e^2\). Students halve or double \(w\) by mixing these; the intensity FWHM is \(w\sqrt{2\ln 2}\approx 1.18\,w\), not \(w\).
- Sign / factor errors in \(1/q\): writing \(1/q=1/R+i\lambda/\pi w^2\) with the wrong sign flips the beam between converging and diverging, or between a physical (decaying) and unphysical (growing) Gaussian.
- Adding \(w\) or \(R\) directly through a system instead of propagating \(q\): only \(q\) obeys the ABCD law. \(R\) and \(w\) combine nonlinearly, so you must go through \(q\).
- Forgetting the Gouy phase: omitting \(\psi(z)\) misplaces resonance frequencies of a cavity and the on-axis phase; the extra \(\pi\) through focus is physically real (measured phase anomaly).
- Using \(R(z)=z\) everywhere: that is only the far-field limit. Near the waist \(R\to\infty\), and \(R\) has a minimum \(2z_R\) at \(z=z_R\).
- Treating the on-axis intensity as constant: peak intensity falls as \(w_0^2/w^2(z)\); the \(w_0/w\) amplitude factor is what conserves total power, not the peak.
- Non-normalised ray matrices: the ABCD law assumes \(\det M=n_1/n_2=1\) in a single medium; using an un-normalised matrix corrupts \(q_2\).
Discussion
The paraxial wave equation \(\nabla_\perp^2\psi=2ik\,\partial_z\psi\) is formally the time-dependent Schrödinger equation of a free 2D particle, with \(z\leftrightarrow t\) and \(\hbar/2m\leftrightarrow 1/2k\). The Gaussian beam is then the exact analogue of a minimum-uncertainty coherent-state wavepacket: it stays Gaussian, spreads with distance, and its "waist" is where the transverse momentum and position uncertainties are minimally correlated. The Rayleigh range is the diffraction analogue of the wavepacket spreading time, and the hyperbolic \(w(z)\) is the ballistic spreading of that packet. This is why the same complex-\(q\) trick reappears throughout wave physics.
The complex beam parameter is powerful precisely because it linearises a nonlinear problem. Spot size and curvature evolve nonlinearly and couple to each other, but their specific combination \(1/q=1/R-i\lambda/\pi w^2\) evolves by a fractional-linear (Möbius) map, and Möbius maps compose exactly as \(2\times 2\) matrices multiply. That algebraic isomorphism is the whole content of the ABCD law: the ray matrix you already computed for geometric optics, unchanged, propagates the wave-optical \(q\). Cavity stability, mode matching into fibres, and telescope design all reduce to iterating or inverting that one map.
The Gouy phase \(\psi(z)=\arctan(z/z_R)\) is a genuinely wave phenomenon with no ray counterpart: a focused beam accumulates \(\pi\) less phase over its length than a plane wave would, an "anomalous" advance concentrated within a few \(z_R\) of focus. Physically it reflects the spread of transverse momenta: components with nonzero \(k_\perp\) travel with reduced axial wavenumber \(k_z=\sqrt{k^2-k_\perp^2}\approx k-k_\perp^2/2k\), and averaging over the mode's momentum distribution yields exactly the Gouy shift. It sets the longitudinal mode spacing in resonators and, generalised, underlies the topological phase of structured light.
The full mode family emerges by the same route: separating the paraxial equation in Cartesian coordinates yields Hermite–Gaussian modes \(H_m(\sqrt2 x/w)H_n(\sqrt2 y/w)\) and in cylindrical coordinates Laguerre–Gaussian modes carrying orbital angular momentum \(\ell\hbar\) per photon. All share the same \(q(z)\), \(w(z)\), \(R(z)\); only their Gouy phase generalises to \((m{+}n{+}1)\psi\) or \((2p{+}|\ell|{+}1)\psi\). Because different transverse orders acquire different Gouy phases, they are non-degenerate in a real cavity — the physical origin of transverse mode structure and of the ability to select TEM\(_{00}\) by aperturing. These sets are complete, so any paraxial field is a superposition with fixed \(q\), reducing free-space propagation to a phase bookkeeping over mode indices.
Common misconceptions. A Gaussian beam is not a ray bundle that focuses to a point — diffraction forbids a spot smaller than \(\sim\lambda\); the waist \(w_0\) and divergence \(\theta\) are reciprocally locked by \(w_0\theta=\lambda/\pi\), a diffraction "conservation law." Nor does the beam "come from" the waist as a point source except in the far field. And \(q\) is not merely a bookkeeping convenience: its imaginary part is a physical, invariant measure (the beam's \(M^2=1\) quality) that no lossless linear system can degrade.
Worked examples
Reading. The beam stays near \(w_0\) for over a metre, then spreads: at 2 m it has nearly doubled. A tighter waist would diverge faster — the reciprocal lock \(w_0\theta=\lambda/\pi\) is unavoidable.
Units check. \(z_R\) in metres, \(\theta\) dimensionless (radians), \(w\) in millimetres. All consistent.
Reading. A collimated beam focuses essentially one focal length behind the lens to a spot \(w_0'\approx f\lambda/\pi w_0\): a larger input beam or shorter \(f\) gives a tighter focus. The tiny \(z_R'=3.4\) mm shows the focused beam is very short — depth of focus \(2z_R'\approx 7\) mm.
Units check. \(1/q\) in \(\text{m}^{-1}\), \(q\) in m, \(w_0'\) in m. \(f\lambda/\pi w_0\): \([\text m][\text m]/[\text m]=[\text m]\). Consistent.
Problems
- A beam has \(w_0=0.2\) mm at \(\lambda=800\) nm. Compute \(z_R\) and the spot size at \(z=0.5\) m.
Solution
\(z_R=\pi w_0^2/\lambda=\pi(2\times10^{-4})^2/(8\times10^{-7})=\pi(4\times10^{-8})/(8\times10^{-7})=0.157\) m. At \(z=0.5\) m: \(z/z_R=0.5/0.157=3.18\), so \(w=0.2\,\sqrt{1+3.18^2}=0.2\sqrt{1+10.1}=0.2\times3.33=0.67\) mm. - Show that the wavefront radius \(R(z)=z(1+z_R^2/z^2)\) is minimum at \(z=z_R\) and find \(R_\text{min}\).
Solution
\(R=z+z_R^2/z\). Then \(dR/dz=1-z_R^2/z^2=0\Rightarrow z=z_R\). Second derivative \(2z_R^2/z^3>0\), a minimum. \(R_\text{min}=z_R+z_R^2/z_R=2z_R\). So the most sharply curved wavefront occurs one Rayleigh range from the waist, with radius \(2z_R\). - A Gaussian beam with \(q_1=0.5i\) m (waist, \(z_R=0.5\) m) propagates 0.5 m in free space. Find \(w\) and \(R\) at the new plane. (\(\lambda=1\ \mu\)m.)
Solution
Free space: \(q_2=q_1+d=0.5i+0.5=0.5+0.5i\) m. Then \(1/q_2=(0.5-0.5i)/(0.25+0.25)=(0.5-0.5i)/0.5=1-1\,i\ \text{m}^{-1}\). So \(1/R=\mathrm{Re}=1\Rightarrow R=1\) m. \(\mathrm{Im}(1/q)=-\lambda/\pi w^2=-1\Rightarrow w^2=\lambda/\pi=10^{-6}/\pi=3.18\times10^{-7}\), \(w=5.64\times10^{-4}\) m \(=0.56\) mm. Check: at \(z=z_R\), \(w=w_0\sqrt2\); waist \(w_0=\sqrt{\lambda z_R/\pi}=\sqrt{10^{-6}\cdot0.5/\pi}=3.99\times10^{-4}\) m, and \(w_0\sqrt2=5.64\times10^{-4}\) m. Consistent. - Derive the far-field divergence \(\theta=\lambda/\pi w_0\) from \(w(z)\), and state the beam-parameter product \(w_0\theta\).
Solution
For \(z\gg z_R\), \(w(z)=w_0\sqrt{1+z^2/z_R^2}\to w_0\,z/z_R\). The half-angle is \(\theta=\lim_{z\to\infty}w(z)/z=w_0/z_R\). Substituting \(z_R=\pi w_0^2/\lambda\): \(\theta=w_0/(\pi w_0^2/\lambda)=\lambda/(\pi w_0)\). The product \(w_0\theta=\lambda/\pi\), a diffraction invariant independent of \(w_0\) — squeezing the waist proportionally widens the divergence. - A beam of waist \(w_0=1\) mm (\(\lambda=633\) nm) passes through a telescope of angular magnification \(M=5\) (ray matrix \(\mathrm{diag}(1/M,\,M)=\mathrm{diag}(0.2,5)\), \(B=C=0\)). Find the output waist assuming input and output planes are at the respective waists.
Solution
At a waist \(q_1=iz_R\), \(z_R=\pi w_0^2/\lambda=\pi(10^{-3})^2/6.33\times10^{-7}=4.96\) m, so \(q_1=4.96\,i\). ABCD law with \(A=0.2,B=0,C=0,D=5\): \(q_2=(Aq_1+B)/(Cq_1+D)=0.2\,q_1/5=0.04\,q_1=0.04\times4.96\,i=0.198\,i\). Purely imaginary, so output plane is a waist with \(z_R'=0.198\) m. Then \(w_0'=\sqrt{\lambda z_R'/\pi}=\sqrt{6.33\times10^{-7}\times0.198/\pi}=\sqrt{3.99\times10^{-8}}=2.0\times10^{-4}\) m \(=0.2\) mm. So the waist scales by \(1/M\): a \(5\times\) angular telescope shrinks the beam \(5\times\) (and expands its divergence \(5\times\)), as expected since \(q\to q/M^2\) rescales \(z_R\) by \(1/M^2\) and \(w_0\propto\sqrt{z_R}\) by \(1/M\).