Field Lagrangian and Gauge Invariance
Statement
Starting from the covariant field Lagrangian density \( \mathcal{L} = -\frac{1}{4\mu_0}F_{\mu\nu}F^{\mu\nu} - J^\mu A_\mu \), with \( F_{\mu\nu} = \partial_\mu A_\nu - \partial_\nu A_\mu \), we derive the inhomogeneous Maxwell equations \( \partial_\mu F^{\mu\nu} = \mu_0 J^\nu \) as the Euler–Lagrange field equations, recover the homogeneous pair as an identity of the potential formulation, and prove that invariance of the action under the local gauge transformation \( A_\mu \to A_\mu + \partial_\mu \chi \) holds if and only if the coupled current is conserved, \( \partial_\mu J^\mu = 0 \).
Why it matters
This derivation is the template for every fundamental interaction we know. Electromagnetism is here recast not as a set of experimental field laws but as the unique consequence of a variational principle plus a symmetry demand: the physics cannot depend on the unobservable phase convention encoded in \( A_\mu \). The same logic, with \( U(1) \) replaced by \( SU(2) \) or \( SU(3) \), generates the weak and strong interactions. Understanding exactly where each Maxwell equation comes from — dynamics for the inhomogeneous pair, geometry (an identity) for the homogeneous pair — is essential to reading any modern field theory.
It also settles a conceptual question of the first importance: charge conservation is not an extra empirical input bolted onto electromagnetism. Local gauge invariance forces \( \partial_\mu J^\mu = 0 \); a theory that coupled \( A_\mu \) to a non-conserved current would have an ill-defined action. Conservation of charge and the masslessness of the photon are two faces of the same symmetry.
Assumptions
Derivation
Result
Reading. All four Maxwell equations emerge from one scalar Lagrangian: the sourced pair (Gauss, Ampère–Maxwell) as genuine equations of motion, the sourceless pair (no monopoles, Faraday) as an identity of the potential description. The third relation is not optional: demanding that the action be blind to the local redefinition \( A_\mu \to A_\mu + \partial_\mu \chi \) forces the current to be conserved, and equivalently the antisymmetry of \( F^{\mu\nu} \) makes the field equation inconsistent for any non-conserved source.
Units check. \( [F^{\mu\nu}] = \mathrm{T} = \mathrm{V\,s\,m^{-2}} \), so \( [\partial_\mu F^{\mu\nu}] = \mathrm{T\,m^{-1}} \); and \( [\mu_0 J^\nu] = (\mathrm{T\,m\,A^{-1}})(\mathrm{A\,m^{-2}}) = \mathrm{T\,m^{-1}} \). Consistent. For the Lagrangian density itself: \( [F^2/\mu_0] = \mathrm{T^2}/(\mathrm{T\,m\,A^{-1}}) = \mathrm{T\,A\,m^{-1}} = \mathrm{J\,m^{-3}} \), and \( [J^\mu A_\mu] = (\mathrm{A\,m^{-2}})(\mathrm{V\,s\,m^{-1}}) = \mathrm{J\,m^{-3}} \): both terms are energy densities, as an \( \mathcal{L} \) must be.
Limiting cases
- Static sources (\( \partial_t = 0 \)): the field equations decouple into electrostatics \( \nabla \cdot \vec{E} = \rho/\varepsilon_0 \), \( \nabla \times \vec{E} = 0 \) and magnetostatics \( \nabla \times \vec{B} = \mu_0 \vec{J} \), \( \nabla \cdot \vec{B} = 0 \).
- Vacuum (\( J^\mu = 0 \)) in Lorenz gauge \( \partial_\mu A^\mu = 0 \): the field equation collapses to \( \Box A^\nu = 0 \) — free waves travelling at \( c \), with two physical polarisations after the residual gauge freedom is spent.
- Global gauge transformation (\( \chi \) constant): \( A_\mu \) is unchanged, \( \delta S = 0 \) trivially; the local demand is strictly stronger and is what carries the physics.
- Point charge at rest: \( \nu = 0 \) equation with \( \rho = q\,\delta^3(\vec{r}) \) reproduces Coulomb, \( \vec{E} = \frac{q}{4\pi\varepsilon_0 r^2}\hat{r} \).
- The theory is linear, so the "weak-field limit" is exact: superposition holds at any classical field strength within the domain of validity.
Breaks when
- Vacuum nonlinearity at extreme field strength. Near the Schwinger critical field \( E_c = m_e^2 c^3/(e\hbar) \approx 1.3 \times 10^{18}\ \mathrm{V\,m^{-1}} \), electron–positron vacuum polarisation adds Euler–Heisenberg quartic terms \( \sim (F_{\mu\nu}F^{\mu\nu})^2 \) to \( \mathcal{L} \): light-by-light scattering appears and superposition fails, though gauge invariance itself survives.
- Quantum regime. The classical action is only the saddle point of the QED path integral; for single photons, spontaneous emission, or Casimir physics the classical field equations are inadequate and \( A_\mu \) must be quantised (with gauge fixing, e.g. Faddeev–Popov, to make the path integral well defined).
- Magnetic monopoles. If \( \nabla \cdot \vec{B} \neq 0 \) anywhere, no single-valued global potential \( A_\mu \) exists; the Bianchi identity (Step 9) fails and the potential formulation needs patching (Dirac strings, fibre bundles) plus the quantisation condition \( qg = 2\pi n \hbar \).
- Massive photon. A Proca term \( \tfrac{1}{2}\mu_\gamma^2 A_\mu A^\mu \) breaks the symmetry of Step 10 explicitly: current conservation must then be imposed by hand, and static fields acquire Yukawa screening \( e^{-\mu_\gamma r}/r \). Experiment bounds \( m_\gamma \lesssim 10^{-54}\ \mathrm{kg} \).
Failure modes
- Losing the factor of 4 in Step 3. Differentiating \( F_{\alpha\beta}F^{\alpha\beta} \) and getting \( 2F^{\mu\nu} \) or \( F^{\mu\nu} \) instead of \( 4F^{\mu\nu} \): you must count both the product-rule factor and the two-delta antisymmetric contraction. The wrong count produces a spurious \( \tfrac{1}{2} \) in Maxwell's equations.
- Treating \( \partial_\mu A_\nu \) and \( \partial_\nu A_\mu \) as the same variable. They are independent components of the derivative array; the antisymmetrisation comes out of the algebra, it is not to be imposed by hand.
- Claiming \( \mathcal{L} \) itself is gauge invariant. It is not: \( \delta \mathcal{L} = -J^\mu \partial_\mu \chi \neq 0 \). Only the action is invariant, and only when \( \partial_\mu J^\mu = 0 \) and boundary terms vanish. Conflating the two hides exactly the physics this derivation exists to expose.
- Deriving the homogeneous equations from the Lagrangian. Students often try to extract \( \nabla \cdot \vec{B} = 0 \) from the Euler–Lagrange machinery. It cannot be done — it is an identity of \( F = \partial A \), true off-shell, independent of any dynamics.
- Signature slips. With \( (-,+,+,+) \) the sign of the \( F^2 \) term flips to \( -\tfrac{1}{4\mu_0}F^2 \to \) same expression but component dictionary changes (\( F_{0i} \) picks up a sign). Mixing conventions mid-derivation flips the sign of the displacement-current term.
- Varying \( J^\mu \) as if it were dynamical. For an external current, \( J^\mu \) is held fixed under \( \delta A \); varying it produces meaningless extra "equations".
Discussion
The deepest lesson is the direction of the logic. One does not start from Maxwell's equations and observe, as a curiosity, that they admit gauge transformations. One starts from the demand that physics be invariant under local redefinition of the potential and asks: what couplings are then allowed? The answer is startlingly restrictive. The kinetic term must be built from the invariant \( F_{\mu\nu} \); a mass term \( A_\mu A^\mu \) is forbidden (the photon is massless because of gauge symmetry, not by accident); and the linear coupling \( J^\mu A_\mu \) is admissible only for a conserved current. Symmetry dictates interaction.
Steps 13 and 14 are two theorems meeting in the middle. Step 13 is an instance of Noether's second theorem: a continuous symmetry parametrised by an arbitrary function (not a constant) yields an off-shell identity among the equations of motion, here \( \partial_\nu(\partial_\mu F^{\mu\nu} - \mu_0 J^\nu) \equiv 0 \). Step 14 shows the field equations enforcing the same constraint from pure antisymmetry. This redundancy is why only four of the eight Maxwell component equations are independent dynamical statements, and why the initial-value problem for \( A_\mu \) requires gauge fixing: the theory literally contains fewer physical degrees of freedom (two) than field components (four).
The template generalises with almost no new ideas. Replace the phase group \( U(1) \) by a non-abelian group, promote \( \partial_\mu \) to a covariant derivative \( D_\mu = \partial_\mu + igT^aA^a_\mu \), and the field strength acquires a self-coupling term \( F^a_{\mu\nu} = \partial_\mu A^a_\nu - \partial_\nu A^a_\mu - g f^{abc} A^b_\mu A^c_\nu \): the Yang–Mills Lagrangian of the strong and electroweak interactions. The photon does not self-interact because \( U(1) \) is abelian (\( f^{abc}=0 \)); gluons do because \( SU(3) \) is not.
A sharper statement of what gauge symmetry "is": it is not a physical symmetry mapping one state to a different state, but a redundancy of description — the physical configuration space is the quotient of field space by gauge orbits. Yet this redundancy has teeth. In the Hamiltonian formulation, Gauss's law \( \nabla \cdot \vec{E} - \rho/\varepsilon_0 = 0 \) is not an evolution equation at all but a first-class constraint that generates the gauge transformations; it is preserved by the dynamics and must be imposed on initial data once, after which Ampère–Maxwell propagates it. Quantum mechanically the redundancy becomes physical at the boundary: the Aharonov–Bohm phase \( \exp\left( \frac{iq}{\hbar}\oint A_\mu dx^\mu \right) \) is gauge invariant and observable even where \( F_{\mu\nu} = 0 \), showing that the holonomies of \( A_\mu \), not the field strengths alone, exhaust the physical content of the theory.
Common misconceptions. (i) "Gauge invariance is empty because it's just redundancy" — the redundancy constrains the allowed interactions and outlaws a photon mass, which is as physical as it gets. (ii) "\( \mathcal{L} = 0 \) for a plane wave means the field does nothing" — \( \mathcal{L} \) is not the energy density; the invariant \( \tfrac{1}{2}(\varepsilon_0 E^2 - B^2/\mu_0) \) vanishing merely says the field is null, while the energy density \( \tfrac{1}{2}(\varepsilon_0 E^2 + B^2/\mu_0) \) is strictly positive. (iii) "Charge conservation follows from Gauss's law alone" — it follows from the antisymmetry of \( F^{\mu\nu} \) (all four inhomogeneous equations together), or equivalently from gauge invariance; a single component equation does not deliver it.
Worked examples
Example 1 — the Lagrangian density of a plane wave is zero. A linearly polarised plane wave in vacuum has \( E_0 = 3.0 \times 10^{2}\ \mathrm{V\,m^{-1}} \) and \( B_0 = E_0/c \). Evaluate \( \mathcal{L} \).
Reading. A plane wave is a null field: its two Lorentz invariants vanish, so every inertial observer agrees the Lagrangian density is zero — while the energy density \( u = \tfrac{1}{2}(\varepsilon_0 E_0^2 + B_0^2/\mu_0) \approx 8.0 \times 10^{-7}\ \mathrm{J\,m^{-3}} \) is emphatically not.
Units check. Both terms carried \( \mathrm{J\,m^{-3}} \) throughout; their difference is a legitimate energy-density-valued scalar.
Example 2 — gauge invariance dictates the charge density accompanying a wave of current. An antenna carries \( \vec{J} = J_0 \sin(kz - \omega t)\,\hat{z} \) with \( J_0 = 5.0 \times 10^{2}\ \mathrm{A\,m^{-2}} \), \( \omega = 2\pi \times 10^{8}\ \mathrm{rad\,s^{-1}} \), and \( k = \omega/c \). What charge density must coexist with it for the coupling \( J^\mu A_\mu \) to be admissible?
Reading. The current wave cannot exist alone: gauge invariance of the action (equivalently, consistency of \( \partial_\mu F^{\mu\nu} = \mu_0 J^\nu \)) forces a charge-density wave locked in phase with it, amplitude \( J_0/c \). An antenna's current and charge distributions are not independent design choices.
Units check. \( [J_0/c] = \mathrm{C\,m^{-3}} \) as required for a charge density; the sinusoid is dimensionless.
Problems
- Show that \( F_{\mu\nu}F^{\mu\nu} = 2\left( B^2 - E^2/c^2 \right) \), then evaluate it for a region where \( E = 1.0 \times 10^{3}\ \mathrm{V\,m^{-1}} \) and \( B = 1.0 \times 10^{-5}\ \mathrm{T} \) (fields perpendicular). Is the field electric-dominated or magnetic-dominated?
Solution
Split the double sum: the six independent components give \( F_{\mu\nu}F^{\mu\nu} = 2F_{0i}F^{0i} + F_{ij}F^{ij} \). With \( F_{0i} = E_i/c \), \( F^{0i} = -E_i/c \), the electric part is \( 2\sum_i (E_i/c)(-E_i/c) = -2E^2/c^2 \). With \( F_{ij} = -\varepsilon_{ijk}B_k \), the magnetic part is \( \sum_{ij} \varepsilon_{ijk}\varepsilon_{ijl}B_k B_l = 2\delta_{kl}B_kB_l = 2B^2 \). Total: \( 2(B^2 - E^2/c^2) \). Numerically: \( B^2 = 1.0 \times 10^{-10}\ \mathrm{T^2} \); \( E^2/c^2 = (1.0\times 10^6)/(8.988 \times 10^{16}) = 1.11 \times 10^{-11}\ \mathrm{T^2} \). So \( F_{\mu\nu}F^{\mu\nu} = 2(1.0\times 10^{-10} - 1.11\times 10^{-11}) = 1.78 \times 10^{-10}\ \mathrm{T^2} > 0 \): magnetic-dominated. There exists a frame in which the field is purely magnetic; no frame makes it purely electric. - Starting from \( \partial_\mu F^{\mu 0} = \mu_0 J^0 \), derive Gauss's law and use it to find the charge density in a region where \( \vec{E} = \alpha x\,\hat{x} \) with \( \alpha = 1.0 \times 10^{2}\ \mathrm{V\,m^{-2}} \).
Solution
Only spatial derivatives contribute (\( F^{00} = 0 \) by antisymmetry): \( \partial_i F^{i0} = \mu_0 J^0 \). The dictionary gives \( F^{i0} = E^i/c \) and \( J^0 = c\rho \), so \( (1/c)\nabla \cdot \vec{E} = \mu_0 c \rho \Rightarrow \nabla \cdot \vec{E} = \mu_0 c^2 \rho = \rho/\varepsilon_0 \). Here \( \nabla \cdot \vec{E} = \partial_x(\alpha x) = \alpha \), so \( \rho = \varepsilon_0 \alpha = (8.854 \times 10^{-12}\ \mathrm{F\,m^{-1}})(1.0 \times 10^{2}\ \mathrm{V\,m^{-2}}) = 8.85 \times 10^{-10}\ \mathrm{C\,m^{-3}} \), uniform throughout the region. - Apply the gauge transformation generated by \( \chi(t) = C t^2 \) with \( C = 2.0\ \mathrm{V\,s} \) to a configuration with scalar potential \( \phi = 10\ \mathrm{V} \) (uniform) and \( \vec{A} = 0 \). Find the new potentials at \( t = 3.0\ \mathrm{s} \) and verify \( \vec{E} \) and \( \vec{B} \) are unchanged.
Solution
In 3+1 language the transformation \( A_\mu \to A_\mu + \partial_\mu \chi \) reads \( \phi \to \phi - \partial_t \chi \), \( \vec{A} \to \vec{A} + \nabla \chi \). Here \( \partial_t \chi = 2Ct \) and \( \nabla \chi = 0 \), so \( \phi' = 10 - 2(2.0)(3.0) = -2.0\ \mathrm{V} \) at \( t = 3.0\ \mathrm{s} \), and \( \vec{A}' = 0 \). Fields: originally \( \vec{E} = -\nabla \phi - \partial_t \vec{A} = 0 \), \( \vec{B} = \nabla \times \vec{A} = 0 \). After: \( \phi' \) is still spatially uniform so \( \nabla \phi' = 0 \), and \( \vec{A}' = 0 \), hence \( \vec{E}' = 0 = \vec{E} \), \( \vec{B}' = 0 = \vec{B} \). The 12 V shift in the potential at \( t = 3\ \mathrm{s} \) is pure description; nothing measurable moved. - A proposed source has \( \rho = \rho_0 e^{-t/\tau} \) (uniform in space) and \( \vec{J} = \frac{\rho_0 z}{\tau} e^{-t/\tau}\,\hat{z} \), with \( \rho_0 = 1.0 \times 10^{-6}\ \mathrm{C\,m^{-3}} \) and \( \tau = 2.0\ \mathrm{\mu s} \). (a) Verify this current may legally couple to \( A_\mu \). (b) Evaluate \( J_z \) at \( z = 0.10\ \mathrm{m} \), \( t = \tau \).
Solution
(a) Continuity: \( \partial_t \rho = -\frac{\rho_0}{\tau}e^{-t/\tau} \) and \( \nabla \cdot \vec{J} = \partial_z J_z = \frac{\rho_0}{\tau}e^{-t/\tau} \). Sum: \( \partial_t \rho + \nabla \cdot \vec{J} = 0 \) identically — the current is conserved, so the coupling \( -J^\mu A_\mu \) yields a gauge-invariant action and the field equations are consistent. (b) \( J_z = \frac{\rho_0 z}{\tau}e^{-1} = \frac{(1.0 \times 10^{-6})(0.10)}{2.0 \times 10^{-6}} \times 0.3679 = (5.0 \times 10^{-2}) \times 0.3679 = 1.8 \times 10^{-2}\ \mathrm{A\,m^{-2}} \). Units: \( \mathrm{C\,m^{-3}\,m\,s^{-1}} = \mathrm{A\,m^{-2}} \). - Add a Proca mass term so \( \mathcal{L} = -\frac{1}{4\mu_0}F_{\mu\nu}F^{\mu\nu} + \frac{\mu_\gamma^2}{2\mu_0} A_\mu A^\mu - J^\mu A_\mu \). (a) Show the field equation becomes \( \partial_\mu F^{\mu\nu} + \mu_\gamma^2 A^\nu = \mu_0 J^\nu \) and that taking \( \partial_\nu \) of it forces the Lorenz condition when \( \partial_\mu J^\mu = 0 \). (b) Show the transformation of Step 10 no longer leaves the action invariant. (c) The experimental bound on the photon mass is \( m_\gamma \lesssim 1.0 \times 10^{-54}\ \mathrm{kg} \). Compute the corresponding minimum range \( \lambda = \hbar/(m_\gamma c) \) and compare with an astronomical scale.
Solution
(a) The mass term contributes \( \partial \mathcal{L}/\partial A_\nu = (\mu_\gamma^2/\mu_0) A^\nu \) to the Euler–Lagrange equation, giving \( -\frac{1}{\mu_0}\partial_\mu F^{\mu\nu} + \frac{\mu_\gamma^2}{\mu_0} A^\nu - J^\nu = 0 \), i.e. \( \partial_\mu F^{\mu\nu} + \mu_\gamma^2 A^\nu = \mu_0 J^\nu \) (here \( \mu_\gamma = m_\gamma c/\hbar \) has units \( \mathrm{m^{-1}} \)). Apply \( \partial_\nu \): the first term dies by antisymmetry, the source term dies by conservation, leaving \( \mu_\gamma^2\, \partial_\nu A^\nu = 0 \). For \( \mu_\gamma \neq 0 \) the Lorenz condition \( \partial_\nu A^\nu = 0 \) is a forced constraint, not a gauge choice — there is no gauge freedom left to choose with. (b) Under \( A_\mu \to A_\mu + \partial_\mu \chi \), the mass term shifts by \( \frac{\mu_\gamma^2}{2\mu_0}\left( 2A^\mu \partial_\mu \chi + \partial_\mu \chi\, \partial^\mu \chi \right) \), which is not a total divergence for arbitrary \( \chi \): gauge invariance is explicitly broken. (c) \( \lambda = \hbar/(m_\gamma c) = \frac{1.055 \times 10^{-34}\ \mathrm{J\,s}}{(1.0 \times 10^{-54}\ \mathrm{kg})(2.998 \times 10^{8}\ \mathrm{m\,s^{-1}})} = 3.5 \times 10^{11}\ \mathrm{m} \) — about 2.3 astronomical units. Any photon mass would exponentially screen magnetic fields beyond \( \lambda \); the observed coherence of planetary and solar-wind magnetic fields over such scales is precisely how the bound is set.