physics2u
Tier
⌕ Search ⌘K
Derivation

Proca Field and Its Degrees of Freedom

D-375 Home PU-402 Threads fields · force · matter Depends on Maxwell's Equations from −¼F²
Statement

Starting from the Maxwell Lagrangian augmented by a mass term, \(\mathcal{L}=-\tfrac{1}{4}F_{\mu\nu}F^{\mu\nu}+\tfrac{1}{2}m^{2}A_{\mu}A^{\mu}\), we show that the mass term breaks the gauge invariance \(A_\mu\to A_\mu+\partial_\mu\lambda\); that the resulting field equation \(\partial_\mu F^{\mu\nu}+m^{2}A^{\nu}=0\) forces the Lorenz condition \(\partial_\mu A^{\mu}=0\) as a dynamical constraint (not a gauge choice); and that each surviving component obeys \((\Box+m^{2})A^{\nu}=0\), leaving a spin-1 field with exactly three physical polarizations.

Why it matters

The photon is massless because gauge invariance forbids a mass term. The moment you insist a vector boson be massive — as the \(W\) and \(Z\) are — you must confront what a mass does to the two-derivative structure of electromagnetism. The Proca field is the minimal, self-consistent answer, and the counting "four components minus one constraint equals three polarizations" is the prototype for every massive spin-1 field in nature.

It also sets up the central puzzle the Higgs mechanism resolves: a naive mass term destroys gauge invariance and, with it, the renormalizability that made QED work. Understanding precisely how the Proca constraint removes the fourth degree of freedom is what lets you later recognise the "eaten" Goldstone boson as that same longitudinal mode restored by spontaneous symmetry breaking.

Assumptions
Flat Minkowski spacetime, metric signature \((+,-,-,-)\).On curved backgrounds \(\partial_\mu\to\nabla_\mu\) and \([\nabla_\mu,\nabla_\nu]A^\nu\) generates curvature terms; the clean divergence identity below picks up a Ricci contribution and the constraint is modified.
Strictly \(m\neq 0\).At \(m=0\) the step "divide by \(m^{2}\)" is illegal, the Lorenz condition reverts to an optional gauge choice, and the degree-of-freedom count drops from three to two.
Free field: no coupling to sources or self-interaction.With a source \(J^\nu\) the constraint becomes \(m^{2}\partial_\nu A^\nu=-\partial_\nu J^\nu\); it only reduces to \(\partial_\nu A^\nu=0\) when \(J^\nu\) is conserved.
The Lagrangian is at most quadratic in \(A_\mu\), so \(F_{\mu\nu}\) is the unique gauge-covariant two-derivative kinetic structure.Adding a term \(\propto(\partial_\mu A^\mu)^2\) changes the constraint structure, un-does the Proca counting, and generically re-introduces a propagating ghost scalar (the Stückelberg/gauge-fixed picture).
Derivation
1
\[\mathcal{L}=-\tfrac{1}{4}F_{\mu\nu}F^{\mu\nu}+\tfrac{1}{2}m^{2}A_{\mu}A^{\mu},\qquad F_{\mu\nu}=\partial_\mu A_\nu-\partial_\nu A_\mu.\]
The Maxwell kinetic term is inherited from maxwell-from-gauge-lagrangian; the only new object is the Lorentz scalar \(A_\mu A^\mu\) of correct mass dimension four. A
2
\[A_\mu\to A_\mu+\partial_\mu\lambda\;\Rightarrow\;\tfrac{1}{2}m^{2}A_\mu A^\mu\to\tfrac{1}{2}m^{2}A_\mu A^\mu+m^{2}A_\mu\partial^\mu\lambda+\tfrac{1}{2}m^{2}\partial_\mu\lambda\,\partial^\mu\lambda.\]
Test gauge invariance. \(F_{\mu\nu}\) is unchanged because \(\partial_\mu\partial_\nu\lambda\) is symmetric and cancels; but the mass term picks up \(\lambda\)-dependent pieces that do not vanish. Gauge invariance is broken by the mass term alone. A
3
\[\frac{\partial\mathcal{L}}{\partial A_\nu}=m^{2}A^{\nu},\qquad \frac{\partial\mathcal{L}}{\partial(\partial_\mu A_\nu)}=-F^{\mu\nu}.\]
Compute the two Euler–Lagrange ingredients. The derivative of \(-\tfrac14 F^2\) with respect to \(\partial_\mu A_\nu\) gives \(-F^{\mu\nu}\) exactly as in the massless case; the mass term supplies the algebraic piece \(m^2 A^\nu\). B
4
\[\partial_\mu\frac{\partial\mathcal{L}}{\partial(\partial_\mu A_\nu)}-\frac{\partial\mathcal{L}}{\partial A_\nu}=0\;\Longrightarrow\;\boxed{\;\partial_\mu F^{\mu\nu}+m^{2}A^{\nu}=0\;}\]
Substitute into the Euler–Lagrange equation. This is the Proca equation — Maxwell's \(\partial_\mu F^{\mu\nu}=0\) with a mass term that couples \(A^\nu\) to itself. B
5
\[\partial_\nu\!\left(\partial_\mu F^{\mu\nu}+m^{2}A^{\nu}\right)=\underbrace{\partial_\nu\partial_\mu F^{\mu\nu}}_{=\,0}+m^{2}\partial_\nu A^{\nu}=0.\]
Take the four-divergence of the field equation. The first term vanishes identically: \(\partial_\nu\partial_\mu\) is symmetric in \(\mu\nu\) while \(F^{\mu\nu}\) is antisymmetric, so their contraction is zero. B
6
\[m^{2}\neq 0\;\Longrightarrow\;\boxed{\;\partial_\nu A^{\nu}=0\;}\]
Because \(m\neq0\) we may divide by \(m^{2}\). The Lorenz condition is now a consequence of the equations of motion — a genuine constraint on every solution — not a freely-imposed gauge fixing. This step is precisely where masslessness would fail. B
7
\[\partial_\mu F^{\mu\nu}=\partial_\mu(\partial^\mu A^\nu-\partial^\nu A^\mu)=\Box A^{\nu}-\partial^{\nu}(\partial_\mu A^{\mu})=\Box A^{\nu}.\]
Expand \(F^{\mu\nu}\) and use the constraint \(\partial_\mu A^\mu=0\) from step 6 to kill the second term. B
8
\[\boxed{\;(\Box+m^{2})A^{\nu}=0\;}\qquad\text{subject to}\qquad\partial_\nu A^{\nu}=0.\]
Insert step 7 into the Proca equation. Each of the four components separately satisfies the Klein–Gordon equation, so all solutions propagate on the mass shell \(p^{2}=m^{2}\); the constraint ties the components together. B
9
\[A^\nu(x)=\varepsilon^\nu e^{-ip\cdot x}:\quad (\,-p^2+m^2)\varepsilon^\nu=0,\;\; p^2=m^2,\qquad p_\nu\varepsilon^\nu=0.\]
Fourier-mode counting. Klein–Gordon fixes \(p^2=m^2\); the constraint becomes the algebraic transversality condition \(p_\nu\varepsilon^\nu=0\), one linear relation removing one of the four \(\varepsilon^\nu\). Four minus one leaves three independent polarizations. C
10
\[\text{Rest frame }p^\mu=(m,\mathbf{0}):\quad p_\nu\varepsilon^\nu=m\,\varepsilon^0=0\Rightarrow\varepsilon^0=0,\quad \varepsilon^{(1,2,3)}=\hat{\mathbf e}_{x,y,z}.\]
Evaluate the transversality condition in the rest frame. It forces the time component to vanish and leaves three spacelike unit polarizations — the three \(m_s=+1,0,-1\) states of a massive spin-1 particle. A canonical (Hamiltonian) analysis reaches the same count: \(\pi^0=\partial\mathcal L/\partial\dot A_0=0\) and its consistency condition are two second-class constraints, removing \(2\) of the \(8\) phase-space variables of \((A_\mu,\pi^\mu)\) to give \(6/2=3\) configuration-space degrees of freedom. C
Result
\[(\Box+m^{2})A^{\nu}=0,\qquad \partial_\nu A^{\nu}=0,\qquad \#\text{polarizations}=4-1=3.\]

Reading. A vector field with a mass term is no longer gauge-redundant. Its equation of motion automatically enforces the Lorenz condition, which acts as a single scalar constraint that eliminates the unphysical time-like component. What remains is a massive spin-1 particle: two transverse polarizations (as for light) plus one longitudinal polarization that only a massive vector can carry. The field is short-ranged, with Yukawa-type behaviour \(\sim e^{-mr}/r\) rather than the Coulomb \(1/r\) of the massless case.

Units check. In natural units (\(\hbar=c=1\)) a vector potential has mass dimension \([A]=1\), so \([m^{2}A_\mu A^\mu]=2+1+1=4\), matching the dimension of a Lagrangian density in four spacetime dimensions — identical to \([F_{\mu\nu}F^{\mu\nu}]=4\). Restoring factors, the mass sets an inverse length \(m c/\hbar\) (the Compton wavenumber), so \(\partial_\nu A^\nu=0\) balances dimensions term-by-term and the range \(\hbar/(mc)\) carries units of length.

Limiting cases
  • \(m\to 0\): the constraint softens to an optional gauge choice, the longitudinal mode decouples from any conserved current, and the count reverts to the two transverse photon polarizations of Maxwell theory.
  • Static, source-driven (\(\partial_t\to 0\), point source): the equation becomes \((-\nabla^{2}+m^{2})A^{0}=\rho\), whose Green's function is the Yukawa potential \(e^{-mr}/(4\pi r)\) — the massive analogue of the Coulomb potential.
  • High energy \(E\gg m\): the longitudinal polarization \(\varepsilon_L^\mu\to p^\mu/m\) grows like \(E/m\); this enhancement is the origin of the Goldstone-boson equivalence theorem.
  • Non-relativistic \(|\mathbf p|\ll m\): all three polarizations become the ordinary three spin states of a massive particle at rest, degenerate in energy.
Breaks when
  • The mass is set to zero. Every step from "divide by \(m^2\)" onward is invalid; the fourth would-be constraint disappears and the theory has a gauge symmetry with only two propagating modes. The polarization sum \(-g_{\mu\nu}+p_\mu p_\nu/m^2\) is singular at \(m=0\), signalling the discontinuity in the naive counting.
  • The field couples to a non-conserved current. With \(\partial_\mu F^{\mu\nu}+m^2A^\nu=J^\nu\) the divergence gives \(m^2\partial_\nu A^\nu=\partial_\nu J^\nu\neq0\); the Lorenz constraint fails, the (would-be) fourth degree of freedom is excited, and unitarity/renormalizability are spoiled — the disease that non-abelian gauge theories cure via the Higgs mechanism.
  • Self-interactions or higher-derivative terms are added. A term \(c\,(\partial_\mu A^\mu)^2\) or a non-abelian field strength changes the constraint algebra; second-class constraints can become first-class or disappear, altering the degree-of-freedom count and often introducing a propagating ghost.
  • Strong gravity / curved spacetime. Replacing \(\partial\to\nabla\) makes \(\nabla_\nu\nabla_\mu F^{\mu\nu}\) pick up a Ricci term \(\propto R^\nu{}_\mu A^\mu\), so \(\partial_\nu A^\nu=0\) is corrected and the clean three-count no longer follows from a one-line identity.
Failure modes
  • Calling \(\partial_\mu A^\mu=0\) a "gauge choice." In Proca theory it is a theorem forced by the equations of motion, not a condition you are free to impose or relax — there is no residual gauge freedom left to fix.
  • Counting four polarizations. Students forget that transversality \(p_\mu\varepsilon^\mu=0\) removes one; a massive vector has three states, not four.
  • Counting two polarizations by analogy with the photon. The extra, longitudinal state is precisely what distinguishes massive from massless — omitting it undercounts.
  • Sign/normalization of the mass term. Writing \(-\tfrac12 m^2 A_\mu A^\mu\) with the wrong sign for the chosen metric gives a tachyonic or wrong-sign kinetic energy; the sign must make the spatial components' energy positive.
  • Assuming the massless limit is smooth for all observables. The number of degrees of freedom jumps discontinuously (3→2); only quantities coupling to conserved currents are continuous.
  • Treating \(A^0\) as dynamical. \(A^0\) has no time-derivative in \(\mathcal L\); it is an auxiliary field fixed by a constraint, not an independent propagating variable.
Discussion

The deep lesson is that gauge symmetry and masslessness are two sides of one coin. For a spin-1 field, the gauge redundancy \(A_\mu\to A_\mu+\partial_\mu\lambda\) is exactly what is needed to remove two of the four components (one by the gauge choice, one by the residual constraint), leaving the two helicity states of a massless particle. A mass term explicitly breaks that redundancy, so only one component can be removed — by the dynamically-enforced Lorenz constraint — and three physical states survive. The counting \(4-1=3\) is not a coincidence but a direct readout of how much symmetry the mass term destroyed.

Physically, the third (longitudinal) polarization is what makes a massive vector fundamentally different from light. A photon has no rest frame and only transverse polarizations; a massive vector can be brought to rest, where all three spatial directions are equivalent and it manifestly carries spin 1 with \(m_s=-1,0,+1\). Boosting the rest-frame states, the longitudinal polarization vector \(\varepsilon_L^\mu=(|\mathbf p|/m,\,E\hat{\mathbf p}/m)\) becomes nearly parallel to \(p^\mu\) at high energy, and its \(1/m\) growth is what makes the massless limit subtle and drives the equivalence theorem in the Standard Model.

At the canonical level the three-count is enforced by a pair of second-class constraints: the primary \(\pi^0\approx0\) (no conjugate momentum for \(A_0\)) and the secondary \(\partial_i\pi^i+m^2A^0\approx0\) obtained by demanding \(\dot\pi^0\approx0\). Their Poisson bracket is non-vanishing and proportional to \(m^2\), which is exactly why they are second-class and why the whole structure collapses at \(m=0\), where the bracket vanishes and the constraints become first-class generators of gauge transformations. Second-class constraints are solved (not gauge-fixed), removing one phase-space pair each: \(8-2\times1=6\) phase-space dimensions, i.e. three configuration degrees of freedom. This is the rigorous statement behind the momentum-space count of step 9.

Common misconceptions. The Proca mass is not the Higgs mechanism — Proca simply postulates a mass by hand and thereby sacrifices gauge invariance and renormalizability. The Higgs mechanism instead generates the same three-polarization massive vector while keeping a gauge-invariant, renormalizable underlying theory, by having a scalar's Goldstone mode supply the longitudinal component. Proca is the low-energy effective description you recover after the Higgs field is integrated out; it is correct as an effective theory but breaks down at energies \(\sim m\) times a coupling, exactly where the missing longitudinal dynamics matters.

Worked examples
1
\[\text{Range of a massive vector force: } R=\frac{\hbar}{m c}=\frac{\hbar c}{m c^{2}}.\]
The Yukawa potential \(e^{-mr}/r\) that solves the static Proca equation has characteristic length \(R=\hbar/(mc)\). Symbols first; insert numbers for the \(W\) boson, \(m_W c^2=80.4\ \text{GeV}\). A
2
\[R=\frac{\hbar c}{m_W c^{2}}=\frac{197.3\ \text{MeV·fm}}{80.4\times10^{3}\ \text{MeV}}=2.45\times10^{-3}\ \text{fm}.\]
Use \(\hbar c=197.3\ \text{MeV·fm}\) and convert \(80.4\ \text{GeV}=80.4\times10^{3}\ \text{MeV}\). A
\[R_W\approx2.5\times10^{-3}\ \text{fm}=2.5\times10^{-18}\ \text{m}.\]

Reading. The mass term makes the weak interaction extremely short-ranged — roughly a thousandth of a proton radius — which is exactly why beta decay looks point-like at nuclear energies. A massless vector (photon) would instead give the infinite-range Coulomb law.

Units check. \(\text{MeV·fm}/\text{MeV}=\text{fm}\), a length; \(1\ \text{fm}=10^{-15}\ \text{m}\) gives \(2.5\times10^{-18}\ \text{m}\).

1
\[\varepsilon_L^{\mu}=\left(\frac{|\mathbf p|}{m},\,\frac{E}{m}\,\hat{\mathbf p}\right),\qquad E=\sqrt{|\mathbf p|^{2}+m^{2}}.\]
The longitudinal polarization is built by boosting the rest-frame state \((0,\hat{\mathbf p})\); it satisfies \(p_\mu\varepsilon_L^\mu=0\) and \(\varepsilon_L\!\cdot\!\varepsilon_L=-1\). Take a vector of mass \(m=1\ \text{GeV}\) with momentum \(|\mathbf p|=100\ \text{GeV}\). B
2
\[E=\sqrt{100^{2}+1^{2}}\ \text{GeV}=100.005\ \text{GeV},\quad \varepsilon_L^{0}=\frac{|\mathbf p|}{m}=\frac{100}{1}=100.\]
Numerically \(E\approx|\mathbf p|\) at high energy, and the time component of \(\varepsilon_L\) is \(|\mathbf p|/m\). Check transversality: \(p\cdot\varepsilon_L=E\varepsilon_L^0-|\mathbf p|\,E/m=(E|\mathbf p|-|\mathbf p|E)/m=0.\) B
\[\varepsilon_L^{\mu}\approx(100,\,100.005\,\hat{\mathbf p})\quad\Rightarrow\quad \varepsilon_L^\mu\to\frac{p^\mu}{m}\ \text{as }E\gg m.\]

Reading. The longitudinal polarization grows like \(E/m\sim100\), while the two transverse polarizations stay \(O(1)\). This factor-of-\(E/m\) enhancement is the physical reason massive-vector scattering amplitudes appear to grow with energy — and why the longitudinal mode "is" the would-be Goldstone boson.

Units check. \(\varepsilon^\mu\) is dimensionless: \(|\mathbf p|/m\) is (GeV)/(GeV), a pure number. Transversality \(p\cdot\varepsilon_L=0\) holds identically, confirming the constraint of step 9.

Problems
  1. Verify explicitly that \(F_{\mu\nu}\) is invariant under \(A_\mu\to A_\mu+\partial_\mu\lambda\), and hence that only the mass term breaks gauge symmetry.
    Solution \(F_{\mu\nu}\to\partial_\mu(A_\nu+\partial_\nu\lambda)-\partial_\nu(A_\mu+\partial_\mu\lambda)=F_{\mu\nu}+(\partial_\mu\partial_\nu-\partial_\nu\partial_\mu)\lambda=F_{\mu\nu}\), since partial derivatives commute. Thus \(-\tfrac14 F^2\) is invariant. The mass term transforms as \(\tfrac12 m^2A_\mu A^\mu\to\tfrac12 m^2A_\mu A^\mu+m^2A_\mu\partial^\mu\lambda+\tfrac12 m^2\partial_\mu\lambda\partial^\mu\lambda\), which is not invariant for \(m\neq0\). Hence gauge invariance is broken solely by the mass term.
  2. Derive the Yukawa potential from the static, source-driven Proca equation for the time component, \((-\nabla^2+m^2)A^0=\rho\), with a point charge \(\rho=g\,\delta^3(\mathbf r)\).
    Solution The Green's function of \((-\nabla^2+m^2)\) is \(G(r)=e^{-mr}/(4\pi r)\), obtained by Fourier transform \(G(r)=\int\frac{d^3k}{(2\pi)^3}\frac{e^{i\mathbf k\cdot\mathbf r}}{k^2+m^2}=\frac{e^{-mr}}{4\pi r}\). Therefore \(A^0(\mathbf r)=g\,\frac{e^{-mr}}{4\pi r}\). As \(m\to0\) this reduces to the Coulomb potential \(g/(4\pi r)\); for \(m\neq0\) the field is screened over the length \(1/m\).
  3. Compute the range of the \(Z\) boson force given \(m_Z c^2=91.2\ \text{GeV}\), and compare with the \(W\) result of the worked example.
    Solution \(R_Z=\hbar c/(m_Z c^2)=197.3\ \text{MeV·fm}/(91.2\times10^3\ \text{MeV})=2.16\times10^{-3}\ \text{fm}=2.2\times10^{-18}\ \text{m}\). This is slightly shorter than \(R_W=2.5\times10^{-3}\ \text{fm}\) because \(m_Z>m_W\); range scales as \(1/m\), so \(R_Z/R_W=m_W/m_Z=80.4/91.2=0.88\), consistent with \(2.16/2.45=0.88\).
  4. Show that the polarization sum for a massive vector is \(\sum_{\lambda=1}^{3}\varepsilon_\mu^{(\lambda)}\varepsilon_\nu^{(\lambda)*}=-g_{\mu\nu}+\dfrac{p_\mu p_\nu}{m^2}\), and verify by contraction with \(g^{\mu\nu}\) that it counts three states.
    Solution The right-hand side must be a symmetric tensor built from \(g_{\mu\nu}\) and \(p_\mu p_\nu\) satisfying transversality \(p^\mu(\text{RHS})_{\mu\nu}=0\): indeed \(p^\mu(-g_{\mu\nu}+p_\mu p_\nu/m^2)=-p_\nu+p^2 p_\nu/m^2=-p_\nu+p_\nu=0\) on shell (\(p^2=m^2\)). In the rest frame it reduces to \(\mathrm{diag}(0,1,1,1)\), the sum of the three spatial projectors \(\hat e_i\hat e_i\), confirming the coefficients. Contracting: \(g^{\mu\nu}(-g_{\mu\nu}+p_\mu p_\nu/m^2)=-4+p^2/m^2=-4+1=-3\). Since each spacelike polarization has \(\varepsilon\cdot\varepsilon^*=-1\), the total \(-3\) counts exactly three physical states.
  5. Using canonical analysis, count the degrees of freedom of the Proca field by identifying its constraints, and contrast with the massless case.
    Solution The conjugate momenta are \(\pi^\mu=\partial\mathcal L/\partial\dot A_\mu=-F^{0\mu}\). Since \(F^{00}=0\), we get the primary constraint \(\pi^0\approx0\). Consistency \(\dot\pi^0=\{\pi^0,H\}\approx0\) yields the secondary constraint \(\partial_i\pi^i+m^2A^0\approx0\). Their Poisson bracket \(\{\pi^0(\mathbf x),(\partial_i\pi^i+m^2A^0)(\mathbf y)\}=m^2\delta^3(\mathbf x-\mathbf y)\neq0\), so they are second-class. Phase space \((A_\mu,\pi^\mu)\) has \(8\) functions per point; two second-class constraints remove \(2\), leaving \(6\) phase-space dimensions \(=3\) configuration-space degrees of freedom. In the massless case the bracket vanishes, the constraints become first-class (generating gauge transformations), and one further removes \(2\times1\) more, leaving \(4/2=2\) polarizations — the photon.