Distribution Theory: Generalized Functions That Work

Distribution Theory: Generalized Functions That Work

Build a working mental model of distributions as continuous linear functionals, then use weak derivatives, Fourier duality, and wavefront sets to reason about singularities and when operations like products are actually defined in PDEs and physics.

The fastest way to stop being confused by distributions is to stop picturing them as weird functions. Distribution theory is a calculus of actions on test functions, where singular objects behave predictably because the rules are built from continuity, not pointwise values. The Dirac delta is the poster child for this shift. It is not an infinite spike you squint at. It is an evaluation map that makes differentiation, Fourier transform, and PDE forcing terms work cleanly when classical functions run out of runway.

The Dirac delta as a functional

The Dirac delta only looks mystical if you insist it must be a function. Treated correctly, Dirac delta δ\delta is a linear functional that takes a test function φ\varphi and returns one number, δ,φ=φ(0)\langle\delta,\varphi\rangle=\varphi(0). That single formula is already the whole object.

Approximations like narrow bumps are useful, but they are a story about how δ\delta acts, not what it is. The bump picture tempts you to talk about height and pointwise values, then you start trying to multiply spikes or cancel infinities. The functional definition keeps you honest. If you cannot describe an operation in terms of what it does to every φ\varphi, you do not yet have a well-defined operation.

See how δ\delta acts as you shrink a family of normalized bumps against different test functions.

What survives the limit is not a graph, it is the invariant action φφ(0)\varphi\mapsto\varphi(0). That is why the width can go to zero while the output remains stable, and why the area normalization matters more than peak height.

Picture trap
When you catch yourself saying infinite spike, translate it into an action on test functions. If you cannot, you are about to rely on a metaphor where the math needs a definition.

Test functions, dual spaces, and why topology matters

A distribution lives on a choice of test space. Change the test functions and you change what continuity means, which changes what objects are allowed. The usual starting point is test functions D=Cc\mathcal{D}=C_c^\infty, smooth functions with compact support. A distribution is then a continuous linear functional on D\mathcal{D}, written D\mathcal{D}'.

Continuity is the whole game because D\mathcal{D} is not normed in the way beginners hope. Its topology encodes control of all derivatives on compact sets. That is why convergence φn0\varphi_n\to0 in D\mathcal{D} means eventually the supports stay in one compact set and every derivative goes to zero uniformly there. A distribution is continuous if it sends such convergent sequences to numbers that go to zero.

Switch the test space and you get a different dual:

  • D=Cc\mathcal{D}=C_c^\infty supports local singularities and compactly supported probes.
  • Schwartz space S\mathcal{S} adds rapid decay and makes Fourier analysis behave.
  • D\mathcal{D}' vs S\mathcal{S}' differs mainly in what growth at infinity you permit in the functional.

Explore the inclusions and dual relationships that explain why S\mathcal{S}' is smaller than D\mathcal{D}', and why that restriction is a feature.

The practical consequence is that statements like the Fourier transform of a distribution depend on which dual you are in. If you skip the topology, you end up treating continuity like an afterthought, then wonder why limits and transforms misbehave.

Weak derivatives make differentiation always possible

Distributions make differentiation an algebraic operation defined by integration by parts. If TDT\in\mathcal{D}', its distributional derivative iT\partial_iT is defined by

iT,φ=T,iφ.\langle\partial_iT,\varphi\rangle=-\langle T,\partial_i\varphi\rangle.

This works because iφ\partial_i\varphi stays in the same test space and the right-hand side is linear and continuous in φ\varphi.

For a locally integrable function ff, you can view it as a distribution TfT_f via Tf,φ=fφ\langle T_f,\varphi\rangle=\int f\varphi. Then iTf\partial_iT_f coincides with the classical derivative when it exists, and still exists when it does not. The archetype is a jump. A function like the Heaviside step has a derivative that is zero almost everywhere, yet its distributional derivative is δ\delta because the jump contributes a boundary term.

Simulate the integration-by-parts shift and watch how a jump discontinuity produces a delta term.

The mental model is that differentiation moves from the rough object onto the smooth test function. Every time you see a boundary term in classical integration by parts, expect a singular distribution to appear when you remove smoothness assumptions.

Support, singular support, and wavefront intuition

Support is blunt. It tells you where a distribution is nonzero, but not how it is singular. Two distributions can share the same support and still be radically different microlocally. Singular support refines this by marking where the distribution fails to be smooth.

The next refinement is the wavefront set (WF). Instead of only locating singularities in space, WF also tracks directions in frequency where the object fails to be smooth. An edge singularity is not the same as a point singularity even if both sit on the same set. WF records that difference by attaching cones of bad directions to points.

Visualize how support can agree while WF differs, and how directionality separates edge-like from point-like singularities.

This is not decoration. WF is the bookkeeping device that decides whether you can multiply distributions, pull them back under maps, or propagate singularities through PDEs without cheating.

Microlocal lens
If your operation depends on how oscillations behave at high frequency, support is too crude. Reach for WF before you start guessing.

Fourier transform on tempered distributions

The Fourier transform is naturally at home in S\mathcal{S}', the space of tempered distributions. Definition by duality is clean. For TST\in\mathcal{S}', define T^\widehat{T} by

T^,φ=T,φ^(φS),\langle\widehat{T},\varphi\rangle=\langle T,\widehat{\varphi}\rangle\quad(\varphi\in\mathcal{S}),

because φ^S\widehat{\varphi}\in\mathcal{S} again.

This is the featured-snippet fact that matters. The Fourier transform extends from Schwartz functions to tempered distributions by duality, and it stays well-defined because S\mathcal{S} is invariant under Fourier transform and controls growth at infinity.

Compare what breaks in D\mathcal{D}' and what works in S\mathcal{S}' when you try to define Fourier transforms via duality.

The restriction to tempered objects is not aesthetic. It is the minimal constraint that rules out pathologies at infinity while keeping the singular objects you actually use, like polynomials, δ\delta, and principal value kernels.

Multiplication and nonlinearity

Addition is always defined in a vector space. Multiplication is not. The product of two distributions is often undefined because there is no consistent way to define ST,φ\langle ST,\varphi\rangle that respects limits and agrees with classical multiplication when both are functions.

The classic warning sign is colliding singularities. Multiplying HH and δ\delta can be made sense of in certain frameworks, but δδ\delta\cdot\delta is the canonical example that fails in standard distribution theory. The obstruction is not philosophical. It is that any attempt to define the product clashes with how approximations converge.

Hörmander’s criterion uses WF as a compatibility check. Roughly, you can multiply uu and vv if their wavefront sets do not contain opposite covectors at the same point, because that is the configuration that makes high-frequency singularities interact in an uncontrollable way.

Explore WF-based compatibility for products and see why some singularity pairings fail immediately.

When someone writes a nonlinear PDE with distributions and manipulates products as if everything were smooth, this criterion is the first place to look for the hidden assumption.

Regularization as a disciplined hack

Regularization is not a guilty secret. It is a controlled way to replace an ill-defined object by a family of smooth ones and extract a limit. The catch is that regularization is a choice, and for nonlinear expressions different choices can produce different answers.

A mollifier ρε\rho_\varepsilon is a smooth approximate identity. Convolving TρεT*\rho_\varepsilon often gives a smooth family that converges back to TT in the distributional sense. For linear operations, this is usually safe. For nonlinear operations like squaring, the limit can depend on the mollifier. That is why principal value distributions and boundary value prescriptions matter. They are attempts to specify a canonical limit, often motivated by symmetry or complex analysis.

Reveal how different regularizations can change nonlinear outcomes, and when a limit is canonical rather than conventional.

Nonlinear warning
If your result changes when you change mollifier shape, you did not compute a property of the distribution. You computed a property of your regularization scheme.

Distributions in PDEs and physics

Distributions earn their keep when you rewrite a PDE so every term makes sense as an element of a dual space. A forcing term like δ\delta is not an exotic source. It is a compact way to encode point loads, impulses, or initial data. The solution concept shifts from pointwise equality to equality after testing against φ\varphi and integrating by parts.

Green’s functions and fundamental solutions are often distributions because the operator’s inverse cannot be represented by an honest function globally. The distribution viewpoint tells you what equation is satisfied, what boundary conditions are encoded, and which function space is the right one for existence and uniqueness.

Work through prompts that translate a PDE into distribution form, interpret δ\delta-type forcing, and choose a natural test space.

Once you think this way, weak solutions stop feeling like a compromise. They become the native language for problems with shocks, corners, point sources, and oscillation.

Design rules for safe use

You rarely get in trouble because you used distributions. You get in trouble because you forgot which test space you were in, or you performed an operation that needs microlocal hypotheses.

Keep three rules in your working memory:

  • Pick the test space that matches the operations you need, especially Fourier transform and growth at infinity.
  • Track singularities with singular support and WF before you multiply, compose, or restrict.
  • Do not cancel infinities. Replace that impulse with a statement about limits in the chosen topology.

The concrete next step is to take one identity you use routinely in smooth calculus, for example commuting differentiation and Fourier transform, and restate it purely as an equality of actions on φS\varphi\in\mathcal{S}. If you can write it cleanly, you can trust it in the singular regime.

Was this lesson helpful?
Dive Deeper

Generate a follow-up sub-lesson on any aspect of this topic

Related content