Hide outline
Feedback

Measure Theory: The Machinery Behind Integration

Measure Theory: The Machinery Behind Integration

Build a working mental model of modern integration by learning why σ-algebras exist, how measurability is forced by countable additivity, and how convergence and change of variables become clean once you treat null sets correctly.

Measure theory feels backwards at first. Instead of starting with a formula for area, it starts by restricting which sets you are even allowed to talk about, then defines integration as a consequence. That restriction is not pedantry. It is what makes countable additivity, limits, and change of variables coexist without contradictions. Once you internalize that the objects are really equivalence classes modulo null sets, many classical headaches in analysis become bookkeeping problems with a standard solution.

Measuring without length

A σ-algebra is a promise. You promise to only ask questions about sets that are stable under the operations analysis actually uses. Complements encode negation, finite intersections encode conjunction, and countable unions encode limits of events or approximations. A measure then assigns sizes to these promised sets in a way that respects disjoint countable unions.

The key move is that measurability is not about geometry. It is about building a calculus of sets where limit operations do not break your notion of size. If you want to talk about lim sup\limsup sets, or about approximating a complicated set by simpler ones, countable closure shows up whether you asked for it or not.

See how σ-algebras sit between the power set and the bare minimum needed for limits.

Once you accept the promise, the usual identities become reliable tools. For disjoint EnE_n in a σ-algebra,

μ(n=1En)=n=1μ(En).\mu\Big(\bigcup_{n=1}^\infty E_n\Big)=\sum_{n=1}^\infty \mu(E_n).

This is not a definition you can safely impose on all subsets of R\mathbb{R} while keeping translation invariance and other niceties. The restriction is doing real work.

Closure pays
Countable unions are not a taste choice. They are the set-theoretic shadow of taking limits, and analysis takes limits constantly.

Why σ-algebras are the price

If you try to assign a length to every subset of R\mathbb{R} that is translation invariant and countably additive, you run into non-measurable sets. The point is not the pathology itself. The point is that countable additivity plus mild symmetry is strong enough to force contradictions unless you restrict the domain.

The Carathéodory extension viewpoint explains the restriction as an engineering constraint. Start with something you can measure consistently on a generating family, such as intervals. Build an outer measure, then declare a set measurable precisely when it behaves well with respect to that outer measure, meaning it splits every set additively.

Explore how the outer measure criterion selects a σ-algebra and why arbitrary subsets break additivity.

What you get is not maximal generality. You get maximal generality compatible with the limit operations you care about. In practice, this is why Lebesgue measurability is the right default in Rn\mathbb{R}^n. It is the completion of the Borel structure by null sets, tuned to make approximation and convergence arguments painless.

Measurable functions are preimage maps

A function is measurable when it preserves the set structure you have decided to trust. Concretely, given (X,A)(X,\mathcal{A}) and (Y,B)(Y,\mathcal{B}), a map f:XYf:X\to Y is measurable if f1(B)Af^{-1}(B)\in\mathcal{A} for every BBB\in\mathcal{B}. The direction matters. Images do not preserve unions and complements nearly as well as preimages do.

Two patterns dominate actual work.

Generators

You rarely check measurability on all of B\mathcal{B}. You check it on a generating class. For Y=RY=\mathbb{R} with the Borel σ-algebra, it is enough to verify f1((,t))Af^{-1}((-\infty,t))\in\mathcal{A} for all tRt\in\mathbb{R}. This is why distribution functions and threshold sets appear everywhere.

Stability under limits

Measurable functions are closed under pointwise limits of sequences when you have measurability at each stage and the limit is taken through operations compatible with σ-algebras. This is one reason analysis prefers monotone or dominated setups. They let you pass measurability and integrability through the limit without surprise.

See the preimage perspective and how generators cut the verification down to size.

The quiet takeaway is that measurability is not an extra property tacked onto functions. It is the compatibility condition between your function and your set calculus.

Lebesgue integral via level sets

The Lebesgue integral can be defined in several equivalent ways, but the most structural slogan is this. It integrates by measuring where the function sits above a threshold, then stacking those measures.

For f0f\ge 0 measurable, one guiding identity is

fdμ=0μ({x:f(x)>t})dt,\int f\,d\mu=\int_0^\infty \mu(\{x:f(x)>t\})\,dt,

whenever the right side makes sense. This is the distribution function view. It explains why convergence theorems are natural. Limits of functions become limits of nested level sets, and measures handle nested limits well.

To build the integral cleanly, you start with simple functions, finite linear combinations of indicator functions of measurable sets. You define sdμ\int s\,d\mu by linearity and the measure. Then you approximate ff from below by an increasing sequence of simple functions snfs_n\uparrow f and set fdμ=supnsndμ\int f\,d\mu=\sup_n\int s_n\,d\mu.

Watch the threshold picture connect μ({f>t})\mu(\{f>t\}) to f\int f.

Three theorems then become the workhorses.

  • Monotone convergence theorem (MCT): fnff_n\uparrow f implies fndμfdμ\int f_n\,d\mu\uparrow \int f\,d\mu for fn0f_n\ge 0.
  • Fatou: lim inffndμlim inffndμ\int \liminf f_n\,d\mu \le \liminf \int f_n\,d\mu for fn0f_n\ge 0.
  • Dominated convergence theorem (DCT): if fnff_n\to f pointwise and fng|f_n|\le g with gL1g\in L^1, then fnf\int f_n\to\int f.

Choose your theorem
MCT is about order, DCT is about control, and Fatou is what you use when you have neither but need an inequality.

Null sets and where analysis lives

A null set is not just small. In measure theory it is ignorable in a way that survives limits and integration. The slogan is that functions that differ only on a null set are the same for integration, and most convergence notions in analysis are designed to respect that.

This is the entry point to LpL^p spaces (L^p). You define

fp=(fpdμ)1/p\|f\|_p=\Big(\int |f|^p\,d\mu\Big)^{1/p}

and then identify functions that agree almost everywhere. The result is a complete normed space for 1p1\le p\le\infty, the natural habitat for Fourier analysis, PDE estimates, and probability.

The subtlety is convergence. Pointwise, almost everywhere, in measure, in LpL^p, and weak convergence are different lenses. Experts switch lenses deliberately, because implications you wish were true often fail without extra hypotheses.

Compare the main convergence modes and the implication arrows that break.

A useful mental model is that LpL^p convergence measures average error, not worst-case error, and almost everywhere convergence is a statement about exceptional sets, not rates. Mixing them up is the source of many false proofs.

Densities and change of measure

The Radon–Nikodym theorem is the statement that a measure can have a density with respect to another measure, provided it is absolutely continuous. If νμ\nu\ll\mu and both are σ-finite, then there exists a measurable function h=dνdμh=\frac{d\nu}{d\mu} such that for all measurable AA,

ν(A)=Ahdμ.\nu(A)=\int_A h\,d\mu.

This is the abstract form of change of variables and the foundation of expectations under different distributions.

Once you have hh, integrals transform by

fdν=fhdμ.\int f\,d\nu=\int f\,h\,d\mu.

No coordinate chart is required. It is a theorem about measures, not about Rn\mathbb{R}^n.

Work a few density transformations and see how sets and integrals change together.

The failure mode is also crisp. If ν\nu puts mass where μ\mu sees nothing, there is no density. The singular part is real information, like point masses inside an absolutely continuous background.

Product measures and iterated integrals

A product measure μ×ν\mu\times\nu formalizes the idea that measuring rectangles A×BA\times B by μ(A)ν(B)\mu(A)\nu(B) should extend to a full σ-algebra on X×YX\times Y. Once you have it, you want to compute

X×Yf(x,y)d(μ×ν)\int_{X\times Y} f(x,y)\,d(\mu\times\nu)

by iterating one-dimensional integrals.

Tonelli is the nonnegative theorem. If f0f\ge 0 is measurable, then

fd(μ×ν)=(f(x,y)dν(y))dμ(x)\int f\,d(\mu\times\nu)=\int\Big(\int f(x,y)\,d\nu(y)\Big)\,d\mu(x)

and both sides may be ++\infty. Fubini is the integrable theorem. If fL1(μ×ν)f\in L^1(\mu\times\nu), then the iterated integrals exist as finite numbers and agree almost everywhere.

What experts actually check is not a philosophical condition. It is one of these.

  • Nonnegative measurable, so Tonelli applies.
  • Absolutely integrable, so Fubini applies.
  • Split f=f+ff=f^+-f^- and verify f+<\int f^+<\infty and f<\int f^-<\infty.

The typical pitfall is to interchange integrals or sums based on pointwise convergence rather than an applicable theorem. If you need to swap limits, make the hypothesis you are using explicit.

How experts wield the machinery

Measure theory is the language of controlled approximation. You model the objects you can compute with, then prove that your target object is a limit of those, with the limit justified by a convergence theorem whose hypotheses you can actually check.

Three places this shows up constantly.

Probability kernels

A probability kernel is a measurable way to assign a probability measure to each point, such as xK(x,)x\mapsto K(x,\cdot). The measurability requirement is again preimage based, because you want xK(x,A)x\mapsto K(x,A) measurable for each measurable AA. This is what makes conditional distributions and Markov transitions composable.

Weak convergence

Weak convergence is about testing measures against bounded continuous functions. The tightness and uniform integrability conditions you see in probability are ways of importing compactness and domination into this test-function world so limits exist and are meaningful.

Pitfalls that bite

Two classics. Assuming every subset is measurable, which is false in any rich setting. Treating almost everywhere statements as uniform ones, which breaks when exceptional sets depend on parameters.

Bring a scenario and pin down the measure space, measurability requirements, and the right convergence theorem.

If you want one concrete next step, take a proof you already know in analysis that uses an ϵ\epsilon and a partition, then rewrite it using simple functions, σ-algebras, and one convergence theorem. The content barely changes. The proof becomes modular, and the conditions under which it works become obvious.

Was this lesson helpful?
Dive Deeper

Generate a follow-up sub-lesson on any aspect of this topic

Related content