Disintegration theorem

From HandWiki
Short description: Theorem in measure theory

In mathematics, the disintegration theorem is a result in measure theory and probability theory. It rigorously defines the idea of a non-trivial "restriction" of a measure to a measure zero subset of the measure space in question. It is related to the existence of conditional probability measures. In a sense, "disintegration" is the opposite process to the construction of a product measure.

Motivation

Consider the unit square S=[0,1]×[0,1] in the Euclidean plane ℝ2. Consider the probability measure μ defined on S by the restriction of two-dimensional Lebesgue measure λ2 to S. That is, the probability of an event E⊆S is simply the area of E. We assume E is a measurable subset of S.

Consider a one-dimensional subset of S such as the line segment Lx={x}×[0,1]. Lx has μ-measure zero; every subset of Lx is a μ-null set; since the Lebesgue measure space is a complete measure space, E⊆Lx⟹μ(E)=0.

While true, this is somewhat unsatisfying. It would be nice to say that μ "restricted to" Lx is the one-dimensional Lebesgue measure λ1, rather than the zero measure. The probability of a "two-dimensional" event E could then be obtained as an integral of the one-dimensional probabilities of the vertical "slices" E∩Lx: more formally, if μx denotes one-dimensional Lebesgue measure on Lx, then μ(E)=∫[0,1]μx(E∩Lx)dx for any "nice" E⊆S. The disintegration theorem makes this argument rigorous in the context of measures on metric spaces.

Statement of the theorem

(Hereafter, 𝒫(X) will denote the collection of Borel probability measures on a topological space (X,T).) The assumptions of the theorem are as follows: [1]

  • Let Y and X be two Polish spaces (i.e. separably completely metrizable spaces).
  • Let μ∈𝒫(Y).
  • Let π:Y→X be a Borel-measurable function. Here one should think of π as a function to "disintegrate" Y, in the sense of partitioning Y into {π−1(x) | x∈X}. For example, for the motivating example above, one can define π((a,b))=a, (a,b)∈[0,1]×[0,1], which gives that π−1(a)=a×[0,1], a slice we want to capture.
  • Let ν∈𝒫(X) be the pushforward measure ν=π*(μ)=μ∘π−1. This measure provides the distribution of x (which corresponds to the events π−1(x)).

The conclusion of the theorem: There exists a ν-almost everywhere uniquely determined family of probability measures {μx}x∈X⊆𝒫(Y), which provides a "disintegration" of μ into {μx}x∈X, such that:

  • the function x↦μx is Borel measurable, in the sense that x↦μx(B) is a Borel-measurable function for each Borel-measurable set B⊆Y;
  • μx "lives on" the fiber π−1(x): for ν-almost all x∈X, μx(Y∖π−1(x))=0, and so μx(E)=μx(E∩π−1(x));
  • for every Borel-measurable function f:Y→[0,∞], ∫Yf(y)dμ(y)=∫X∫π−1(x)f(y)dμx(y)dν(x). In particular, for any event E⊆Y, taking f to be the indicator function of E,[2] μ(E)=∫Xμx(E)dν(x),which shows that the family {μx}x∈X is a regular conditional probability.

Applications

Product spaces

The original example was a special case of the problem of product spaces, to which the disintegration theorem applies.

When Y is written as a Cartesian product Y=X1×X2 and πi:Y→Xi is the natural projection, then each fibre π1−1(x1) can be canonically identified with X2 and there exists a Borel family of probability measures {μx1}x1∈X1 in 𝒫(X2) (which is (π1)*(μ)-almost everywhere uniquely determined) such that μ=∫X1μx1μ(π1−1(dx1))=∫X1μx1d(π1)*(μ)(x1), which is in particular[clarification needed] ∫X1×X2f(x1,x2)μ(dx1,dx2)=∫X1(∫X2f(x1,x2)μ(dx2∣x1))μ(π1−1(dx1)) and μ(A×B)=∫Aμ(B∣x1)μ(π1−1(dx1)).

The relation to conditional expectation is given by the identities E⁡(f∣π1)(x1)=∫X2f(x1,x2)μ(dx2∣x1), μ(A×B∣π1)(x1)=1A(x1)⋅μ(B∣x1).

Vector calculus

The disintegration theorem can also be seen as justifying the use of a "restricted" measure in vector calculus. For instance, in Stokes' theorem as applied to a vector field flowing through a compact surface Σ⊂ℝ3, it is implicit that the "correct" measure on Σ is the disintegration of three-dimensional Lebesgue measure λ3 on Σ, and that the disintegration of this measure on ∂Σ is the same as the disintegration of λ3 on ∂Σ.[3]

Conditional distributions

The disintegration theorem can be applied to give a rigorous treatment of conditional probability distributions in statistics, while avoiding purely abstract formulations of conditional probability.[4] The theorem is related to the Borel–Kolmogorov paradox, for example.

See also

References

  1. ↑ Bogachev 2007, §10.4, §10.6
  2. ↑ Dellacherie, C.; Meyer, P.-A. (1978). Probabilities and Potential. North-Holland Mathematics Studies. Amsterdam: North-Holland. ISBN 0-7204-0701-X. 
  3. ↑ Ambrosio, L.; Gigli, N.; Savaré, G. (2005). Gradient Flows in Metric Spaces and in the Space of Probability Measures. ETH Zürich, Birkhäuser Verlag, Basel. ISBN 978-3-7643-2428-5. 
  4. ↑ Chang, J.T.; Pollard, D. (1997). "Conditioning as disintegration". Statistica Neerlandica 51 (3): 287. doi:10.1111/1467-9574.00056. http://www.stat.yale.edu/~jtc5/papers/ConditioningAsDisintegration.pdf. 
  • Bogachev, Vladimir Igorevich (2007). Measure theory. Springer Berlin, Heidelberg. ISBN 978-3-540-34513-8.