Hilbert projection theorem

From HandWiki
Short description: On closed convex subsets in Hilbert space

In mathematics, the Hilbert projection theorem is a famous result of convex analysis that says that for every vector x in a Hilbert space H and every nonempty closed convex C⊆H, there exists a unique vector m∈C for which ‖c−x‖ is minimized over the vectors c∈C; that is, such that ‖m−x‖≤‖c−x‖ for every c∈C.

Finite dimensional case

Some intuition for the theorem can be obtained by considering the first order condition of the optimization problem.

Consider a finite dimensional real Hilbert space H with a subspace C and a point x. If m∈C is a minimizer or minimum point of the function N:C→ℝ defined by N(c):=‖c−x‖ (which is the same as the minimum point of c↦‖c−x‖2), then derivative must be zero at m.

In matrix derivative notation:[1] ∂‖x−c‖2=∂⟨c−x,c−x⟩=2⟨c−x,∂c⟩ Since ∂c is a vector in C that represents an arbitrary tangent direction, it follows that m−x must be orthogonal to every vector in C.

Statement

Hilbert projection theorem — For every vector x in a Hilbert space H and every nonempty closed convex C⊆H, there exists a unique vector m∈C for which ‖x−m‖ is equal to δ:=infc∈C‖x−c‖.

If the closed subset C is also a vector subspace of H then this minimizer m is the unique element in C such that x−m is orthogonal to C.

Detailed elementary proof

Proof by reduction to a special case

It suffices to prove the theorem in the case of x=0 because the general case follows from the statement below by replacing C with C−x.

Hilbert projection theorem (case x=0)[2] — For every nonempty closed convex subset C⊆H of a Hilbert space H, there exists a unique vector m∈C such that infc∈C‖c‖=‖m‖.

Furthermore, letting d:=infc∈C‖c‖, if (cn)n=1∞ is any sequence in C such that limn→∞‖cn‖=d in ℝ[note 1] then limn→∞cn=m in H.

Consequences

Proposition — If C is a closed vector subspace of a Hilbert space H then[note 3] H=C⊕C⊥.

Properties

Expression as a global minimum

The statement and conclusion of the Hilbert projection theorem can be expressed in terms of global minimums of the following functions. Their notation will also be used to simplify certain statements.

Given a non-empty subset C⊆H and some x∈H, define a function dC,x:C→[0,∞) by c↦‖x−c‖. A global minimum point of dC,x, if one exists, is any point m in domain⁡dC,x=C such that dC,x(m)≤dC,x(c) for all c∈C, in which case dC,x(m)=‖m−x‖ is equal to the global minimum value of the function dC,x, which is: infc∈CdC,x(c)=infc∈C‖x−c‖.

Effects of translations and scalings

When this global minimum point m exists and is unique then denote it by min⁡(C,x); explicitly, the defining properties of min⁡(C,x) (if it exists) are: min⁡(C,x)∈C and ‖x−min⁡(C,x)‖≤‖x−c‖ for all c∈C. The Hilbert projection theorem guarantees that this unique minimum point exists whenever C is a non-empty closed and convex subset of a Hilbert space. However, such a minimum point can also exist in non-convex or non-closed subsets as well; for instance, just as long is C is non-empty, if x∈C then min⁡(C,x)=x.

If C⊆H is a non-empty subset, s is any scalar, and x,x0∈H are any vectors then min⁡(sC+x0,sx+x0)=smin⁡(C,x)+x0 which implies: min(sC,sx)=smin⁡(C,x)min(−C,−x)=−min⁡(C,x) min⁡(C+x0,x+x0)=min⁡(C,x)+x0min⁡(C−x0,x−x0)=min⁡(C,x)−x0 min(C,−x)=min⁡(C+x,0)−xmin(C,0)+x=min⁡(C+x,x)min(C−x,0)=min⁡(C,x)−x

Examples

The following counter-example demonstrates a continuous linear isomorphism A:H→H for which min⁡(A(C),A(x))≠A(min⁡(C,x)). Endow H:=ℝ2 with the dot product, let x0:=(0,1), and for every real s∈ℝ, let Ls:={(x,sx):x∈ℝ} be the line of slope s through the origin, where it is readily verified that min⁡(Ls,x0)=s1+s2(1,s). Pick a real number r≠0 and define A:ℝ2→ℝ2 by A(x,y):=(rx,y) (so this map scales the x−coordinate by r while leaving the y−coordinate unchanged). Then A:ℝ2→ℝ2 is an invertible continuous linear operator that satisfies A(Ls)=Ls/r and A(x0)=x0, so that min⁡(A(Ls),A(x0))=sr2+s2(1,s) and A(min⁡(Ls,x0))=s1+s2(r,s). Consequently, if C:=Ls with s≠0 and if (r,s)≠(±1,1) then min⁡(A(C),A(x0))≠A(min⁡(C,x0)).

Iterated projections

For any closed convex nonempty subset C⊂H, let PC:H→C be the projection function.

If there are multiple closed convex subsets C1,C2,…,Cn, then one can approximate the projection operator PC1∩…∩Cn by applying PC1,PC2,…,PCn in sequence, then do it again and again. That is, one can approximate (PCn…PC2PC1)k→PC1∩…∩Cn as k→∞. The Kaczmarz method is a commonly used special case. Such methods can be computationally effective. For example, if C is a complicated shape, then projecting directly to C may be difficult. However, C can be approximated as an intersection of simple objects like half-spaces, hyperplanes, finite-dimensional subspaces, or cones.[4]

If C is a closed subspace, then it is convex. In this case, the projection function P:H→C is an orthogonal projection (a continuous linear operator that is self-adjoint). A classic theorem states that, if C1,…,Cn are closed subspaces, then[5] limk→∞‖(PC1⋯PCn)kx−PC1∩…∩Cnx‖=0,∀x∈H

See also

Notes

  1. ↑ Because the norm ‖⋅‖:H→ℝ is continuous, if limn→∞xn converges in H then necessarily limn→∞‖xn‖ converges in ℝ. But in general, the converse is not guaranteed. However, under this theorem's hypotheses, knowing that limn→∞‖cn‖=d in ℝ is sufficient to conclude that limn→∞cn converges in H.
  2. ↑ Explicitly, this means that given any ϵ>0 there exists some integer N>0 such that "the quantity" is ≤ϵ whenever m,n≥N. Here, "the quantity" refers to the inequality's right hand side 2‖cm‖2+2‖cn‖2−4d2 and later in the proof, "the quantity" will also refer to ‖cm−cn‖2 and then ‖cm−cn‖. By definition of "Cauchy sequence," (cn)n=1∞ is Cauchy in H if and only if "the quantity" ‖cm−cn‖ satisfies this aforementioned condition.
  3. ↑ Technically, H=K⊕K⊥ means that the addition map K×K⊥→H defined by (k,p)↦k+p is a surjective linear isomorphism and homeomorphism. See the article on complemented subspaces for more details.

References

  1. ↑ Petersen, Kaare. "The Matrix Cookbook". https://www.math.uwaterloo.ca/~hwolkowi/matrixcookbook.pdf. 
  2. ↑ Rudin 1991, pp. 306–309.
  3. ↑ Rudin 1991, pp. 307−309.
  4. ↑ Deutsch, Frank (2001), "The Method of Alternating Projections", Best Approximation in Inner Product Spaces (New York, NY: Springer New York): pp. 193–235, doi:10.1007/978-1-4684-9298-9_9, ISBN 978-1-4419-2890-0, http://link.springer.com/10.1007/978-1-4684-9298-9_9 
  5. ↑ Netyanun, Anupan; Solmon, Donald C. (August 2006). "Iterated Products of Projections in Hilbert Space" (in en). The American Mathematical Monthly 113 (7): 644–648. doi:10.1080/00029890.2006.11920347. ISSN 0002-9890. https://www.tandfonline.com/doi/full/10.1080/00029890.2006.11920347. 

Bibliography

Template:Hilbert space