Convergence of Random Variables
What does it mean for a sequence of random variables to converge to another random variable ?
Suppose we have a sequence of random variables Z 1 , Z 2 , … Z_1, Z_2, \dotsc Z 1 , Z 2 , … defined on some ( Ω , F , P ) (\Omega, \mathcal{F}, \mathbf{P}) ( Ω , F , P ) , what does it mean to say that { Z n } \{Z_n\} { Z n } converges to Z Z Z as n → ∞ n \to \infty n → ∞ ?
One notion that we have already seen, though we have not said it in such words, is the notion of pointwise convergence.
That is, for all ω ∈ Ω \omega \in \Omega ω ∈ Ω , lim n → ∞ Z n ( ω ) = Z ( ω ) \lim_{n\to\infty} Z_n(\omega) = Z(\omega) lim n → ∞ Z n ( ω ) = Z ( ω ) .
However, we can also come up with a weaker notion of convergence, which essentially says the convergence is pointwise almost everywhere --- everywhere except on a set of measure zero.
We also call this convergence almost surely or with probability one , which means that P ( lim n → ∞ Z n = Z ) = 1 \mathbf{P}(\lim_{n\to\infty} Z_n = Z) = 1 P ( lim n → ∞ Z n = Z ) = 1 .
To aid in establishing this notion of convergence, consider the following proposition, which mirrors the almost everywhere notions of measure theory.
Proposition.
Let Z 1 , Z 2 , … Z_1, Z_2, \dotsc Z 1 , Z 2 , … be random variables .
Suppose that for each ϵ > 0 \epsilon > 0 ϵ > 0 , we have P ( ∣ Z n − Z ∣ > ϵ i.o. ) = 0 \mathbf{P}(|Z_n - Z| > \epsilon \text{ i.o.}) = 0 P ( ∣ Z n − Z ∣ > ϵ i.o. ) = 0 .
Then P ( Z n → Z ) = 1 \mathbf{P}(Z_n \to Z) = 1 P ( Z n → Z ) = 1 , i.e. { Z n } \{Z_n\} { Z n } converges to Z Z Z almost surely.
In other words, P ( { ω ∈ Ω , : lim n → ∞ Z n ( ω ) = Z ( ω ) } ) = 1 \mathbf{P}(\{ \omega \in \Omega, \,:\, \lim_{n\to\infty} Z_n(\omega) = Z(\omega)\}) = 1 P ({ ω ∈ Ω , : lim n → ∞ Z n ( ω ) = Z ( ω )}) = 1 .
Proof.
It follows from the discussion about the complementarity of infinitely often and almost always that
P ( Z n → Z ) = P ( for all ϵ > 0 , ∣ Z n − Z ∣ < ϵ a. a. ) = 1 − P ( there exists ϵ > 0 , ∣ Z n − Z ∣ > ϵ i.o. ) . (1) \htmlId{eq-1}{\begin{aligned}
\mathbf{P}(Z_n \to Z) &= \mathbf{P}(\text{for all } \epsilon > 0, |Z_n - Z| < \epsilon \text{ a. a. })\\
&= 1 - \mathbf{P}(\text{there exists } \epsilon > 0, |Z_n - Z| > \epsilon \text{ i.o.}).
\end{aligned}} \tag{1} P ( Z n → Z ) = P ( for all ϵ > 0 , ∣ Z n − Z ∣ < ϵ a. a. ) = 1 − P ( there exists ϵ > 0 , ∣ Z n − Z ∣ > ϵ i.o. ) . ( 1 ) Countable sub-additivity tells us that
P ( there exists ϵ > 0 , ϵ rational, ∣ Z n − Z ∣ > ϵ i.o. ) ≤ ∑ ϵ > 0 , rational P ( ∣ Z n − Z ∣ > ϵ i.o. ) = 0. (2) \htmlId{eq-2}{\begin{aligned}
\mathbf{P}(\text{there exists } \epsilon > 0, \epsilon \text{ rational, } |Z_n - Z| > \epsilon \text{ i.o.}) \leq \\
\sum_{\epsilon > 0, \text{ rational}} \mathbf{P}(|Z_n - Z| > \epsilon \text{ i.o.}) = 0.
\end{aligned}} \tag{2} P ( there exists ϵ > 0 , ϵ rational, ∣ Z n − Z ∣ > ϵ i.o. ) ≤ ϵ > 0 , rational ∑ P ( ∣ Z n − Z ∣ > ϵ i.o. ) = 0. ( 2 ) Note that for any ϵ > 0 \epsilon > 0 ϵ > 0 there exists an ϵ ′ > 0 \epsilon' > 0 ϵ ′ > 0 which is rational and where ϵ ′ < ϵ \epsilon' < \epsilon ϵ ′ < ϵ .
For this ϵ ′ \epsilon' ϵ ′ , we have { ∣ Z n − Z ∣ ≥ ϵ i.o. } ⊆ { ∣ Z n − Z ∣ ≥ ϵ ′ i.o. } \{|Z_n - Z| \geq \epsilon \text{ i.o. } \} \subseteq \{|Z_n - Z| \geq \epsilon' \text{ i.o. } \} { ∣ Z n − Z ∣ ≥ ϵ i.o. } ⊆ { ∣ Z n − Z ∣ ≥ ϵ ′ i.o. } .
It follows that
P ( there exists ϵ > 0 , ∣ Z n − Z ∣ > ϵ i.o. ) ≤ P ( there exists ϵ ′ > 0 , ϵ ′ rational, ∣ Z n − Z ∣ > ϵ ′ i.o. ) = 0 (3) \htmlId{eq-3}{\begin{aligned}
\mathbf{P}(\text{there exists } \epsilon > 0, |Z_n - Z| > \epsilon \text{ i.o.}) \leq \\
\mathbf{P}(\text{there exists } \epsilon' > 0, \epsilon' \text{ rational, } |Z_n - Z| > \epsilon' \text{ i.o.}) = 0 \\
\end{aligned}} \tag{3} P ( there exists ϵ > 0 , ∣ Z n − Z ∣ > ϵ i.o. ) ≤ P ( there exists ϵ ′ > 0 , ϵ ′ rational, ∣ Z n − Z ∣ > ϵ ′ i.o. ) = 0 ( 3 ) thus giving the result.
■ \blacksquare ■
In combination with the Borel-Cantelli Lemma, we see that
Corollary.
Let Z 1 , Z 2 , … Z_1, Z_2, \dotsc Z 1 , Z 2 , … be random variables .
Suppose for each ϵ > 0 \epsilon > 0 ϵ > 0 , we have ∑ n P ( ∣ Z n − Z ∣ ≥ ϵ ) < ∞ \sum_n \mathbf{P}(|Z_n - Z| \geq \epsilon) < \infty ∑ n P ( ∣ Z n − Z ∣ ≥ ϵ ) < ∞ .
Then P ( Z n → Z ) = 1 \mathbf{P}(Z_n \to \Z) = 1 P ( Z n → Z ) = 1 , i.e. { Z n } \{Z_n\} { Z n } converges to Z Z Z almost surely.
Another notion of convergence involves only probabilities.
Definition. Convergence in probability
We say that { Z n } → Z \{Z_n\} \to Z { Z n } → Z in probability if for all ϵ > 0 \epsilon > 0 ϵ > 0 , lim n → ∞ P ( ∣ Z n − Z ∣ ≥ ϵ ) = 0 \lim_{n \to \infty} \mathbf{P}(|Z_n - Z| \geq \epsilon) = 0 lim n → ∞ P ( ∣ Z n − Z ∣ ≥ ϵ ) = 0 .
We now consider the relationship between convergence in probability and convergence almost surely.
Proposition.
Let Z , Z 1 , Z 2 , … Z, Z_1, Z_2, \dotsc Z , Z 1 , Z 2 , … be random variables .
Suppose that { Z n } → Z \{Z_n\} \to Z { Z n } → Z almost surely.
Then { Z n } → Z \{Z_n\} \to Z { Z n } → Z in probability.
That is, if a sequence of random variables converges almost surely, then it converges in probability to the same limit.
Proof.
This mirrors the proof of uniform convergence almost everywhere.
Fix ϵ > 0 \epsilon > 0 ϵ > 0 , and let A n = { ω : there exists m ≥ n , ∣ Z m − Z ∣ ≥ ϵ } A_n = \{\omega \,:\, \text{ there exists } m \geq n, |Z_m - Z| \geq \epsilon\} A n = { ω : there exists m ≥ n , ∣ Z m − Z ∣ ≥ ϵ } .
Then, { A n } \{A_n\} { A n } is a decreasing sequence of events and, by assumption lim n → ∞ A n = A \lim_{n\to\infty} A_n = A lim n → ∞ A n = A is such that P ( A ) = 0 \mathbf{P}(A) = 0 P ( A ) = 0 .
More precisely, for ω ∈ ⋂ n → ∞ A n \omega \in \bigcap_{n\to\infty} A_n ω ∈ ⋂ n → ∞ A n , then Z n ↛ Z Z_n \not\to Z Z n → Z and hence P ( ⋂ n A n ) ≤ P ( Z n ↛ Z ) = 0 \mathbf{P}(\bigcap_n A_n) \leq \mathbf{P}(Z_n \not\to Z) = 0 P ( ⋂ n A n ) ≤ P ( Z n → Z ) = 0 .
By continuity of probabilities, we have that P ( A n ) → P ( ⋂ n A n ) = 0 \mathbf{P}(A_n) \to \mathbf{P}(\bigcap_n A_n) = 0 P ( A n ) → P ( ⋂ n A n ) = 0 , and so
Hence, P ( ∣ Z n − Z ∣ ≥ ϵ ) ≤ P ( A n ) = 0 \mathbf{P}(|Z_n - Z| \geq \epsilon) \leq \mathbf{P}(A_n) = 0 P ( ∣ Z n − Z ∣ ≥ ϵ ) ≤ P ( A n ) = 0 , as desired.
■ \blacksquare ■
On the other hand, the converse of the previous proposition is false.
Convergence in probability is a weaker notion of convergence than almost sure convergence.
The counter-example here is very weird to me.
Laws of Large Numbers
The words “weak” and “strong” here should evoke the relative weakness and strength of the two forms of convergence we just saw.
Consider the weak law of large numbers.
Theorem. Weak Law of Large Numbers
Let X 1 , X 2 , … X_1, X_2, \dotsc X 1 , X 2 , … be a sequence of independent random variables , each having the same mean m m m , and having variance ≤ v < ∞ \leq v < \infty ≤ v < ∞ .
Then for all ϵ > 0 \epsilon > 0 ϵ > 0 ,
lim n → ∞ P ( ∣ m − 1 n ∑ i = 0 n X i ∣ ≥ ϵ ) = 0. (4) \htmlId{eq-4}{\lim_{n \to \infty} \mathbf{P}\left( \left| m - \frac{1}{n}\sum_{i=0}^n X_i \right| \geq \epsilon \right) = 0.} \tag{4} n → ∞ lim P ( m − n 1 i = 0 ∑ n X i ≥ ϵ ) = 0. ( 4 ) In words, the partial averages converge in probability to m m m .
Proof.
Take S n = 1 n ( X 1 + ⋯ X n ) S_{n} = \frac{1}{n} \left( X_1 + \cdots X_n \right) S n = n 1 ( X 1 + ⋯ X n ) .
We have to bound the probability that ∣ S n − m ∣ ≥ ϵ |S_n - m| \geq \epsilon ∣ S n − m ∣ ≥ ϵ , which is possible with Chebyshev’s inequality.
Recall that
P ( ∣ S n − m ∣ ≥ ϵ ) ≤ 1 ϵ 2 V a r ( S n ) ≤ v n ϵ 2 . (5) \htmlId{eq-5}{\mathbf{P}(|S_n - m| \geq \epsilon ) \leq \frac{1}{\epsilon^2}\mathrm{Var}(S_n) \leq \frac{v}{n\epsilon^2}.} \tag{5} P ( ∣ S n − m ∣ ≥ ϵ ) ≤ ϵ 2 1 Var ( S n ) ≤ n ϵ 2 v . ( 5 ) Now, that immediately implies that
lim n → ∞ P ( ∣ S n − m ∣ ≥ ϵ ) = 0 , (6) \htmlId{eq-6}{\lim_{n \to \infty} \mathbf{P}(|S_n - m| \geq \epsilon) = 0,} \tag{6} n → ∞ lim P ( ∣ S n − m ∣ ≥ ϵ ) = 0 , ( 6 ) as required.
■ \blacksquare ■
Solutions to Selected Exercises
Exercise 5.5.11
Consider the sequence of random variables on the standard Lebesgue probability triple
{ Y n } = 1 [ 0 , 1 2 ) , 2 ⋅ 1 [ 1 2 , 1 ] , 3 ⋅ 1 [ 0 , 1 3 ) , 4 ⋅ 1 [ 1 3 , 2 3 ) , … (7) \htmlId{eq-7}{\{Y_n\} = \mathbf{1}_{[0, \frac{1}{2})}, 2 \cdot \mathbf{1}_{[\frac{1}{2}, 1]}, 3 \cdot \mathbf{1}_{[0, \frac{1}{3})}, 4 \cdot \mathbf{1}_{[\frac{1}{3}, \frac{2}{3})}, \dotsc} \tag{7} { Y n } = 1 [ 0 , 2 1 ) , 2 ⋅ 1 [ 2 1 , 1 ] , 3 ⋅ 1 [ 0 , 3 1 ) , 4 ⋅ 1 [ 3 1 , 3 2 ) , … ( 7 )
(a) Clearly, we have that lim n → ∞ P ( 1 n ∣ Y n ∣ ≥ ϵ ) = 0 \lim_{n \to \infty} \mathbf{P}( \frac{1}{n}|Y_n| \geq \epsilon) = 0 lim n → ∞ P ( n 1 ∣ Y n ∣ ≥ ϵ ) = 0 .
(b) Moreover, we have convergence of Y n / n 2 Y_n / n^2 Y n / n 2 with probability one.
Let ϵ > 0 \epsilon > 0 ϵ > 0 , then for n > 1 ϵ n >\frac{1}{\epsilon } n > ϵ 1 , we have
{ ω ∈ Ω : 1 n 2 Y n ≥ ϵ } = ∅ , (8) \htmlId{eq-8}{\left\{\omega \in \Omega \,:\, \frac{1}{n^2} Y_n \geq \epsilon \right\} = \emptyset,} \tag{8} { ω ∈ Ω : n 2 1 Y n ≥ ϵ } = ∅ , ( 8 )
which implies that P ( lim n → ∞ 1 n 2 Y n = 0 ) = 1 \mathbf{P}\left( \lim_{n \to \infty} \frac{1}{n^{2}} Y_n = 0 \right) = 1 P ( lim n → ∞ n 2 1 Y n = 0 ) = 1 .
In other words, for every ω \omega ω , Y n n 2 → 0 \frac{Y_n}{n^2} \to 0 n 2 Y n → 0 , which implies convergence almost surely.
(c) For any ω ∈ Ω \omega \in \Omega ω ∈ Ω , the sequence has Y n ( ω ) / n = 1 Y_n(\omega) / n = 1 Y n ( ω ) / n = 1 infinitely often.
Hence, we do not have pointwise convergence anywhere, and therefore no convergence almost surely.
Exercise 5.5.15
Prove that if { X n } \{X_n\} { X n } converges to X X X almost surely, then
P ( ∣ X n − X ∣ ≥ ϵ i.o. ) = 0. (9) \htmlId{eq-9}{\mathbf{P}(|X_n - X| \geq \epsilon \ \text{ i.o.}) = 0.} \tag{9} P ( ∣ X n − X ∣ ≥ ϵ i.o. ) = 0. ( 9 )
Proof.
For the sake of contradiction, suppose not.
Define
E = { ω ∈ Ω : ∣ X n ( ω ) − X ( ω ) ∣ ≥ ϵ i.o. } . (10) \htmlId{eq-10}{E = \left\{ \omega \in \Omega \ : \ |X_n(\omega) - X(\omega)| \geq \epsilon \ \text{ i.o.} \right\}.} \tag{10} E = { ω ∈ Ω : ∣ X n ( ω ) − X ( ω ) ∣ ≥ ϵ i.o. } . ( 10 ) By assumption, P ( E ) > 0 \mathbf{P}(E) > 0 P ( E ) > 0 .
But then E E E is precisely the set where X n ( ω ) ↛ X ( ω ) X_n(\omega) \not\to X(\omega) X n ( ω ) → X ( ω ) , and its complement is the set where we do have convergence.
P ( lim n → ∞ X n = X ) = P ( E c ) = 1 − P ( E ) < 1. (11) \htmlId{eq-11}{\mathbf{P}\left(\lim_{n \to \infty} X_n = X\right) = \mathbf{P}(E^c) = 1- \mathbf{P}(E) < 1.} \tag{11} P ( n → ∞ lim X n = X ) = P ( E c ) = 1 − P ( E ) < 1. ( 11 ) ■ \blacksquare ■