Sometimes, the subscript is omitted and we simply write ϕ(t) if the random variable is clear from context.
Moreover, we can immediately see that the characteristic function depends only on the distribution of X, by the change of variables theorem.
On the surface, the characteristic function closely resembles the moment generating function.
However, the addition of the imaginary unit i=−1 gives it different properties.
For starters, the characteristic function is always finite, since ∣eitx∣=1 and so ∣ϕX(t)∣≤1<∞, in contrast to the moment generating function which may be infinite for any s=0.
Like for moment generating functions, we have ϕX(0)=1 for any X, and if X and Y are independent, then ϕX+Y(t)=ϕX(t)ϕY(t).
Letting μ=L(X), we have
Now, as h→0 the bounded convergence theorem, (given that eihX−1≤2), states that the last expression goes to 0.
We can therefore conclude that ϕX is always a (uniformly) continuous function (since the last expression does not depend on t).
Similarly, we can define the derivatives of ϕX(t).
It will be analogous to the derivation for the moment generating function, but we’ll need fewer constraints thanks to the facts we just showed.
The Continuity Theorem
The whole point of this chapter is to get a new tool.
This tool is the fact that characteristic functions are unique, and you can (usually) back out the corresponding distribution.
The implication, which is what has to be proven, is that sequences of characteristic functions correspond to sequences of distributions, and the limits match.
In other words, we can work entirely in terms of characteristic functions, compute the limit, and then back out the corresponding distribution.
In order to prove this result, we’ll need some machinery.
For the result to hold, we need to know that the function limit actually corresponds to the characteristic of some distribution.
For this to happen, the limit of distributions has to be well-behaved, and the probability can’t escape to infinity.
In this case, the limit of characteristic functions may exist, but not actually correspond to a distribution.
Also, the limiting function has to correspond to a unique distribution.
We’ll cover the notion of a “tight” sequence of distributions, Fourier inversion, uniqueness, and eventually the central limit theorem and the method of moments.
First, the following theorem
We need lemmas to prove this:
Moreover,
The proof follows from standard results from complex analysis, namely contour integration.
These results then allow one to prove the inversion theorem purely mechanically by substituting the exponential by its sine and cosine representation.
The proof is ommitted for brevity.
However, we note an important corollary:
This shows that at least the notion we want to show is well-defined.
It’s plausible that the limits in characteristic functions ensure corresponding limits in distribution.
However, this does not ensure that the limit actually exists.
To prove this, we need a new fact, called the Helly selection principle.
The proof is tedious, so rather than state it formally, we’ll give a coarse statement with the proof idea.
If we have a sequence of cumulative distribution functions{Fn(x)}, then there’s always a subsequence with limit function F.
The limit function F is non-decreasing and right-continuous, and points x where F(x) is continuous have limk→∞Fnk(x)=F(x).
However, the function may not start from 0 and go to 1.
The general idea of the proof is to pick a countable set of support points, say the rationals, where we will define an auxiliary function G.
It can be shown by a diagonalization argument that you can find a subsequence such that Fnk converges for all points in the countable set.
We call the limit function G(q) for q∈Q, and then we pick
F(x)=inf{G(q):q∈Q,x>q}.(7)
This function satisfies the desired properties, which the reader should verify.
The additional constraint on the sequence we need is that of “tightness”.
As foreshadowed, we have to disallow sequences of distributions where the probability mass escapes to infinity for the continuity theorem to hold.
We then have the following desireable theorem:
The proof consists of showing that the limit of CDFs F is a proper CDF using the tightness condition.
Then we can set μ((a,b])=F(b)−F(a) which is indeed the weak limit.
To get a proper result for the continuity theorem, we need to know that the sequence of distributions is tight.
We can now finally prove the continuity theorem.
The Central Limit Theorem
Let’s pause and figure out what we’ve learned and where we should go from here.
First, we acquainted ourselves with characteristic functions of random variables.
They have some nice properties, like being unique and invertible.
Moreover, manipulating a characteristic function is often easier than manipulating the underlying measure.
As a result of these niceties, we’d really like it to be true that a sequence of distributions converges to μ if and only if the corresponding characteristic functions converge to μ‘s characteristic function.
In other words, we can work in characteristic function space, and go back at the end.
It turns out this is true, as we showed.
With this new tool at hand, we’d do well to apply it.
This is what we’ll do in this section, where we use the previous section’s results to prove the central limit theorem.
Loosely, it states that a sum of i.i.d. random variables tends to the normal distribution.
We’ll want to show that the characteristic functions of the partial sums tend to the characteristic function of the normal.
The result then follows by the continuity theorem.
There are several generalizations to the classical central limit theorem.
For example, we can drop the i.i.d. assumption, or get additional results on the rate of convergence.
We don’t give the results here, since they are digressions.
However, we mention that with additional conditions on the third moment, it is possible to drop the assumption that the variables are identically distributed.
See the Lyapunov Central Limit Theorem, or the Lindeberg Central Limit Theorem.
Method of Moments
We saw that we could look at the characteristic functions to check or find limits of distributions.
However, there are other ways to do that and prove the central limit theorem.
One such method, called the method of moments, looks only at the moments of the distributions.
Some questions arise in this approach, similar to those we had about convergence in the characteristic functions.
How do we know that a convergence in the moments yields a unique distribution? When is a distribution uniquely defined by its moments?
And, when is the limit and the distribution interchangeable in the moments?
It turns out that not all limits have the right properties.
We’ll say that a distributionμ is determined by its moments if all its moments are finite and no other distribution has identical moments.
The proof is not particularly informative, but since the distributions have finite moments, they are tight.
Since some subsequence must converge to some distribution, say ν, then ν=μ since we can identify all the moments of ν with those of μ.
For this fact, use Skorohod’s theorem.
Finally, we specify which distributions are uniquely defined by their moments.
Exercises
11.5.1
a) No. The probability mass escapes to +∞.
Formally, for every ϵ>0, if you pick an a,b∈R, at some point in the sequence the mass escapes from [a,b].
Specifically, starting at n=⌊b+1⌋.
b) No, no tightness so no theorems hold. The Helly Selection principle says nothing about the distributions converging to anything.
c) Yes, the Helly selection Principle says so. The function is given by
F(x)=0.(13)
d) Same, but F(x)=1.
11.5.6
It must be tight.
11.5.9
Yes, it does. Intuitively, we know it’s the Dirac delta distribution at 0.
Can we show it rigorously?
The distributions are tight, that’s straightforward, so the continuity theorem holds.