Let’s get right into it.
We’ve already defined probability triples, now let’s start talking about random variables.
Notice that before, we didn’t have any randomness.
Actually, it will still take a bit of time before we encounter any stochasticity.
This is really just a technical definition to be on common ground.
Not much intuition can be or should be given to it.
In essence, this just says that X is measurable: the σ-algebra contains the relevant sets so that, along with P, we can determine the probability of events.
One can almost think of the function X as defining the cumulative density function, but that’s not quite the case.
This gets a lot out of the way. In particular, we can perform basic algebra with random variables.
Can we extend this notion to continuous functions of random variables?
It turns out we can.
Essentially, a measurable function is one where the pre-images of Borel sets are still Borel sets.
That is, we can apply a measure to the resulting transformation based on the size of the set that yielded the result.
Note that the Borel sets is the smallest σ-algebra containing all the intervals of R.
Back to our question about taking functions of random variables.
For a measurable function f and a random variableX, it turns out that f(X) is still a random variable,
since compositions of measurable functions are measurable.
For example, if f(x)=xk, then f is Borel-measurable.
Hence, for if X is a random variable, then so is Xk for all k∈N.
The book makes an additional remark about the random variable possibly being a function that maps to some other measurable space.
This is nice for completeness, but in my experience, mapping to the Lebesgue measure is more than expressive enough.
Independence
Now for a classic word: independence.
We know what this means intuitively: two events are independent if they do not affect each other’s probabilities.
In other words, A and B are independent if P(A∩B)=P(A)P(B).
Three events are independent if they are pairwise independent as above andP(A∩B∩C)=P(A)P(B)P(C).
You need all these conditions, it may be good to think about why.
The definition extends naturally to random variables by associating the events A and B to the measurable sets X−1(S1) and Y−1(S2).
That is, the random variables are independent if, for all Borel sets S1,S2, the events X∈S1 and Y∈S2 are independent.
All these notions can be extended to collections of events or variables by considering all finite subsets of events in the collection.
Given a standard probability triple and events A1,A2,…∈F, we write {An}↗A to mean
that A1⊆A2⊆⋯, and ⋃nAn=A.
In words, the events An increase to A.
Similarly, we write {An}↘A to mean that {Anc}↗Ac, or equivalently that
A1⊇A2⊇⋯ and ⋂nAn=A; the events decrease to A.
Limit Events
Given events A1,A2,…∈F, we define
nlimsupAn={An infinitely often }=n=1⋂∞k=n⋃∞Ak(7)
and
nliminfAn={An almost always }=n=1⋃∞k=n⋂∞Ak.(8)
Since F is a σ-algebra, these are well defined events.
They correspond to those events in the collection of An that happen either infinitely often, or the complement.
The theorem is striking: if the events are independent, the limsup is either 0 or 1 and never anything else.
The key theme here is that if we can show that the sum of probabilities of a bunch of events diverges, then these events must occur infinitely often.
Tail Fields
For example, consider infinite fair coin tossing.
Let Hn denote the event that the n-th coin comes up heads.
Then τ includes the event limit event limsupnHn, liminfnHn, and many more interesting events.
Before we can prove it, we need a corollary on independence.