Transcription
Welcome back. Okay, so we're starting to get towards one of the most important sets of results in all of probability and statistics, which is the central limit theorem and also the law of large numbers. This essentially tells you how probability distributions limit in the limit of large n. For example, if I add up a bunch of probability distributions, they will limit towards a normal distribution. The law of large numbers states that the sample mean, if I sample a distribution and average that sample, that will converge to the expected value of the distribution.
There are some really, really important, uh, kind of mathematical theorems—Markov's inequality and Chebyshev's inequality—that we're going to need to prove, uh, those, those limiting theorems. So Markov's inequality is a really intuitive one, really simple, and I'm just going to walk you through it. So it states that if you have a random variable X that is non-negative, so it only takes on positive values, then, uh, essentially the probability that X is greater than or equal to a, some value a, the probability that it's bigger than or equal to a, is less than or equal to the expected value of x divided by that value a.
Now, this seems a bit strange. Maybe it's true, maybe it's not. Why is it useful? Um, I'll write down just an example to build your intuition, and then we'll prove it. So the example: let's say that my expected value of x is equal to 10. Then what this means is that the probability of X being greater than or equal to 20—so here 20 is a—is less than or equal to my expected value 10 over 20, which is 1/2.
And so, in words, what this means is that if I have some probability distribution function—let's, let's draw, you know, my normal probability distribution—remember, X is strictly non-negative, so it only takes on positive values from zero. This is X. What this means is that if my, if my mean, expected value is 10, at most half of my distribution can be to the right of 20. Okay, so if I have, um, I can't have more than half of my distribution to be at the right of 20 and still have enough distribution left over here to balance it out to have this expected value equal 10. That kind of makes sense. This is very intuitive. Like, if my average value is 10, I can't have that much of my distribution too far to the right of some value because I wouldn't have enough probability density over here to kind of balance it out and make it equal to 10. And this formalizes that in a super useful, simple formula.
In fact, my probability being greater than 20 being less than or equal to 1/2, if half of my probability is to the right of 20, it actually all has to be stacked up here at 20, and then the counterweight has to all be stacked up at x equals 0. So that's kind of the, the limiting case. The only way I can have half of my distribution to the right of 20 is to have it be kind of all at 20 and the other half at zero. If I have less than half of my distribution to the right of 20, I can have it be distributed a little bit, and I can have my counterbalance be distributed a little bit. Okay, so let's walk through the proof of this. It's relatively simple.
Um, so the expected value of x, we're just going to write down this formula: the expected value of x, um, and I guess I'm going to write this down for a discrete random variable X, but you could do it for a continuous random variable too, no big deal. Um, this is the sum over all possible values that this variable can take times of that value times the probability that big X equals that specific value of x. Okay, and this is going to always be, this, this total sum here, because X is positive and the probabilities are non-negative, this is always going to be greater than or equal to if I started my sum for X's bigger than a. So if I take the exact same probability, uh, the exact same expression here, but I only limit my sum to being values of X bigger than a, then this is always less than or equal to this bigger sum. There's more things I'm adding up over here, and they're all positive. Okay.
Um, and this expression here, if X is bigger than or equal to a, this is equal to, uh, this is equal to—checking my notes here—okay, this thing, if X is bigger than or equal to a, this is also greater than or equal to the same sum X greater than or equal to a, if I replace this X with an a probability of x equals X. I'm going to write this out, and then we're going to double-check that every step makes sense. The expectation of X, this is just the definition; this is always bigger—I'm adding up over all possible values, little x here—if I restrict myself to only the values of X that are to the right of a, this here is a, this is my a value. If I only restrict to adding up values to the right of a, this sum is less than this sum because I'm adding up less of these positive values, and I could replace this x with a constant a, and up here x is always bigger than or equal to a, so this is always going to be bigger than or equal to this. If I replace X within this constant a, this is always going to be less than or equal to this expression here. And now I can pop this a outside, and I have this equals a times the probability that X is bigger than or equal to a. So that last step you might have to pause and think about it for a minute; it's not that complicated. I popped out my a, and then the sum of the probability x equals little x for all X bigger than a is just the probability that big X is greater than or equal to a. That's exactly—like, this is the definition of that sum—and then this expression here, I can take this and essentially get, uh, this expression. Okay, so, um, the probability here is less than or equal to—a * this probability is less than or equal to the expectation—so I can divide both sides by a, and I get this probability is less than or equal to my expected value divided by a. This is the proof. Okay.
Very, very useful theorem or inequality. We're going to use this to prove Chebyshev's inequality in the next video, and then we're going to use Chebyshev's inequality to prove the law of large numbers. So this is a really nice kind of backpocket theorem, and it's super intuitive. It says that if you have a non-negative distribution and you have a mean value, you can't have too much of your distribution too far to the right of that mean because you wouldn't have enough distribution left to counterbalance. Makes perfect sense, and this is kind of how you codify or go through the proof of something like Markov's inequality. And there are tons of theorems like this, so get comfortable with what we're doing here because it's going to get a little more sophisticated, uh, in the next examples. Okay, thank you.