📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Chebyshev's Inequality in Probability: Second Order Estimates

Steve Brunton9:44

Transcription

Welcome back. Okay, we last time derived Markov's inequality, which is a really intuitive, simple expression for how much of a probability density can be how far to the right of its expectation value if uh the random variable is non-negative. And today we're going to derive uh and state Chebyshev's inequality. This one is super useful, and we're going to use this specifically to prove the law of large numbers, uh one of the most important central results in probability and statistics in the next lecture. Okay.

So Chebyshev's inequality um is a little bit more, more sophisticated um than Markov's inequality. So remember Markov's inequality uh states that for a non-negative random variable, the probability of that variable being greater than some value a is less than or equal to its expected value divided by a. And this basically means you can't have too much mass in the distribution too far to the right of the expected value because there wouldn't be enough room left um you know, to balance it out on the other side of the expected value, roughly speaking. This only uses the expectation value though. Chebyshev's inequality is going to use the variance; it's actually going to give you a result about how much the variance um kind of tightens or spreads. So this is, I think in my estimation, a little bit more useful.

So, for—I'm going to state it for any positive number a, for any a greater than zero—then uh then the probability of x minus mu, the absolute value—I'm going to again write it down and then we're going to talk about it—the probability of the absolute value of x minus mu, this is the deviation of x from its mean, being bigger than or equal to a, this is less than or equal to the variance of my distribution, Sigma squared, divided by a squared. Okay.

So essentially what this is saying is that, you know, I have some distribution—I'm just going to actually draw—I think I'm going to draw a distribution here because I like having a a picture in my mind of what I'm talking about—I have some distribution here and I have some mean value mu, and my distribution has a standard deviation Sigma. Okay. Then the probability of finding x minus mu, of sampling X and finding it, you know, uh some distance away from mu, the probability that that distance being greater than or equal to a for some reason has to be less than or equal to the variance of my distribution divided by a squared. This is a little less easy to say intuitively why this is the case than Markov's inequality, but we're going to reason through it. And this is true for lots of distributions, not just a normal distribution. This would also be true if my distribution had some weird, you know, some weird bumps in it, um presumably. Okay, this could be a a weird distribution that is not normal, but if it has this variance and this expectation, then this is going to be true. Okay. This is an important result, uh and so we're going to work through proving it, and then we're going to try to talk through understanding a little bit more about why this might be, might be the case.

Okay, so let's prove this thing. Um so we're going to introduce a new variable y equals uh x minus mu, x - mu squared. Okay, y equals x - mu squared, and we're going to define B = a squared. There's going to be this a squared popping out here, so we're going—this proof we're going to go through it and we're going to make some assumptions that were very convenient and non-obvious. It's not obvious why I would do this um to prove this. Okay, so Y is going to be this random variable, and the probability we're going to try to relate this probability here to some probability in terms of Y and B. Okay, that's what we're going to try to do. So the probability that this is true, that uh x minus mu absolute value is bigger than or equal to a is the same as the probability that x—I can square both sides of this—this is an interesting property; you really need to like slow down and convince yourself—you know, Jerry Marsen used to say, you know, you need to go to a quiet room or like sit in the dark and think about why this is true and convince yourself—I'm going to state something that's true: the probability that the absolute value of x minus mu is greater than or equal to a is the same as the probability of x - mu squared being greater than or equal to a squared, to a squared. So I can square both sides of this, and this inequality—this probability is still—these are equal—this is an a squared—I just really botched it with my bad writing skills—this is an a squared. This is true, and now this is equal to the probability that my random variable y—this is just my random variable Y—is greater than or equal to my rand—to my constant B. This is the probability that Y is greater than B, um and this I can use Markov's inequality. So I know uh something about the expected value of y; that's going to be basically—the expected value of y is the variance; that's the definition of variance of X. So I I don't want to go too fast; I want to remind you the expected value of y is the expected value of x minus mu squared. This is the definition of the variance of X, which in our case is Sigma squared. So we're going to use this property. Um so the probability of Y being bigger than or equal to B, the probability of some variable being bigger than or equal to some other number is less than or equal to the expected value of y, to the expected value of y divided by this value B, divided by B. So we used—we just used Markov's inequality right here. This is from Markov's inequality, and now I know that the expected value of y is equal to Sigma squared, and so all of this equals—this is all less than or equal to Sigma squared over B, which is just a squared. Okay.

Okay, so the probability—this thing I'm talking about here—this probability of x minus mu absolute value greater than or equal to a has to be less than or equal to this, to the variance of x divided by a squared. Okay. Um it's an outfall of Markov's inequality in this new variable Y, which is this uh squared deviation of x from its mean. And this is a really, really interesting, interesting result. It's kind of like Markov's inequality, which says you can't have too much mass of the distribution too far away from the center uh from from the mean. This is saying if I have a variance Sigma, I can't have too much of my probability density, too many of my points X spreading too far from my mean. If I have, you know, X—if I have too much of my distribution too far from my mean—this means too far—a far away—the probability of that happening has to be back-bounded to still have this variance um equal to Sigma squared. That's roughly what it's stating. These are both stating that to have this variance, I can't have too much of my distribution too far away from my mean, um or else I can't have that variance. That's essentially what this is stating. These are both super, super useful theorems, inequalities; we use them all the time in probability and statistics, and soon we're going to use this to prove the central limit theorem, one of the cornerstone results in in all of probability and statistics, which states that if I randomly sample from this distribution X over and over and over and I average those samples, that average will converge to the expectation value of x, and convergence means that my variance—that that my uncertainty—will tighten and tighten and tighten; my my sample mean, my average over random samples of X will get closer and closer and closer to this expected value. So we're going to do something to kind of bound the variance of the deviation of that sample mean from the expectation value, and that's how we're going to prove the uh law of large numbers. Okay, thank you.