📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Variance and Standard Deviation

Steve Brunton12:59

Transcription

Welcome back. Okay, so in the last couple of lectures, we have defined the expected value of a distribution X, this kind of uh expectation of a random variable X, which measures the center of mass of that distribution. For nice, well-behaved distributions like a Gaussian, um it is actually the most probable and center of the distribution. Today, we're going to introduce the variance and standard deviation of a random variable X.

So, if the mean μ is the measure of center of mass of the distribution, then the variance of X, and I'll define it in a minute, the variance in X measures the average squared deviation of x from that mean value. This measures uh the average square deviation of x from this mean value μ, and the standard deviation is just the square root of the variance. So let me draw a really quick picture here because I think this will help. Uh maybe I'll start in uh in yellow. So let's say I have some probability distribution function like this nice Gaussian. Then the mean, the expected value, is the kind of center of this distribution μ. The variance uh is how much spread, how what is the expected squared deviation of x from this mean value uh μ. And maybe I'll also write down the standard deviation. Um the standard deviation of x, uh SD of X, this is um essentially measures the spread, and spread is a non-technical word; we can kind of define it as the standard deviation if you want, the spread of the histogram of the PDF of X. So, for example, in the Gaussian uh normal distribution, the standard deviation are these plus or minus kind of Sigma points, plus or minus Sigma, where Sigma is the standard deviation, inside of which about 68% of the distribution lives within plus or minus one Sigma, one Sigma. And then you have two Sigma, two standard deviations, 3, 4, 5. So μ measures the center of mass; standard deviation measures the spread; and the standard deviation, um SD of X, is just the square root uh of the variance of X.

Okay, so I'm going to write down now what the formula is for the variance. Uh we're going to show kind of how it works and how to compute it, um and then zoom back out and talk about it with respect to these distributions. Okay, and this is a really, really easy idea here. So, um let's start with this variance uh of X, and maybe I'll just do it over here for a minute. So the variance of X is defined—this is kind of the definition, so maybe I'll do like a little triangle equals—this is defined as the expected squared deviation of x from μ. It's a weird mouthful; I'll write it down in math; it's actually going to be more sensible in math. It's the expectation of my variable x minus the mean quantity squared. So we know that μ is the mean value; it's the average value. But if my distribution has some spread, if I randomly sample values of X, they're probably not going to be exactly equal to μ; there's going to be some x minus μ. So what is the expected value of the square of that spread, the square of the difference between X and μ? What's the expected value of the difference between X and μ squared? And that gives me an idea of, you know, if I have really, really, really long tails or a really wide distribution, this is going to be bigger because I'm expecting x minus μ to be pretty large most of the time that I randomly sample X. Good, that makes a lot of sense. And we can essentially derive a very useful formula. So this is also equal to—and this is something we're going to have to derive, so this is not obvious; this is something we would have to derive—that this is equal to the expectation value of my variable x squared, the square of my random variable X, the expectation of x squared minus the expected value of x quantity squared.

Okay, this is interesting. I'll I'll derive this for you in a minute. This is a very useful quantity. This thing is hard to necessarily work with; it's a little messy. This thing is a lot easier to work with; expectation of X is just μ, so this is just minus μ². If I know the mean, I don't even have to compute the second term, and now I only have to compute this term, which is the second moment of my uh probability distribution over X. Okay, so this is a useful property. So maybe I will uh derive this now here. Okay, so the variance of X is this quantity here. So uh Var(X) = the expected value of (x - μ)², and I can actually expand this thing out; I can do math on this. This is a function of my random variable; this is just some G(X), and we know actually how to compute E[G(X)] um from an earlier video. But let's actually work this out. So this is the expectation of x² - 2μx + μ², and because of a really important property of expectation values, the expected value of the sum of three terms is the sum of the expectation value of each of those three terms. So this equals expectation of X² um I can actually pull this two out, but I'm not going to yet uh plus expectation of -2μx plus expectation of μ². Okay, now the expectation of a constant is just that constant. Okay, um that that's really simple. The expectation of that constant—you can actually um do this using the formula for expectation; you plug in the expected value of a constant; it's just the sum of that constant times the probability over all of the states, and those probabilities add up to one, so you just recover the constant. This is kind of an exercise; maybe I'll switch—I'm just going to write some things down here. So this is just μ². The expected value of a constant times x is just that constant times the x; expectation value of x, so this is -2μ; expected value of x is -2μ², that's this term; and this expectation of x² is just expectation of X². So all of this adds up to the expectation of x² - μ², which is minus the expectation of x quantity squared. Okay, so that proves this useful formula here that we just derived. So we essentially derived from this this useful formula. So if you have the mean, all you have to compute is the second moment, um this expectation of X², and we know how to compute the expectation of a function of X. This guy here, I'll just write it down, uh the expectation of a function of X like X², let's say we're dealing with a continuous random variable like a a Gaussian here; this is going to be the integral from negative infinity to infinity of x² times my probability density dx. So this is an easily computable quantity; um you just plug in, you know, x² here; this is the second moment of this probability density function, uh second moment. And remember I said that there are more moments; the first moment was μ; the second moment is this quantity; there's a third, a fourth, a fifth; there's an infinite series of moments that characterize funky distributions that have weird behavior and asymmetries and things like that. And it's kind of like the Taylor series; it uniquely identifies your probability density. But if you have something nice and well-behaved like a Gaussian, it's completely characterized by these two numbers; it's first and second moment; it's mean and its standard deviation or variance; it's, you know, average value and the spread of the function uniquely determines the Gaussian or normal distributed function. So really, really useful, and this is something we can compute. So in in a next example, we're actually going to compute the expected value and the variance and the standard deviation for the exponentially distributed random uh variable, for a Gaussian, for a normal. So if x uh is a normally distributed random variable with μ uh and σ² as the parameter, um which I think the PDF for this would be um f(x) = 1 / √(2π)σ e^-(x - μ)² / 2σ², I hope. Okay, if this is the PDF, then you can actually kind of go the other way from this PDF for this Gaussian; you can show that the expected value of x is μ; this is the center of the distribution; and the variance of X is σ²; that's how much spread or deviation from μ you expect, uh squared deviation from μ you expect to have in your distribution. Okay, is there anything else I want to tell you? Um kind of a homework problem; I think you should know—you should do this yourself: why is the variance always greater than or equal to zero? What in this—I'm taking something minus something—why is this always going to be greater than or equal to zero? I want you to think about that; why is the variance always a positive number? That's kind of a cool question for you to ponder on.

Um, and then one other fact I want to point out—this is very, very important—is if X and Y are independent—again, I always want to write down what happens if you have independent variables; it's super important—then the variance of x + y is equal to Var(X) + Var(Y). Okay, this is not obvious; you'd actually have to show that this is true. But for independent random variables X and Y, this is true. This is very much not true if X and Y are dependent. So, for example, um you know, Var(2X) is absolutely not equal to 2Var(X) or Var(X) + Var(X); I'm pretty sure it's equal to 4Var(X). Okay, so this is definitely not true if X and Y are not independent, but it is true if they are independent. Okay, um so taking a step back, mean and either variance or standard deviation are very, very useful ways to quantify the behavior of well-behaved distributions like a normal distribution. Um, in fact, they uniquely characterize a normal distribution, and they are the first of many moments that will characterize kind of generic um, you know, funky, weird distributions. So this distribution might need more moments to uniquely characterize it, um that kind of generalize the notion of mean and standard deviation. Okay, thank you.