📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Random Variables and Probability Distributions

Steve Brunton21:56

Transcription

Welcome back. So today we're going to introduce a really, really major concept in probability: the concept of a random variable. So, just like in, uh, algebra and calculus, you can have a variable X that takes on values two, or three, or Pi, or whatever values, a random variable in probability and statistics is also a variable that can take on a given set of values, and then you can assign a probability to each of those values of X. Okay. So I'm going to give you some examples. Um, in fact, let's just start off with an example and then kind of define what we mean by X and give a more formal definition. So, um, a good example is let's say that I am flipping my fair coin. Okay. Remember, uh, I have my fair coin here. So I'm going to flip a coin four times; flip a coin four times. And we've already talked about the sample space Omega of all of the things that could happen. You could have heads heads heads heads, or heads heads heads tails, etc. A random variable is going to be some real-valued kind of number that you associate to each of those experiments in that sample space. So let's say I, you know, flip a coin four times. X, my random variable, could be the number of heads that occur. This could be the number of heads that occur. So, um, we—we know—I'm actually just going to write this in orange—we know our sample space, um, Omega for four, four coin flips is just—I'm just going to list it out—heads heads heads heads, heads heads heads tails, heads heads tails heads, uh, dot dot dot dot all the way to tails tails tails tails. Okay. Um, and I think there are 16 elements of the sample space; there's 16 different sequences of heads and tails I can flip—2 to the 4—and this random variable X is the number of heads that could occur in these coin flip sequences. And so X, we would say—we would say that X is uh takes values between 0, 1, 2, 3, and 4. It could be an integer between 0 and 4. I could get zero heads; I could get four heads; I could get one, two, three heads as well. Okay. So X is a variable, and it has, um, you know, this kind of set of numbers that it could—it could take. So this is a discrete random variable. My random variables could be discrete; um, they could be discrete like the number of heads in N coin flips, or they could be continuous. They could be continuous random variables. A good example of a continuous random variable would be the height of an average—the height of an American—the height of Americans, um, height of Americans. You could make it the height of Brazilians, the height of French, the height of people worldwide; it doesn't matter. The height of Americans. This could be in feet; it could be in meters; it could be in centimeters, but this is a variable that could be kind of any continuum of values, you know, within some reasonable range. Discrete random variables can only take on a discrete set of values. Okay. Now, importantly, I can assign a probability to each of these values of X. That's really, really important. So to each value of x, I can assign a probability. So the probability of x equals zero—that I get no heads—there's only one out of these 16 cases where I get zero heads, so this is a 1/16 probability. Um, the probability that x equals 1—I'm just going to use shorthand—x equals 1, the probability—we're going to use the binomial distribution; we know that this satisfies or follows the binomial distribution—remember Pascal's triangle and N choose R—so the probability that x equals 1 is going to be, uh, I think 4/16. I think the probability that x equals 2 is 6/16. Then x equals 3, the probability again is 4/16, and the probability that x equals 4 is 1/16. So this just is a nice example of a binomial distribution, and what we would say—I'll get into this in a minute—but we would say x is distributed as a binomial random variable. So it is a random variable that adheres to the binomial distribution of probabilities that we can calculate here. And one of the cool things, um, that you can do then when you have this random variable and you can compute the probabilities of each of these states of the random variable is this gives you a very nice concise way of plotting or summarizing those probabilities in a histogram, in a plot. So let's actually do that. So we're going to plot, um, the probability of X—number of heads out of four coin flips—so this is my variable X, um, this is going to be 0, 1, 2, 3, and 4, and I'm going to roughly try to do this; it should go from 1, 2, 3, 4, 5, 6. So the probability of getting zero heads—the probability of X being zero—is 1/16. The probability of X being one—of having one head in this coin flip—is 4/16. 1, 2, 3, 4/16. Each of these ticks is a 1/16 of the probability. The probability of X equaling 2 is 6/16; that's this one here. The probability of X equaling 3 is again 4/16, and the probability of X equaling 4 is 1/16. And you should always label your axes. This is 0, 1, 2, 3, and 4 heads, and this is probability in 16ths; this is 1/16. Okay, good. So this is a really nice concise way of summarizing this relatively big set of things that could happen—all of these 16 possible cases of four coin flips—can be summarized in this distribution. This is called a probability distribution over my random variable X.

Now, when I was learning algebra, I remember this idea of a variable being kind of a tough abstraction. I think this is something that kids—I had a hard time with this; I remember my mom telling me she had a hard time with this—um, um, you know, my kids—I remember there was a day or two where they had a hard time with this idea of an abstraction of a variable. So you're used to doing math with numbers: 3 + 5 = 8, you know, and then when you start introducing x + 5 = 8, what is X? That takes people some time to get used to. Now we're all super comfortable with that abstraction now, but this might take you a little while. So it's okay if it takes you a minute for this abstraction to sink in and for you to have intuition. We're going to do probably 10 lectures just on examples of random variables, how to compute functions of random variables, what are some of the most common examples of random variables, and you're going to build this intuition just like you did in algebra and calculus. So it's going to be okay. Okay. So, um, a random variable essentially is just like a normal variable, except I can assign a probability to X being in each of its possible states. That—that's really important. Um, and maybe I'll write that down here. So, um, we essentially have P(x)—is a function; you can plot—is a function called a distribution—called—it's called a pro—a distribution or a probability distribution—called a distribution. And roughly speaking, what you need—you need the values of X; you need to define kind of what is the domain of this probability space and what is the range of this probability distribution—so the values that X can take—and then you need to have, um, essentially the probability, uh, P of each of each value. Okay, probability p of each possible value of x. So here we know that our number of heads—our random variable is the number of heads in four coin flips—X can take five different values, and here are the probabilities of X taking each of those five values. So we have defined a probability distribution over our random variable X, and that's a really, really useful concept. Um, this is the kind of thing that you're going to really be glad that you have access to this abstraction of probability distributions. For example, let's say I flip a coin 100 times. I don't want to list 2 to the 100 possible coin flips; it's—it's kind of incalculably large, the number of possible sequences of coin flips if I flip a coin 100 times. I don't want to ever write that down on a whiteboard, um, or try to enumerate that. But the number of heads that occur in a 100 coin flips is a binomial distribution; it has a name, and there is a function for the probability density of each of, you know, for there being 20 heads or 21 heads or 22 heads. And so, to some extent, you can write down the probability distribution—sometimes we call it a probability density function—you can write down that probability distribution for much more complicated random variables, for much more complicated processes where you could never enumerate all of the possibilities. So this becomes really useful, um, and we have this number of heads in N coin flips that is going to be—we're going to say X follows a binomial distribution; we're going to define what this is in one of the next lectures—it's related to those binomial coefficients that we looked at earlier. If we have a continuous variable like the height of Americans, this is actually going to—in this case—follow a normal distribution, a normal or a Gaussian distribution. So you've seen this before; it's kind of the bell curve, a normal distribution here. And you'll notice actually the binomial distribution—if I have N gets very large—if I have 100 coin flips or a thousand coin flips—this will start to approximate or converge to a normal distribution. So really, really kind of useful stuff. And you can do things like—you can calculate what is the probability—if this is my distribution of heights—so this is x in, let's say, feet—and let's say this, you know, average height—I don't know what the average height of Americans is; let's say it's like 5'8" or something like that—and maybe this is six feet—and I don't know, I'm just making up numbers, right?—let's say 5, 6, 7, 4, 3, 2, whatever—you can compute the probability that someone is less than six feet tall. You can compute what is the probability that someone is less than six feet tall just by integrating all of these probabilities up to that point. I can compute the probability that someone is between seven and eight feet tall, um, you know, by computing—adding up all the probabilities between x equals 7 and x equals 8. So these random variables and these probability distributions on those random variables—you can do all kinds of things; you can do calculus on those variables; you can integrate probabilities and find what's the probability within some range that I'm within some range of values of X. Um, I can compute the expected value: What do I expect if I just pick someone off the street? What is the expected height of that person? And how much spread do I have in that expectation? How would I be surprised if someone was two feet shorter or taller than that expected value? These distributions have all of that information, and you can define functions on these random variables. The probability is one function; you can also define the expected value of x or the variance of X or the standard deviation of X and all kinds of other functions of this random variable. Okay, good. Um, so we're going to have a bunch more examples. We're going to talk about binomial, normal, Poisson, exponential, gamma—a lot of the most useful random variables that are useful for things like: How long do you expect to wait at the DMV if there's five people ahead of you and the average wait time is two minutes? Um, how many emails do I expect to get in the next 10 minutes given a certain rate of emails? And would I be surprised if I got no emails in that time? Like these are the kinds of questions you can ask and answer now that we have this abstraction of random variables. So maybe I'll just write down some of the why. So I've already said most of this, but like why do we want random variables? Um, first, it's a concise summary—a concise expression of all of your probabilities—a summary of all probabilities. Good. It's a concise summary of all of the probabilities. Um, sometimes it's hard to count these probabilities. Remember, if I have 100 coin flips, I don't want to count all of the ways I can get 50 heads. My probability distribution over that random variable—it's a function that I can write down, so I don't have to count all of those possibilities. Okay. So sometimes we get a function for P(x), so we don't have to count. So no counting. Remember, the older I get, the harder it is for me to count. I have to have everyone be quiet while I'm counting to 20, and so I don't like counting. If I have a function for these probabilities, I want to use that function most of the time. Uh, two—the probability density function—this probability distribution—sometimes we're going to call it a PDF—which is a probability density function—and usually I think of probability density functions for these continuous variables. A PDF—this PDF—this is just shorthand I'm going to use to define my probability density. This PDF is a model of a random process, and this is a really, really important idea here that the real world never is exactly Bernoulli or binomial or normal or any of the distributions. These are mathematical approximations, just like if I throw a ball, you know, and I write down F = ma, I might neglect wind resistance; I might neglect rotational effects. There are all these things—these approximations I make to get my simple, you know, ballistic trajectory or my simple F = ma descriptions or my simple, um, you know, Galileo's Tower of Pisa constant gravity—like all of these things are approximations. That's also true in probability. The PDF is just a model of a random process. Um, if I flip this coin really, there's wind resistance, probably. It's not perfectly 50%; probably the way I flip a coin might be like a tiny bit biased. There's all of these things that make it not perfectly random, but this is a good model of that random process. And an important point of that model is that this allows you—so we're talking about probability for the most part here, but this dual notion of statistics is—let's say I collect data; I should be able to test my hypothesis that my system is binomially distributed or normally distributed. I should be able to test hypotheses. This allows me to test hypotheses with data. So, for example, if I flip 10 coins in a row and I get heads all 10 times, does that—do I actually think that my PDF is binomial, um, with, you know, equal probabilities? That would be a hypothesis testing problem, and I could get a probability of how likely that sequence of 10 heads is given that my coin is binomial—a fair binomial or Bernoulli coin—and so you can do hypothesis testing. You can also do parameter estimation, um, parameter estimation, uh, again with data. So this is a machine learning or a data statistics problem. If I know that my distribution of Americans' heights is normal, we know that the normal distribution has two parameters that completely define it: the mean and the standard deviation. So I can take my data; I can sample 100 people off the street, and I can get a really good estimate of that mean and standard deviation, and then from that small sample, I can say something important about the much larger distribution of, you know, whatever 300 million people. Good. Okay. Three: You can visualize data in a much easier way. You can visualize—you can visualize and compute. So we can compute our probability density functions; we can—we can plot—maybe I'll switch to yellow—we can compute things like the probability that X is less than 6 and greater than 4. Let's say, you know, what's the probability that a random person is greater than four feet and less than six feet? This is something you can compute once you have this probability distribution. Um, I can compute how unlikely is event X? How unlikely is it that I am greater than seven feet tall? All kinds of things like that. Uh, and you can do calculus. Um, in this case, this probability is going to be the integral from 4 to 6 of my probability density function; let's just call this P(x) dx. Okay. So you can do calculus on these continuous distributions. If I had a discrete distribution, I would be doing a sum, you know, from k = 4 to 6 of P(x) equals k, something like that. Okay. You can compute these things using sums or integrals on these distributions. You can do computations with these. Um, last thing—and there's a lot more dot dot dot dot dot—but the last thing I'm going to mention is that you can build functions; you can build functions on X. You can compute the distribution of X squared; that can be useful sometimes. You can compute the distribution of X squared; that's a function of X. You can compute the expected value of x; that's another function, or the variance of X; that's a function. Um, you can take an uncertainty—a normal Gaussian distributed uncertainty—and you can propagate it through some engineering process, some manufacturing process, or some dynamical system. I can take an uncertainty in my initial condition of a differential equation and I can propagate that uncertainty and see how it spreads. Uh, in chaotic systems, it gets spread around the chaotic attractor. All kinds of interesting things you can build functions on X; you can propagate X through dynamical systems; uh, you can propagate uncertainty through dynamics. This is what Gaussian processes do is they propagate uncertainty through dynamics—super, super useful ideas. Um, okay, that's probably all I want to show you right now. Um, one thing I'll just very, very briefly mention is that if we had this probability density function that tells me the probability of X being at a certain value, there is also this notion of a cumulative probability: It's the probability of X being less than a certain value, um, and that would just be the integral of this probability distribution up to that point x. That's called the cumulative density function—the CDF. We'll talk about that later; I just wanted to mention that that's also a useful function of X is the cumulative distribution function. Okay. Um, a lot more coming up soon. I'm going to show you examples of binomial, normal, Poisson, exponential, gamma, um, you know, a bunch more. We're going to work out examples; we're going to use these to compute really intuitive things that you're going to be able to use both in your daily life and, um, you know, to do better engineering. Okay, thank you.