📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Joint Probability Distributions: Marginal and Conditional Densities

Steve Brunton9:36

Transcription

Welcome back. So in the last lecture, we introduced this notion of a joint probability distribution between two random variables, X and Y. Um, essentially, you can define a probability of X happening and Y happening. Okay? This is kind of like what we had before where it's the probability of X and Y, uh, from conditional probability. And that notion is actually going to help us uh use these joint distributions to compute conditional densities and something called the marginal density. So these are important concepts you should know.

Um, one of the reasons I really liked the uh probability in stats course I took, which was kind of a senior undergrad class, is because it allowed you to do calculus. So calculus is a super powerful way of handling functions like probability densities. And the marginal and conditional densities are essentially things that you get if you do clever calculus on these joint distributions, and they're related to the notion of conditional uh probability from before, things you know that we use to derive Bayes' theorem and things like that. So we've already seen that you can have joint distributions like this, uh kind of two-dimensionally symmetric Gaussian in X and Y where each of X and Y is itself distributed as a Gaussian. And I hinted that if you take this two-dimensional uh Gaussian probability density and you just average out the X variable, you'll get a Gaussian in Y, and if you average out the Y variable, you'll get a Gaussian in X. I just want to formalize that here. Those are called the marginal density functions.

So if um, in kind of a continuous random variable X and Y, we know that our probability density function—I'll just write this down—our uh probability density function uh is given by this function f(x, Y). And I can compute the probability of my random variable X and Y living in some 2D area by just integrating this thing up over all of those little infinitesimal dX dY in that area. Okay? That's the PDF. And the marginal density is defined in the following way: the marginal density. So we've heard, you know, marginal all the time in like economics and statistics. Marginal density essentially allows me to take this PDF in X and Y, this joint PDF—let's call this a joint PDF—and it allows me to write a PDF just in terms of X by averaging out the Y variable. So I can get the marginal density F, uh, just in terms of the X variable, F(x). And this is essentially what I would get if I take this uh joint distribution and just integrate out the Y variable. I'm basically saying, uh, what is the probability of X conditioned on something in Y happened? Any, you know, Y can take on all of these values. I'm just going to integrate over all the probabilities of all of the things that could happen in the Y direction and get rid of that Y variable. So this equals the integral from minus infinity to infinity of my joint probability distribution f(x, y) dy. Good. Um, and that's that's it. It's a really, really simple definition. You literally just—sorry, not dX dY, just dY—it's a really, really simple definition where essentially what you're doing is you're just integrating the Y variable to get a function that only depends on uh on the X variable. And again, roughly speaking, we remember the the law of total probability: something has to happen; Y has to take on one of the possible values that this random variable could take on. So if I integrate out all of those possible, you know, possibilities of Y, then I'm left with just a probability distribution of what X is going to be, um, kind of averaged over all of those things that Y could have been. And you can do this again uh for Y. That's pretty easy. You can build the marginal density in Y; it's kind of exactly the same thing, but now uh we're integrating out the X variable. Okay.

Um, in discrete time or sorry, discrete random variables—these are continuous; in discrete random variables, it's kind of the same thing. So if this is a Bernoulli random variable or a Poisson random variable, you can do the same exact thing, um, where now if I have this P(X,Y), I can derive a probability just in X, um, that essentially, you know, of little x, and what it is is I'm going to take this distribution here, this uh x = x, y = y, and I'm just going to average out all of these Y variables. So I'm going to say, uh, I'm going to add up all of the possible probabilities over all of the possible states that my Y variable could take, and I'm going to essentially average out this Y variable to get something that's just a function, a function of X. I have too many parentheses here, but that doesn't really matter. Okay? So that's a really simple idea of this marginal density function, and it's just something you can kind of define. If you have a joint distribution, you can average out one of those variables to get just the distribution in X. Things I want you to do is to verify that this is actually a well-defined PDF. If you integrated this uh from negative infinity to infinity, it had better equal one. Um, so make sure that you actually believe that these really are PDFs, and make sure that you think you can go back uh backwards and forwards. So you can actually um look up the formula for a two-dimensional Gaussian. I'm going to write down what I think it is: e to the minus let's say x² + y^2 / 2 * 1/(2π). There may very well be an integration factor I'm missing here, but let's say that this is f(x, y). You could easily write this as f(r, θ). This is just r squared, so you could write this as f(r, θ). And I want you to go through the exercise of going back and forth between uh these marginal densities and this probability distribution. Convince yourself that this makes sense, that you can manipulate these things, and then try it on some simpler distributions too. Okay.

One last thing I want to point out: um, there was this notion that was super important earlier of a conditional probability. Um, so all of Bayes' theorem uh and kind of inverse statistics is based on this conditional probability and figuring out the probability of X given that we know Y happened or vice versa. And so I just want to write down how this looks using these joint distributions. Um, so the probability—and I'll do this in discrete random variables first—the probability that X um takes on some value given that Y takes on a little value y is going to be my joint probability distribution, probability of um X, y divided by the marginal probability of Y, which I didn't write down here, but it's exactly the same where you just average out X divided by the marginal probability of of Y. And I think you can actually, you know, I want you to go back a couple of lectures to the conditional probability and the Bayes' theorem, and I want you to write down a page, you know, on, you know, a white sheet of paper. I want you to write down kind of that version of this math where now the probability of X and Y happening is probability of X and Y, kind of probability of X and Y divided by the probability of Y. This is almost identical to what we wrote down in conditional probability earlier. This is now just using these distribution functions to make it a little bit more uh more formal. This is a function over all values of X and Y, um, which is a little bit more general. And similarly, we can do this in continuous time. I'll just write this down because it's pretty cool. Um, the the probability density of—here I did X given Y, down here I'll write Y given X just to make it more interesting—and you can again just flip the variables and you get Y given X, no big deal. So the marginal—sorry, the the the conditional probability distribution of Y given X for a continuous random variable—this is of little y given little x—is just equal to my probability distribution x, y divided by my marginal density for X, this f<sub>X</sub>(x). Okay. Um, and again, you can convince yourself this is really like the probability of f(x, y) divided by the probability of X. Okay? So this is very much like what we did before in conditional probabilities. Now we're defining these conditional densities functions. Okay. Um, that nothing here was complicated. It's a lot of information, but I think it's it's all useful information. From your con—from your discrete and continuous joint probability densities, you can derive marginal densities where you integrate out one of the variables, or you can write down conditional densities where it's the probability distribution of one variable given that you know another variable exists and or sorry, is takes this value. And this is essentially creating a distribution out of those conditional probabilities that we wrote down earlier. Okay. Thank you.