Transcription
Welcome back. Okay, so we've introduced the concept of random variables and probability distributions over those random variables. Now it's time to talk about joint probability distributions. So this is not how you would hand out 100 joints at a fish concert; this is how two random variables may or may not depend on each other and jointly affect some probability of both of those events happening.
Given two random variables, X and Y—so X and Y are two random variables; they don't have to be the same distribution—two random variables, I can define a joint probability distribution as the probability, um, little x, little y. This is the probability that my random variable x equals X and my random variable y takes on the value little y. Okay, this is a really simple idea. We've already talked about conditional probabilities—you know, what is the chance of X happening given that y happens? This is very, very closely related. And here I've drawn this, or I've written this in kind of a discrete random variable form, but you can also do this in continuous random variables, and I'll do that in just a minute. Okay, so I just want to give a couple of examples to motivate why we're introducing this new concept. So a number of examples—in fact, there's a ton of examples.
One of the ones I think about a lot is if you have a turbulent fluid, then the velocity components—the X, the Y, and the Z velocity components—are jointly distributed random variables for a turbulent fluid. So in a turbulent, in turbulence, the U, V, and W velocity components in the X, Y, and Z directions are jointly distributed random variables, and that joint distribution would depend on the fluid flow of interest. If I have like kind of random isotropic turbulence where direction doesn't matter, maybe these would be independent; I don't know. If I have a boundary layer where the flow is going from left to right in the X direction, there will be a very specific structure to this joint probability distribution of how V and W and U kind of co-relate. Okay, so that's one cool example.
Another big example is in things like population health and medical outcomes. So if you have patients' kind of biometrics and demographics—so let's say health and patient demographic, but also biometrics—there will be correlations, for example, in the probability of heart disease given that I am, you know, a 40-year-old male living in the US, okay, in Washington. Like those demographics and biometrics, and you know, let's say I'm 6 feet tall and a certain number of kilograms—that taken together can inform a joint distribution of different health outcomes, you know, that that might be relevant to make actionable decisions based on. And actually, this is super, super closely—when you have these joint distributions of things like this—this is one of the underlying assumptions that goes into principal components analysis. So lots of you have actually already seen PCA, principal components analysis, before. PCA—this is how you take high-dimensional data that you collect from a system. Maybe I just measure the demographics and biometrics and health outcomes of a thousand or 10,000 people, and I do principal components analysis on that data to extract kind of approximations of that joint distribution. I have a whole lecture series on PCA; this is kind of an advanced topic in statistics. We'll get to that sometime soon. And this is essentially assuming—sometimes we neglect this assumption, or we kind of forget casually that there's this assumption—that the joint distribution for PCA is a joint Gaussian distribution, that these would be kind of normally distributed random variables. But sometimes we can kind of forget that. And there's a lot more for examples. I want you to be thinking of joint distributions for discrete variables, for continuous variables, and things you can do with that.
One of my favorite examples—and this is actually how I'm going to introduce the continuous version of a joint distribution—one of my labmates in grad school, in one of his follow-on jobs, worked at a sports analytics company, and I heard that they collected a bunch of data, like camera data of a basketball court, of the basketball court during games, and they could follow players around. And so you can actually—it's a really, really simplified court—and what you can do is you can actually follow a player around through an entire season, and you can build a probability density of where you are most likely to see that player. So maybe you have a player that hangs out here more often, and sometimes they're here, and very rarely they'll be here. That would be a probability distribution for where that player—let's say, you know, LeBron James—is going to be across a season. And that again, we're starting to get into this notion of building these distributions from data that you actually collect. This is a data-driven approximation to a probability distribution; you're modeling where this person is going to be as a probability density function. And so again, in continuous time, we often denote this PDF as this function f(x, y), and roughly speaking, it's the probability of finding them at an infinitesimal little dx by dy kind of teeny tiny little infinitesimal section here. And so you can compute the probability that, um, you know, my my random variable or my my person is going to be in a region. So let's say I define some region here, some region A; the probability of X, Y being in that region A is just the integral of this PDF; it's the integral over that domain A of f(x, y) dx dy. So it's exactly how we do a single random variable, a PDF of a single random variable, but now you can integrate over. And you could do this for three-dimensional random variables; you would integrate over volumes. Really simple idea to calculate the area of being in some some, you know, finite region of this court given this PDF. Good.
What are some other things I want to tell you? I want to tell you what happens if these two variables are independent; that's pretty important, and connect it to separation of variables. Maybe just for a moment I'll go back to this Gaussian example here. So what if I have two variables, X and Y, and they are both normally distributed random variables? They're both Gaussians. Then this can actually set up a new distribution that we're going to call f(x, y), where these are kind of jointly distributed, where X is a Gaussian and Y is a Gaussian, and you could basically build another two-dimensional Gaussian. I'm not going to write out the PDF, but I'm going to draw a picture for you if I can; it's a little bit of a hard picture to draw. So if it's a Gaussian in X and let's say it's a Gaussian in Y, then you get this kind of radially symmetric Gaussian in X and Y. So we're going to say this is my my X direction, let's say this is my Y direction, and you'll notice that this, you know, joint PDF is kind of itself a two-dimensional Gaussian. Again, that's the underlying assumption of PCA, principal components, is that your high-dimensional data is a high-dimensional Gaussian, kind of this high-dimensional Gaussian structure. And you'll notice that if I kind of average out all of the X variables, I should recover a Gaussian PDF in Y, and if I average out all of the Y variables, I should get a Gaussian distribution in X. And these are called the marginal distributions, where you average out all Y to get the marginal distribution in X, or you average out all X to get the marginal distribution in Y. And we'll see this more later; I just wanted to kind of paint this picture for you that if you have, for example, two variables that are in that are themselves Gaussians, you can build a joint distribution that's a two-dimensional Gaussian where each of the marginal probabilities are themselves Gaussians. Okay, good.
The last kind of major thing I want to show you is this notion of independence. This is a really important property, and we're going to use this over and over and over again when we compute the expected value of a joint distribution, when we compute the variance in PCA, things like that. So we have already seen this notion of independence when we looked at conditional probabilities. So two variables, X and Y, are independent if knowing about Y doesn't change my probability of X and vice versa. And you can also write it for these joint distributions as well. So independence, independence—these variables X and Y, these random variables are independent if the probability of X equaling some specific value and Y equaling some specific value is the product of the two independent densities; is the product of the prob—the probability of X being this value times the probability of Y being this specific value, little y. And you'll notice this is almost exactly like the notion of separation of variables when you solve a partial differential equation. So if I'm solving the heat equation on a two-dimensional rectangle and I'm looking at the heat distribution, I can often separate that solution into a function of x times a function of y. That's the same exact idea here. So this is just like separation of variables. And actually, in that heat example, the separation of variables for the heat distribution we actually also involve convolution with a Gaussian; the solution to the heat equation involves a Gaussian heat kernel, but that's that's neither here nor there. So this of independence is really, really important; it allows you to multiply these probability density functions to get the joint distribution if X and Y are totally independent, independent random variables. Now, sometimes this doesn't work; separation of variables, we know that sometimes this works and sometimes this doesn't work. So let's come up with an example. Let's say I'm throwing darts at a board. Okay, so let's say I have a dartboard; usually they're circular, and usually they're divided into a bullseye and then rings and sectors. So I could try and—actually, if I'm, you know, any good at throwing darts, you'll notice that the the distribution will start to look like a Gaussian ideally. If you're actually good at darts, you'll start to get a Gaussian distribution about where you're aiming for. So if you aim for the bullseye, you'll have kind of this 2D Gaussian distribution around that bullseye. In fact, in one of our old houses, we had a dartboard on on a wall, and occasionally we would miss the board because we weren't that good, and when we took the board off, you could actually see this, you know, tapering Gaussian distribution of density away away from the board. Okay. Now, if I tried to write this PDF in X, Y variables, it wouldn't be separable because if I go out in the X direction, my Y probability really does depend on where I am in the X variable. So I can't separate this PDF into X and Y; it's not separable in X and Y, just like the heat equation. If I hit a blowtorch on this circular disc at the center, the heat distribution is not going to be separable in X and Y, but it is separable in theta and r, in the the radial and angular coordinates. So there often are good coordinates in which I could represent my probability density that would kind of separate due to independence, and there's coordinates that are bad, just like in differential equations. There's good coordinates; in probability, there's often good coordinates also. Now, of course, gravity does bias this and kind of breaks the symmetry, so maybe we're throwing darts in zero gravity in the International Space Station, but this is kind of a useful idea that you should know of independence, and we're going to use this all the time, this property of independence. If two processes are independent, you can build a joint distribution by multiplying their densities, or you can kind of compose a density into two products. Okay, good.
One of the things I'm going to do in a follow-up video is talk about something called the marginal distribution. I already hinted at it here; it's this idea that if I have a joint distribution, I can kind of average out Y and get a marginal distribution in X or vice versa, to do that in Y. And I'm going to relate that to the notion of conditional probability, this kind of, you know, conditional idea that we use in Bayes' theorem earlier, and we're going to relate it to independent and joint distributions, I think in the next video. All right, thank you.