Transcription
Welcome back. So I want to take a little tiny aside right now because I made a statement in one of my earlier lectures about the moment generating function, uh, and I want to clarify. So I said this moment generating function is a very useful transformation of your probability density function; in fact, it's the Laplace transform of the PDF of your random variable X. And I said that this moment generating function uniquely determines the cumulative probability distribution, the probability that your variable is less than or equal to some value X.
Now this is weird. Why didn't I say that this uniquely determines the PDF, the probability density function? Why did I say the CDF, the cumulative distribution function? And this is a really subtle but important point about functions in general, but especially probability distributions, that it turns out that the cumulative distribution function is actually more general than and and easier to work with than the PDF. So the CDF, the integral of the PDF, is sometimes easier to work with and more general than the PDF. So remember that this cumulative distribution function, the probability that X is less than some value, is the integral from negative infinity to that value of my probability density function f(x) dx. So for a normally distributed variable, we get that nice sigmoidal error function for the cumulative distribution, that F function.
And what I want to point out here—this is just a brief aside, not a big deal, and we'll talk about this later—is that often times my probability density function can actually be pretty nasty; it can have delta functions and discontinuities. So let me do an example here. So one of the things my probability distribution can can have is it can have weird spikes. Uh, this pink marker is dead; it can have weird spikes in this PDF. So this is like a delta function. Okay. And where could this come up? I mean, why am I making up some weird PDF, you know, f(x) that has like this normal with a delta function? Well, it turns out if you look at the statistics of the heights of men in the US, um, the heights of men in the US, either on, you know, driver's licenses or dating apps or whatever, you get a pretty normal distribution; people's height is fairly, you know, Gaussian or fairly normal. But people seem to pile up at 6 feet tall. There seems to be this preponderance of people that's exactly 6 feet tall. Now, of course, we know statistically that's because there's social pressure to be a certain height, and if you're close, if you're like 5'11 1/2", you're going to bump yourself up; even 5'10 1/2" people might bump themselves up to 6 feet. And so there is this kind of big spike or discontinuity at a specific value. So even though height is a continuous random variable, Gaussian distributed, in reported heights you get this big spike at 6 feet, okay, for men in the US. I don't know about men in, in, you know, Sweden; it might be different. Okay.
And so if I look at the cumulative distribution function, this is a nasty function; this has a delta function; um, this is kind of a generalized function; it's a pain in the butt. But the cumulative distribution function, this big F(x), is a little bit easier to work with. So what this cumulative distribution function does is it just integrates from left to right, so it finds the value, you know, of the integral to the left of some little x, and that's what it reports. And for a regular Gaussian, we would get a sigmoidal error function, and we get essentially the same thing here, a sigmoidal error function, but at this discontinuity we essentially get a little jump in the CDF. So I can have a discontinuous, a discontinuous cumulative distribution function, and that discontinuity is is nasty, but it's workable. I can write down this function; I can do stuff on it; I can compute it. It's much harder to work with these abstract kind of generalized delta functions. And lots of times you actually have this; you have this kind of continuous distribution and this discrete kind of point spectrum of delta functions.
Now there's a whole field of math, um, of functional analysis where you can handle these kinds of functions; you can integrate these. The basic idea—it's called Lebesgue measure theory—and again, measure theory because typically we're measuring probabilities; so Lebesgue measure theory and Lebesgue integration, it's an alternative to Riemann integration for these nasty functions. But the basic idea is that often times we're going to work with cumulative distribution functions instead of probability distribution functions. So when we want to prove the central limit theorem, we're going to prove that the sum of a bunch of independent, identical random variables approaches a normal distribution; we're going to prove that by showing that the moment, the moments of that sum converge to the moments of a of a Gaussian. But we're going to do that essentially proving that the cumulative distributions converge, not the probability distributions. It's harder to show that things converge when you have these weird properties; it's easier to show that things converge in this integrated cumulative distribution function because it's better behaved; it's, it's, you know, for every discontinuity here, for every discontinuity here, it becomes even worse here. So integration smooths things out, makes things better behaved; just like computing derivatives of noisy data makes it worse, but integrating noisy data makes it better; integrating your messy kind of probability distributions makes them easier to analyze. So often times we're going to work with cumulative distribution functions, and this is just at least a sketch of why CDFs often come up instead of PDFs. We'll dig into this more later; we'll talk about Lebesgue theory; we'll talk about measure theory and measure spaces and like kind of the the more abstract theory, um, but I just wanted to give you like a little hint of why this came up the way it did. Okay, thank you.