📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Probability and Statistics: Overview

Steve Brunton29:43

Transcription

Welcome back. I'm Steve Brunton, a professor at the University of Washington in Seattle, and I am super excited to introduce a new short course on probability and statistics. This is one of my absolute favorite topics in all of mathematics because it is one of the most powerful tools we have to describe the complexity we observe in the real world. So probability and stats is up there with differential equations, calculus, and linear algebra in terms of the most kind of foundational tools, and it's important in the modern era now that we're thinking about data and machine learning. This is really one of the foundational mathematical topics.

So I've been looking forward to teaching this class for over a decade. Um, I was actually fortunate to learn probability and statistics when I was a teenager. I was a 17-year-old kid in Texas, uh, at the University of North Texas in Denton, and I learned probability and stats from one of the true greats, kind of all-time greats, Dr. John Quintanilla. Uh, in fact, I still have his uh lecture notes here. So this is my uh bound copy of his lecture notes; my wife uh has told me that I've had this since before I met her. This is one of my most prize possessions. Uh, and Dr. Quintanilla's notes were really, and lectures were incredible and brilliant and changed my life. And so a lot of what I'm going to tell you about in this course is going to kind of channel, um, hopefully some of that inspiration I got when I learned it from Dr. Q, uh, 20 years ago. Okay.

So I'm super excited to launch into this. This is going to be a short course, about 10 hours of probability and about 10 hours of Statistics, starting with pretty introductory, um, essential material and then getting into special advanced topics pretty quickly. So about half and half intermediate and advanced topics, probability and stats, about 10 hours each. Um, and I'll probably keep adding videos and lectures along the way because there's so many interesting topics you could just, you know, there's there's no end to the interesting uses um and properties of probability and stats. So in this introductory overview video, I'm going to give you some examples uh of how we use probability in the real world. I'll give you an outline of what this course is going to look like, especially the first half on probability. Um, again, I hope you are as excited about this as I am. This is a new tool set in your arsenal if you haven't seen this before that's going to open up a ton of possibilities for modeling the complex real world around us. So let's jump in. Um, let's start with some examples. I think examples are how I always, uh, I always think about this, so let's talk about some examples of what we model in the world with uh with probability. And again, I'm going to keep using this word uncertainty. The real world is complex; it's uncertain, and it's uncertain specifically in my ability to measure everything and to model all of the physics and and behavior. And so that's what I mean by uncertainty. If something is too complex to model or measure, then it's a very good candidate for building a probability model of.

So maybe the first uh kind of classic example I think of is going from, you know, gas molecules to descriptions like temperature and entropy. So going from something like a gas, you know, we have 10 to the 23 um gas molecules in Avogadro's number in a mole of, you know, of air or nitrogen or oxygen, and that's too many degrees of freedom first off to measure. I can't measure the position and velocity of every molecule of air in this room, um, and I don't want to, and it's also too complicated to simulate all of those particles. And so we have built this incredibly efficient, elegant, beautiful description of this very complex physical system, this gas, in terms of a few thermodynamic properties: the temperature, uh, the entropy, and so on and so forth, you know, density. And so this is one of the successes of all of modern physics, one of the kind of real uh cornerstone successes of modern physics is this thermodynamic description of gas in terms of these simplified statistical quantities. Just a few numbers characterize this, you know, astronomically complex system. Okay, that's a great example of something to model probabilistically. You know, you could model this with a Boltzmann distribution or Maxwell's distribution, things like that, to get um these simple statistical quantities. Um, and you know, fast forward about a hundred years, we're still trying to build these kinds of thermodynamic closures for systems like turbulence. Um, so a turbulent fluid, um, like you know, the the air over the wing of an airplane or in your coffee when you stir it, that is also a very, very complex system with so many degrees of freedom, so many particles, um, that we can't measure all of them with certainty, and we can't describe their motion, their behavior; it's too complicated. And so turbulence is a great candidate again for probabilistic and statistical models.

Um, one of the classics again, even before uh our thermodynamic description of gases is the notion of measurement error. So we have been modeling measurement error, um, measurement error. So uh if you do some experiment, you're trying to measure, you know, uh the acceleration of gravity, that's a great one, or the speed of light, or the mass of an electron, or the charge of an electron. If you're doing some physical measurement, there's going to be error associated with that. So if you repeat that measurement 30 times, you're going to get 30 slightly different values, and it turns out that often times that measurement error behaves according to this normal or Gaussian distribution, which is a probability distribution. Um, so this is actually, in my mind, one of the genesis points of modern probability and statistics, Laplace. Um, so our our great French mathematician Laplace—any of you who know me know that I, you know, am a huge fan of Laplace from differential equations—Laplace was also essentially one of the pioneers, one of the founders of modern probability and statistics. So he developed a lot of the theory that we use today, and the nomenclature, the perspective, the philosophy, um, really to understand this notion of measurement error in physical systems, to start quantifying the uncertainty in our models of the real world. Uh, and in fact, you know, Bayesian statistics is one of the most important topics in statistics. It was discovered by Bayes, and a couple of years later, Laplace independently discovered it and massively extended it, generalized it, popularized it. I think Bayesian statistics should be called Bayes-Laplace statistics, um, but again, modeling measurement error, which tends to be a Gaussian distributed random variable, was a huge kind of cornerstone um event in probability. Laplace brought this into the kind of modern era.

Um, other things that are important: if you like control theory like I do or dynamical systems, you'll remember the Kalman filter. The Kalman filter um is a great example of merging kind of dynamics and control with a probabilistic perspective. Again, if I'm taking measurements of a system to do feedback control, those measurements have error; we typically assume a Gaussian. We also assume that my system is being kind of externally forced by things that are beyond my ability to model: uncertainty. Maybe I'm trying to build a cruise controller for an automobile, but I don't know if it's windy outside or if it's raining or if I'm going uphill or downhill. Those uncertainties are often modeled probabilistically in things like a Kalman filter, and this is going to very naturally segue into the topic of of stochastic differential equations, where my differential equation itself is forced with a probability distribution or a random variable that has a distribution. And so SDE is a topic I'm going to cover a lot later; this is one of those special topics for later, but super, super important and interesting for dynamics and control. Um, and the list goes on and on; there's so many great examples. Um, I think weather and human behavior are really good ones. Um, so weather and let's just say human behavior is a catch-all um for tons of complexity, behavior, um, sometimes good behavior, sometimes bad behavior, typically interesting and complex, um, and it's interesting that if you think back to the very, very, you know, earliest history of of human thought and communication, the, you know, earliest oral traditions and written traditions, people have always tried to wrestle with the uncertainty in the complex real world, including human behavior. Weather phenomena, you know, is the weather going to be favorable for tomorrow's battle? Is the weather going to be favorable for planting my crops for this year's crop yield? Um, you know, things like that; those are the kinds of things people have been wrestling with this massive uncertainty for millennia, from the dawn of human um kind of tradition. We have been modeling uncertainty, and I think it's fascinating that one of the ways we model uncertainty, have modeled uncertainty historically, we call this divination, essentially—are things like astrology or augury or tea reading—reading the tea leaves, stirring up your tea and reading the tea leaves, um, rolling the bones of animals and seeing how they scatter to predict something about our uncertain real world. It's quite fascinating that one of the ways humans have grappled with uncertainty and the probabilistic nature of the complex real world is to actually try to analyze it and understand it with another random process, like rolling the bones of animals or how tea leaves dry in a cup of tea or things like that: pyromancy, um, how the flame patterns move. There's hundreds of words for different types of divination. It's one of the most common themes throughout all of human history is trying to grapple with our observed uncertainty of the real world, of the complexity of the real world like weather and human behavior, by trying to simplify it into a simple or random process that we can maybe try to analyze or and understand, like rolling the bones or something like that. Fascinating. I think of this, you know, all the time when I think about probability and statistics. There's a rich history and tradition, um, and again, Laplace did a a huge service kind of making this a mathematical, quantifiable science, and a lot of the the thought leaders at that time really changed kind of our our world from this astrology to, you know, scientific astronomy, kind of alchemy to chemistry, uh, you know, divination to probability and statistics. Huge sea change in how we think about the world, which is inherently uh chaotic; it's a chaotic dynamical system. And this is one of the things I really want to to point out is that these probability and statistics systems often times actually are fundamentally deterministic; they're not actually random. If I think about a coin—this is my fair coin—if I flip this coin, this technically is not a random system; this is governed by force equals mass times acceleration, F=ma; it has wind resistance, mass, inertia; it's under the effect of gravity. Um, this is a deterministic physical system that technically I could model on a computer with equations and do a pretty good job of predicting, but as far as I'm concerned as a human observer, there is too much uncertainty in my measurements, in my ability to model this in my head. So for me, a coin flip is a random or uncertain uh event that I have to model probabilistically. Okay, and I can use this uh to generate random numbers kind of um because it it does have this kind of chaotic randomness in the the wind resistance and its motion. And so a lot of these systems—turbulence, gas dynamics, weather phenomena—those are actually fundamentally deterministic dynamical systems; they are predictable. Often the chaos in that dynamical system means that my prediction horizon is is finite, and after some point, for all intents and purposes, I have to model it with probability models. So there's this fine line between deterministic uh systems that I could model if I had enough information and us actually having to model them probabilistically because there's so much uncertainty in our physics models, in our uh measurements of the initial conditions of things like that. Good. Okay.

Um, so now I'm going to give you an overview of what we're actually going to learn in this first kind of 10-hour block on probability, and I'll hint at some of the ways this is going to tie to statistics. These are intimately connected; so you can't really separate these. Um, so I'm going to give you the outline of the topics here. What I really want to do first is is point out just kind of a definition here. So probability is essentially assuming that you have a model of your system. So you assume um that you have kind of a known probability distribution; you assume a known uh probability distribution, and I'll tell you what a probability distribution is in a minute. It's thing; it's something like a Gaussian, something where uh you expect to find your variable at this value with some probability. It's a probability distribution. Uh, so you assume your probability distribution is known, and you try to say something about future data you might observe. So you essentially uh don't know the samples from the future; samples are unknown. In statistics, we often call our data samples. So the data from the future is unknown. In this fair coin, I have a model of its probability; there's a 50% chance of flipping heads, but if I but I don't know what will actually happen, what the exact sequence will be if I flip 10 uh coins in a row. So probability will give me some way of quantifying the likelihood of seeing, let's say, five heads out of 10 flips, or seven heads out of 10 flips, or 10 heads out of 10 flips. So I don't know the future, and I'm trying to say something quantifiable about what the future might look like based on a known probability distribution. Statistics is the flip side of that, where we actually um the data is known; the samples are known. So let's say I flip the coin 10 times or 100 times, and I measure, I observe that system, and now we're trying to say something about the probability distribution. So the probability uh distribution is unknown. And so these are really flip sides of the same coin; they're dual problems. Um, in probability, we're going to start with probability because that it gives us the family of models that we're going to then use when we have data. So we're going to learn about probability and how to model probabilistic systems, and then once we have that mathematics under our belt, then when we collect data, we can say really, really precise things about that data using uh kind of these probability models we learned earlier. So let's get into it. Um, so things you're going to learn in this class: I'm going to break this into a few key chunks or sections that I thought were nice break points. We're going to start with essentially um introductory probability. So this is really kind of uh introduction, intro to probability, and introduction to probability involves things like um building intuition and doing a lot of examples. What is the chance of of flipping seven heads in 10 coin flips? We'll be able to precisely quantify that. How many types of poker hands can I deal off of a 52-card deck? Um, if I roll three dice, what's the chance that they add up, the numbers add up to 13? Okay, those are the kinds of things that we're going to to look at initially. This might seem like a slow introduction if you already have some background in probability; you can probably skip this or, you know, watch this at one and a half speed. But really, this is going to be examples and intuition, examples and intuition, and especially this notion that probability is really, if you boil it down, it's all about counting sets of things that can happen, counting sets; it's it's an advanced way of counting uh and grouping events into things that can happen in different ways and counting them. So we're going to build a lot of intuition and give a lot of examples here. This will be pretty rapid actually. If you think about the thermodynamic closure of gases when you learned this in school, you might have learned about the, you know, canonical ensemble or the Boltzmann distribution; that's essentially a fancy way of counting the different ways molecules can be arranged and and uh and turning that counting into a notion of entropy. Okay, so a lot of intuition here, and then very quickly we're going to get into the meat; we're going to get into an abstraction that's essential: the abstraction of a random variable. So uh maybe I'll do this again in blue. So this is the abstraction of a random variable; a random variable uh we're going to call this X and functions of random variables uh and distributions um let's say and distributions. This is the key point. Okay, good.

And so a random variable, just like in regular mathematics, algebra, and calculus, you have a variable that can take a specific value. In probability, we now have random variables that have a probability of taking a certain value. So this random variable, we now say um we would essentially say that we have some random variable X, and it has a probability of taking on some specific value given some parameters um that that specify that system. So if I have a Gaussian, if I have a normally distributed random variable that I'm using to represent my measurement error, it probably has a mean, an average value, a mean, and maybe a standard deviation. Okay, those are two numbers, two parameters, theta, that I need to specify this probability distribution, and that tells me the probability of finding my random variable at a particular value V. Okay, very useful generalization of this notion of a variable and a function from calculus. Now you have variables and functions in probability and statistics, and those variables have a probability associated with finding them in a certain value, but you can do still do things like you can take the square of this, you can take the log of x, you can plot, you can do calculus on this. Very, very useful notion of random variable, and we're going to build up a bunch of distributions, probability distributions that describe different kinds of events. So, for example, um we're going to have the Bernoulli uh random variable. This is a probability distribution that describes the probability of, let's say, coin flips, systems that have two outcomes, you know, heads or tails, success or failure, up or down. Bernoulli random variables is the distribution, the the the probability distribution to describe those kinds of events. If I have a bunch of coin flips, if I flip, you know, 10 or 12 coins in a row and I want to know how many heads do I get, that's something called a binomially distributed random variable, binomial. Very, very, very, very important um distribution; we'll use a lot. It's the sum of a bunch of independent Bernoulli random variables, and if I have a large sample size, a large number of coin flips in my binomial distribution, it starts to tend towards a normal distribution, a normally distributed random variable. And that's actually kind of where this notion of Gaussian measurement error comes in. Gaussian and normal are the same thing; they they're the same name for the same distribution—sorry, different names for the same distribution—is if my measurement error comes from a bunch of different factors, you know, like the temperature and the wind and all of these different factors that kind of add up to give me a measurement error, it turns out that if you add up a bunch of random variables, very often their sum starts to look normally distributed, even if individually they have a weird distribution. Themselves. So measurement error, there's this like uncanny uh fact that measurement error tends to often look Gaussian, and that's a very fundamental probability property um that we'll talk about later called the central limit theorem. So large uh n limit, the limit of of a large number of events in a binomial distribution tends to be normal. If they are rare events, then they tend to be Poisson. These are both limits of the binomial distribution in the large n limit; Poisson for rare events like um light bulbs failing or the radioactive um emissions, you know, alpha particle emissions from a radioactive element; that's a Poisson distributed random variable. These are two of the most important uh probability distributions; we'll talk about them. They're used everywhere across the board in probability and statistics all the time; super foundational uh stuff. And then a bunch of other ones; we'll talk about things like um the exponential distribution, um exponential distribution, exponential distribution is tells me the probability of waiting times between Poisson events, like radioactive decay. If I have alpha particles being emitted, what's the probability that it'll be 0.5 seconds until the next one? That is an exponentially distributed random variable. And dot dot dot dot dot dot; there's dozens of these; I'll probably tell you about 10 of them, the 10 most useful ones that you'll use all the time in probability and statistics. Um, good. And so this is a model of the likelihood of finding my random variable at a value given some parameters that characterize that system. Statistics is going to flip it on its head, and given data, we're going to find the probability of the parameters of my distribution. Again, that's the that's the statistics problem, but but today we're talking mostly about probability. Good. Um, and sometimes the distribution is known; I'll make a little note of this. Uh, sometimes sometimes the distribution is unknown. For lots of complex real-world systems, the probability distribution—we believe there is one, like in turbulence—but it might not be known. That is essentially a machine learning problem. So when we don't know the distribution but we have data and we're trying to learn the distribution from data, that's essentially a machine learning problem. So that's how this really connects to machine learning is we're learning these distribution functions from measurement data for much more complex systems like turbulence and weather and human behavior, behavior, things like that. Okay, good. Uh, let's keep going. So uh topic three, big, big topic here is um now we're going to talk—now that we have random variables and distributions—we're going to talk a lot more about functions of random variables, um things like the expectation value. So the expected value, um the expected value, what like literally what value do I expect my variable to take if I have this distribution? What's the expected value? What's the variance um or standard deviation of that value, like how much spread do I have in my prediction? Uh, things like the median; these are robust predictions that are related to the mean. Essentially, how do I quantify this probability distribution in a few numbers that tells me the essential things I need to know? That's very much like this gas distribution going to temperature and entropy; there is a distribution for my gas, but there are a few numbers like the expected value and the standard deviation that I really need to know um to say a lot about that system. So this is also going to be super important to tie back to statistics because when we have data, these are the things we're actually going to be estimating about our probability distribution. Often we're going to be estimating the mean, the variance, the median, things like that, from data to say things about our probability distribution. Okay.

Um, good. And then I'm going to kind of end the course um on something that is really near and dear to my heart, which is the central limit theorem. Central limit theorem, and it's one of the most important topics again in all of probability, and it's a cornerstone of statistics. So this is really where we start transitioning from probability to statistics. It says that if I have a bunch of random variables that have the same distribution and I add up those random variables or I average those random variables, that new average quantity starts to look like a normally distributed random variable, regardless of what the original distribution was. We use this all the time in statistics. When you take large samples of data, you can estimate things like their average value using the central limit theorem. Um, it's used in every single one of these examples in some way or another. So the central limit theorem—it's so important—I'll actually just give you kind of an idea. If I have a bunch of random variables, um, let's say x<sub>i</sub>, i can be 1 to n, so I've got like a bunch of data, then the average of this—literally we call it x̄, it's called the the sample mean, the average—it's, you know, 1/n * the sum of all of these variables or all this data—this quantity, this sum of random variables, tends to be normally distributed, normally distributed with some mean uh and variance; that's what we call the mean and variance. This tends to look like a normally distributed random variable, regardless of what these distri—how this data was distributed. This data could be distributed as a Poisson, binomial, exponential, you name it, some weird machine learning distribution, and if I add up those variables, if I take their average, it starts to look like a normal distribution. Again, that's where this Gaussian measurement error kind of comes from is the central limit theorem, and this is the…

Cornerstone, uh, of kind of modern probability theory and modern statistics. And in fact, this is where we're going to very naturally segue into the next set of lectures on statistics. I'll just give you a tiny sneak peek of what we're going to see in statistics. This is all kind of flipping the probability on its head.

So instead of having a model of a coin and then saying things about what's likely to happen in the future, in statistics I'll have data from coin flips. I'll have 10 coin flips or 100 coin flips, and I'll start asking questions: Is this a fair coin? What's the probability of heads? Can I estimate that from the data I have?

So there's going to be things like hypothesis testing. If I run a drug trial, let's say I have a clinical trial for a new cancer treatment, does that cancer drug work or not? That's a hypothesis that I would test with data. I collect data from a trial. And I use some models from probability to say very quantifiable things like, "I'm 95% confident that this drug does work or doesn't work." That's what we mean by hypothesis testing.

Other things we'll be able to do are things called survey sampling. If I have a large population, like the population of the United States—300+ million people—if I take a small sample of maybe a thousand people and I ask them questions or I measure their height, can I get information about the larger population from that small sample? That's very, very closely related to these topics: variance, standard deviation, expected value, and central limit theorem.

But again, this is based on data. I'm trying to say something about the probability distribution. And so here we had a probability of my data given some parameters of the distribution, like the mean and the variance. Statistics is all about finding the probability of the parameters given the data. What's the most likely model? What's the most likely model parameters given the data? This is the statistics problem.

And this "theta" doesn't have to be the parameters of a named distribution. These could be the parameters of a neural network that I'm using to model an unknown distribution. In machine learning, we model probability distributions using things like neural networks; the weights of the nodes, the weights of that neural network are these parameters I'm trying to estimate from my data. I'm trying to learn my parameters from my data—that's fundamentally a statistics problem.

And then the list goes on and on and on—fitting distributions, estimating parameters, and so on and so forth, eventually getting to this machine learning idea. So how likely are our parameters given our data? That's a statistics problem. It's very intimately related to Bayesian statistics. There is this very useful notion of Bayesian statistics where you can incorporate prior information and data to get a better statistical or probabilistic model. And that's really, really commonly used in machine learning today as this kind of notion of statistics, and especially Bayesian statistics, where you mix data and some prior knowledge about what your distribution might look like.

Okay, this is the thumbnail sketch, the mile-high overview of the first half of this new course on probability and statistics. We're going to start with probability, build our intuition, build the mathematical models—these are the tools we're going to use when we then have data and we're trying to say statistical things from that data. I'm super excited to walk you through this. I've been waiting for this for literally over a decade, um, and I hope you get as much value out of this as I do. Thank you.