Transcription
Welcome back. So we're introducing this, uh, notion of probability, or the chances of some event occurring out of a larger set of possible things that can happen. We've looked at things like dice rolls, poker hands, um, coin flips, things like that. And we've seen very quickly that one of the issues is how to enumerate or count all of the possible outcomes.
Before I go much further, I am going to need to introduce some basic set theory, some kind of formalism that allows us to say things about these different events and probability spaces and sample spaces and things like that. This is going to seem a little bit technical and dry, but it's really, really important to set up the basic language of how we think about probabilities in a way that generalizes and gives us shorthand so we can communicate much more rapidly together. Okay, so, um, again, we know this basic idea that a probability is measuring the the chances or likelihoods of some event happening. So, um, probability measures the likelihood or chance of some event A happening, some event A occurring or happening. Good.
And I'm going to start defining some sets, um, like you know, we we we've learned about set theory before. There's the set of integers and the set of the real numbers and things like that. So we know that probability is measuring the likelihood. I'm going to start relating this to set theory. Okay, so we're going to define a sample space. A sample space, Omega, is the set of all possible things that can happen in my experiment. So basically, I'm going to run either thought experiments or real experiments. We're going to talk about flipping 10 coins and what is the chance that at least five of them are heads? So the sample space, Omega, is the set of all possible things that could happen in that experiment. So the sample space, Omega, is the set of all possible outcomes of a kind of random experiment, of some of some random experiment I'm doing, of a random, let's say random experiment. Good. And this could be me flipping a coin 10 times. This could be, uh, dealing a hand of poker. It could be playing, you know, rolling the dice in in backgammon. It could be any number of of experiments, either real or thought experiments. So this is the sample space, and it's going to be a set of all possible outcomes. And we're going to say that a specific element—let's put this in in a different color—a specific element, Omega, we say that this is in the set big Omega. A specific little Omega is basically one instance or one actualization, one realization of this experiment, is one—we're going to say that this is one outcome or realization of our experiment. So you know, lots and lots and lots and lots of examples of what this sample space could be. This could be, um, the number of possible ways that I could flip five coins. Okay, um, so I have five coin flips. Omega could be the set of all possible ways I could flip five coins. Omega could be, um, the number of heads in five coin flips. Omega could be the number of emails I get in an hour. Okay, that's a a large number. Um, Omega could be the, uh, um, you know, a distribution of heights sampled from 10 people. There's all kinds of things Omega could be. Um, it could be today's high temperature. It could be, um, you know, the number of hands of poker, things like that. And a specific element could be: I rolled five heads in a row—heads, heads, heads, heads, heads. That would be a specific element of the space of five coin flips. Good.
Um, and we're going to start to get into this notion of a random variable soon, but for now let's just talk about the space of all possible things that can happen. So let's make this concrete. Let's say that I have—um, I'm going to go to green, keep it colorful—let's say as an example I have, uh, three coin flips. So three coin flips, and for this I can actually enumerate this set Omega. So Omega equals heads, heads, heads; heads, heads, tails; heads, tails, heads; heads, tails, tails; tails, heads, heads; tails, heads, tails; tails, tails, heads; tails, tails, tails. So there are eight, uh, possible elements of of this set Omega. And we're going to define these events, A. We're going to say them in words, but then they're also going to define subsets of my sample space. So if I have this event A, let's say, um, the first flip is a head. Let's say first flip is ahead. That defines an event A where A is a subset of Omega. So if the first flip is ahead, then I can take all of the elements of of Omega where the first coin flip was ahead, and you'll notice that the first four of these are all of the instances, all of the specific elements where the first flip was ahead. So A is the subset HHH, HHT, HTH, HTT. And again, I can compute the probability of A by just counting how many elements are in A divided by how many elements are in Omega. That's a really simple way because these are all equally likely to happen. If these were different probabilities of these happening, I'd have to compute my probability slightly differently. Okay.
Um, I can do other things. I can say my second flip, uh, second flip is a tail. So let's call that event B, and let's just, you know, we're going to do this the slow way. Let's enumerate all of the possible ways my second coin flip can be a tail. So I have, uh, heads, heads, tails; heads, tails, tails; tails, heads, tails; tails, tails, tails. So that's my set B, and these are two different events. There's one event where my first flip is a head, and there's another event where my second flip is a tail. It turns out these are completely independent; the first flip and the second flip are independent. We'll talk about what that means later in terms of probabilities, but these are just two outcomes that could happen, um, and we can compute probabilities of the first flip being a head, the second flip being a tail, and so on and so forth. We can start to think about the sample space and the probabilities in terms of pictures like Venn diagrams. Okay, so I'm trying to keep my sample space kind of blue here. So let's say that I have, um, you know, my sample space Omega, and I'm just going to draw kind of abstractly what this means, and let's say I have event A and event B. Technically here, these events are independent, um, so it's a little different. I'll show you what that means in in a minute, but I'm just going to define kind of roughly these different sets, a set A, a set B, and a set Omega, and there are things we can define, um, in terms of set theory that are going to be really useful in probability theory. So we can, I, um, we can define this notion of a union. So A, uh, or B is any element with everything in A or B. So either A happens or B happens. So either the first flip is a head or the second flip is a tail. And the way we think about that is everything in the set A and, uh, in B. So we basically take everything in set A and everything in set B, and we, uh, we combine them together. And I really should be careful and not use the word "and" because that is like the exact opposite of what I mean. Everything in A or in B, anything that is in either A or B, and that is going to equal—I could literally just enumerate everything in A or B is—HHH, HHT, HTH, HTT, TTH, TTT. So this set is bigger than A and and B. So this one has six elements because it's anything in either of these two. Okay, so that's what A or B means. So I'm going to just going to draw that here like A or B is the set here of all of anything in A or in B. Very common sense.
Um, I also have the set A and B. So A and B is, uh, A and B is everything that is in both A and B. This is called the intersection, and it means that for an element to be in the intersection, it has to be in both A and in B. So it would be this intersection region here, this region here that's A and B. And we can compute A and B just as easily. So we have to go through and look at which elements are in both A and B, and very quickly we see that heads, tails, heads and heads, tails, tails. And that makes sense. These are the elements that the first flip is a head and the second flip is a tails. So this is a smaller set. The intersection generally is more restrictive; the union is more of a big umbrella, more permissive, more inclusive. So union is everything that's either in A or B; intersection is everything that is in both A and B. And then there's another one which is the complement that's kind of useful too. Maybe I'll just pick orange. So, uh, A complement is everything not in A, and I don't want to enumerate it, but it's all of the things not in A. So since A is the first flip is a head, then everything not in A is everything where the first flip is a tail. Okay. And pictorially we would draw that as the set of everything outside of A. It's literally everything that's not in this set A. And these are like the three most useful kind of concepts that we're going to use all the time. So we're going to compute the probability of A and B or the probability of A or B or the probability of A not happening, everything but A, and we're going to to use enumerations. We're going to count how many elements are in the union, in the intersection, in the complement to compute those probabilities. Okay, very, very useful. There are some basic properties here like there's, um, things like commutativity: A or B is the same as B or A; A and B is the same as B and A; A or B or C, it doesn't matter what order you do the or; A and B and C, doesn't matter what order you do that; it's the same answer. Okay, so that's, um, commutativity and associativity. These are commutative and associative operations on these sets, is the mathematical way of saying that, but it's very common sense in terms of these pictures.
Um, if I really wanted to be clever about this or if I wanted to be really, really accurate, what I'd actually draw—maybe forget three coin flips, let's just say there's two coin flips, make my life a lot easier—so I have heads, heads; heads, tails; tails, heads; and tails, tails. What I could do is I could say that the first flip is a head would be all of these. So say this is the first flip being heads, this is the first flip being a tail, and let's say that the second flip being a heads is this half, and the second flip being a tail is this half. That would be element B. So this would be B, the second flip is a tail; this set would be A; this is A; this is B; and it's really easy now to see that the probability of A and B is one quarter; the probability of A or B is three quarters; the complement of A is one half. You can do this really, really easily. You can actually make this diagram make sense with the problem, uh, at hand here. Okay, good. Uh, okay, cool.
So last thing I want to say is again we have this probability is the measures the likelihood of some event happening, and we have these sets which determine the sample space of possible outcomes, and the sample the space of of how many of those outcomes, uh, are consistent with some event that I'd like to keep track of, that I'd like to compute the probability of, like getting a head on the first, uh, the first coin flip. Now here we have in all of our examples we've dealt with events that have equal probability. It's equally likely to get a head or a tail, but I could have a biased coin. So before I shaved today, I was thinking about bringing a bringing a biased coin where it would have had a, you know, more likely chance of getting heads than tails, and then it would be a little bit different than just counting the number of outcomes of A divided by the number of outcomes of Omega. I'd have to weight those by their by their likelihood, by their their probability. So a probability is actually a measure. So this probability P is specifically a measure on these subsets. It says how likely is a given subset? So it is a map from subsets of Omega to the real numbers, to the reals, and specifically to a number between 0 and 1 in the s. Okay, so this is mathematical language that says: probability. For every single set that is a subset of Omega, for every subset A, P is going to assign a value to that subset, which is the probability between zero and one of that subset of of of an of a random element being in that subset. And this sounds abstract, but it allows us a lot more flexibility. Now I can have, uh, the probability of heads and tails be different. I can have the probability of my second coin, I can have one fair coin and one unfair coin, and I can define the probabilities of each of those events with this probability measure P. So this is abstract, and if this is taking a minute to sink in, that's totally fine. We're going to come back to this, and we're going to develop this more, but this probability measure has some properties that I want to tell you. The first most important property is that P of Omega always equals one. This is essentially a statement that the probability of something happening is one. If I, you know, if I flip two coins or three coins, one of these things is going to happen. The probability of Omega is equal to one. So this is, uh, property one. Property two is is that if A is a subset of Omega, this is A, and it's a subset of Omega, then the probability of A is greater than or equal to zero. There is some chance that A will happen. There is some some chance. It could be zero, but it's going to be, you know, if A is a subset, then it has a probability. And three, the third property: if A and B are disjoint—disjoint means that they are completely separate—that would mean, um, A and B have no overlap. So this means A and B equals the empty set; they have no elements in common—then the probability of A or B is equal to to the probability of A plus the probability of B. Okay, these are the basic basic axioms of probability, and from this you can build almost all of the properties of probability. Now this is pretty abstract. We don't really think about, you know, sets and measures every day on simple problems, but for complicated problems and for theorems you're absolutely going to need to understand what a set is, how to think about these things, these basic axioms, and how to use them. And I'll point out, um, this idea here of like I drew this picture of these probabilities. I literally am thinking that the area of A divided by the area of Omega is the probability of A, and that's a cartoon that helps me think about this like if I'm throwing darts at this board and I'm uniformly, I've got equal chance of hitting anywhere in Omega, the chances of me hitting A are is the area of A divided by the area of Omega. So that's how I think about this pictorially, um, for these problems. And there's a lot of of of follow-on from these axioms. You can derive some other things: the probability of A complement equals 1 minus the probability of A. That means the probability of A not happening is 1 minus the probability of A. Very common sense, and you can think about it in terms of these, uh, these notions. Another one: the probability of nothing happening, of the empty set, is zero. Um, if A is a subset of B, if it's inside of B or equal to B, then the probability of A is going to be less than or equal to the probability of B. Uh, and then the last one, and then I'm going to stop—this is already getting long—the last last property—this is actually a pretty important property, and I want to I to show—I'm just going to do it over here so I don't get into too much, uh, of a space issue—is that if I have the probability of A or B—this is A or B here—the probability of A or B is equal to the probability of A plus the probability of B minus the probability of A and B. And again, you can think of this pictorially almost as the area of A or B equals the area of A plus the area of B minus the area of the intersection of A and B because this piece gets double counted if I just add those up.
Okay, let's zoom out for one minute, and, uh, we'll come back to this occasionally, but I just want you to know this actually goes deeper. Um, there are, um, you know, like pretty big fields of calculus and, uh, and probability theory called measure theory about how do you measure exotic sets. Here these are discrete sets. How do I measure, um, you know, the size of natural images in image space? That's a weird thing to measure. How do I measure, uh, the Lorenz attractor in three space? How do I measure fractals? How do I measure Cantor sets? You can go really deep down the rabbit hole, and you can assign probabilities to things that are on weird sets that I'm not talking about here. This is like page one of a book on on measure theory, but the basic idea is that formally we define a space of all the things that can happen, and we define events to be subsets of that sample space, and for every possible set we can define a probability, um, of that set, of that event happening, and as long as long as these three things are true, that is a well-defined probability mapping. As long as these things are true for all sets A and B, then it is a well-defined probability, and so you're going to be building these probabilities for different sets, for different random experiments—coin flips, how long you're going to wait at the DMV before getting helped, um, probability of failure of a part, all kinds of things. We are going to to define a sample space and a probability measure, and as long as these are true, then it's a well-defined probability, and I can do these common sense things with that probability measure. Okay, thank you.