Transcription
Here are the top 10 most important things to know about probability.
Must know number one: Experimental probability. Let's start off by doing an example where we conduct an experiment where we flip a coin 10 times and then we calculate what the experimental probability of flipping tails is. So, let's conduct this experiment. I'll flip 10 coins. And then, before we write down what the experimental probability of flipping tails was from this experiment, let me give you a couple quick definitions about probability and experimental probability.
Probability is simply how likely something is to happen. The value of a probability always falls somewhere on a scale between 0 and 1. A probability of zero means an event is impossible, and a probability of one means an event is certain to happen. And any probability, let's say a probability of 0.9, can be expressed as a decimal or a fraction. So, in this case, 9 out of 10, or a percentage, 90%.
Now, an experimental probability is a probability that is determined based on the results of an experiment. And let me make some room to add the formula for how we calculate an experimental probability. Let's say we wanted the experimental probability of some event happening, we'll call it event A. To calculate the probability of event A happening, you do the number of times event A occurred and divide that by the total number of trials of the experiment. And there's a shorter notation for writing that, where N(A) stands for the number of times event A occurred, and N(S) means the total number of elements in the sample space, which in this case means the total number of trials of the experiment.
Now, back to our experiment where we flipped a coin 10 times. We wanted the experimental probability of flipping tails. To calculate the experimental probability of flipping tails, I would need to do the number of times that tails occurred divided by the total number of elements in the sample space, or the total number of trials of this experiment. So, based on our experiment, we got three tails out of 10 total flips of the coin. So, our experimental probability of flipping tails is 3 out of 10, or we could write that as a decimal, 0.3, or as a percentage, 30%.
And let me add a note before moving on. The more trials you do of an experiment, the closer the experimental probability is likely to be to the theoretical probability. So, if we were to continue flipping this coin hundreds and thousands of times, you would notice that the relative frequency, or experimental probability, of flipping tails would be approaching 0.5.
And that leads us right into our next must know.
Must know number two: Theoretical probability. Theoretical probability uses reason or math instead of an experiment to calculate the likelihood of an event occurring in the future. And the formula for calculating the theoretical probability of any event, we'll say event A, is you do the total number of possible favorable outcomes and you divide that by the total number of possible outcomes. And in a condensed notation, we could write that as the number of ways that event A can happen divided by the total number of elements in the sample space, which just means the total number of possible outcomes.
Let's now do three quick examples that are common types of simple theoretical probability questions. We'll do questions about rolling a die, drawing a card, and picking a marble. Let's start with rolling a die. For any rolling a die question, we'll assume that we have a standard six-sided die. And let's say we roll that standard six-sided die one time. Let's calculate first the probability of rolling a four.
To calculate the theoretical probability of rolling a four, we would have to do the number of fours that are on the six-sided die and then divide that by the total number of elements in the sample space, which means how many total possible outcomes are there when we roll the die. Well, on the six-sided die, there is only one side that has a four on it, so there's one way we can roll a four. And there are six different outcomes when we roll the die, so the total number of elements in the sample space is six. So, the theoretical probability of rolling a four is 1 out of 6.
Let's also calculate the probability of rolling an odd number. We would have to do the number of outcomes that are an odd number and divide that by the total number of possible outcomes. Well, there are three odd numbers on a six-sided die: 1, 3, and 5. And six total possible outcomes. So, the theoretical probability of rolling an odd number is 3 out of 6, which reduces to 1 out of 2, or we could write that as 0.5 or 50%.
Let's move on to drawing a card. For any question involving playing cards, we'll assume it's a standard deck of 52 cards, where there are the numbers 2 through 10, Jack, Queen, King, and Ace for each of the four suits: hearts, diamonds, spades, and clubs. If we were to draw a single card from a deck of cards, let's calculate the probability that that card is the seven of hearts.
The theoretical probability of being a seven of hearts, we would do the total number of seven of hearts that are in the deck of cards divided by the total number of possible outcomes, so the total number of cards in our sample space. Well, there's only one card that is a seven of hearts, and there are 52 total cards in the deck of cards. So, there is a 1 out of 52 chance of being dealt the seven of cards.
What about calculating the probability that the single card you draw is a king? We would have to do the number of kings that are in a deck of cards and divide that by the total number of elements in our sample space, the total number of cards in the deck of cards. Well, in the deck of cards, there are four kings, one for each of the four suits. And the total number of possible outcomes, there's still 52 total cards. 4 out of 52 could reduce to 1 out of 13.
And let's move on to the last theoretical probability question I'll do with you: picking a marble. Suppose you have a bag containing three red, two blue, and five green marbles. If you draw one marble at random from the bag, let's calculate the probability that you draw a red marble.
To calculate that theoretical probability, we would do the total number of red marbles in the bag and divide that by the total number of marbles that are in the bag. There are three red marbles, and the total number of marbles in the bag, 3 + 2 + 5, is 10. So, your chance of drawing a red marble is 3 out of 10.
Let's do one more question. It's going to look a little bit different. Let's calculate the probability that you don't draw a blue marble. This notation here, this apostrophe after blue, means the complement of the probability of drawing a blue marble. That means the probability of not drawing a blue marble. So, I'll make a little note of that, and how you calculate the probability of an event not happening is by doing 1 minus the probability that it does happen. So, for this question, that would be 1 minus the probability of blue. Well, the probability of drawing a blue marble, two of the 10 marbles are blue, so that'd be 1 minus 2 out of 10. And 2 out of 10 can reduce to 1 out of 5. And 1 minus 1 over 5 is 4 over 5, which we could instead write as 0.8 or 80%.
Must know number three: Probability using sets. Let's say a survey was conducted where three questions were asked: Do you like hockey? Do you like soccer? And do you like basketball? The results of that survey could be shown in a Venn diagram. This Venn diagram helps us understand the relationship between the different answers to this survey. Inside the rectangle is the entire sample space, so everyone who answered the survey is represented inside that rectangle. We see all of the numbers in this circle represent all of the people that said they like hockey, and this one, everyone that said they like basketball. This five in the very middle would be the people that said they like all three sports, and this four on the outside would represent the four people that said they don't like any of those three sports.
What I want to do with this Venn diagram is help you understand two important concepts: the intersection of sets and the union of sets. Let's start with the intersection of sets. So, let's say we had two sets, set A and set B. The intersection of those two sets is the set of elements that are in common to both A and B. So, in the Venn diagram, that would be the overlap of the two circles. And the notation that represents the intersection of the two sets would look like this: A ∩ B. We read that as "and," and it means the intersection of the sets.
So, let's do a probability example where we look at the intersection of two sets. The example says to calculate the probability that a respondent likes hockey and basketball. So, that word "and," remember, means the intersection of the two sets. So, we are looking for the probability that a respondent likes hockey and basketball, where this upside-down U means "and," it's the intersection of the sets. And then, when calculating the probability, you do the number of elements that are in the event space and then you have to divide that by the total number of elements in the sample space, so the total number of people that took the survey.
Let's start by finding the intersection of hockey and basketball. So, the number of people that like both hockey and basketball would be the overlap of those two circles, which is right here. I see that there are 5 + 3, so there are 8 people that like both hockey and basketball. So, my probability calculation, it would be 8 divided by how many ever total people took this survey. And to find the total number of people that took the survey, we would have to add every single number that we see in the Venn diagram. And you'll see that that equals 100. So, the probability of a person liking hockey and basketball is 8 out of 100, which could reduce to 2 out of 25, or we could write it as a decimal as 0.08.
Now, let's do an example for the union of sets. Let's start off with what is the union of two sets. I'll draw two sets once again. I'll draw set A and set B. And then, if I wanted to represent the union of those two sets, I would need to show the elements that are inside of A or B or both. So, that would be anything anywhere inside of this shaded region. And the notation for the union of two sets looks like this: A ∪ B. And for the union of sets, we read this as "A or B." So, the elements are anywhere inside of A or B. And there's actually a formula for counting the number of elements that are inside of A or B. What we do is we take the number of elements that are inside of A, and to that, we add the number of elements that are inside of B. So, we take everything that's inside of A, and to that, we add everything that's inside of B. But notice, if we do that, we would have double-counted all of the elements that are in the intersection of those two sets. So, for that reason, we need to subtract the number of elements that we double-counted, which are the number of elements that are inside of A and B, the intersection of the two.
Let's now see if we can do an example involving the union of two sets. The example says to calculate the probability that a respondent likes basketball or soccer. That word "or" tells me that we're looking at the union of two sets. So, I'll go ahead and write the probability formula. The probability of a respondent liking basketball or soccer would equal the number of elements that are in the event space and divide that by the total number of elements inside the sample space. Now, there are a couple different ways we could figure out the number of people that like basketball or soccer. We could use this formula. If I did that, it would equal the number of people that like basketball plus the number of people that like soccer. But when adding those two together, I would have double-counted the people that like both basketball and soccer. So, I have to then subtract the number of people that I double-counted, the number of people that like basketball and soccer, divided by the number of elements in the sample space.
And if I do this formula, let's look back up at our Venn diagram. The number of people that like basketball, inside of the basketball circle, I see 48 elements, plus the number of people that like soccer, inside of the soccer circle, I see 42 elements. And then I have to subtract the number of people that like basketball and soccer. So, that would be the intersection of the basketball and soccer circles. I see 15 people inside that intersection. Those are the people that I counted as both the basketball circle and the soccer circles, so they're double-counted. So, I have to subtract that 15. And then the size of the sample space was still 100. If I simplify this, I get 75 over 100, which reduces to 3/4, or as a decimal, 0.75.
Now, it might have been easier to get this 75 people that like basketball or soccer just by using the Venn diagram, looking at those two sets, the basketball and the soccer set, and then just adding together one time all of the numbers that fall anywhere inside of those two circles, right? Because the union of two sets is anything anywhere in those two. And if I added all those numbers together, I would get 75.
Must know number four: Conditional probability. Conditional probability is the probability of an event occurring given that another event has already occurred. Now, let me show you the formula for how we calculate a conditional probability. Let's say we wanted the probability of event B happening given that event A has already occurred. In this notation, this vertical line, we read it as "given." To calculate this conditional probability, before I give you the formula, let me give you a visual representation of what this conditional probability would look like. So, I'll draw two sets, set A and set B, inside of a sample space. So, let me zoom in on this. We've got set A, we've got set B, and then the overlap of those two sets is the intersection of the two of them, which we can write like this: A ∩ B. The set of elements that are in A and B.
Now, if we were just finding the probability of event B happening, I would simply do the number of elements that are inside of event B divided by the total number of elements inside of the sample space. But we don't just want the probability of B, we want the probability of B given that A has happened. So, I'm going to have to change both of these. My sample space is no longer everything inside of this rectangle, because we're given that A has occurred. I can shrink my sample space to only being the elements inside of event A. So, I'll change my denominator to the number of elements inside of A. And I'll shade in that new sample space. So, if this is our new sample space, what are the number of elements inside of set B? Well, that would be just these elements here that are in the overlap of A and B. So, I'll change my numerator to the number of elements that are in both A and B. So, this is my conditional probability formula. Let me zoom out and let's try an example to see how it works.
In this example, we have a table that shows the results to the question, "Do you like school?" and then it asks us to determine the probability that a respondent likes school given they are female. This word "given" tells me this is a conditional probability. I'm interested in finding the probability that they like school given they're female. Since we are given the respondent is a female, I can shrink the sample space to just be the respondents that are female. So, in my probability calculation, I'll be dividing by just the number of females. And since we're given that they're a female, the only way that they could like school is if they like school and are a female. So, my numerator is the number of people in the survey that like school and are a female. So, basically, all the question is saying, of the 14 females, what's the probability that they like school? Well, 10 of the 14 females like school, right? The number of students that like school and are female are these 10 people right here. And that's divided by the total number of females, which is 14. And that could be reduced to 5 over 7.
Must know number five: The multiplication law. If we're trying to find the probability of multiple events happening in sequence, the formula we use will depend on if the two events are independent or if the events are dependent. Now, what would make two events independent would be that if the occurrence of one event does not affect the occurrence of the other. So, let's say we had two events, event A and event B, and I wanted to find the probability that event A happens and then event B happens. If the events are independent of each other, meaning the occurrence of event A does not affect the occurrence of event B, then I can calculate the probability of the two of them happening in sequence just by doing the probability of event A happening and then multiplying that by the probability that event B happens. And sometimes, instead of using the intersection of set symbol, which means "and," sometimes you'll see it written as a comma between the two events: P(A, B), which means the same thing, just multiply the probabilities of each event in the sequence.
Now, let me just shrink this a bit and let's make some room for an example of how to use this multiplication law for independent events. The example says, suppose you roll a die and flip a coin. What is the probability of rolling a three and flipping tails? Because I want two different events to happen, I want to roll a three and flip tails, I know I need to use the multiplication law to find the probability of event A and event B happening. So, if I write a formula for this, to find the probability of rolling a three and then flipping tails, I would need to do the probability of rolling a three multiplied by the probability of flipping tails. The probability of rolling a three, well, one of the six sides of a die are a three, so you have a 1 out of 6 chance of that happening, multiplied by the probability of flipping tails. One of the two sides of a coin are a tail, you have a 1 out of 2 chance of that happening. So, the probability of both of those things happening is the product of those two: 1/6 * 1/2. That's just 1/12.
In this example, we just did rolling a three on the die did not affect the probability of flipping tails. But let's look at what would happen if the two events in our sequence are dependent on each other. Dependent events are two events where the occurrence of one of them affects the occurrence of the other. So, the formula for finding the probability of event A happening and event B happening, if the two events are dependent on each other, we would find the probability of event A, and then multiply that by not just the probability of event B, but the probability of event B given that A has occurred, because that affects the probability of B.
So, let's do an example where we need this multiplication formula that involves a conditional probability. The question reads, what is the probability of drawing two kings in a row without replacement? So, we're trying to figure out the probability of the first card you draw being a king and the second card you draw is also a king. In probability, when you see a symbol for "and," you want to think multiplication. So, we'll have to multiply the probability of the first card being a king by the probability of the second card being a king given that the first card was a king. The reason why this second probability has to be a conditional probability is because it says when you draw the cards, you do it without replacement. So, after you take the first card out of the deck, that affects the probability of the second card being a king. So, the events are dependent on each other.
And now, let's figure out those probabilities. What's the probability from a standard deck of 52 cards that the first card you draw is a king? Well, there are four kings in a deck of cards out of the 52 cards, so the probability of that happening is 4 out of 52. And that gets multiplied by the probability that the second card you draw is a king given the first card you drew was a king. So, we've already removed a king from the deck of cards, meaning there are only three kings left in the deck. And the deck doesn't have all 52 cards anymore, you've already taken out one of the kings, so there are only 51 cards left to draw from. And if we do this multiplication and simplify, we would get 1 out of 221 as the probability of drawing two kings in a row without replacement.
Must know number six: Permutations. Sometimes, when counting the number of elements in a sample space or event space for a probability, we need to consider the number of permutations of objects there are. Where permutations are just ordered arrangements of objects. The number of permutations of n distinct objects is just equal to n factorial. And in case you don't know what a factorial is, n factorial would just mean the sum of all the positive integers up to the value of n. So, for example, 5 factorial would be 5 * 4 * 3 * 2 * 1.
Let's do an example now where we calculate how many different orders can the letters A, A, B, and C be arranged in. Since we're doing the number of ordered arrangements of three objects, that just means how many permutations of three objects can we make. Since the number of permutations of n objects equals n factorial, I could say the number of permutations, or ordered arrangements, of three objects would just be equal to 3 factorial, which means 3 * 2 * 1, which equals 6. So, there are six different orders I can rearrange those three letters into. And let's take a look at why that makes sense. Let me write those letters A, B, and C, and let's say we're placing them into these three spots. We need to start by choosing a letter to go into the first spot. Let's say we choose B. Well, how many options do we have for what could go there? There were three options. Now, we need to choose something for what goes in the second spot. Since we've used a letter in the first spot, there's only two options remaining. Let's say we choose A for that spot. And now there's only one spot left and only one letter we can choose for it. So, we only have one option for what goes in the last spot. And hopefully, you can see the relationship between what I wrote here, 3, 2, and 1, and how we count the number of ordered arrangements of three objects. We just multiply the number of options we had at each step in the sequence when choosing the order of the letters. So, BAC was one of the six possible permutations. I'll just write the other five quickly, just so you can see them.
Let me make a little bit of room. And let's look at what happens if we're just doing ordered arrangements of part of a set of objects. The number of ordered arrangements of n items taken r at a time could be calculated using this formula. You say the permutations of n items taken r at a time would equal n factorial divided by (n - r) factorial. And let's see how that works with an example. Let's say there are 10 people in a race. How many different ways could they finish first, second, and third? So, basically, we only want to find how many ordered arrangements of three I can make from 10 people. To calculate that, I do the number of permutations of 10 taken three at a time would equal the formula tells me to do 10 factorial, but then divide that by 10 - 3 factorial. And I'll simplify 10 - 3 to 7. And then, if I were to start expanding the 10 factorial in the numerator, that'd be 10 * 9 * 8 * 7 and then all the way down until I get to 1. But I can stop the expansion of a factorial by putting a factorial symbol. That means that it continues all the way down to one. And the reason I'm stopping at 7 is because there is a 7 factorial in the denominator. So, what happens is those 7 factorials cancel, and what I'm left with is just 10 * 9 * 8. And why does that make sense? Well, from the 10 people, we're only filling three spots. We're only making an ordered arrangement of three from the 10 people. So, how many choices do we have for who could come first? We had 10 choices. How many options for then who could come second? Well, there's nine people left that could come second. And for third, there would be eight people left. So, we do 10 * 9 * 8, and we figure out that there are 720 different ways the people in the race can finish first, second, and third.
And let me do one last example where we see how this could come into play when calculating a probability. This example says that a lock opens if the right order of three numbers from 0 to 59 is input. Numbers can't be repeated. What's the probability of guessing the correct passcode on your first try? So, for this example, I want to calculate the probability of guessing the correct passcode. I need to do the number of correct passcodes. Well, there's only one actually correct passcode that will open the lock. And then I have to divide that by the total number of possible different passcodes. Well, there are 60 numbers between 0 and 59, including 0 and 59. Because I'm doing an ordered arrangement of items that can't be repeated, I know this is a permutations problem. So, I can just use the permutation formula to figure out how many permutations of 60 objects there are if I take them only three at a time. And if I use the permutation formula for that, it would simplify to just 60 * 59 * 58. And that's because you have 60 choices for the first number, then 59 for the second, then 58 for the third. And that would mean our total probability is 1 over 25,320.
Must know number seven: Combinations. A combination is a selection of all or part of a set of objects where the order of the objects does not matter. And you can calculate the number of combinations of n items taken r at a time using this formula. We see the number of combinations of n items taken r at a time, which we can write like that, or we can write it like this: C(n, r). And we usually pronounce this as "n choose r." And that's equal to n factorial divided by (n - r) factorial, and then there's another r factorial in the denominator. So, it looks really similar to the permutations formula, but there's this additional r factorial in the denominator.
So, let's see how this formula works and when to use it. In this example, it says, how many groups of three can be made from five people? Based on this question, it doesn't seem like the order of the people in the group would matter. So, we're just wondering how many combinations of three can we make from five people. Using the combinations formula, we would say that 5 choose 3 would equal 5 factorial divided by (5 - 3) factorial times 3 factorial. This 5 - 3 changes to 2. And then I'll simplify this by expanding the 5 factorial to 5 * 4 * 3 * 2. Instead of writing times 1, I'll just stop the expansion with a factorial symbol. And I'll do that because I see that it cancels with the 2 factorial that's in the denominator. And now, what I have in the numerator, 5 * 4 * 3, will give me the number of ordered arrangements of three I can make from a group of five. So, that would be 60. So, why isn't that our answer? Well, let me show you with letters. Let's say from the five people, uh, one of the groups of three we make is person A with person B with person C. We wouldn't care what order those people are in. It could be ABC, or ACB, or any of the six different permutations I could make with those three people. I wouldn't want to count those six different permutations as different groups. I would want to count them as just one combination of people, which is why we divide by the number of different permutations of those three people in the group that we can make. We divide by 3 factorial. That divides this answer by six to give us the total number of unique groups where the order does not matter. And that would tell us that there are only 10 different groups of three people that we can make from five people.
And let's do one more example where we have to calculate a probability that's going to involve combinations. This question says, from a group of seven kids and 11 adults, if you're making a team of six, what's the probability there is only one kid on the team? To calculate the probability that there's exactly one kid on the team, I would have to do the number of teams that have exactly one kid and divide that by the total number of teams. So, divide that by the number of elements in the sample space. Well, the total number of teams possible, there are 18 people to choose from, and I'm making a team of six where the order doesn't matter. So, I use combinations. I can just do 18 choose 6. And in the numerator, I need to figure out the number of groups that have one kid. Well, if I'm making a group of six that has one kid, that means of the seven kids, I need to choose only one of them. So, I can do 7 choose 1. And from the seven kids, if I'm only choosing one of them for my group of six, that means from the 11 adults, I have to fill the other five spots for the team. So, I would have to do 11 choose 5 to figure out the number of different groups of five adults I can make from the 11 adults. And then, if I do that multiplication, I get 3,234. That's the total number of groups that have exactly one kid on the team. And the total number of different teams of six that could have been made is 18,564. If I evaluate that as a decimal, it's about 0.1742. So, there's about a 17% chance that there's exactly one kid on the team.
Must know number eight: Continuous probability distributions. Now, there are lots of different types of continuous probability distributions. A few types are the normal distribution, exponential, chi-squared, and uniform. To explain continuous probability distributions to you, I'm just going to focus on normal distributions. And before I go into a full explanation, let me give you a sketch of what the distribution of a set of normally distributed data would look like. This function that I drew right here, this function f(x), this is what we call a probability density function. And what it does is it models the probabilities of outcomes of some continuous random variable X.
Now, there's a couple things I should comment on in more detail. First of all, a continuous random variable is a variable that can take on, within a certain interval, it can take on an infinite number of uncountable values. So, for example, somebody's height. Let's say we look at the interval between 1.6 and 1.7 meters. There's an infinite number of heights somebody could be within that interval. They could be 1.6543 meters. There would be no way to count all the possible heights within that interval. So, that makes it a continuous random variable. So, because there's an infinite number of values a continuous random variable can take, the probability of any one specific value happening is going to be zero. But what we do is use this probability density function to calculate probabilities over an interval by using the area that's underneath the function.
Now, the PDF function of any normally distributed data is calculated solely based on its mean and its standard deviation. So, depending on the mean and standard deviation of the data, the position or shape of this function might change a little bit, but it's going to follow the same group of properties. And let me show you what those properties are. Let me make a little bit more room here. The mean of the PDF function is always going to be right in the middle, right in line with the highest peak of the function. And then 68% of the data, so 68% of the area under this curve, is going to fall within one standard deviation of the mean. And because this function is always symmetrical, that would mean that that 68% could be broken in half into these two sections. And then, if I move two standard deviations out from the mean, within two standard deviations of the mean lies 95% of the data. And then within three standard deviations of the mean, you would find 99.7% of the data, assuming it's normally distributed. So, some approximate values for these little areas here, uh, I know this isn't going to add up to 99.7 exactly because there's been a lot of rounding happening here, um, but it's about 2.25% in each of those sections. One other important piece of information that I haven't written down yet is that the total area underneath a PDF curve is equal to one. That would mean that the definite integral of the PDF function f(x) between negative infinity and infinity would be 1.
Now, there's a very special normal distribution. It's called the standard normal distribution. Let me make a bit of room and we'll write about that. A standard normal distribution has a mean that is equal to 0 and a standard deviation that equals 1. So, down here on my graph of my PDF function, 0 would be in the middle. And then each unit I increase or decrease by would be one. So, to the right would be 0, 1, 2, and 3. And to the left, -1, -2, -3. Because the mean is zero and the standard deviation is one, the actual value of the variable communicates how many standard deviations you are to the right or left of the mean. This value of 2 means you are two standard deviations to the right of the mean, and this value of -1 means you are one standard deviation to the left of the mean. So, these values have special names. These are called z-scores. And there is something called a z-score table where we could look up z-scores, and the table would tell us what percentage of the data is less than or equal to the z-score that we're looking up. So, it will give us the area under the curve to the left of the value we look up. And we'll use that in our next example.
In this example, it says, the lifespan of regular smokers follows a normal distribution with a mean of 68 and a standard deviation of 10. What percentage of smokers will live beyond 76? Let me start by drawing a rough sketch of a normal distribution. The question tells me that the mean is 68, so I know that goes in the middle. And then I'll label these three units to the right going up by the standard deviation each time, and then to the left, label these three spots by going down by the standard deviation of 10 each time. And what the question is asking, it's saying, what percentage of smokers will live beyond 76? To find what percentage of smokers live beyond 76, we would have to find the area under this curve that is to the right of 76. There are a couple different ways to do that. The first way I'll show you is using a z-score table. So, what we have to do is we have to figure out what is the z-score of 76. And what a z-score is, is just telling you how many standard deviations this value is away from the mean. So, it's telling you this point's relative position over here on the standard normal curve. And you can calculate a z-score, right, the number of standard deviations from the mean, just by doing the x value minus the mean and dividing by the standard deviation. So, the z-score for 76, I would do 76 minus the mean of 68 and then divide by the standard deviation of 10. And I figure out that the z-score is 0.8. And notice, 0.8 on this standard normal graph, 0.8 would be right about here. If I found the area to the right on the standard normal graph, it would match exactly the area to the right of 76 on this graph. Now, like I said, we want the area to the right of it. But what a z-score table does, a z-score table is only capable of telling you the area to the left of the z-score that you look up. So, what we'll first do is find the probability of the z-score being less than or equal to 0.8. If we look up 0.8 in the z-score table, it'll tell us that the area to the left is about 0.7881. So, we figured out that this area to the left is 78.81%. But we want the area over here on the right. Well, because the total area, remember, equals 1 under the curve, if I do 1 minus this area, it'll give me that area. So, the probability of a z-score being greater than 0.8 would just be 1 minus the probability that's less than or equal to 0.8. So, 1 minus that probability, which is 0.2119. We can now answer our final question, which says that the probability that a smoker lives beyond 76. So, this area is about 21.9%.
And now there are calculator functions that you can use to not have to use a z-score table. For example, on a graphing calculator, the TI-84, you could find the option for the normal CDF function, which is the cumulative density function. And then within that option for normal CDF, you input a lower boundary, which is 76, an upper boundary for the area we want to type infinity, but we can just type, um, a really big number, so like 1 times 10 to the power of 99, and then input the mean and standard deviation, and it'll give you the area within the interval that you asked for.
Must know number nine: Binomial probability distributions. A binomial probability distribution describes the probability of the number of possible successes in an experiment. Referred to be a binomial probability distribution, this experiment that I'm talking about has to follow a set of criteria. There has to be a fixed number of trials. For each of the trials, there's only two possible outcomes: either a success or a failure. The probability of success stays constant, which means each trial is independent. And there's a formula for calculating the probability of having K successes in N trials. The formula is: P(X = K) = C(n, k) * p^k * (1 - p)^(n - k). And I'll explain to you what each of those variables stands for. P is the probability of success, so 1 - P would then be the probability of failure. K is our number of successes, and N is the total number of trials. And then I suppose I could add, but I'm out of room, N - K would be the number of trials minus the number of successes, so that would just be the number of failures. And this function that we generated here, that can find us the probability of K successes in N trials, is called a probability mass function. It can find us the probability of any number of successes happening. And because our variable is number of successes, that's very countable, right? There's a discrete number of successes we could have. So, this is called a discrete probability distribution. This is different than the continuous probability distribution we did in the last section.
So, let's do an example of a binomial probability distribution. The example says, suppose you roll a six-sided die four times. Create a theoretical probability distribution for the number of threes rolled. Now, because this is a discrete probability distribution, we can typically communicate the probability of any number of threes being rolled in a table. And in the table, our variable is the number of threes that you roll. So, how many threes could you roll when you roll the die four times? Well, you could roll no threes, or 1, or 2, or 3, or 4. And now, what we need to do is find the probability associated with each of those values of X. And I should mention, in this experiment, rolling a three is what we consider a success. So, we could have no successes, or one, or so on.
Let's figure out first of all the probability if we are going to have, let's say, two successes. If we are going to have two successes, let's try and use this formula. In the formula, we do "n choose k" first, the number of trials choose the number of successes, four rolls of the die, and we want two of them to be successful. That gets multiplied by the probability of success to the exponent of the number of successes. Well, the probability of rolling a three is 1 out of 6, and we want that to happen two times. And that gets multiplied by the probability of failure to the exponent of the number of failures. Well, the probability of not rolling a three would be 5 out of 6. And if we're going to have two successes out of four rolls, that means we're going to have two failures as well. So, if we have a look at this formula, if we specifically look at this part, it does the probability of success times the probability of success times the probability of failure times the probability of failure. So, what that considers is the probability of having a success, then a success, then a failure, then a failure. But that's not the only way you could have two successes. Four choose two tells you the number of different ways we could choose two of the four rolls to be successful. So, finding the product of all three of those will give us the probability of having two successes out of four rolls. And if we calculate that, it's about 0.157.
Let's now find the probability of having three successes. That means we roll a three three times. Using the probability mass function for a binomial distribution, I would do 4 choose 3, right? I have four rolls, I want three successes. The probability of success, I want that to happen three times. And the probability of failure, well, I want to fail only one time. If I'm succeeding three times, calculating that would be 0.154. So, hopefully, you can see how this formula works. I'll just fill in these other three spots for you so we have the complete distribution.
So, there's the complete probability distribution. For any probability distribution, the total of all the probabilities should equal one. And also, if we wanted to calculate what's called the expected value of the number of threes we would roll, there are two ways you can do that for any probability distribution. You can just find the sum of each x value with its probability. So, we do 0 times its probability plus 1 times its probability and so on. But for a binomial probability distribution, all you have to do is the number of trials times the probability of success. So, there are four trials, and the probability of success on any one trial was 1 out of 6. So, that equals 2/3, or 0.67. So, what does this mean? 0.67. An expected value is just if we were to conduct this experiment over and over and over again, the average number of threes that we would roll, we would expect the average to be 0.67.
Must know number ten: A geometric probability distribution. A geometric probability distribution models the probability of the number of trials needed to achieve the first success in an experiment. Now, just like a binomial distribution, the experiment has some conditions. There can only be two possible outcomes for each trial: a success or a failure. There has to be a constant probability of success, meaning the trials are independent. Now, even though it shares all three of these conditions with a binomial experiment, our variable of interest is different. We're not finding the number of successes in a fixed number of trials. We're instead finding out how many trials we need to do until we get the first success. So, the formula that can calculate that, the probability mass function for a geometric distribution, is the probability of the number of trials until the first success, which we call the waiting time, is equal to k, is P(X = k) = (1 - p)^(k-1) * p. And let me tell you, uh, what these variables stand for and why this makes sense. Our variable of interest, X, is the number of trials till first success. P is the probability of success. So, that would mean that 1 - P would be the probability of failure. And K is equal to X, so it's the trial number of the first success. So, what's happening in this formula? The probability that K is our first success, we would do 1 - P, right? That's probably a failure. We'd have to fail one less than K times, and then after that, we'd have to have our success.
Let's now do an example of a geometric probability distribution question. If you continue rolling a pair of dice until you get doubles, part A says, what is the probability that it takes four tries? This is a geometric probability question because you keep rolling the dice until you get doubles, so it just keeps going until you get your first success, and the probability of getting doubles is constant on each trial. The probability of rolling doubles, well, if you roll two dice, there's six options on the first die, six options on the second die, so there are 36 total different outcomes that could happen. So, the probability of doubles, the denominator, the sample space is 36. And there are six different doubles that you can get: doubles of ones, twos, threes, fours, fives, or sixes. So, our probability of success in this experiment is 6 out of 36, or 1 out of 6. So, in part A, when I want to know what's the probability that the waiting time is equal to 4, based on the probability mass function for a geometric distribution, I would do the probability of failure, so 1 minus the probability of success, that would be 5 out of 6, to the exponent of k minus 1, to the exponent of 4 minus 1, which is 3, right? If my first success is going to happen on the fourth try, I'd have to fail three times, and then on the fourth try, I'd have to succeed. And the probability of success is 1 out of 6. And evaluating this, I get 0.0965.
Let's do a part B. If we want the probability it takes fewer than three tries to get doubles, so probability the waiting time is less than three, well, that would mean that the waiting time is either one or two. Either happens on your first try or your second try. So, if I add the probability that the waiting time is one with the probability that the waiting time is two, that will answer this question. And for the waiting time to be one, that would mean that we would fail zero times and succeed right away on the first try. And the probability the waiting time is two, we'd have to fail on the first try and then succeed on our second try. And if we evaluate this, it's about 0.3056.
And the last question we'll do is we will calculate the expected waiting time. The expected value of a geometric distribution, the expected waiting time before you get your first success, it's always 1 divided by the probability of success. So, you just have to do 1 divided by 1 over 6, which equals six. So, on average, if you continued this experiment multiple times, it would average out that it would take you six times to get doubles. Sometimes it'll take you less, sometimes more, but on average, six.
So, that's the end of the top 10 must-knows for probability. Make sure to stay tuned to the channel because I'm going to put out in my next video a sequence of 10 probability questions that get increasingly harder where you can test out all of this knowledge.