📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Making Predictions

Ally Daubs Math10:31

Transcription

All right, so we've done all this work to see if there's a linear relationship between these two variables. The whole point of this is to create a linear equation to model that relationship. So, the equation that we create is called the regression line. This is the linear equation that summarizes the relationship of the data set, and we want to do this so that we can make predictions for values that are not included in that original data set. Remember, the data set is going to most likely come from a sample, and it's not going to include all possible values. So, we want to be able to come up with an equation that best fits the data and can be used to make predictions for values not included in that original data set.

So, the equation for the regression line follows this general formula: y-hat equals ax + b. And it hopefully looks familiar. It's not that different from y = mx + b because it is still a linear equation. However, in this case, we are using y-hat because we are making predictions, and then because we are in the world of regression and statistics, we use 'a' to represent the slope and we use 'b' to represent the y-intercept. You might also see equations that look like this: y-hat equals this funky looking B with a zero plus another funky looking B with a one, and then that is multiplied by the X. In that case, this B with the zero is the y-intercept, and this B with the one is the slope. That's an alternative way that you could see this.

We can flip-flop where the ax and the B are at any point, same with the b0 and the B1. The important thing to remember is that the slope is always the number being multiplied by the X, and the y-intercept is always the number by itself. So, we do have formulas for how we calculate the slope and the y-intercept. And just like with the correlation coefficient, I have these formulas here, but we're going to use technology to actually find the slope and the y-intercept for everything that we do. But the formulas are there. They involve the correlation coefficient, the standard deviations, and the means.

But once we have this formula, we've taken our data set, we've determined we have a linear relationship, we've calculated the slope and the y-intercept, and we have an equation. We can use that equation to make predictions, and that's really what we want to do. So, that's what I have an example of down here. So, in this example, we created an equation to model the relationship between the number of hours a student studied and the score they received on the exam that they were studying for. And that equation is y-hat equals 5.43x + 76.4. So, 5.43 is the slope, X represents the hours that they studied, 76.4 is the y-intercept, and then y-hat is the prediction for the score that a student would get.

So, we're going to use this formula to predict the score that a student would receive if they studied for 2.5 hours. So, I'm not going to come up with a number out of thin air. I'm going to use the equation. So, I'm going to say, "Okay, y-hat, that means predict. My slope is 5.43. The number of hours that I'm looking at is 2.5, plus 76.4." Order of operations says we do the multiplication first, so that gives me 13.575 plus 76.4. So, that is a predicted score of 89.975 for this student.

So, now predict the score a student would receive if they studied for 1.75 hours. So, we're doing the same thing, but now instead of 2.5, I'm plugging in 1.75. So, I'm still predicting using the equation. Still going to do my multiplication first, and that multiplication is 9.5025. And this time, my y-hat is 76.4 added on. This prediction is an 85.9025. So, based on whatever the original data set was that gave us this equation, if a student studied for 1.75 hours, which would be an hour and 45 minutes, this equation predicts that they would score an 85.9 on this exam.

So, a quick word of caution with making predictions. Extrapolation is something we want to be careful about. Extrapolation occurs when we start making predictions for X values that are not in the original range of explanatory values, so they're outside of the original range. So, if we only had students that studied between zero and four hours, and I said, "Well, what if I studied for six hours? What would I've gotten then?" That would be extrapolation. I don't have data for anyone that studied six hours. I only have data that goes up to four hours, and six is quite a bit different from four when we're talking about how long you studied for.

So, we want to be careful about this because it can lead to inaccurate results because we don't know if the pattern that we observed would continue outside of that original data set. If a student who studied for four hours got a 99, let's say, how much more would I get for studying six hours? Or would I eventually hit a point where I've overstudied and now my score starts to go down if I studied for six hours? We don't know that. So, we don't want to start making predictions for data values that are too far outside the original data set. You want to stay in, or really close to, the original range that we have. And that is how we can use these couple formulas. They're not too bad. We didn't actually use them, but you could. We have our regression line formula and then how we can use that formula to make predictions within our data set.