Transcription
Welcome to PSYC2317 Lecture 10, where we are going to finally introduce the concept of Bayesian hypothesis testing. I'm extremely excited to give you this video today, so let's start, as we always do, with an example.
Let's suppose that we are interested in assessing the effectiveness of some mathematics instruction program. After implementing the program, we measure mathematical ability using something called the Scale for Advancing Mathematical Ability, the SAMA. It's this fictional national assessment that has a known mean score of 50. Okay, so that's what we're going to be comparing our treatment group to.
Now, our sample—these are the people that had the training—had a mean of 54.4 on this test with a standard deviation of 10. Our question is: Did the training work? This is a very typical type of problem that we would attempt to solve this semester, and to do it, we're going to do hypothesis testing.
Of course, you've been watching all semester, and you know how hypothesis testing works. But at this point, it's really important for us to review what this basic statistical framework says. So let's review.
The basic statistical framework that we use is as follows: First, we define two competing models. We call these hypotheses in general; that's why we call it hypothesis testing. These models are statistical models that could possibly generate the observed data that we have.
One model that we do is a model where this training worked; we call that an alternative hypothesis. We define another model where it didn't work, that there was no effect of the training; we call that a null hypothesis. So we have these two models now that we pit against each other, and essentially, we assess the fit of these models against the observed data.
So how exactly does that work? How do we proceed with a hypothesis test? Well, the first thing we should do, of course, is define some hypotheses. Now, we've done that this semester; we've defined hypotheses about the specific means that we expect to see. But this time, I'm going to do something slightly different. Instead of defining a hypothesis about the population mean, I'm going to define a hypothesis—or two, actually—about the effect size.
This symbol that you see, it's kind of a little squiggly line; that is a Greek letter, d—it's a delta. Okay, so we're going to be using that Greek letter a lot today. The two hypotheses are this: The null hypothesis is that the effect size is zero, so we would write H0 as delta equals zero. The alternative hypothesis is going to be that the treatment worked; that is, that it had a positive effect on these math scores, and mathematically, we would write that as delta is greater than zero.
So those are our two hypotheses. The next step, of course, is to collect some data, which, in the setup to the problem, we've done that. Okay, so what's next? Well, what's next—at least earlier this semester—is we would have said, "Okay, now that we've collected the data, let's compute the p-value." That is, let's compute the probability of observing our data if the null hypothesis is true.
Once we've done that, we interpret the p-value as follows: We say if the p-value is small—typically, we say less than .05—then that data is rare under the null. It's very unlikely to have occurred if the null were true. Our action then is we reject the null in favor of the alternative. That's how we get support for the alternative hypothesis; that is, support for the notion of a positive effect.
Okay, so that's one way we could do it. You've had lots of practice with this, but there's another way that we can do it, and that's what I want to introduce you to today. Instead of computing the p-value, today we're going to learn to compute something called a Bayes factor.
A Bayes factor looks like this; it looks like a fraction—almost—of two p-values. If you look at the mathematical form of this p-value here, that's kind of what we've written here. Now, if you are watching this video and you know something about Bayes factors, you'll know that what I said was absolutely false. But intuitively, I think the beginner can begin to think about a Bayes factor in this manner.
What the Bayes factor does is take the probability of the data under one hypothesis—say the null—and compare it to the probability of the data under the other hypothesis. So you get the symbol BF01, which tells you the likelihood of your data under the null compared to the likelihood of your data under the alternative.
So it's a fraction that tells you the relative support of your two hypotheses. It's interpreted as follows: If this BF01 number is bigger than one, well then remember this is a fraction, so that tells you that this number is bigger than this number. That tells you the data are more likely under the null than they are under the alternative.
On the other hand, if this fraction is less than one, that tells you that the denominator is bigger, and so the data are more likely under the alternative. So this is a fundamentally different way of assessing the fit of our models. We're still looking at the likelihood of our data, but instead of simply looking at the likelihood of the data under the null and not even considering the alternative—like the p-value does—we can then instead look at the relative likelihood of our data under both models.
So this is what's called a Bayes factor, and this is what we're going to use to do Bayesian hypothesis testing. Now, before proceeding with exactly how that works, it's worth pointing out a couple of differences between doing inference with p-values versus doing inferences with Bayes factors.
So let's just recall their definitions. The p-value is the likelihood of your data or more extreme under the null. As I mentioned earlier, it only considers the fit of the null as a potential model. There's nothing in this formula about the fit of the data under H1, so it completely ignores that.
What that leaves us with is this slightly, I think, less than stellar position where our support for the alternative is only indirect. It's only support for the alternative because the null doesn't fit. Ah, maybe the alternative doesn't fit either, right? But we never assess that with a p-value.
On the other hand, a Bayes factor, which looks like this, is the relative fit. Right? It considers the relative adequacy of both models as predictors of our data, and as such, it can directly index support for either model. It can index support for H0 or it can index support for H1. That's something we absolutely do not get with p-values.
With p-values, you might remember we can reject the null, and that sort of leaves us with support for the alternative. But if we fail to reject the null, we're sort of left with nothing, right? The Bayes factor does a little bit better and lets us directly support either hypothesis. That's really a nice thing.
So let's talk about a couple of examples of how you would interpret a Bayes factor, and then we'll actually compute it for the data that we had. So let's suppose that you got BF10 equals 9. Now, you'll notice I read that as BF10, not BF10. That's because these are actually two different numbers; these are the two hypotheses, H1 and H0, in order.
So this says the BF, or the Bayes factor, for H1 over H0 is nine. What that means literally is that the observed data are nine times more likely under H1 than they are under H0. Now, it could be instead that you get something like this—looks almost the same—but instead, you'll notice the subscript is 0 1. So that's a Bayes factor for the null over the alternative.
If that was equal to 9, it would say the same thing, except that the observed data are nine times more likely under the null hypothesis than they are under the alternative. So whichever one of these you get gives you a measure of relative support for either hypothesis over the other. It's really kind of a nice thing.
So how do we interpret Bayes factors? Well, I've already given you one interpretation with those two sentences, and that's an interpretation that I call the relative predictive adequacy of the models. So when you say something like the observed data are nine times more likely under H0 than H1, you're saying that H0 does a better job predicting the data. In fact, it does a nine times better job predicting the data than H1 does.
So we call this the relative predictive adequacy interpretation, and many times this is the one you'll see written up in papers that use Bayesian hypothesis testing. The other—it's mathematically equivalent but conceptually a little bit different—is we consider the Bayes factor as an updating factor.
So how does that work? Well, I'll show you exactly what that means, but we would write something like this: If we got BF01 equals 9, again we would then say that after observing the data, my prior odds for H0 and H1 have been increased by a factor of nine.
Okay, so what do I mean by that? I think the best way to explain what I mean here is to draw a picture. So let's look at an example. Before seeing any data, you might consider H0 and H1 to be equally likely, right, as predictors of your data. You don't have any strong belief—or let's suppose you don't have any strong belief—in which one is the correct model, which one best explains your data.
So you might set what we call the prior odds for these two models as one to one. That is, one-half probability for H0 being the true model and a one-half probability for H1 being the true model. Okay, so this is what are these are—this is my prior belief in the two models; this is before seeing data.
Now we observe the data, and then again suppose we got a BF01 of 9. So what does that do? Well, this notion of updating my prior odds says I'm going to change my odds from 1 to 1 to 9 to 1. I'm just going to simply multiply by nine, and that would look like this: That would mean that now nine-tenths of that circle is shaded in favor of H0, whereas only one-tenth of the circle remains for H1.
We call these posterior odds. This word "posterior" means after seeing the data. So a lot of Bayesian inference is all about the change from prior to posterior. The Bayes factor—this nine that we're seeing—is the multiplying factor or the updating factor from these prior odds to these posterior odds. Okay, so that's what we mean by that.
So let's do our working example. Let's go back to it and see how we use this notion of Bayesian hypothesis testing to solve this problem. So remember our problem was we've got this test, and we want to assess its efficacy.
So we take a sample, and I don't think I mentioned this earlier, but let's say we had 65 participants. From those 65 participants, we observed a sample mean of 54.4 with an estimated standard deviation of 10. So how do we assess our models using these data?
Well, the first thing that we do, just like when we were dealing with p-values, is the first thing we do is convert our observed data into a test statistic. This is where our t-score comes in. This t-score indexes where our data is on the distribution of all possible sample means under the null hypothesis.
If we do that—and remember the form of a t-score looks like this—we put in all these numbers: our sample mean of 54.4, the hypothesized sample mean—or sorry, the hypothesized population mean of 50—and then our estimated standard deviation and the sample size. This ends up giving us a t-score of 3.55.
So now what do we do with this? Well, earlier we would have said, "Okay, let's go and compute the probability of observing that score or bigger on a t-distribution." That would give us our p-value, and we would make a decision about the null.
Well, now we're going to do something slightly different. Instead of converting this t-score to a p-value, we're going to convert this t-score to a Bayes factor. Now, the way you can do this is you need some software. I'm going to give us two options, and I'm going to show you one of them today.
The first one I'll mention is actually the second one on my list, and that is if you're familiar with JASP, you can just use the summary stats module, which is a little add-on to JASP that you just click, and it will let you put in your t statistic and sample size. It'll pop out a Bayes factor to you.
But because many of my students don't have computers that they can install JASP on or might just be using their phones to sit in a lecture, I have made an interactive web app that you can open on your phone or an iPad or just a computer. It's at this address. You don't have to copy all that down; actually, in the description below, there's a link to this.
But it's just my typical Shiny apps address that we've used this semester, and the name of the calculator is called Bayes Factor Calc, with a capital F and a capital C. I'm going to demonstrate that for us now.
So I'm going to go over to Google Chrome, and I've already got it pulled up here. I just need to reload it and give us just a second. Here is the Bayes factor calculator, and so it gives us lots of things here. I think the best thing to do is let's put in our data, and that's going to be done over here on the left-hand side.
Then we'll look at what happens on the right-hand side in terms of things that we can interpret. So this is a single sample, so I'm going to click that as my design. You can see it can also handle independent samples designs as well. This example, though, is single sample.
Is there a predicted direction? Well, we did predict that this is a positive effect. So what I want you to notice is up here at the top for model definitions, I have written here the model definitions: H0 is that the effect size is zero; H1 is the effect size is not zero. That's if there's no direction.
If I click on positive effect, this updates to say that your effect size is greater than zero, and these indeed are the two models that we want to test. Now down here, we can just put in our test statistic. When we did the t-score a second ago, we got 3.55.
So when I type in 3.55, you'll notice that this picture changed. Okay, but it's not quite correct yet because I don't have the right sample size. The default here is 10, but I want to put in my sample size of 65. You'll see now that we've got a graph here.
It looks similar to that graph that I drew earlier in the lecture. This is a graph of my predictive adequacy. This tells me that my data are much more likely under the alternative than they are under the null. Exactly how much? Well, just scroll down here, and you'll see there's an output that tells you exactly how much.
The Bayes factor for the alternative is BF10 equals 69.37, and under that, I put an interpretation. This means that the observed data are approximately 69.37 times more likely under H1 than under H0. You can see this is exactly what that would look like.
If your data are almost 70 times more likely under H1, the vast majority of this pie here—or pizza, as my friend EJ Wagonmaker likes to call this pizza plot—here, the vast majority is going to reflect the alternative, with only a very small slice reflecting any sort of support for the null.
Below this, even—I haven't talked about this—but you get posterior probabilities, and we'll talk potentially more about posterior probabilities in another lecture. But these are, after seeing the data, what is the probability that H1 is the correct model versus H0? You can see that for H1, the probability is 0.9858. That's a 98.6% chance that H1 is the right model and a very small probability of H0.
These are not p-values, okay? I just want to be very, very clear. It's tempting to interpret p-values and posterior probabilities as the same, but they are not the same. But I don't want to get into all of the technicalities there; they are just different. Just separate them in your brain, okay?
So with just very little input on our part, we get a lot of information here. Now, the question is: How do we interpret all of this? Well, let's go back to our notes and solve our problem and write down a little results section in which we can interpret what we have done.
So let's go back and do that. So what are the elements to report? Well, first, let's define H0 and H1 and then ignore this part right here for now about specifying a prior. We'll talk about that in another lecture. For now, all I want to do—in fact, let me just erase this.
Okay, so we're going to define H0 and H1. We might say something like this: Under the null hypothesis, we expect an effect size of zero, and so we would define H0 as delta equals zero. Under the alternative, we would expect a positive effect, and that is we would write H1 as delta greater than zero.
Okay, so nothing new here. We now are ready to report and interpret our Bayes factor. Okay, so we got a Bayes factor of—let's just remind ourselves real quick—we got a Bayes factor of 69.37, or say 69.4.
So we would write something like this: We found a Bayes factor of BF10 equals 69.4, which means that the observed data are approximately 70 times more likely under H1 than H0. So that's just a direct interpretation of the Bayes factor using that predictive adequacy interpretation.
Next, we might want to calculate and report these posterior model probabilities. Okay, so that we got from the output below the Bayes factor—this was down here. I'm going to report the posterior model probability for this winning model, H1, and I might say, "Assuming prior odds of one to one for H1 and H0," okay, which is what we did here.
Our observed data updated these odds to 69.4 to 1 in favor of H1. So that is that updating interpretation, and then that is equivalent to a posterior model probability of this equals 0.986.
Again, just remember the 0.986 comes directly from this output on the Shiny app. Okay, and let's see, that's all that I might want to write there.
So real quick, I want to show you what would happen in a situation where we would have support for the null instead. So suppose that we did a t-score and we got a very, very small t-score, like maybe 0.1. Okay, so in this case, you can see that much of the plot is in favor of H0, and we got a Bayes factor in favor of the null of almost 7.
Then these things would all update appropriately. Okay, so it can happen here, and it can happen that you can get support for either one of the models. The nice thing about Bayes is you can interpret them directly. You don't have to say, "Well, you know, I failed to reject something." There's none of that going on here; you get a direct measure of support using the Bayes factor.
Okay, so let's finish up here. Just a couple of things that I want to mention: What about this notion of statistical significance? So all along the course so far, we've used that phrase quite a bit. Well, we usually don't use this phrase in a Bayesian context; instead, we speak of model evidence from the data.
In terms of what is sufficient model evidence, there are guidelines for this. Some of the most popular go back to Jeffries in the 60s, who proposed this scheme. So if you get a Bayes factor between 1 and 3, we tend to think of this as anecdotal evidence. 3 to 10 gets us up to moderate, 10 to 30 strong, 30 to 100 very strong, and beyond 100, maybe we would say extreme.
I do want to caution you; these are only guidelines. These are not cut-offs. If you get a Bayes factor of 9.8, that's not any more moderate than 10.2 is strong. They're still about 10 to 1 in favor of that specific hypothesis. But for a beginner to know, you know, what is sufficient evidence if you're not used to thinking about odds in favor of things, these are some pretty good guidelines to use.
So using these guidelines, our data might be written up this way: We might say, "In summary, these data provide very strong evidence." Well, that's because we had a Bayes factor in this range; it was almost 70, right? These data provide very strong evidence in favor of H1, and then, of course, interpreting it: That is, the training program had a positive effect on math performance.
So that is a quick, quick, quick introduction to using Bayesian inference to answer questions that we would traditionally use p-values to do hypothesis testing with. One of the big things I hope you get from this is what a Bayes factor is.
So go back and look at those definitions of the Bayes factor as predictive adequacy and as an updating factor. And also, importantly, we've got a nice little web app for you to use to calculate Bayes factors from your t statistics, and it works for a variety of designs.
So that's all for now, and I hope that you find this to be a helpful app and a helpful thing to do with your own data. So until next time, that's it for now.