📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Philip Tetlock on Superforecasting 12/19/2015

EconTalk1:00:00

Transcription

Welcome to Econ Talk, part of the Library of Economics and Liberty. I'm your host, Robert. From Stanford University's Hoover Institution, our website is econ talk dot org, where you can subscribe, comment on this podcast, and find links and other information related to today's conversation. You'll also find our archives, where you can listen to every episode we've ever done, going back to 2006. Our email address is feedback at econ talk dot org. We'd love to hear from you.

[Music]

Today is November 30th, 2015, and my guest is Philip Tetlock, the Annenberg University Professor affiliated with the Wharton School and the School of Arts and Sciences at the University of Pennsylvania. He is the author, along with Dan Gardner, of Superforecasting: The Art and Science of Prediction, which is the subject of today's episode. Philip, welcome to Econ Talk.

Well, thank you. So, you start with a lot of criticisms, or throughout the book, I'd say, you have a lot of criticisms of pundits. Some of those have PhDs, and some of them are journalists, and some are just so-called experts who make predictions. But it turns out a lot of those, you can't really hold their feet to the fire when it comes time to judge where their predictions are accurate or not. Are they good forecasters or not? And why is that? What's the challenge with our sort of day-to-day world, where people claim that something is going to happen and print it in the newspaper?

Well, the pundits, of whom you think, let me say, we're critical, probably thinking of people like Tom Friedman or Niall Ferguson, and people on the left or people on the right. We, we identify all sorts, but they're all pretty uniform. They're pretty uniformly very smart people. They're, they're very articulate, they're very knowledgeable. They offer, make many observations about world politics and economics that seem very insightful. It is extremely difficult, however, to gauge the degree to which their assessments of possible futures, of the consequences of going down one policy path or another, are correct or incorrect. Because they rely almost exclusively on what we call vague verbiage forecasting. They don't say that there's a 20% likelihood of something happening or an 80% likelihood of something happening. They say things like, "Well, it's a distinct possibility that there'll be a global deflation in 2016." When you ask people what "distinct possibility" could mean, it could mean anything from about 20% to 80% probability, depending on the mood they're in when they're listening.

And I didn't mean to suggest you're critical of them, although you sometimes are, but you're critical of our culture that takes these vague pronouncements and then there's a "gotcha" game that gets played by people on the other side. But of course, there's always a way to weasel out of it because there's usually some hedging in that, in that verbiage, correct?

Well, that's right. If you, if you exist in a blame game culture in which people are going to pounce on you whenever you make an explicit probability judgment that appears to be on the wrong side of maybe, it's pretty rational to retreat into vague verbiage. So, we talk in the book about a brilliant journalist, the New York Times journalist David Leonhardt, who created The Upshot, a quantitative column in the New York Times. And he wrote a piece back in, I guess it was 2011 or 2012, and the Supreme Court narrowly upheld Obamacare by a 5-4 margin. And the prediction markets had been putting a 75% probability on the law being overturned. And David Leonhardt, who doesn't have any grudge against prediction markets, as far as I know, concluded that the prediction markets got it wrong. Now, that's, that's a harsh judgment on the prediction markets because they make hundreds of predictions on hundreds of different issues over years, and they're not bad. When they say there's a 75% likelihood of something happening, it's pretty close to a 75% likelihood, which means that 25% of the time, it doesn't happen. So, if you're going to throw out a very well-calibrated forecasting system every time it's on the wrong side of maybe, you're not gonna have any well-calibrated forecasting systems at your disposal.

I would say that's a second problem, really, which is that even when you do quantify your prediction, by definition, you're allowing the possibility that it doesn't happen. And then the question is, how do you assess the accuracy or judgment of the person who makes a statement like that?

Yes, exactly. And that then requires some, some understanding of probability and some willingness, some patience, and some willingness to look at track records over time. So, let's begin with your particular track record. You've done a lot of research in this area, this question of whether prediction is possible, how accurate is it, or experts good at forecasting. Talk about your background. We're going to get to the tournament that's at the heart of your book, but I want to start with your research history and what you found in the past and how people reacted to it.

Well, I guess that's another I ask, is it just exactly how old must I be? Because I've been doing longitudinal forecasting tournaments a long time. So, let's just put on the table, I'm 61 years old, and I got started at this right after I got tenure at the University of California, Berkeley. And I was a little, little more than 30 years old, in 1984. And the Soviet Union still existed. Gorbachev had yet to become General Secretary of the Communist Party of the Soviet Union. And we did our initial pile of studies back in the mid-1980s when people were hawks and doves were arguing about the best ways of dealing with the Soviet Union. And now we're doing forecasting tournaments as hawks and doves arguing about the best ways of dealing with the Iranian nuclear program, or for that matter, for dealing with Russia and the Ukraine. So, we've been running forecasting tournaments off and on for 30-plus years. The first big set of forecasting tournaments were done in the late '80s and the early '90s and were reported in a book, Expert Political Judgment, that came out in 2005. And the second wave of forecasting tournaments were much larger, involving many thousands of forecasters, a million-plus forecasts, and were sponsored by the US intelligence community. And they ran from 2011 to 2015. And in fact, they're still running. So, if your readers are interested in signing up for an ongoing forecasting tournament, they should consider visiting the website at gjopen.com.

Going back to the earlier work that you did before the fall of the Soviet Union, what were some of the main empirical takeaways from that work?

Well, one big takeaway was that liberals and conservatives had very different policy prescriptions, and they had very different conditional forecasts about what would happen if you went down one policy path or another. And that nobody really came close to predicting the Gorbachev phenomenon. Nobody, for that matter, came really close to predicting the disintegration of the Soviet Union later on. But everyone, after the fact, seemed to have an explanation that either appropriated credit or deflected blame, and it was consistent with their worldview, I'm sure, and meshed perfectly with their prior worldview. So, so it was as though we're in an outcome-irrelevant learning situation. It didn't really matter what happened. People would be in an excellent position to interpret what happened as consistent with their prior views. And this ends the idea of forecasting tournaments, which was to make it easier for people to remember their past dates of ignorance.

Well, this isn't a side of sorts, but it's just, it's just a wonderful insight into human nature. And it's a theme here at Econ Talk, which is, when you went back and asked people to give their, what they remember as their probability of, say, the Soviet Union falling, what did they say?

Well, they certainly thought they assigned a higher probability to the dissolution of the Soviet Union than they did. And there were a few people who assigned really very low probabilities, who are being on assigning higher than a 50% probability. So, people really pumped up those probabilities retrospectively. The psychologist calls that hindsight bias, or the "I knew it all along" effect. And we saw that in spades in the Soviet forecasting tournament.

Yeah, I think that's an incredibly important thing that we all tend to do. We tend to think we had much more vision than we actually had. And we usually don't write those things down. You haven't written them down. So, that was awkward that they actually had their original forecast. But most of us, the "I knew it all along" problem is a bigger problem for most of us because we don't write it down. What we truly remember it differently. Even if you think the person on the other side of the table knows what the correct answer is, you still tend to misremember it. Yeah.

So, this more recent tournament was rather remarkable. Give us the background of who competed and your role in it, and how it was set up, and what some of the questions, for example, were that people are competing on.

Sure. This was work I did jointly with my, my wife, research collaborators, Barb Mellors, and we were faculty then at the University of California, Berkeley. And we didn't leave for the University of Pennsylvania until about 2010. But we were visited by three people from the Office of the Director of National Intelligence when we're at Berkeley, I guess, late in 2009. And at least two of them were quite enthusiastic about the idea of the US intelligence community using some of the techniques that were employed in my earlier work, Expert Political Judgment, for keeping score on the accuracy of intelligence analysts' judgments. And that was the core idea behind the, what became known as the IARPA forecasting tournaments. IARPA is the research and development branch of the Office of the Director of National Intelligence, which is the umbrella organization for all intelligence agencies, CIA and DIA and so forth, all 16 of them. And the idea would be, would be they would have a competition, and major universities and consulting operations would would apply for large contracts to assemble teams whose purpose would be to assign the most realistic probability estimates to possible futures that the US intelligence community deemed to be of national security relevance. So, those can turned out to be questions on everything from Sino-Japanese clashes in the East China Sea to Greece leaving the eurozone and Spanish bond yield spreads, to Russian relations with the near abroad, so, Ukraine, Georgia, of course, conflicts in the Middle East, Ebola, H5N1 issues, just an enormous range of issues, 500-plus questions over about four years. And the goal would be of each of the research operations would be to come up with the best possible ways of designing probability estimates. Now, they, they spent everybody for their academic bona fides, so they want to make sure that everybody was legit, weren't using Ouija boards or anything like that. But the US intelligence community was simply interested in who could generate the most accurate probability estimates for these extremely diverse questions. And they didn't really care whether we took a more psychological approach, or a statistical approach, or a composite approach. What they, what they cared about was accuracy. And that was it: accuracy, accuracy, accuracy. So, our group, Barb and my wife and I, put together this group called the Good Judgment Group, which is an interdisciplinary consortium of wonderful scholars. And we went out, we tried to recruit good forecasters who, and we tried to give them the best possible training in principles of good probabilistic reasoning. And we assembled some of them into teams, and we gave them guidance on how teams can work effectively together. We put some of them into prediction markets, and we wanted to see how well prediction markets would work. We did, we experimented with a lot of different approaches. We also have really good statisticians who experimented with different ways of distilling wisdom from crowds. So, our approach is very experimental. I think some of the other approaches were experimental as well, but our experiments worked out better than their experiments. So, we won the tournament by pretty resounding margins in the first two years, sufficiently resounding that the US intelligence community decided to funnel the remaining money into one big group, which would be the Good Judgment Project, which could hire some of the best researchers from other teams who were competing against.

Well, we originally, we were competing against different, different competitive benchmarks here. Originally, we were competing against the other institutions that received contracts from the government, like, oh gosh, MIT and the University of Michigan and George Mason University, places like that. Then later, and we were competing against a prediction market that we ourselves were running, by, but the firm known as Inkling. And also against internal benchmarks, US intelligence analysts themselves generating probability estimates and competing against them. Although that was classified because, of course, the US intelligence analysts were classified. But David Ignatius at the Washington Post leaked some of that information. And I think at the end of the second year or third year.

But after two years, your team trounced everybody. And then what happened going forward after that?

Well, we were able to absorb resources from the other teams because the government was obviously saving a lot of money by by suspending the funding of the other teams. We were able to consolidate some resources, and we were able to compete all the more aggressively against the, the other remaining benchmarks. Which the key benchmarks for us to be were an external benchmark, the prediction market run by Inkling, and the more confidential one inside the US government.

Now, you mentioned, and this is just by actually, I'm going to read a quote from the book which I loved, which is relevant, which is from Galen. They, uh, early, early physician. And what time period of Galen? Roughly, I guess, second century after Christ? Years ago, okay. I thought it was later than that. So, so he wrote a long time ago. And you write the following, that he wasn't into experiments. And you wrote the following, that he, here's the quote: "All who drink of this treatment recover in a short time, except those whom it does not help, who all die. It is obvious, therefore, that it fails only in incurable cases." So, what could be better than that? I mean, that's phenomenal. And I was reminded, I think you work, that's where you apply the quote, of even the pundit who puts it in miracle value on a on a certain event happening. As a 63.7% chance that this will happen. Whether it happens or not, if it does happen, he says, "I told you, it was 63.7%." And if it doesn't happen, he can say, "Well, I said there was a 36.32% chance that it wouldn't happen. It was so, when it didn't happen, I'm still right." So, the question then becomes, when you say you trounced the other teams, there has to be a way to evaluate probabilities. And in the book, you present the Brier score, or try to give us the flavor of how you measured success in prediction.

Oh, that's an excellent point. It really is impossible to measure the accuracy of a probability judgment of an individual event unless the person, the forecaster, is rash enough to assign a probability of zero and it happens, or a probability of 1.0 and it doesn't happen. Otherwise, the forecaster can always argue that something improbable happened. So, assessing the accuracy of individual events is impossible, except in those limiting cases. But it is possible to assess the accuracy across many events and many time periods. So, good judgment in world politics means you're better than other people at assigning higher probabilities to things that happen than things that don't happen, across many events, many time periods. So, the example would be, let's talk about the second particular example. We're going to, we're going to try to forecast the probability of Greece leaving the eurozone. So, I say it's 0.51, and you say, "Seize maxillae fit," okay, 0.49. Because I think they're not likely. I'm going guys, because it's below 0.5. I said 49, and you say 0.1. And it doesn't happen. So, my argument is that you did a better job than I did.

You don't know that for sure. You don't respect to Grexit. That's correct. You do know it probabilistically, across the full range of questions posed in the ARPA tournament. Now, insofar as you've been predicting 0.49 consistently over several years, and I've been predicting 0.1, and it doesn't happen, you might be tempted to draw the conclusion, even with respect to Grexit, that I've been closer to the truth. You might be.

Is that one of the things I found troubling about the the setup and the way of assessing good judgment? And it, it, one of the things your book makes one ponder is just how hard it is to assess whether someone has good judgment. It's, it's, that's absolutely true. I couldn't agree more. It's a very difficult concept to operationalize. Yeah. So, this particular way, even though, so let's take this case. Let's say there's 10 things where I tended to predict 0.45, and you predicted 0.1, and none of them happens. So, that we were both "quote right," and that we both thought it was below 1/2, it was less likely. But you were "quote more right" than I was. Because, because what? And here's where my, here's what I want you to respond to. It seems to me you could argue you just had were calm than I did. You were more strategic in how you picked your number. You didn't have any more accurate knowledge of the actual probability.

Well, how many times did you have to flip that coin before you decided that the person who claims the coin is biased is closer to correct than the person who claims the coin is very close to equilibrium? Well, that's a challenging question. I thought while I was reading the book, I thought of Bill Miller of Legg Mason. So, Bill Miller beat the S&P 500, I think, for at least 15 years in a row, maybe more. And a lot of people concluded he had to be a genius because, well, he beat the S&P 500 one year, not so impressive. But 15 years, that's so unlikely. But of course, we know that that doesn't prove he's a genius. It doesn't even prove he's smart. It might merely mean he was lucky. Of the thousands and tens of thousands of managers of mutual funds, he was the one who happened to beat the S&P 500 15 years in a row. And we know that over enough time and enough managers, that's gonna happen. And so, we know nothing about his ability going forward. And in fact, he didn't do particularly well after his streak was broken. Did he get less smart? Did he get overconfident? We have no way of knowing. So, I find myself, even though I, I found many things in the book that are useful in thinking thoughtfully about looking into the future, the fundamental measurement technique strikes me as a challenge. What do you say to that?

I think that is a great question. Really, really deep and a good question. People in finance argue, of course, about whether there is such a thing as good judgment. If you're a very strong believer in the efficient market hypothesis, you're going to be very skeptical. If you toss enough coins enough times, a few of them are bound to wind up heads 76, 60, 70, 80 times in a row. You can just keep, keep doing that. And there are skeptics who argue that Bill Miller, or for that matter, Warren Buffett, or George Soros, were just one of those lucky sequences of coin flips, and then we anoint them geniuses. We are very sensitive to the possibility that superforecasters could be super lucky. And we're always open to the possibility that any given superforecaster has been super lucky. We're always looking for patterns of regression toward the mean. The more chance there is in a task, the greater the regression toward the mean effect. And that's just something we're continually looking for. Our best estimates are that the geopolitical forecasting tournament sponsored by ARPA had about a 70/30 skill/luck ratio, based on the regression toward the mean effects we're observing. Which means there's a big element of skill, and there's a significant element of luck. And, and based on other factors, like we introduced experimental manipulations that reliably improve accuracy. If it were pure noise, but if it would be possible to do that, so it wouldn't be possible to develop training modules or teaming mechanisms that improve accuracy if we were dealing with a radically noisy, dependent variable. It is possible to do that. So, various converging lines of evidence, with individual difference evidence among the forecasters, and experimental evidence, suggest that we're not dealing with a radically indeterminate phenomenon here. There, there is such a thing as good judgment, but there is certainly a significant element of luck as well.

One of the challenges when you read Warren Buffett or Charlie Munger's partners' analysis of the market, they're really smart, they're full of interesting insights, right? So, it reinforces your view that it may be it's not luck. Their challenge, of course, is that you don't know whether those particular insights really matter. And that we are in complete agreement on this subject. Let's take, let's take an example from the book, which I found really illuminating, which is an example of how there, there is a role for skill, at least in some forecasting problems and some estimation problems, which is, you give the example of, you're told that there's a family, their last name is Renzetti, they have an only child. What are the odds that they have a pet? And talk about how you might think about that more thoughtfully than just saying, "Well, I don't know," or worse. Well, if they have an only child, and that's important, and you, the inside-outside view, I found very illuminating.

Well, it's part of a more general discussion in the book about what distinguishes superforecasters from regular forecasters. And that is the tendency of the superforecasters to start with the outside view and gradually work in. So, you would start with, you're in it. You would start your initial estimate, whether it's, you know, trying to estimate the number of piano tutors in Chicago, or whether a particular family has a pet, or whether a particular African dictator is likely to survive in power another year. All those kinds of examples, you would start by saying, well, what's the base rate of survival? Or what's the base rate of pencil? What's the other example is this African dictator problem. We might ask you a question about whether dictator X in country Y is likely to survive in power for another year. And you might shrug and say, you know, I barely heard of the country, let alone the dictator. But you do know a couple of things. You know more than you think you know. And one of them is that once a day, if the dictator's been in power more than a year or two, the likelihood of the dictator being in power another year is very high. It's a 90, 90% plus. So, you could, even though you know nothing about the dictator or the country, you can say, well, I know that when someone has established a power base within the country, it's difficult to dislodge them. Now, if say, you would start, you would start your estimation process with a high probability because of that fact. It's just a simple, demonstrable statistical fact. And then you would say, well, now I better do a little bit of research and find out a little bit about this guy and this country. If you discovered that this particular person is 91 years old and has advanced prostate cancer, you might want to modify your probability. If you discovered there's fighting in the suburbs of the capital, you might want to modify your probability. So, these are the, it captures part of the distinctive working style of the superforecasters, is that they try to get as much initial statistical leverage on the problem as they can before they delve into the messy historical details.

And I, I think all of us like the idea of evidence-based medicine, evidence-based forecasting. And your book is certainly a tribute to the potential for data and statistics to help improve our ability to anticipate events that are important. I guess the challenge is, which evidence, and how we incorporate the other factors. I mean, one of the, you tell a lot of really interesting stories of the way the different forecasters, many of them were just quote amateurs, which is beautiful. They're not burdened by the PhD that I have, and that others have, who tend to try to predict things. So, you talk a lot about how they weigh evidence. Part I find intellectually challenging and accepting these results are two issues. One is, and I want to make it clear, these amateur teams that you put together, along with experts, and the aggregation of folks into teams with advice on how to, how to work together and how to avoid groupthink, which is a large part of it, are very, very interesting and very useful, I think, to anybody. In all of these examples, they dominate. It's not like they do three percentage points better than the others. I just want to make that clear, right? It's, they really did a lot better than just some of the more educated folk and the so-called experts, correct?

Well, when you throw everything together, the cumulative advantage does get to be quite staggering over the ordinary folks in the tournament. That's true. But you were talking about different components here. It certainly helps to have talent and to get the right people on the bus. So, individual differences among super-superforecasters are not just regular people. They are different in certain measurable ways. They score higher on measures of fluid intelligence. They're more politically knowledgeable. They're more open-minded. But most important, I don't think. I think they have all those advantages over regular folks, and I don't know, and those matter. But I don't think they have those advantages over professional intelligence analysts. I don't think they have greater fluid intelligence, greater, definitely, own expertise of college. And I don't really think they're more evenly more open-minded, although they are pretty open-minded. I think what really distinguishes the superforecasters from the seasoned professionals in the intelligence community, whom they were able to outperform, and that was really, I thought, the most difficult of all the benchmarks. I think what really distinguishes them is that they believe that subjective probability estimation is a skill that can be cultivated and is worth cultivating. I think many of the sophisticated analysts, like many of the sophisticated pundits, when they when they see a question like, "Well, how likely is Greece to leave the eurozone?" or "How likely is it that Putin will try to export Ukrainian territory?" though they'll shrug and they'll say, "But this is a unique historical event. There's no way we can assign a probability to this." You should have learned this is statistics 101. You know, you can, you can make problem, you can learn to make refined probability judgments in poker and things like that. You can learn to distinguish 60/40 bets from 40/60 bets in poker because in poker, you're having, you have repeated play and a well-defined sampling universe. And indeed, the frequentist statistics everybody learns in stat 101 apply. Those statistics just don't apply here. So, you're engaging in an exercise instead of precision. So, you've got people with really high IQ, that strikes people with really high IQ saying really smart things like this, and it blocks them from exploring the potential of learning to do it better, which is what, which I think is what the ARPA tournament proved as possible.

Now, I want to try to take that criticism, operate, phrase it a little differently. If, if you asked me, let's predict there's a football game. So, we're recording this on a Monday. There's a football game tonight. It's the Cleveland Browns against the Baltimore Ravens. If I remember correctly, it's not a very interesting game. And we won't try to figure out the probability that Baltimore is going to win. I think they're probably favored. Okay, so they're supposed to win. But we know that they might not. So, we'd like to know, though, what the probability is. Now, there are many ways to go about this question. The way the people in your book go about it is they take a base rate, or this is one of the ways they would take a base rate. Like we talked about a minute ago, how many dictators who've been in office X years, rate offices, and an additional year? Or in the case of the pet example, you didn't mention it, but in the book, you talk about what's the proportion of households that have pets. That would be a great starting place. You start with that, then you dig deeper and you try to find out more stuff. It's, first of all, it's really hard to know what the base rate is. Because it's, at the base rate of underdogs on a Monday night, is the base rate of teams that have lost two games in a row. So, what that people started to do, and they can do it in football, it's harder to do with Greece exiting the euro, is they try to accumulate statistical evidence that, you know, in a systematic way. They run multivariate regression, and they're pretty good at that. Because for the nature of football, we're pretty good at narrowing down. We can look at past performance, we can take account of injuries that can mess things up. We'll never know, as Hayek pointed out in a different context, we'll never know if the quarterback had an unsettling argument with his wife the night before, a bad meal at lunch, that's that's affecting his play. But in football, we can get pretty good at predicting probabilities. But did I have those tools? And in fact, when we look at say, Greece exiting, and worse than that, when we do have those tools, we often can't do it very well. So, we try to say, estimate in epidemiology, the effect of drinking a lot of coffee, whether you're more likely to get cancer. We can't measure that. So, how are these people somehow absorbing all this information, which you talked about, how they, they read a lot, and they, they, they talk and they share ideas, and they bounce ideas off each other when they were doing this? How are they able to somehow hone in with accuracy without, without using a formal statistical model that people we use for almost statistical models we can't do very well?

Yeah. Well, they were opportunistic. And sometimes they do find statistical models in unlikely places. And they know, one of the first places they would go for your football game is they look at, look at Las Vegas. And although what the odds are. They go, you know, they would do it. They would do what you might superficially consider to be cheating. They would say, well, the, there are some very efficient information aggregators already out there, like there's Nate Silver and 538, and there's this, and there's that. And I'm going to take a look at each of those, and I'm going to average those. My gonna take does my initial estimate. And then if I know something about the quarterback's relationship with the spouse, I might factor that in too. But I probably not going to get very much weight to it. And that, that turns out to be a pretty good strategy. You're raising a deep philosophical question, what the limits of precision. And, and I think it's just wonderful. This is one of the best interviews I've had. I think it is a deep question. And why don't we just say, I don't know the answer to what the liberal elements of precision are? You don't know what the answer is. Why don't we run studies like the ARPA tournament and find out where they are? And that's essentially what I did. It's a very pragmatic attitude and said, how we could, we could, we could hunker down in philosophical positions, and you could, I could say, on Bayesian, you're a frequentist, and I think we can make probability, make a difference. You think, no, this is, it's, there's just too much noise, and it's not working. Opportunities, we could argue about that until the cows come home. But what really the right thing to do here is to run forecasting tournaments and explore what the limits of decision are. And, and the little bit, there are real limits of precision. The ARPA tournament, I mean, at its very best, forecasters on average are not doing much better than assigning 75% probability to things that have happened and 25% probabilities to things that haven't. So, there's still a lot of residual uncertainty here. There are big pockets of irreducible uncertainty. And then the tournament, there's lots of room for error. They make lots of errors. We, what we simply showed is that it's possible to do something that people previously supposed was pretty impossible.

Let me ask a different question. I guess there's two issues related to that empirical finding. One would be, could you do it again? Right? It would be a question of replication. Could you replicate the success with the same team? Do you think they would continue to outperform the benchmarks? That'd be the first question. The second question is, what do you do with it? So, well, that chance of the first one first. Do you, is there any plans to try to replicate, replicate these results?

Well, we're doing that. ARPA is doing exactly the right thing. They're setting up a mechanism for exploring how replicable results are. So, we're going to be running more forecasting tournaments. One of the reasons I invited your readers to participate in GJ Open is they can explore their skills. And who knows, there might be some superforecasters listening right now. You're here, I guess. And then the next question would be, is this really valuable? So, you have to have contingencies. Often, you want to have contingency plans for the contingencies that could happen. So, you, you, you, you want to know whether it's really likely that Greece will leave the eurozone, or you want to know whether China's going to do such and such militarily. You want to know if there's going to be a coup in this or that country. But does it affect our actions to know that it's really 73% rather than 58%? So, what's the consequence of this improved? Do you believe that improving those probabilities are, are going to lead to better policy?

You know, again, I, I don't like, I reread the risk of flattering the interviewer. That's just a superb question. It depends on the domain. If we were talking about pricing futures options on oil, I think Wall Street professionals would say, "Yeah, I would really want to know the difference between a 60/40 probability and a 40/60 probability." Brown, the chief risk officer of AQR, said as much when we interviewed him for the book. And I know that's a common attitude among people in the hedge fund world. So, in that world, options pricing and finance, I don't think there's much question about it. Poker, I don't think there's much question about it. Now, if we had an opportunity to talk to a senior official in the US intelligence community about the project, say, two years ago, and we asked, "Well, if you had known that the probability of a Russian incursion into the Ukraine was not 1%, but was 20% during the Sochi Olympics, would you have done something different?" And we got an interesting reaction. It was, "But I've never, I've never even heard a question like that before."

[Applause]

And it is a fascinating problem, and it raises very deep questions. I think the short answer is, everybody would agree you're not worse off with better probability estimates than worse probability estimates in the long run. I don't think there's any. The question would be, are the increments, improvements in accuracy we're able to achieve, do they translate into enough better decisions in a given domain to justify the cost of achieving those improvements? I think that would be your question in the intelligence context. And that is, I think, something that the intelligence community is quite sensibly exploring right now. But I think the other question is, you might lead yourself down a path of thinking you've got more certainty than you actually do. So, there's a downside, a downside risk also of using a more organized method, even though we all want to, we all think that's got to be better. It doesn't have to be, unfortunately, right? But, but, but the opposite error is also possible. I mean, psychologists, you're right, tend to emphasize the dangers of overconfidence. But there's also the danger of underconfidence. Yeah. So, we talked about the situation in which President Obama was making the decision about whether to go after Osama bin Laden. And he was probably, he'd drawn underconfident conclusions, we think, from the probability judgments that were offered to him in that room. When you have people with different expertise and different points of view, all offering probabilities, almost all of them offering probabilities on the about 50%. What's the right way to process those probabilities? Should you simply take the median, or should you take something more extreme than the median? And that was one of the issues our statisticians wrestled with. I mean, it's treated as a thought experiment. In a thought experiment, you're the President of the United States. You're around, you have a table of the elite advisors, each of whom offers you a probability estimate that Osama bin Laden is residing in a mystery compound in Abbottabad, Pakistan. And each one says, "Mr. President, I think there's a 0.7 probability." The next one, "Oh, 0.7." "Point seven." All around the table, uniform 0.7. What conclusion should the President of the United States draw about whether Osama is there and whether to consider going to the next step of launching a Navy SEAL attack?

Well, the short answer is, it depends on whether the advisors are clones of each other or not. If they're clones of each other, the answer is 70%. They're all drawing on the same information, the processing at the same way, 70%. There's no incremental information divided by each 70%. But if they're drawing on different types of evidence and different processing in different ways, is one guy with satellites, but information, and others a codebreaker, and other human intelligence, and so forth. If they're drawing on different sorts of information, processing in different ways, each of them still arriving at 70%, but not knowing all the information the other people had when they reached their 70%. Now, what's the correct probability? And the answer as I've to the question I just posed is mathematically indeterminate, but it's statistically estimated. And we did statistically estimate it over and over again during the course of the ARPA tournament. It was one of the big drivers of our forecasting success. And typically, in the ARPA tournament, you would, you would extremize. You would move from 70% to 85 or 90%. You know, more you get there, you know more than you think you do.

Explained when you said you would move. I did not understand you saying that's what you would discover.

Well, in that, it's a question of who your advisors are. If your advisors are all drawing on the same information and reaching 70%, the answer is, when you average their judgment, you're just going to be 70%. Only of really of one estimate. You're fooling yourself if you think.

Give it ten. That's right. But if the advisors are all saying 70%, but they're drawing on diverse sources of information, it's a little bit counterintuitive, but the answer is going to be quite a bit more extreme than 70%. Now, how much more extreme is going to be a function of how diverse the types of information are, and how much expertise is in the room. And these are difficult to quantify things for sure. What our statisticians did is they used an extremizing algorithm. It simply, when the weighted average of the best forecasters was tilted in on one side of maybe or another, they extremized it. They moved from 30% down to 15, or from 70% up to 85.

You're saying they, when they had that average estimate of 70%, they actually, they were, they pretended it was higher. They gave it more confidence than than just the 10, because they drew on different information.

Well, our statisticians, when we submitted forecasts at 9:00 a.m. Eastern Time every day during the forecasting tournament. So, there's no wiggle room here. I mean, this is very carefully monitored research, right? This doesn't have the problems of some research where, you know, people can have wiggle room. There's no wiggle room here. It is being run like a bank, with the transactions every day. And we were betting on those aggregation algorithms. And that particular extremizing algorithm I'm describing in verbal terms right now was essentially the forecasting tournament winner. It was more accurate than 99% of the individual superforecasters from whom the algorithm itself was largely derived.

Yeah, I just want to emphasize again that going back to our earlier discussion, is that didn't really mean that the number was 0.8, 0.9. I'm not sure that's a meaningful number. Just that you were more confident that it was Osama bin Laden, say, than the 0.7, the point 7 number suggested. That would give some come to the president in launching a SEAL attack, knowing another way to say it would be that even though they all thought it was more likely than not, 0.7, if it came from different sources of information, you could be more confident it was more likely than not, it was even closer to certain.

Exactly right. Yeah. I want to come back to the wisdom of crowds in a minute, and an aggregation issue. But well, since we talked about a president making a decision, you have some interesting thoughts on how a leader balances humility with confidence. And I find, I've, we talked about this on the program a lot. I'm, I'm skeptical, but I said, as I'm too skeptical. I have, I need to be more skeptical about my skepticism because I have trouble accepting things that might be true that go against my skeptical beliefs. So, you deal with that in the book. That a leader, most leaders are not very skeptical. They, they seem to be bold. Winston Churchill, I'd be a quintessential example. They don't say, "Well, it could be 73 or 80." It's, they say, "Well, we know the truth. We got to move forward." Talk about this issue of balancing humility and overconfidence and confidence that they say. This is a topic I, if I were to write a sequel book, it's one that I would very much want to have feature prominently in the book.

Well, let's use a sports analogy. I'm not a big sports fan, but my co-author, Dan Gardner, is a big hockey fan, and he's a Canadian. And I was actually born in, originally in Canada myself, but I'm US naturalized. But the Ottawa Senators were apparently in a Stanley Cup final one year, and they were down three to one in, in, in the series. Best-of-seven. Sands best-of-seven series, right? And some reporter thrust the microphone in front of the coach's mouth and said, "Hey, coach, you think I got a chance?" And the coach, instead of doing what coaches are supposed to do, is they, of course, we're gonna kick butt, we can do it, we're gonna win. He went into superforecaster mode and said, "Well, what's the base rate of success for teams that are down three to one? That doesn't look very good, does it? It's, this is not what coaches or leaders are supposed to do." And it raises the question about what, what are the conditions under which leaders are supposed to be liars?

Yeah. And what's the answer?

Well, that's why I said, the next book would wrestle with this. We do talk about it in Superforecasting at some length, and, and, and we have some interesting, I think, military examples and some other examples of where, you know, the need for leadership and a need for confidence are in some degree of tension with the need for circumspection and the need for confidence. And our intention with each other.

Yeah, it's always fascinated me how hard it is for a leader to say, "Ex post, I made a mistake." Or a pundit to say, "I made a mistake." Most of them don't. They hedge. They say, "Well, I had that in mind. Here's the word that suggests that I knew that. Or I didn't know this piece of information. If I'd known that, of course, I wouldn't have done it." Or just my favorite, "It wasn't a mistake. Everyone else thinks they're wrong. It was a great thing." So, get the whole range. But I think there's a cycle. You know, you're a psychologist. I think the psychological challenge of admitting a mistake and being, you know, it's one thing to say, "Well, I know it's a long shot, but I won't say it because that would be bad for the team." I think I'm afraid to say it. I suspect a lot of great leaders don't even think that it's a long shot. They just say, "We're gonna win," and they actually believe it. Right? And, and there's a question about whether you would prefer to have a leader who believes, who is capable of self-deception, or a leader who is capable of being two-faced and has one set of private, private numbers and another.

A set of public numbers. Yeah, for sure. Let's talk about the wisdom of crowds, which you refer to a few times in the book, and you've just, we implicitly talked about it a minute ago. Talk about how you aggregated folks and how you avoided the cloning problem or the groupthink problem in your, in your estimates.

Well, there was a big argument in our research group after early on about whether it would be good, better for our forecasters to work as individuals or work as teams. And the anti-team faction correctly pointed to the dangers of groupthink and all the other dysfunctions of groups. Anyone has ever worked in a team knows how bad teams can be: bullying, all of the above. And then there was another group that said, "Look, there are conditions under which teams can be more than the sum of their parts, and if we give them the right guidance on how to work as a team, they can, they can deliver some great stuff." And we resolved it by running an experiment. And it turns out, at the end of the first year, the teams were better, significantly better. How much better? 10%. You know, what we're talking about is many small factors and a few big ones that cumulatively produce a really big, big advantage for the superforecaster teams.

So, superforecasters do better because they have certain natural or acquired talent advantages. They do better partly because they work in a cognitively enriched environment with other superforecasters. They do partly, they do better partly because they're given them a lot of training and guidance on how to do probability estimation. And then they taught each other. And for that matter, the superforecasters have taught us things. So our training has become better by virtue of the feedback from the superforecasters. And then finally, they do better because of the algorithms.

Well, one thing I want to make clear, which we didn't stress enough: these folks who are doing this, and as you said, there's a lot of forecasts and they're doing it on an ongoing basis, these are not people doing this as a full-time job. You know, these are not, and these aren't professors of political science forecasting what's gonna happen in the South China Sea. These are just really smart, everyday people who were doing this on the side.

Correct. Well, I wish we had more professors of political science. We have a few, but we don't have as many. Advice or the problems? Some of mine too, right?

Well, who are the superforecasters? So the media likes to focus on the superforecasters who are the most counterintuitive. So there's Ann Kilkenny, who is a housewife in the Civil Air Patrol. And there's, like, there's a social work caseworker in Pittsburgh. And there is a person that works as a pharmacist in Maryland. Other superforecasters work as analysts on Wall Street, or were previously analysts in the intelligence community, or were in Silicon Valley, or are really adept software programmers who we developed interesting tools for helping people decide which problems to focus on and how to, how to winnow media sources and so forth. So superforecasters are really quite varied. Some, some of them, you know, which fit more the stereotype of what you'd expect a superforecaster to look like: some Silicon Valley, Wall Street, IQ of 180 type. And others have looked a lot more like intelligent, thoughtful citizens you've run across in everyday life.

And it's interesting how well they get along. It's actually a wonderful dynamic to behold. In the super teams, they are, they are a diverse group. They have, and they're very clever at working out what their sources of comparative advantage are and in dealing with problems and allocating labor. They create, in effect, many organizations. They create, in effect, what they created with many intelligence agencies that were generating probability estimates more accurate than those that were coming out of many intelligence analysts.

But you said that you gave them advice. You didn't just throw them into teams and say, "Good luck, hope you work it out." You did some very thoughtful things that the book describes to get them to perform effectively in these teams rather than as clones or groupthink.

We did. We did. We did. Because there's always a tension in groups to get to the truth. And in groups, you often have to ask questions that might offend people a little bit. So mastering the art of disagreeing without being disagreeable, and mastering the art of what some consultants in California called precision questioning. We found that to be a very useful tool to transfer to other super, to our, to all of our teams, with regular or testing teams and superforecasting teams. You know, because we have thousands of forecasters and many experimental conditions.

Give us a two-sentence description of precision questioning. What is that?

Well, when someone makes a claim like, "Soccer is declining in popularity as the world's most popular pastime," you would want to figure out what exactly we mean by "pastimes," what do they mean by "declining." And you want to get them to be much more specific than people normally are. And when you start probing people, they're often unable to become more specific. And they often, when they then feel when they're being probed, they get irritated and they'll say, "Quit bugging me about this." So superforecasting teams have learned to push the limits of precision but maintain reasonable etiquette inside the group. And I think that's crucial for getting at this, you know, more underlying question you've been, you've been raising throughout the conversation, which is, what are the limits of precision? When does precision become pseudo-precision?

When you start, you need an, you don't know until you test it. Well, the simple answer is when you start using decimal points. Seventy-three point du-pit. Well, you know, I think that's right. I think when you look at, well, how many degrees of uncertainty? Let's just say, for the sake of argument, we treat the probability scale as having a hundred points along it, rather than being infinitely divisible. As to say, it's 100-point probabilities, go, which is the one we actually used in a tournament. How many degrees of uncertainty were superforecasters collectively usefully distinguishing then they make their forecasts? You can estimate that statistically by rounding off their forecast attempts and so forth. And I think our best estimate is they can distinguish somewhere between about fifteen and twenty degrees of uncertainty. A lot of probability still. Most people distinguish about five.

Before I make a point, I, we're almost out of time. I want to get to an economics issue that you raised in the book, which is often on my mind. Which is, I'm going to couch it the way listeners here would expect, which is, we passed this enormous, seemingly enormous, you can debate whether it was enormous or not because there's always a debate after the fact whether the pace conditions even held. We passed a seemingly enormous stimulus package to fight the recession. And there were some predictions that were made sort of about what that would achieve. And there were people on one side of the fence who said it was gonna, you know, unemployment over a certain period of time. Other people said it's gonna make things worse. A lot of people just said, "Oh, I really like it," or "I really don't," without making any kind of even beginnings of a quantitative prediction. But then the dust settled, and afterward, everybody said, "I was right." On either side of the debate. And I'm familiar, I find it deeply troubling that in economics, in particular, but it's elsewhere, that there's no accountability. And if there's no accountability, why do we even begin to pay attention? If there's no authoritative way, even mildly authoritative way, to assess whether a prediction is accurate, whether a model is accurate, whether a policy prescription is fulfilled, how can we make any progress? I don't, and I don't see it in my profession. And you suggest the possibility of some, what some ways we might hold people's feet to the fire and at least have some accountability. How might that work?

I think you might be referring to the proposal of adversarial collaboration tournaments. And we use the example of Niall Ferguson and Paul Krugman and the debate over quantitative easing. Yep. Right. Well, I think it's a great, it's a great model. It has some utility in science. I think it has some utility in public policy debates for running an Iranian nuclear accord tournament. Now, on GJ open. Here's, here's one of the key things you would do. Each side would have a chance to nominate five or ten questions that it thinks it has a comparative advantage in answering. The questions have to be relevant to the underlying issue and have to pass the clairvoyance test, which means that to be rigorously scorable for accuracy after the fact. And victory has a pretty clear meaning in this kind of context. It means you not only can answer my questions, you're not, I can answer your questions, and I can, you can answer my questions better than I can. And that leaves me in an awkward situation because I can't simply say, "Well, you posed some stupid questions." I have to say, "All my questions were stupid too." It's much more awkward. Now, of course, pundits are going to be very reluctant to engage in a game like this. I mean, why would you want to engage in a game where the best possible outcome is it's not very, it's not nearly attractive enough to justify the risk? The only way we're ever going to induce high-status pundits to agree to participate in level playing field forecasting tournaments in which they pit their predictions about the future against their competitors is if the public demands it. And if there is a groundswell demand for that. If pundits feel that their credibility is beginning to suffer because they're refusing to offer more precise and testable predictions visa vie their competitors, then I think that would be, that would be the only force on earth capable of inducing them to do it. I think there's a shame factor. I think an external source, maybe this program could shame some people into participating.

But it's an interesting question. One of the challenges, I think, in economics, and it's somewhat of a question, I think, when you think about it carefully, and in other fields as well, is what are you measuring? So if what we really care about, say, it's not the whole picture, but if we're worried about whether the minimum wage, say, causes unemployment, one reason people would say, "Well, I can't, I can't participate in that because there's so many other factors besides that increase the minimum wage. I can't guarantee that they're not going to be in place." And then I was not a probability. What if you can't? Then you should shut your mouth because you're just talking. I mean, there's no, and I do it too. I should just make it, I shouldn't say just them. I don't pretend mine is scientific in the way that they sometimes do with empirical data. I just, I'm trying to rely on my principles, which I think are pretty reliable. But I'm probably at risk of fooling myself there too.

Oh, I think the minimum wage would be a wonderful example of where adversarial collaboration could work because there are so many states and municipalities take independent actions on that front. Yeah, maybe something I can enable if I play my cards right.

My guest today has been Philip Tetlock. Philip, thanks for being part of EconTalk. And listeners, thanks for putting up with some of that electronic noise. We're trying to make it better.

[Music]

This is EconTalk, part of the Library of Economics and Liberty. For more EconTalk, go to econtalk.org, where you can also comment on today's podcast and find links and readings for today's conversation. The sound engineer for EconTalk is Rich Goyer. I'm your host, Robert. Thanks for listening. Talk to you on Monday.

[Music]