📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

What Is Statistics: Crash Course Statistics #1

CrashCourse13:00

Transcription

Hello, I'm Adrian Hill, and this is Crash Course Statistics. Welcome to a world of probabilities, paradoxes, and p-values. There will be games, thought experiments, coin flips, and lots and lots of coin flips. Statisticians love talking about coin flips. By the time we finish this series, you'll know why we use statistics, how we use statistics, and what questions you should be asking when statistics are used in the world. They're everywhere.

Statistics can help you guess whether or not you'll get into Harvard. Marketers use it to sell us gold and flared trousers. Netflix uses statistics to predict the shows we might want to watch later. You use statistics when you look at the weather forecast and decide what to wear – a dress or jeans. Policymakers use it to decide whether or not they want to invest in early childhood education or whether or not they should spend more on mental health services. Statistics is all about understanding data – and figuring out how to use and employ that information. Today, we're going to answer the question, "What is statistics?"

Legend says that during an English tea party in the late 1920s in Cambridge, a woman claimed that a cup of tea with the milk added afterward tasted different than tea where the milk was added first. The brilliant minds present that day started thinking of ways to test her claim. They arranged eight cups of tea, all sorts, to see if she could really tell the difference between the milk-first and the tea-first cups. But even after seeing her guesses, how could they really decide? Because she'd get half the cups right just by random guessing – either milk or tea. And even if she could actually tell the difference, it's entirely possible she'd miss a cup or two. So how do you know if this woman was actually a tea connoisseur? What's the line between a lucky tea guesser and a tea taster?

As fate would have it, the super speedy statistician and part-time potato scientist Ronald A. Fisher was present. During his lifetime, Fisher began the work that paved the way for much of the statistics that is the focus of this series. Statistics can help us make decisions in uncertain situations – tea tastings and beyond. Fisher's insights on experimental design helped transform statistics into its own scientific system. And although Fisher never published the results of this tea test… the story goes… in the end, the woman correctly sorted all the cups of tea. Just in case you were curious at this point, it's worth noting that there are two related – but separate – meanings to the word "statistics." We can refer to the field of statistics… which is the study and practice of collecting and analyzing data. And we can talk about statistics as in facts about… or summaries… of data.

To answer the question "What is statistics?" we must first… ask the question "What can statistics do?" Let's say you woke up at your desk after a long evening studying for finals with a cheese-steak sub covering your face. And you're wondering to yourself… "Why did I eat that? Does fast food control my life?" But then you say to yourself, "Nah, it's just super convenient." But you're concerned. You're thinking about how great it is that McDonald's serves breakfast all day. But maybe that's normal. This week is finals, so you Google the question "fast food consumption," and you find fast-food survey results. The first thing you can do is start asking questions that interest you. For example, you could ask, Why do people eat fast food? Do people eat fast food on weekends more than weekdays? Does fast food cause me stress? Now, we have some interesting questions, we need to ask ourselves a more important question: Can these questions be answered via statistics?

As I mentioned before, statistics are our tools, but they can't carry the whole load. To answer a question about why people eat fast food, you can ask them to fill out a questionnaire, but you can't know if their answers represent the truth of what they think. Maybe they'll answer dishonestly because they don't want to admit that they're putting on a McDonald's onesie because they're too exhausted to cook dinner, or because they're ashamed to admit they think a Del Taco is delicious, or because none of the answers provided represent their true reasons. Or they might not really know why they eat fast food. With a survey's results, you could tell you that the most common reason people reported for eating fast food is convenience, or that the average number of meals they eat each week is five. But you're not truthfully measuring why people eat so much fast food. You're measuring what we call a "proxy," which is related to what we want to measure, but not exactly what we want to measure.

To answer whether people eat fast food on weekends, or whether eating it more than twice a week increases stress, we'll need more than just how much fast food people eat, which our survey addressed, and what days they eat it. We'll need an additional measure of "stress." You can use statistics to give a good answer about whether you'll order more takeout on holidays, but even the question of whether fast-food consumption is related to higher stress levels is difficult to answer directly. What is stress, and how can we measure it? Do people eat fast food because they're stressed? Or does eating all those calories increase their stress? Often, some of the most interesting questions are the ones that cannot be directly answered by statistics – like why people eat fast food. Instead, we find questions we can answer – like whether people who eat fast food often work more than eighty hours a week.

The tools we use to answer these questions are numerical statistics – and there are two main types: descriptive and inferential. Descriptive statistics, well… describe what the data shows! Descriptive statistics usually involve things like where the middle of the data is, what statisticians call measures of central tendency – and measures of how the data is spread out. They take massive amounts of information that might not be intuitive to us and work to compress and summarize it… hopefully giving us more useful information. Let's go to the Thought Bubble. You've worked for two years at your local waffle factory. Day in and day out, you create the golden-brown, crispiest frozen waffles ever. The holes are perfectly spaced. They scream for syrup. And now you want a raise. You deserve a raise. Nobody can make a waffle like you. But how much to ask for? An extra thousand dollars? An extra five thousand dollars? You know you're valuable, but you have no idea what other waffle makers are making. So you search online, and you find a whole subreddit dedicated to waffle makers. And the username "waffleleaks" has posted a salary table for waffle makers. Now, just by glancing at this huge list of numbers, you can figure out if the woman working a similar job at the competing frozen waffle company is making more than you. You can figure out how much you make more than the new guy, who's just learning how to mix the batter. But you still don't know much about waffle company salaries as a whole, or the industry as a whole. It turns out there are thousands of waffle makers out there, and all you see is a list of data points, not the patterns that could help you learn more about how much you can convince your boss to pay you. This is where descriptive statistics come in. You can calculate the average salary at your company as well as how everyone's salaries are distributed around that average. You'll be able to figure out whether the salaries of the top executives are relatively close to entry-level batter-makers or incredibly far apart, and how your salary compares to both. You can calculate the average salary for everyone in the industry with your job title and look at the high and low end of that pay. Then, armed with these descriptive statistics, you can walk confidently into your waffle-company boss's office and ask to be paid for your talents. Thanks, Thought Bubble.

Although descriptive statistics can be great, they only tell us the basics. Inferential statistics – fancy terms used by those statisticians – allow us to make inferences that go beyond the limits of the data we have available. Imagine you have a barrel of candy full of taffy. Some of it's pink, some white, some yellow. If you want to know how many of each color you have, you could count them one by one. That gives you a set of descriptive statistics. But who has time for all that? Or, you could grab a huge handful of taffy and just rely on the ones you pulled. That uses descriptive statistics. If your candy is, in fact, evenly mixed throughout the barrel, and you grabbed enough, you can use inferential statistics on that "sample" to estimate the contents of the whole taffy stash. We call on inferential statistics to do all sorts of more complicated work for us. Inferential statistics lets us test an idea or a hypothesis, like answering whether people in the United States under thirty eat more fast food than people over thirty. We don't survey everyone to answer this question.

Let's say someone tells you that the new brain vitamin – Smartie-vite – works to improve your intelligence. Would you rush out and buy it? What if they told you that the average increase in IQ for group A – twenty people who took Smartie-vite for a month – was two points, and group B, twenty people who took nothing – was one point. What do you think now? Still unsure? It's a pretty small difference, right? Inferential statistics give you the ability to test how likely it is that the two groups we sampled had different increases in IQ. However, the decision is up to you, as an individual, to decide whether or not it's convincing. And don't be alarmed if the bar you set isn't the same in every situation. It's perfectly acceptable to have different standards for questions like "Does my cat like Fancy Feast more than Meow Mix?" versus "Does this drug cure lung cancer?" It would take far more evidence to convince you to take a new drug supposedly curing cancer than to switch cat food brands. It would take more evidence to convince you to take a new drug supposedly curing cancer than to switch cat food brands. With inferential tests, there will always be a degree of uncertainty because they can only tell you how likely something is to happen or not happen. Your job is to take that information and use it to make a decision *despite* that uncertainty. If statistics were a superhero, its kryptonite would be uncertainty, and its motto would be "When you don't know for sure, but doing nothing isn't an option."

Statistics are tools. Statistics help us make sense of the overwhelming amount of information in the world. Just like our eyes and ears filter out unnecessary stimuli to just give us the best and most useful things, statistics help us filter the amounts of data that come at us every day. Descriptive statistics make the data we get more digestible, although we do lose information about individual data points. Inferential statistics can help us make decisions about data when there's uncertainty – like whether Smartie-vite will actually increase your IQ. But statistics can't do all the work. They're here to help us think, not do the thinking for us. They help us see through the uncertainty, but they don't get rid of the uncertainty. To push our tool analogy a step further, statistics, like chainsaws, are pretty dangerous until you understand how they work. We need to know how to use them and when not to use them. And as we'll see later in subsequent episodes, statistics done poorly can lead us to some silly conclusions. A chainsaw used poorly leads to about 36,000 injuries in the United States every year, 81% of which are lacerations. Did you know that almost nobody dies from chainsaw injuries? Fatal injuries are incredibly rare. 95% of people injured by chainsaws are male. This doesn't necessarily tell us that males are significantly worse at using chainsaws.

Statistics can help us plan a vacation to Bali in December. They can help us improve our chances of winning at soccer. They can help us budget our energy in college. Statistics can help us determine whether that extended warranty a guy is trying to sell us at Best Buy on our new blender is worth it. Statistics can also help us determine whether or not you should get open-heart surgery. Statistics can help NGOs improve the amount of food aid they send to refugee camps. They can help policymakers decide whether or not they should spend more or less money on helping students pay off their student loans. And they can help you determine how much money you should be comfortable borrowing for college in the first place. There's a lot that statistics can help us do, but some things statistics can't do. Statistical thinking means knowing the difference. So when your brother says he used statistics to prove your mom loves him more, you can rest easy knowing that the only question he answered is whether she'll give him more ice cream every night. And he got data suggesting she gives you extra candy sprinkles. Translation by: Shwan Hamid, Twitter: @shwan_hamid