📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

U1 L3: Critical Thinking

OCA Statistics12:49

Transcription

In today's lesson, we're going to discuss critical thinking. The purpose of this lesson is to improve skills in interpreting information based on data. Success in the introductory statistics course typically requires more common sense than mathematical expertise. The lesson is designed to illustrate how common sense is used when we think critically about data and statistics.

You need to think carefully about the context, the source, the methods, the conclusion, and the practical implications when we're looking at data. Sometimes, data can be deceptive. Here are some misuses of statistics: first, the evil intent on the part of a dishonest person, or second, the unintentional errors on the part of the people who don't know any better. We should learn to distinguish between statistical conclusions that are likely to be valid and those that are seriously flawed.

Let's talk about graphs. To correctly interpret a graph, you must analyze the numerical information given in the graph so it's not to be misled by the graph's shape. So, for example, graph A and graph B both represent the same data, but they both look very different. So, we read the labels and the units on the axis. For example, in graph B, the vertical axis begins at 30,000, whereas in graph A, it starts at 0. You can see the less of a jump between the purple and the blue bars versus when you only show a zoomed-in portion of the graph, it makes it look like a much larger gap. This exaggerates the difference between the bars.

Pictographs are drawings of objects using areas or values, and they can also be misleading. So, again, graph B exaggerates the difference by increasing the dimensions and proportions. The actual amounts of oil consumed. So, in graph A, we can see the daily oil consumption of the U.S. is 20 million barrels versus Japan, 5.4 million barrels. But over here, this is actually almost about less than or more than a fourth of this cylinder that represents the United States. So, let's take that image for a second here. So, this pictograph is deceiving because, like I said, 5.4 is more than one-fourth of 20. So, not quite four of these should fit inside the 20s barrel. But when I put four of these in here, notice how it is nowhere near only visually filling up the U.S. barrel. So, this is misleading, and the bar graph would be a better visualization. See how this bar looks a lot more like close to one-fourth of the bar? If I take four of the blue bars, they should pile on top of each other over the U.S. bar just a little bit more than it, because 5.4 is a little bit more than one-fourth of the U.S. bar. The graph is much more clear.

As discussed in the last lesson, the voluntary response sample is one in which the responders themselves decide whether they can be included or not. In this case, a valid conclusion can be made only about specific groups of people who agreed to participate, not about the entire population.

Now, correlation and causality occur when a statistical association between two variables is found, and then it is concluded that one variable causes or directly affects the other variable. It may seem as if the two variables, such as smoking and pulse rate, are linked. This is called correlation. But even if we found that the number of cigarettes was likely linked to a pulse rate, we could not conclude that one variable actually caused the other. So, correlation, seeing a correlation between two variables, does not imply causation. It's just that they are linked in some sort of way.

Now, let's talk about samples that might be too small. So, small samples. The conclusion should not be based on samples that are too small. For example, facing the school suspension rate on a sample of only three students. The percentages can be misleading or unclear. Sometimes, when used, for example, if you take 100% of a quantity, you take it all. If you have improved 100%, then you are perfect. Are you perfect? 100%? 110% of an effort does not make sense. You can't give more than 100%.

Be careful with surveys and loaded questions. If survey questions are not worded carefully, the results of the study can be misleading. Survey questions can be loaded or intentionally worded to elicit a desired response. So, do you agree with the following statements: "Too little money is being spent on welfare" or "Too little money is being spent on assistance to the poor"? So, using the word "welfare" would elicit less of a bias than seeing the words "assistance to the poor." Those are strongly worded questions and it's loaded. See how the percentages of agreeing result based on the same question worded differently? Sometimes, even the ways that you order your questions unintentionally create some kind of bias. So, in the example below, the order of the words "traffic" and "industry" are switched. Notice how the factor people chose most. So, would you say traffic contributes more or less to air pollution than industry? Versus, would you say that industry contributes more or less to air pollution than traffic? So, whichever word comes first. "Traffic" came first in the first one, a higher percentage favored traffic. "Industry" came first in the second question, a higher percentage preferred industry. That wording or felt industry was based on the wording that came first.

A non-response occurs when someone either refuses to respond to a survey question or is maybe unavailable to answer the question. People who refuse to talk to pollsters have a view of the world around them that is markedly different than those that will let poll takers into their home.

Missing data can dramatically affect results as well. Sample value data can be missing because of random or special factors. Subjects may drop out for reasons unrelated to the study. People with low income are less likely to report their incomes, and U.S. Census suffers from missing people, and they tend to be homeless or low income.

Self-interest studies come from parties with an interest to promote or sponsor a specific study. Be wary of those in which their sponsors can enjoy monetary gain based on the results.

Precise numbers. Because of figure, a figure is precise, many people incorrectly assume that it is also accurate. Precise numbers can be an estimate, and it should be referred to that way. Now, when collecting data from people, it is better to take measurements yourself instead of asking people or asking subjects to report results themselves. So, be careful with that.

Biased deliberate distortions. Some studies are or surveys are distorted on purpose. The distortions can occur within the context of the data, the source of the data, the sampling method, or the conclusion. Let's look at a couple of practice problems to help us understand these definitions.

During a study, one-third of participants drop out. Which type of problem is this for the study? Well, if they dropped out, then we must have missing data, right? We're missing data from particular people. The information gathered will be incomplete and can dramatically affect our results.

Researchers from local universities determined that they needed results from at least 200 subjects in order to conduct a certain study. They mailed out 5,000 surveys and received only 354 responses. Is the sample of 354 responses considered to be a good sample? Now, remember, they needed at least 200 subjects, and they got over that 200 subjects. But with a mailed-out survey, it was voluntary. So, because it's a voluntary response, the respondents themselves decided if they were going to be included, and they only received about 7% of the surveys that were sent out. So, the voluntary response sample, in which people with special interests are more likely to respond.

A local high school has a student population of about 1,000. During any given year, student school board members were interested to know the percentage of students suspended from school two or more times during their high school career. So, they chose four students and followed them over four years. They reported that 50% of students were suspended from school two or more times during the four years. What's wrong with the study? The sample size is extremely small, only four students out of a thousand. The conclusion is only based on those four students.

A study uses statistical methods to conclude that there is an association between electrical usage and the number of light bulbs in a home. The study then concludes that removing light bulbs from your home will reduce your electric bill. What's wrong with the reported results of the survey? Association or correlation does not imply causation. Remember, it's because there's an association, it doesn't mean that one has caused the other. A correlation was made between two variables, but then it was concluded that one variable caused the other.

A study produced by local farmers was quoted in a news article to indicate that eating fresh vegetables improved general health. What's wrong with this particular study? Obviously, for farmers, they are going to be growing the vegetables, so they have a self-interest in this particular study. "Eat my vegetables, and you'll be giving me money because so that I can grow my crops again and make money." So, this study was produced by farmers who are likely selling fresh vegetables.