📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Psych 74 Week 12 Lecture Video

Dr. GRS29:18

Transcription

Hi class. So last week, we discussed Survey Research. And you know, as a mode of primarily quantitative data collection, it can also be qualitative as well. But as a mode of quantitative data collection, we would need to use statistical methodologies to analyze any of the data that is collected on our surveys. So that brings us to this week's lecture, where we're going to be reviewing some of the basic concepts in statistical analysis, which is really the seat of any quantitative assessment of survey data.

Now, it's worth noting that this lecture will not be as exhaustive as your previous statistics classes, but this is really just to refamiliarize all of you with some of the core concepts that we would need to understand to be able to interpret quantitative data. And in particular, of relevance to the second part of your research papers, to choose a type of quantitative analysis if that is your your preferred methodology of assessment.

So, the very first thing that's worthwhile to do is to have a working definition of statistics. So, what are they? Well, statistics are a set of mathematical procedures for summarizing and interpreting observations. So obviously, these, when we talk about quantitative data, these observations are typically numerical or categorical facts about specific people or things. And these facts, these this information is usually referred to as data, which is which are facts derived from experience.

Now, when you have statistics that you are making inferences about a larger population, you are going to be collecting data from your target population. Population is all members of your group of interest. And you're going to be deriving a subset from that population. Now, that subset is referred to as a sample. And by using inferential statistics, you're making an inference about what's going on in the broader population. Why would we do this? Well, the biggest reason, as I mentioned in previous lectures, is that it may be nearly impossible to survey the entire members of a population. With the exception of gargantuan undertakings like the US Census, getting data from an entire population of interest is oftentimes cost-prohibitive. So we can use inferential, or as they're also referred to, parametric statistics, to be able to make estimates about the target population from a smaller subset of that population.

Now, it's important when we're looking over some of the basics of any kind of statistical analysis to understand that we need to summarize our data in some way to make it meaningful. And this is where we get to central tendency. It's a single number that's used to represent a form of an average score in a distribution of scores. So when we say distribution of scores, we're talking about the total data set that we have of our collected data from our sample. And so there are three, when we talk about average, there are really three different measures of average. There's mean, which is the common average, which when we think about averages, is the most common representation of it. It's computed by adding all the scores in a distribution and then dividing them by the number of scores in that distribution. But there's also median, which is the middlemost score in the distribution, and mode, the most commonly occurring score in that frequency distribution. So again, that is the most commonly occurring individual score.

And when we, when we look at ways to describe the data, we're also interested in to in the differences between individual data values. And that's where we get into dispersion and variability. So the traditional choice for measuring dispersion and variability is called the standard deviation. It's the typical distance from the average score to the data values. It really indicates the extent of randomness of individuals of around their common average or about their common average. And it's the typical size of these deviations, ignoring positives or negatives, and is a number in the same unit of measurement as the original data. So if you collect data is in the form of dollars or miles per gallon or kilograms, the distance, the standard deviation is also going to be measured in that that same value, whether it be dollars, miles per gallon, or kilograms. And when we're talking about population values, the standard deviation is commonly represented in the notate scientific notation of the Greek letter of Sigma. And this is how you might see this on a data table in a psychological study. It gives you a sense of how spread out the data is.

So we, the first things that we're looking at are measures of central tendency and measures of dispersion. But the simplest measure of variability is simply the difference between the highest and lowest scores within a distribution. And this is referred to as the range.

A third statistical property of a set of observations is a little more difficult to quantify than measures of central tendency or dispersion. This third statistical property is the shape of the distribution of scores. One useful way to feel for a set of scores is to arrange them in order from lowest to highest and graph them pictorially. So the taller parts of the graph represent more frequently occurring scores, or in the case of a theoretical or ideal distribution, more probable scores. Different kinds of distributions include a rectangular distribution, a bimodal distribution, and a normal distribution.

The scores in a rectangular distribution are all about equally frequent or probable. An example of a rectangular distribution is the theoretical distribution representing the six possible scores that can be obtained by rolling a single-sided dice. In the case of a bimodal distribution, two distinct ranges of scores are more common than any other. A likely example of a bimodal distribution would be the heights of athletes attending the annual sports banquet for a very large high school that has only two sports teams: women's gymnastics and men's basketball. If this example seems a little contrived, it should. Bimodal distributions are relatively rare, and they usually reflect the fact that a sample is composed of two meaningful subgroups or sub-samples. In this case, women's gymnasts and men's basketball players. And this would kind of look like two rising groups, right? And you'll see this in a graphical representation in a handout that is posted to web access. Whereas a rectangle, a normal distribution would be shaped kind of like a rising hill. And this is the most important type of distribution. This is referred to as the normal distribution. It's a symmetrical bell-shaped distribution in which most scores cluster near the mean, and in which scores become increasingly rare as they become increasingly divergent from this mean. Many things can be quantified as normally distributed. Distributions of height, weight, extraversion, self-esteem, and age at which infants begin to walk are all examples of approximately normal distributions.

The nice thing about the normal distribution is that if you know the set of observations are normally distributed, this further improves your ability to describe the entire set of scores in the sample. More specifically, you can make some very good guesses about the exact proportion of scores that fall within any given number of standard deviations or fractions of standard deviations from the mean. So one thing that we know consistently about normally distributed data is that about 68% of a set of normally distributed scores will fall within one standard deviation of the mean. About 95% of a set of normally distributed scores will fall within two standard deviations of the mean. And well over 99% of a set of normally distributed scores, 99.8% to be exact, will fall within three standard deviations of the mean. For example, those of you who've taken a developmental psychology course or might be somewhat familiar with this, that modern intelligence tests, such as the Wechsler Adult Intelligence Scale, or the WAIS, are normally distributed in their scores. They have a mean of 100 and they have a standard deviation of 15. This means that about 68% of all people have IQs that fall between 85 and 115, and that an IQ of 100 is an average IQ. Similarly, more than 99% of all people, again 99.8% of all people to be more exact, should have IQs that fall between 55 and 145.

This kind of analysis can also be used to put a particular score or observation into perspective, which is a first step towards making inferences from a particular observation. For instance, if you know that a set of 400 scores on an astronomy midterm approximates a normal distribution, has a mean of 70, and a standard deviation of exactly, you should have a very good picture of what this entire set of scores looks like just from those little pieces of information. And you should know exactly how impressed to be when you learn that your friend, let's call her Amanda, earned an 84 for that exam. She scored 2.33 standard deviations above the mean, which means that she probably scored in the top 1% of her class. How could you tell this? By consulting a detailed table based on the normal distribution. Such a table would tell you that only about 2% of a set of scores are 2.33 standard deviations or more from the mean. And because the normal distribution is symmetrical, half of the scores that are 2.33 standard deviations or more from the mean will be 2.33 standard deviations or more below the mean. Amanda's score was in the half of that 2% that was well above the mean. Translation: Amanda did really well. She kicked butt on that exam.

The basic idea behind inferential statistics, kind of taking all that we've discussed to the next level, and in terms of statistical testing, is that decisions about what to conclude from a set of research findings need to be made in a logical and unbiased fashion. One of the most highly developed forms of logic is mathematics. And statistical testing involves the use of objective mathematical decision rules to determine whether an observed set of research findings is real, for lack of a better way of putting it. The logic of statistical testing is largely a reflection of the skepticism and empiricism that are crucial to the scientific method.

When conducting a statistical test to aid in the interpretation of a set of experimental findings, researchers begin by assuming that the null hypothesis is true. That is, they begin by assuming that their own predictions are wrong. In a sample to group experiment, this would be assuming that the experimental group and the control group are not really different after the manipulation of that other variable of your independent variable, and that any apparent difference between the two groups is simply due to luck or to a failure of random assignment. After all, random assignment is good, but it's rarely perfect. It is always possible that any difference the experimenter observes between the behavior of participants in the experimental and control groups is simply due to chance.

In the context of an experiment, the main thing statistical hypothesis testing tells us is exactly how possible it is, or how likely it is, that someone would get results as impressive as, or more impressive than, those actually observed in an experiment if chance alone, and not an effective manipulation, were at work in the experiment. The same logic applies, by the way, to the findings of all research, survey or interview alike. If a researcher correlates a person's height with that person's level of education and observes a modest positive correlation, such that taller people tend to be better educated, for example, it's always possible, out of dumb luck, that the tall people in this specific sample just happened to have been more educated than the short people, that the relationship really has no significance. Statistical testing tells researchers exactly how likely it is, given a research finding, would occur on the basis of luck alone if nothing interesting is really going on. Researchers conclude that there is a true association between these variables they've manipulated or measured only if the observed association would rarely have occurred on the basis of chance.

Because people are not in the habit of conducting tests of statistical significance to decide whether they should believe what a salesperson is telling them about a new line of athletic shoes, whether there is intelligent life on other planets, or whether their friend's tastes in movies is statistically significantly different from their own, the concept of statistical testing is pretty much foreign to most laypeople. However, anyone who has ever given much thought to how American courtrooms work should be extremely familiar with the logic of statistical testing. This is because the logic of statistical testing is almost identical to the logic of what happens in an ideal courtroom. With this in mind, our discussion of statistical testing will focus on the simile of what happens in the courtroom. If you understand courtrooms, you should have little difficulty understanding statistical testing.

As mentioned previously, researchers performing statistical tests begin by assuming that the null hypothesis is correct. That is, that the researchers' findings reflect chance variation and are not really significant. The opposite of the null hypothesis is the alternative hypothesis. This is the hypothesis that any observed difference between the experimental and control group is real. The null hypothesis is very much akin to the presumption of innocence in the courtroom. Jurors in a courtroom are instructed to assume that they are in court because an innocent person had the bad luck of being falsely accused of a crime. That is, they're instructed to be extremely skeptical of the prosecuting attorney's claim that the defendant is guilty. Just as defendants are considered innocent until proven guilty, researchers' claims about the relation between variables they've examined are considered incorrect unless the results of the study may suggest otherwise. No until proven alternative, you might say.

After beginning with the presumption of innocence, jurors are instructed to examine all evidence presented in a completely rational and unbiased fashion. The statistical equivalent to this is to examine all evidence collected in a study on a purely objective mathematical basis. After examining the evidence against the defendant in a careful, unbiased fashion, jurors are further instructed to reject the presumption of innocence, to vote guilty only if the evidence suggests beyond a reasonable doubt that the defendant committed the crime in question. The statistical equivalent of the principle of reasonable doubt is the alpha level, agreed upon by most statistics and statisticians as the reasonable standard for rejecting the null hypothesis. In most cases, the accepted probability value at which alpha is set is 0.05. And you study statistics, you're going to be hearing this quite a bit.

A final parallel between the courtroom and a psychological or sociological laboratory is particularly appropriate in theoretical fields such as psychology or sociology. In most court cases, especially serious cases such as murder trials, successful prosecuting attorneys will usually need to do one more thing in addition to presenting the body of logical arguments and evidence pointing to the defendant: it will need to identify a plausible motive, a good reason why the defendant might have wanted to commit the crime. It's difficult to convict people solely on the basis of circumstantial evidence. A still similar state of affairs exists within the social sciences. No matter how statistically significant a set of research findings is, most psychologists will place very little stock in it unless the researcher can come up with a plausible reason why one might expect to observe those findings. In psychology, these plausible reasons are called theories. It's quite difficult to publish a set of statistically significant empirical findings unless you can generate a positive theoretical explanation for them.

Having made this friendly pass through a highly theoretical and technical subject, we will now try to enrich your understanding of inferential statistics by using it by discussing things in a little bit more detail. Analyzing and interpreting data from most real empirical investigations requires more extensive calculations than those you like in a statistics class at the undergraduate level. But of course, these labor-intensive calculations are usually carried out by computers. In fact, a great deal of your training, if you can decide to continue with them in the field of any social science focus, is going to involve getting a computer to crunch numbers for you and using statistical software packages such as SPSS or Stata. Regardless of how extensive the calculations are, the basic logic underlying inferential statistical tests are almost always the same, no matter which specific inferential test is being conducted, and no matter who or what is doing the calculations.

Now, as suggested in the thought experiment with American courtrooms, all inferential statistics are grounded firmly in the logic of probability theory. Probability theory deals with the mathematical rules and procedures used to predict and understand chance events. For example, the important statistical principle of regression towards the mean, the idea that extreme scores or performances are usually followed by less extreme scores or performances from the same person or group of people, can easily be derived from probability theory. Similarly, the odds in casinos and predictions about the weather can be derived from straightforward consideration of probabilities.

So, what is probability? Well, from a classical perspective, the probability of an event is a very simple thing. It's the number of all specific outcomes that qualify as the event in question, divided by the total number of all possible outcomes. The probability of rolling a three on a single roll with a standard six-sided die is 1 out of 6, or 0.167, because there is one and only one roll that qualifies as a three, and exactly six equally likely possible outcomes. For the same reason, the probability of rolling an odd number on the same die is 3 out of 6, or 0.50, because three of the six possible outcomes qualify as odd numbers. It's important to remember that the probability of any event or complex sets of events, such as the observed results of an experiment, is the number of ways to observe that event divided by the total number of possible events.

So there are a few things that can go wrong when the as experimenters and social science researchers try to conduct a hypothesis test or a statistical significance test using quantitative data. First of all, it's important to remember that when a researcher conducts a statistical test and obtains a significant result, this does not always mean that his or her hypothesis is correct. Even if a study is performing executed with no systematic design flaws, it's always possible that the researchers' results were due to chance. In fact, the p-value we observed at 0.05, the observing an experiment tells us exactly how that it is that we would have to have obtained results like ours, even if nothing but dumb luck were operating in our study. Statisticians refer to this worrisome possibility, incorrectly rejecting the null hypothesis when it is in fact correct, as the Type 1 error. The likelihood of making a Type 1 error is a direct function of where we set our alpha level. As suggested earlier, if we think it would be a practical or scientific disaster to reject the null hypothesis in error, we might want to set the alpha at a very conservative level, such as 0.01. This would be taking only one chance in a thousand of falsely rejecting the null hypothesis.

So why not only set the alpha at 0.01 or even lower all the time? Well, because we have to strike a balance between being cautious and being so cautious that we become downright foolish. In statistical terms, if we always set alpha at an extraordinarily low level, we would decrease the likelihood of committing a Type 1 error at the expense of increasing the likelihood of committing what's referred to as a Type 2 error. A Type 2 error occurs when we fail to reject an incorrect null hypothesis. That is, when we fail to realize that our study has revealed something meaningful, usually that our hypothesis is correct. The reason it is useful to know about Type 1 and Type 2 errors is that there are things we can do to minimize our chances of making both of these troubles. As I suggested previously, one of the easiest ways to minimize Type 1 errors is to set alpha at a pretty low level. Over the years, most researchers have pretty well agreed that 0.05, or 5%, is a reasonable level for the alpha, a reasonable risk for making a Type 1 error. And of course, we want to be a little more cautious, but we don't want to risk anyone. Or we don't want to ask anyone to adjust any alpha levels. We can always insist on seeing replication. So, in the grand scheme of things, replications are what tell us whether our an effect is real. Whether we do the research and the statistical analysis again and again on different samples on the same population and come up with the same results. And although no one wants to make a Type 1 error, no one really wants to make a Type 2 error either. Several things that influence the likelihood that a researcher will make a Type 2 error and fail to detect a real effect are broad. Some of these are things that are things over which researchers have little or no control, and some of them are things over which researchers have almost complete control.

One thing researchers can't do too much about is their effect size. What's an effect size? Well, that's the magnitude of the effect in which they happen to be interested. If you collected a sample of 20 people and measured their heights and their foot sizes, you could probably expect to observe a statistically significant correlation between height and foot size, even though your sample was pretty small. This is because there is a pretty robust tendency for big people to have big feet. Of course, there are exceptions, but they're relatively rare. We doubt you'll ever meet a gymnast who squeezes into a size 14 or an NBA center who slips comfortably into a size 9. On the other hand, if you gave a sample of 20 people a measure of extraversion and a measure of self-esteem, you might not necessarily observe a significant correlation. Although self-esteem and extraversion do tend to go hand-in-hand, this correlation is much more modest than the substantial correlation between height and foot size.

So finally, as we close this very brief survey on the concepts behind statistics and the social sciences, in terms of research methods and quantitative research, this slide is going to become something that could be very useful to you. I'll leave we go over some of the highlights. These are important statistical notations that you'll see if you read through and quantitative research and as you go through any classes that utilize statistics. And probably one of the most important to note is the population size, referred to as big N, but this is distinct from the sample that is drawn from that population, which is referred to as little n. When we try to understand average, we're referring to, we often use mu when we're talking about the average of some sort of occurrence within the population, but we'll use X bar here as the mean for an average of sample data. When we talk about the sum total of numbers, we will use the summation sign or Sigma. And when we talk about the standard deviation, we're referring to population data versus sample data. We'll use one of these two symbols, with samples being represented as s-x, referring to sample standard deviation. Okay, so printing out this slide and saving it for use in any statistics class might be useful.

All right, so please also watch the supplementary video lectures posted for this module and respond to the discussion question. As always, if you have any questions, post them to the forum, the forums, or send me an email. Have a great week, guys.