📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

PSYCH 74 Week 13 Lecture

Dr. GRS18:58

Transcription

Hi class. For today's lecture, we're going to be finishing up our discussion of statistics, and in particular, our focus has been and will continue to be on parametric statistics. That is, statistics that use the normal distribution, using samples collected from the larger population to estimate what is really truly going on in our population of interest.

So, when we talk about threats to statistical conclusion validity, you may say to yourselves, "Didn't we talk about threats to validity already?" Well, there is a distinction. We talked about threats to internal validity, that is, the extent to which our study establishes true cause-and-effect relationships, whether our various internal validity, our study is valid in and of itself, our measures are valid in and of itself. This lecture is going to be focusing on what some of the threats to our statistical conclusions are, what gets in the way of having our data accurately analyzed and actually reflecting what is really going on, based on the way we conduct statistical analyses and based on the way we interpret those statistical outcomes. Because there are a number of things that we need to keep in mind when engaging in our own research. This will be especially helpful for those who are planning on taking their studies to the next level after this course, whether it be at a four-year institution or under faculty supervision.

So, the first threat that comes to mind is low statistical power. And so, what this has to do with is the ability of our statistical analysis to determine an accurate assessment of what's really going on. And so, an insufficiently powered experiment may incorrectly conclude that the relationship between the, say, for example, in the case of psychology, treatment and outcome is not significant. And this is usually because of a low sample size. So, if a study does not have a large enough sample size, or "n" as the statistical notation would be, then the study is considered to be underpowered, which means that it's impossible to accurately detect an effect or a significant result with the analysis that we choose to run.

So, I would make this argument: when a study does not have significant findings, and you notice that there's also a small sample size, so if you're looking through scholarly articles and you find that, you know, they, they fail to reject their null hypothesis, as would be the determination in statistics, that they fail to find a significant effect, but you also know that their sample size is pretty small, maybe under 30, if they really didn't know what they were doing, and maybe in the, maybe under a hundred, if they still met the, the assumptions to do some statistical analyses, they would still, you might want to call into question whether low statistical power was the result of their lack of findings, or whether it was actually that their results were reflected in the population.

And so, one of the things we mentioned was assumptions. Now, each type of statistical analysis has related assumptions that go with it. And one of the most important ones for a lot of parametric statistical analyses is, for example, the assumption of normality, that the sample that we're looking at is normally distributed, right? That it follows that bell-shaped curve where the majority of individuals' scores are somewhere in the middle, right, with extreme scores at the upper and lower quartiles, or the extreme ends of that bell curve. So, as it is supposed to be administered, you're supposed to check statistically to see if your distribution, if your sample is normally distributed or meets these assumptions. But in the case of some assumptions, say, for example, homogeneity of variance, which is something that goes beyond the scope of this course, but may be touched on in statistics courses you've either taken or will take in the future, or are taking concurrently, this assumption is often just assumed to have been met. Well, this is a pretty egregious violation when you're conducting an experiment or a statistical analysis to just assume that one of the assumptions is met. Because if it's not met, and the data can't be in any way transformed to meet that assumption, then you can't run that type of analysis.

So, it stands to reason that violated assumptions of statistical tests are one major threat to statistical conclusion validity. And so, violations of statistical test assumptions can lead to either overestimate or over it, or overestimating or underestimating the size and significance of an effect. For this threat, we can think of these statistical tests as being either overlooked or inappropriately assessed. And so, if you think when reviewing through articles or conducting your own research, that for any reason that an assumption for a test has been violated, then you can argue that this is really a threat to validity, to the validity of the research.

The next threat is known as fishing and the error rate problem, right? So, when we, oftentimes when researchers in the social sciences are attempting to find significant results, the kind of proper channel for how to engage in social science research is to generate a theory, and then a hypothesis based off that theory, and then to really test it once. However, also something that you might find in your statistical classes or statistics classes is that if you continuously fish through the data to find a significant result and test it again and again and again statistically, you're increasing the likelihood that error will occur. The error will enter into your data, and that it will cause a, it will cause a result to be seen that might not be reflected in the population. For example, normally, we, as kind of an accepted and somewhat unofficial standard within social science research, we accept that a margin of error of 5% is one that is acceptable when engaging in social science research that is quantitative in nature. But this is 5% that gets inflated each time we run an analysis. So, if we're running an analysis again and again and again on the same data, that 5% margin of error that keeps coming into play, odds are, you are going to enter into, you are going to have a result that is artifactually inflated, right? So, what this means is, is that by engaging in these analyses again and again and again, what you end up coming up with is a result that is more likely to include error than if you engaged in simply one analysis. And this is based primarily on probability theory.

Then another major issue that comes into play is unreliability of measure. So, measurement error weakens the relationship between two variables and strengthens or weakens the relationship among three or more variables. This issue won't come up with widely validated measures like tests of intelligence, the way that tests IQ among adults, the Beck Depression Inventory or BDI, which is one of the major psychiatric screeners for depression, or the Minnesota Multiphasic Personality Inventory or MMPI, which tests for personality disorders and personality dysfunction. These have been widely validated and normed on thousands upon thousands of people. And so, we can say that these are reliable within an acceptable margin of error. But when generating your own measure, your own survey, it's very important to validate that measure and make sure it's reliable first before including it in your study. And so, this will have, this would necessitate engaging in many kinds of pre-analyses on your metric to make sure that that metric is actually reliable and valid. But as we've seen far too often in social science research, and one of the major things that comes up among researchers, both in, in psychology and sociology, especially at the graduate level and undergraduate level, when one creates their own metric or adapts a previously reliable metric to answer a new question or changes that metric in some way, it may no longer be reliable and valid. Therefore, you can't generalize the results to the population that you're studying unless you show that with an acceptable margin of error, that it is actually reliable and valid. And unfortunately, this is often overlooked, especially for either undergraduate and graduate level researchers who are trying to make their first mark in the academic and scholarly publication world. And so, it becomes very important to make sure that one's metric, if they generate it or if it's not widely used, is both reliable and valid.

Often times, as you may have noticed when writing your term papers for the course, or when assessing, for you're assessing articles for your term papers for the course, that you may find a tremendous amount of statistical jargon that relates to Cronbach's alpha. And what does this mean? Well, this is looking at the internal consistency and reliability of a metric, right? If Cronbach's alpha is assessed on the result, and if we engage in an assessment of test-retest reliability, where we look at the consistency between two different samples tested on the same metric, then we can with some degree of assuredness or certainty determine that a measure that is deemed to, that is not tested previously, can be deemed to be reliable presently. And so, this becomes a very important part of research and a very important statistical data point to report.

Then our next question relates to restriction of range. How is this a threat to in statistical conclusion validity? Well, reduced range on a variable usually weakens the relationship between it and another variable. This often occurs when all the participants of the study have very similar scores on a measure, for example, everyone having low self-esteem that's within one standard deviation of each other. This may be a threat to validity because the distribution of scores are just so similar to one another. Think about how this would affect the normality of a variable. If the range is restricted, even if the sample size is very high, or there's a high "n" as it were, this may still be a problem, even with a large sample size. So, it's important to consider how similar are each individual data point's scores.

Then we have, very similar to unreliability of measures, we have unreliability of treatment implementation. So, if a treatment that is intended to be implemented in a standardized manner is implemented only partially for some respondents, effects may be underestimated compared with full implementation. So, this is similar to the unreliability of a measure. Think back to what I just mentioned about adapting a measure to a different population, changing the questions around from a previously reliable measure into something new. Well, if you implement a treatment in a way that it wasn't designed to be implemented, then it's not necessarily going to be reliable, is it? Well, if the measure is not standardized or is administered in a standardized way, then it may be unreliable, which affects the measure's validity, and thus the variable's validity. If you make the argument that instrumentation is a threat to internal validity, you likely also want to make the argument that this is a threat to statistical conclusion validity as well.

Then there's also extraneous variance in the experimental setting. So, this has to do with those confounding variables that we've talked about throughout the semester. So, some features of an experimental setting may inflate error, making detection of an effect more difficult. If you make an argument that instrumentation is a threat to internal validity, or that interaction of causal relationships with settings is a threat to external validity, then this is also likely a threat to statistical conclusion validity. Basically, if there's something about the environment that the study was conducted in that makes you think you may change the participants' outcome, then this is also a threat to statistical conclusion validity. So, think about research conducted in the field, right, not in a controlled laboratory setting. There are all sorts of extraneous variables that can come into play when you can't control all of the aspects of the study. So, this could introduce extraneous variance into the experimental setting, thereby changing the relationship between variables and changing the results of your study, such that if you were to engage in the same study in a laboratory setting, you could have completely different results. So, when you think of confounding variables, think of extraneous variance in experimental settings.

The next one is heterogeneity of units. So, what does this have to do with? Well, increased variability on the outcome variable, or the dependent variable as it were, within conditions increases error variance, making detection of a relationship more difficult. In essence, you're increasing the likelihood of error and making it more likely that you won't be able to detect an accurate relationship. This is sort of the opposite of restriction of range. This is going to mean the assumption of homogeneity of variance, which is so important in so many parametric statistical tests, is likely violated. So, if we're assuming that homogeneity of variance is met and we're not testing for it, then if this is the case, then homogeneity of variance is likely violated, meaning that the results of your statistical analysis are likely inaccurate. So, think about it this way: if scores are all over the map, there's high scores, there's low scores, there's scores in the middle, and there's no uniformity to these scores, there's no relationship between these scores, then it's impossible to find a significant result because scores are just too spread out. So, if you notice that scores seem to be very discrepant from one another when conducting research, this is likely a threat to validity.

So, in order to talk about this final threat to statistical conclusion validity, we need to define two statistical terms that hopefully are already familiar to you from your statistics courses. The first one, the first one is effect size. The second one is statistical power. Effect size is a quantitative measure of the magnitude of the experimenter's effect. The larger the effect size, the stronger the relationship between the two variables. Statistical power is the probability of a hypothesis test finding an effect when one indeed exists in the population. So, now that we have these two defined, how do we understand the threat of an accurate effect size measurement? Well, according to a study conducted in 2011 by Brand and colleagues, the reporting of exaggerated effect size estimates may occur either through researchers accepting statistically significant results when power is inadequate, and or from repeated-measures approaches, that is, aggregating and averaging multiple items or multiple trials of an event that's being tested. So, this could lead you to inaccurately understand the, the effect, and seeing an effect that either does exist when none actually exists, or not seeing an effect that exists when one actually exists. So, this really cuts to the core of accurate statistical analysis and accurate statistical results.

Alright class, so that concludes our lecture for this week. If you have any questions, as always, please do feel free to reach out to me either by direct messaging me by inbox or via my La Vallée College email. And also, be mindful of the discussion questions that are due for this week and the previously posted prompt for the second part of the term paper. Alright guys, take care.