📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

U1 L4: Collecting Sample Data

OCA Statistics9:56

Transcription

Today's lesson is about collecting sample data. We'll learn about types of studies and experiments, controlling the effect of variables, randomization, types of sampling, and sampling errors.

Basics of collecting data. Statistic methods are driven by the data that we collect. The first question you need to consider is whether you want to make observations only or somehow modify the subjects. So take these two examples. With a typical survey, we observe subjects by asking them questions, and we don't modify the subjects in any way. That's called an observational study. Versus with a typical clinical trial of a drug, we modify subjects by giving them the drug or a placebo. This is considered to be an experiment. We are modifying our subjects.

Types of observational studies are based on the time period. So a past period of time can be studied in a retrospect or case control study. We go back in time and collect the data over some past period. For one specific point in time, there's a cross-sectional study. This, the data are measured in one point in time. An example is an opinion survey conducted this particular week. And looking up at a prospective study or a longitudinal or cohort study. Go forward in time and observe groups sharing common factors, such as smokers and non-smokers.

As we said before, in an experiment, you're going to be modifying your subjects in some sort of way. You want to first control the effects of the variable by using a completely randomized experimental design, a randomized block design, a regular rigorously controlled experiment design, or a matched pairs design, which will explore all of those in depth. And we want to use a replication by repeating the experiment on a sample size that is large enough to let us know the true nature of its effects. And third, we want to use randomization to assign subjects to different groups.

The different types of methods of sampling include random samples, simple random samples, systematic sampling, convenience samples, stratified samples, and cluster samples. A random sample is when each member of the population has an equally likely chance of being selected. Computers are often used to generate random telephone numbers. So, um, in a simple random sample, a sample of n amount of subjects is selected in such a way that every possible sample of the same size n has the same chance of occurring. See a picture of a computer here developing all of these phone numbers. It's like a random number generator. There's no rhyme or reason to it, no system to it. It just starts to randomly select numbers.

Systematic sampling means that we are going to follow a system to sample each case element in a population. So we would start at a specific value, say we start at the third value, and we decide that we want to sample every third person after we started number three. So that means we would take the third person, and the sixth, and the ninth, and the twelfth, and so on, every third person in a population or element, whatever you're studying, in a systematic way, the same numbered term each and every single time.

A convenience sample is just that. It's very convenient to use. You see this person just yelling out their window, "Hey, do you believe in the death penalty?" just to a random walker. This random walker represents the people that actually get out and walk a dog. So you don't have a method to being very random. What about the people who don't have dogs? How are you going to pull them on the question of, "Do you believe in the death penalty?" So it's very convenient for you. You're just picking people that are close by.

Stratified sampling divides the sample into two or more strata or parts, subgroups, at least two of them, so that the subjects share the same characteristics, and then you draw a sample from each subgroup. So let's say that we separate Democrats and Republicans, and then I want to take a random sample from the Democrats and a random sample from the Republicans. So two subgroups of the population divided into two, and a sample from each subgroup.

Stratified and cluster can often be accidentally interchanged when they're not supposed to be. Cluster sampling means we divide the population into sections or clusters, and then randomly select some of the clusters. So stratified, we're taking elements from each different subgroup. Versus clusters, we're dividing it into clusters like history classes, by science classes, and then polling, picking one of those specific clusters of choice and choosing all the members from that particular cluster. So polling all students in a randomly selected class.

No matter how well you plan to execute the sample collection process, there's likely to be some sort of error in the results. And the sampling error is the difference between a sample result and the true population. Such as error results from chance sampling fluctuation. Non-sampling errors. This occurs when the sample data is incorrectly collected, recorded, or analyzed. So it's not about the sample, but it's about the data that was collected.

Let's practice using these definitions and see if we can apply them to some problems. A survey was taken to determine how many deer were killed in Michigan during hunting season. 10 counties in Michigan were randomly chosen, and all used deer tags were counted. Which sampling method was used? So basically, we divided Michigan into counties, and then from those counties, we picked a couple. We picked 10 of the counties, and then every single deer tag was counted from those counties. So we clustered them together. We took out, picked some clusters. Remember, with cluster sampling, the population is divided into sections, and then the sections are randomly chosen, and all members of that section are chosen. So again, the counties were considered to be our clusters.

In a table tennis ball production line, 20,000 table tennis balls were produced. The first 100 table tennis balls off the line were taken and tested for roundedness. Does this sampling method result in a random sample? Since we're only taking the first 100 table balls, we do not have an equally likely chance of each ball being selected. Not all marbles have the same chance of being selected. The first 100 are guaranteed to be selected, and the rest have no chance of being selected at all.

A completely randomized experimental design is whereby randomness is used to design, assign subjects to the treatment group and a placebo group. Complete randomness. Blinding is a technique in which the subject doesn't know whether he or she is receiving a treatment or a placebo. Blinding allows us to determine whether the treatment effect is significantly different from the placebo effects, which occurs when an untreated subject reports an improvement in symptoms. Double-blinded means that blinding occurred on two levels. First, the subjects don't know whether they're getting the treatment or placebo, and the doctors who gave the treatment and evaluated this result didn't know either.

Replication is the repetition of an experiment. A study was set up in which participants with chronic headaches were given a new medication. Some participants received a sugar pill without their knowledge. The experiment, the element of experiment design is known to be which part of this study? That would be blinding. They didn't know what they were getting. They could have received the treatment or the placebo.