📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

U3A L3: Histograms

OCA Statistics9:57

Transcription

In this lesson, we're going to discuss histograms and how they show distribution of data. We'll be able to understand and interpret those results.

When constructing a histogram, we need to take into account a horizontal scale and a vertical scale. So, the horizontal scale is going to use class boundaries or class midpoints to construct each of the bars, and the vertical scale will use the class frequency.

The relative frequency histogram is going to look very similar to the regular frequency histogram. It's just going to have relative frequency as the vertical axis instead, meaning our percentages.

Now, how to interpret a histogram? Well, if we have what looks to be more like a bell-shaped characteristic, and we call this a normal distribution. So, we would have our histogram, and it looks pretty much uniform, a mirror image of each other about the median. So, this is what we mean by a bell-shaped curve. This is where the frequency increases to a maximum and then decreases again, and it is symmetrical. As I said, this is a roughly bell-shaped curve.

So, which of these is a criteria that can be used to determine whether the data depicted in a histogram have a distribution that is approximately normal distribution? The answer would be A: the distribution bunches up in the middle and tapers off symmetrically on either end. So, you're always looking for those keywords: symmetric about that center.

Um, if the distribution bunches up at either end and tapers off in the center, we would have the opposite of a bell curve. It would look like a U. That is not a normal distribution. The distribution is relatively even from one end to the other. Again, it has to taper off on either end and be symmetrical, and be more stacked up in the center. The distribution bunches up on one end and tapers off to the other end. That would be skewed either left or right, not a normal distribution.

The population of ages at inauguration of all U.S. presidents who had professions in the military is 62, 46, 68, 64, and 57. Identify the false statement about the data.

Well, let's start with B: the histogram will be a repeat of individual numbers with such a small set of data. If we were to draw a histogram, it would just be one of each number, not really a great way to make classes here.

The data set is small enough that the individual eight individual ages can be examined. That is absolutely true. We can examine just five data points just by looking at those data points.

D: The data set is not large enough for a histogram to reveal the true nature of a distribution. That is true. Um, the data set is not big, it's very small, it's only five points, and if we were to plot those, we just those five points, like we said, it would just be a repeat of the individual numbers and would not give us a visual of the of a representation of the numbers.

So, the answer is D or A. The data set is too small to be made into a histogram. So, we could create a histogram, but the data is small enough to be examined by itself. It's possible to make a histogram, and it could look like this.

As one example, listed below are amounts of strontium 90 in a simple random sample of baby teeth obtained from Pennsylvania residents born after 1979. Construct a histogram with the data.

First, let's acknowledge that we have 40 data points total, and our minimum value of 114 and a maximum value of 188. What we're going to do is create a frequency table. So, the amounts of strontium 90 will be on the left, and the frequency of those would be on the right.

We need to come up with a number of classes. Well, if my minimum is 114 and my maximum is 188, um, if I make my class, let's say that I want to go from, start at 110 and I will go in 10 increments. Um, so our minimum of 114 minus 188 and divide that by, let's say I want eight classes. Um, sorry, make that maximum 188 minus 114. If I want eight classes, remember our class width should be somewhere between 5 and 20. Are the number of classes? So, if I chose 8, I get 9.25. So, if I round that up, 10 should be my class width.

So, if I go from 110, then to 120, 130, 140, 150, 160, 170, 180 are my one, two, three, four, five, six, seven, eight lower bounds, and those are gonna go to 119, 129, and so on. And then I'm going to tally up each of these numbers to find out which class do they fall into to help me create my frequency table, which will help me create, in turn, my histogram.

So, in my class of 110 to 119, I see two data points fall into that class. From 120 to 129, I see one, two more. And I'm going to continue this process until I've gotten through all of my classes. 20. And ideally, these are going to look and be equal widths. I'm doing this by hand, so it's going to be a little not perfect, but you get the idea. And we'll go all the way to 200.

Okay, now our frequency just goes up by, we could say it starts at 0, goes up by 2, 4, 6, 8, and 10. And in my first class from 10 to 120, I have only two. From 20 to 29 or 30, we have another two. 30 to 40, I have five. 140 to 149, I have nine. 150 to 159 or 160, I have 12. I need to add on more up there. Then we go back down to six, then two, and then one.

So, this is a relatively bell-shaped curve, approximately bell-shaped curve. And once I've done that, now it's time to create the histogram. Remember, the frequency is your vertical, and the horizontal will tell us, in this particular case, how much strontium we are measuring.

And now we're going to make our bars for each of the classes. In the first class, there's two. And the second class, there's two. Third class, there's five. Then nine. Then 12 or 13, sorry, 13. Six, two, and one. And that's how you make a histogram.