Transcription
Hello! Welcome to another Test Oil live stream. This one is on the lubrication-related failure of a bearing. I'm Evan Zabowski, I'm the senior technical advisor here at Test Oil, and the goal of today's presentation is to answer one of these questions we often sometimes, at least, face; maybe not often, but also show you some snippets again of what's covered in our comprehensive training.
So, what I'll cover today is I'll show you a little bit about what we cover in lubricant analysis, and then I'll show you the report in question. After that, I'll go over a few things that we cover in our "How to Read a Report in Two Minutes or Less" section, before showing you the final answer.
Now with this section here, the lubricant analysis section, we tend to go into a lot of detail regarding the tests and the test results, so I'll only show you a couple of those for the sake of today's presentation. But I'll begin with one of the most common ones, and that is elemental spectroscopy. In our comprehensive course, we do go through why this test is done and how this test is performed.
One of the things we cover there is that your samples are diluted with a solvent and injected as an aerosol into a plasma. When your sample enters this plasma, the electrons in it get excited; or, another word for this is ionized. What we do is we measure off the light that is emitted when those electrons return to their normal states.
So, essentially what happens is, for lack of a better word anyways, your sample's kind of like it's burned. As you know, when we burn things, we give off energy; we give off light. Well, it's kind of the same thing from that perspective. The neat thing about this is every element on the periodic table gives off a unique spectrum of light, so that is what we're actually measuring.
Just to give you an example of what this instrument might see, here are three different, very similar, physically appearing metals: aluminum, tin, and lead—all kind of grayish-silver colored metals—but to this instrument, as those elements are ionized, they would give off radically different signatures, and we can analyze for those.
Now one of the reasons we mention all this detail when talking about how the test is performed is because the way the test is performed centers mostly on zero to five micron particles. It measures them very accurately, but anything larger than that, and the accuracy takes a bit of a nosedive. Somewhere around 10 microns, this instrument is no longer able to detect those particle sizes accurately, and that's an important detail you'll need to remember for some of the other tests here.
But anyways, when we talk about this test, we talk about the results, and two of the results I'll draw your attention to are calcium. Calcium is one of the elements we measure because we expect to find it as an additive if the oil is advertised with detergents. The other element that we often measure for, from a similar perspective, is phosphorus because it is also an additive.
But in this case, as you can see below, one of the things we mentioned is that phosphorus is one of those tell-tale elements that can tell you what basic additive chemistry is present. So anti-wear or AW oils tend to have both zinc and phosphorus in relatively equal proportions. EP oils, on the other hand, only have phosphorus present; they do not contain high levels of zinc.
Then R and O, or rust oxidation oils, tend to contain very, very low to no discernible levels of most of the additive elements that we're testing for. So funnily enough, they almost look like they don't have additives, even though they do; it's just they're made of chemicals that don't contain the elements that are typically measured.
Now with phosphorus, one of the other things to note, though, is that phosphorus is also used as an EP additive in greases. It's one of those things that we're always looking for: Do you expect to find an element present? And what do you do if you don't expect to see it?
So when you don't expect to see an additive element present, you must, of course, locate its source. And often, we're told it's really just another lubricant that's provided it. But the other test I wanted to make sure we covered in some detail here was particle counting.
Now particle counting can be done one of two ways, and most people are quite familiar with optical particle counting, or what's sometimes referred to as laser particle counting. In most presentations you've seen, probably have explained that a particle in the oil, as it's traveling past a light source, casts a shadow, and the instrument counts and sizes the shadows. That is the basic principle behind optical particle counting.
But one of the details we go through in our comprehensive course is to mention how that shadow is sized. It is not sized by the shortest measurement, the longest measurement, or even the average measurement in any one dimension. What it does is something called spherical equivalency. When particles are measured, they're assumed to be spheres, and therefore when their shadows are measured, the shadow is assumed to be a circle.
So when the shadow is being measured, it is being measured by area. What we do is we have to use the instrument to figure out, based on the area, what size of circle or what size of sphere that represents. To give you a quick example, if you had a 1 by 10 micron particle go by, it would, of course, cast a 10 square micron shadow—that's easy math. But then the question becomes, "10 square microns, how big of a circle would create 10 square microns in area?"
To save you scratching your head too long at that math because it's not math you typically do in your head, it's about 3.6 microns. So, in effect, a 1 by 10 micron particle going past the light source, casting the 10 square micron shadow, the instrument would therefore count it as a 3.6 micron particle. So, as you can see, optical particle counting is not directly measuring particles as you may have once thought, and this is an important detail that we like to point out.
Another detail I like to point out, though, is that particle counting doesn't report anything below 4 microns. So that 3.6 micron particle is actually not being reported, even though it was detected; it's not going to appear on the report.
One of the other key things to note about optical particle counting is not just the four micron and up limitation to its measurement, but the interferences, because this instrument relies on counting and sizing shadows. Anything else that can count or show up as a shadow—I should say—can count as a particle; so water molecules, air bubbles, and even soft contaminants like oxidation byproducts, you know, that eventually form varnish, have an ability to cast shadows unless the sample has been properly treated. Typically, that means it's been diluted with a solvent and then has been degassed or vacuumed or sonicated to try and remove as many of these non-particles from the sample.
Now, most labs do actually perform all these steps, but it is extra labor. My point in mentioning this, though, is because at Test Oil we prefer to use a different type of particle counting as our default method, and that is one called pore blockage.
Pore blockage doesn't rely on the same technology at all; in fact, it uses what appears to be a very, very fine screen. But what we're doing is we're not just using the screen to sift out the particles; what we're doing is using the screen to act as resistance. So, what we do is we push the sample through the screen at a constant pressure. What we're actually measuring is the decay in the flow.
When you think about this screening, which is going to be calibrated with all identical size holes, and they're either 5, 10, or 15 microns, imagine that it's a 10 micron screen. Well, what size particle is going to resist going through the screen? It's probably longer than 10 microns in any one direction. But when you think about these larger than 10 micron particles not going through the screen, you have to imagine that they're not necessarily going to block a hole completely.
What happens is the holes will get smaller and smaller and fewer and fewer. Therefore, how the flow decays tells us what size particles are piling up on the screen as it restricts itself. So from this measurement, again, we're not necessarily measuring every single individual particle in the oil, like I said neither does optical particle counting. But here's why we favor it: When you look at the limitations to this test, they pretty much are the same, that it doesn't report any particles smaller than four microns, but as for interferences, it doesn't have any.
That's really one of the best benefits of doing a pore blockage particle count; it doesn't require all the special preparatory steps to the sample, so it can be tested neat. We don't need to dilute it; we don't need to do anything that might modify the particles that are in there. Because when you take a sample and you put it through sonication, if you sonicate it too long, you can actually break up larger particles into smaller particles and give a slightly different view of what's going on.
Now, one last test I want to mention that we do cover in our training is analytical ferography. At Test Oil, this is one of the tests that we automatically add to your report anytime we detect high wear metals or high ferrous wear concentration. What analytical ferography is, is really one of the most forensic-type tests we can do on an oil sample.
It begins like this: we take your oil and we flow it down a glass slide—basically a microscope slide. The slide's at an incline so the oil will flow downhill on its own. Now particles will separate out from largest to smallest based on fluid dynamics. However, when you look at this picture, you're probably noticing they form these cute little straight lines, and you're thinking there's no way that happens naturally. It doesn't, which is why we actually have magnets on either side of the slide.
As the oil is flowing down, these magnets will align all the ferromagnetic particles into straight lines, while all the non-ferrous particles will just randomly distribute. So with this preparation, what we call creating a ferrogram, we can now take that ferrogram to a microscope and look at it through the microscope. We can do a human-based analysis on it based on morphology; the size, the shape, and even the color of the particles can tell us how severe the particles are—being larger than clearance size or smaller. We can tell what mechanism created them, so we know what kind of wear or other degradation mechanisms may have been taking place. The color can give us some hint as to the actual sources; we can tell high carbon steel from low carbon steel, or copper versus a brass or bronze alloy, things like this.
Now, just to be well-rounded in our presentation here, the limitations to this test is it is a subjective human-based test, but the nice thing is it does not have any significant interferences.
So now let me go through the actual report that I'd like you to see. This report here, I will make note, is from a failed bearing. So this one has actually failed already. We're not here to decide if there's a problem; we know there's a problem. We know this bearing has failed. What we're trying to do is see the report from the perspective of can we figure out why it failed and how it failed.
So with this, I'd like to show you a snippet from one of our other sections of the comprehensive course, the "How to Read a Report in Two Minutes or Less." In this section here, we cover quite a bit, and I'm going to reiterate some stuff I showed yesterday in the live stream. That was basically what are the main steps for reading the report.
Step one is simply to read the report from top to bottom—the whole report, every last piece of information. Step two was, as you get to the data section and you look at the analysis results, not to dwell on any interpretation while you're doing that; just decide if the value is abnormal or normal. When the value is abnormal, all I'm asking that you do is classify it as either a cause or an effect. When we get to the end of the report, all of our causes should be able to be grouped together as either one or two causes, and everything else in the report is merely an effect.
What the last step entails is actually looking at those causes and trying to see if we can figure out what to do next. The "what to do next" is often either a confirmation of a reasonable thought that this is probably why that is happening or sometimes it's elimination—we're looking to eliminate a possibility so we can focus our efforts elsewhere. On rare occasions, we sometimes know exactly what's going on that's wrong and we know how to fix that.
So with those three steps in mind, let me walk you through that report in its entirety using this technique. When we say "read the whole report," we do mean read the whole report. So we start at the very top left corner of this report and we start reading. We see here it says it's a plain bearing; it doesn't have some size. We don't really know how big this is, other than there's five gallons. So five gallons of oil, you know, it's not a small amount; it's a reasonable amount—takes a pail of oil.
So in your head, maybe now you've got an idea of what size of bearing this is. But when you look at the oil type, it says it's an R and O 68. Well R and O, from what we discussed previously in the phosphorus section, means that we're probably not expecting any additive elements that are visible, and any that there are will be in very, very low concentrations. The 68 merely means that the viscosity at 40 degrees will be 68 centistokes.
Now if we keep reading the report, you'll see that it has some comments to it, and it identifies the problems. It says this one's got excessive wear and a high particle count. That is probably not that much of a surprise considering this is a post-failure sample—that there is a lot of wear.
So with this one, we now go down to the next section of the report, but still above the oil analysis data. If you look here, we've got these sample dates which are very inconsistent. There's 14 months between the first sample and the second sample; there are about 18 months again, and then just a matter of a few weeks.
So with this type of erratic sampling, it's hard to predict what kind of trends we're going to expect to see. But just as a reminder, when you look at an oil analysis report, there is a new oil column. This is your baseline information—what oil looked like in its virgin state, basically as best as possible. The most current result is this one here that we're looking at, and this is the one that is post-failure. This is the one that has been taken after the bearing has already failed. So it should be no surprise to us that we're going to see a great deal of wear metals because this bearing did mechanically fail.
But as we go down through the rest of the report, now we look at some of the things like lube hours and machine hours and see that they've been inconsistently filled out. We're not getting a lot of good data here, so once I get into the actual oil analysis data section, I've got no easy way of predicting what I'm expecting to find other than at least the last sample is going to look terrible.
So once we start into the oil analysis data section, we look at the data here and we see one of the first things that comes up is, "Wow, that is really, really high iron." But is that a cause or an effect? Well again, wear is usually just an effect; something is making it wear. So why did it fail? Well, we'll have to keep reading. So we'll just log iron down as an effect. Same sort of logic goes with the copper. Copper is just, again, another wear metal, so this is something that we look at and say just an effect—keep going.
We keep going down; we can see that lead is also quite high and identify that as an effect. Luckily, we can skip aluminum; that seems to be a fairly flat trend. But once we get down to the tin here, whoops, the tin is very high, but tin, yet another wear metal, is yet another effect. So we're not really seeing anything that identifies to us that there’s a definitive cause yet.
But if we keep reading, eventually we'll get down here to calcium. Calcium will stand out because it does not match the new oil line and it seems to be increasing in the last couple of samples. Now that is something that typically indicates mixing with a different lubricant. So that is one that we can probably say this is a cause. If we keep going down the report, we'll see that phosphorus also looks very, very similar in that it is appearing when it shouldn't appear at those levels and it is increasing, and again this will be another cause.
As we keep going through the report, we'll eventually need to scroll on down to the bottom half of the first page here, and what's going to jump out to us as we get down through the data is that most things look pretty good until you get to the actual particle count; and once you see the particle count here is very, very high—in fact, it has maxed out the instrument—so this means that we've got a particle count that's so high it actually can't be counted.
Now the tricky question here is whether it's a cause and effect, and that's hard to discern because we know there's a pile of wear metals in there, but we also know that wear metals, as they're measured, are typically in the zero to five micron range. Particle counting is four microns and up, so they don't necessarily overlap that significantly. So it's hard to say that that particle count is entirely wear metals; maybe it's something else.
So if we keep reading the report, we'll see nothing else that's really amiss on this page, and that's what's kind of curious. The viscosity is almost exactly 68—it's almost exactly where we expect to find it. Yet we identified calcium and phosphorus as additives that shouldn't be there, so that seems to be a fairly clear indicator that we've mixed a different lubricant into the system, yet we have not affected its viscosity.
Now once we get to the last page of the report, we will see the analytical ferography results, and we'll see that quite a few things were noted here at very, very high levels and in very, very large sizes. So this is very severe, but again, not really that much of a surprise considering this has already failed.
So when we look at the rubbing wear as an example, again we assign that it's an effect; we can do the same thing with sliding wear—sliding wear, it's just an effect. But when we get down to the bottom, the last one here, dust and dirt—that is when we can probably group as a cause.
Now what's interesting about this one is it might not be related to the calcium in the phosphorus. Dust and dirt typically show up as silicon or silicon and aluminum, so this has nothing to do with the calcium in the phosphorus. This causes us to have to return back to that part of the report where the particle count was found—right here where we said we identified it being very, very high.
Now knowing that there's very high levels of dust and dirt, we can probably classify the particle count as another cause—that that's related to the dust dirt levels. So the final tally, when we get to the end of the report, would say to us that we've got calcium and phosphorus, plus dust and particle count. But everything else we noted on the report was simply an effect.
As I said, we group these together, and we can group like things together. So we'll throw a particle count on this list, but we'll group the calcium and the phosphorus as one thing, and the dust dirt and the particle count as something else.
Now one of the reasons that we group things into causes and effects is we don't focus on the effects; we focus on the causes. This is the section we want to pay attention to. This is the section we must investigate.
One of the stories I like to tell during training is the story of Occam's razor. Occam's razor is just this basic principle that when looking for an answer, you often look at the simplest solution because the simplest solution is often the most correct. You don't try and get into a complicated "if-then" statement for all the sequences of events that happen to make something occur; you look at the simplest explanation for, "Well, how do we think we got away? You know, got there?"
So with this, there's a very famous quote here from Sherlock Holmes that many people paraphrase over the years, but basically stating that if you eliminate every other possible explanation, whatever remains is probably the answer. Where I like to take this a step further is to tell you another bit of an anecdote, though; one to make you think outside of the box.
I have a variation on this tale called Occam's toothpaste. Right? If he's got a razor, then he probably has got some toothpaste. As the story goes, there was a toothpaste factory that had a big issue with shipping empty boxes—right? Product not being in the boxes, so that was costing them money.
As the story goes, the CEO spends millions of dollars upgrading the facility, and what he did was he installed some scales on the production line. After the box is supposed to be filled, the boxes weighed, right? If it weighs something, then there's a tube of toothpaste in the box. If the box is empty, though, the light boxes will trigger an alarm, and therefore, somebody can go over, remove the box, and restart the production line.
So this seems to solve the issue; the toothpaste company is no longer shipping empty boxes. Everybody's happy. But once they start looking at their metrics, their KPIs, they realize that yes, they're catching hundreds of empty boxes, and things are looking better, but after a few weeks, though, there's no more alarms.
So the CEO is concerned now that somebody has circumvented his great system and is not, you know, allowing the system to work the way it should, which is why we're not counting anymore empty boxes. When he goes down to the shop floor, though, to see why things are doing what they're doing, he sees there's no boxes in the reject bin. So this is now very puzzling.
So what does he do next? Well, he looks a little bit up the line and what he notices is there's a very cheap desk fan just blowing across the line. What he sees is that the empty boxes are so light, the breeze from the fan just blows them right off the production line before they hit the trigger to stop the line and, you know, sound the alarm.
So he asks an employee, "Hey, what's this fan doing here?" And the employee says, "Well, you know what? I just got sick and tired of having to restart the line and go remove a box and all these other things." Right?
So sometimes the simplest solution is, you know, deadly obvious. But the other thing is, sometimes there is no box.
But anyways, to return to this example, I've tried to give some prompts along the way, and there's reasons I mentioned some of the testing just before we get into this. But you have to think, "What would cause all those events? Calcium, phosphorus, dust, dirt, particle count? What would cause all those to appear and, of course, lead to a failure?"
So, you know, just to reiterate: calcium, phosphorus—you need the, you know, to think that probably means a different lubricant got in there, but no change in viscosity. Remember that—that's the key. We saw the dust, dirt, we saw the particle count; we know that that's going to lead to wear.
So you have to think whatever the lubricant is, it's contaminated.
Well, where most people go with this much information is they think then, "Well, you know what? We have to think of something that would cause a very sudden and catastrophic failure," and it's probably not this. Right?
Many, many times, if you tell somebody, "Look, there are signs of the wrong lubricant; there are signs that it's dirty," people go, "Wow, they probably grabbed the wrong jug; they used a dirty funnel; there’s cross-contamination; you know, they pulled from the wrong drum." They can come up with a million excuses to come up with a wrong lubricant, coupled with dirty lubricant.
But what you need to appreciate is two things. One is it has to be the wrong lubricant of the same grade because it did not cause a viscosity change. And it has to be so dirty somehow that it did cause a sudden and catastrophic failure.
So the suggestion of grabbing the wrong jug or a dirty funnel or anything along those lines doesn't seem to add up. It's not enough to cause, within a few weeks, a bearing to outright fail.
One of the other snippets I'd like to quickly show you from our comprehensive training, again from the "How to Read a Report in Two Minutes or Less" section, is simply the discussion about who should interpret the reports.
Of course, a lot of people feel like labs should interpret reports because that's what you're paying them for. But as I point out, the lab's expertise has some limitations. Now they know all about test limitations; they know what a test can and cannot tell you, how much variance, you know, the precision, the bias; they know all these things. They know about metallurgy of equipment; they know what wear metals to expect from what sources— you know, that's pretty common—and they also know how to trend. They know what looks normal versus what doesn't look normal; that part they've got.
So they can do this. However, my argument is that there’s somebody else who can do an even better job, and that is typically the end user. What does the end user bring to the table? Well, they probably have many of the same traits that the lab does, but they know about wear and how the sample was pulled.
So any anomalous results they can discount as just being a bad sample and move on. They know how the equipment is configured. Right? They know if it shares lubricant with another device, they know if it has filtration on it—anything that might skew the results.
And here's the key for this particular example: they understand its operating conditions. Right? They know if it works in a dusty environment, they know if it's running hotter than normal; they know these things.
On top of that, though, another key thing to this particular example is they know about other maintenance activities. They know when the oil was changed; they know if there's been filtration done to it; they know a lot of things about the piece of equipment.
So when you revisit this example report and you think about, "Okay, if I was the end user and I knew things about this, would I come up with a different answer?"
So imagine that, of course, the end user knows everything the lab does. They know that calcium and phosphorus indicate a different lubricant; they know that dust and dirt and particle count indicate contamination; they know there was no change in viscosity; they know it caused a sudden catastrophic failure—that part is shared knowledge.
What the end user can add to this equation, though, is other than everything that was previously described, we can add that this thing operates in a very dry environment, but it's very, very dirty—lots of airborne contamination—and this bearing housing has seals on it. Those seals are greased.
We can be very specific and say this is even the grease they’re using. Now I know you can't read the fine print on this, so I'll just zoom in here real quick. But if you look closely, it says it's a calcium sulfonate grease, and it's an EP.
That explains the calcium; that explains the phosphorus. The fact that grease mixes in with the oil would explain where calcium and phosphorus would come from, and if it's only a few drops of grease kind of thing, it could explain why the viscosity didn't change.
But you're wondering, "How does this all mean that the bearing failed?" Well, the last thing I need to show you is a picture of what the bearing housing would actually look like.
So there is the exact same brand of bearing that was used on the report; it's not the bearing. It's an exploded view from the manufacturer so you can see things a bit more clearly. But you can see where the five gallons of oil goes inside the bearing housing, but you can also see at the top edge on both sides a zerk fitting for where the grease goes to grease the seals.
So, in effect, probably what happened with this bearing is it was over-greased. Over-greasing blew out the seals; therefore a little bit of grease migrated into the oil. That didn't cause the catastrophic failure. What did, though, was the lack of seal.
Now, because the seals were compromised, this dirty, dusty environment was allowed to get into the bearing and caused extensive abrasive wear and eventually wiped the bearing.
So, there you have it—just another quick example of some of the stuff we cover in our comprehensive class. I'll now open it up for any questions. If you're watching this live and want to throw a question into the chat, I should be able to see that. If not, I do thank you for your time and for watching this live stream or the recording later.
So, again, thank you!