📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Clase 6/10

Giordano Pérez Gaxiola40:21

Transcription

Alright, let's begin. We're going to change channels. Last week, we were looking at treatment issues, part of the introduction to all of this, and today we're going to start with diagnosis. I always like to show this initial image because if you look at these three cars, these three trucks, they appear to be of different sizes. If we don't consider the context, what's around these images, we might think they are different, that they have three different sizes. But once we put them together, once we realize that all three are the same size, why do they look different? Because of the context. I put this here because many times the individual context of the patient, the probability of a patient having a disease, is ignored, and people try to interpret the result of a diagnostic test while ignoring this context. So, if we think about why we want to make a diagnosis, generally we want to make a diagnosis to increase the certainty of the presence or absence of a disease. We want to make that diagnosis to support a specific therapeutic management, perhaps as an adjunct to prognosis, perhaps to monitor the clinical course of a disease, perhaps to measure the capacity of one or more systems in a given individual, perhaps for epidemiological reasons. And when we want to make a diagnosis, by interrogating the patient, our brain automatically starts to organize the differential diagnosis, or all the diagnostic possibilities, into three parts. On one hand, we have the possibilities, well, with these patient responses, we can think of this, this, and this diagnosis. But at the same time, we start to organize these diagnoses, this differential diagnosis, into probabilities. Well, perhaps of all these, this one is more probable due to frequency than this one and this other one. But at the same time, we start to organize them, to organize this list in a pragmatic way, by priorities. In what sense? Perhaps this one is more frequent than the other, but the other one has a treatment, and the first one is something trivial that will resolve on its own. Therefore, we want to investigate or diagnose the other one because, from a practical point of view, we can treat that one, and the other one, nothing will happen if we don't diagnose it or treat it. So, this is how our mind organizes diagnostic possibilities. And there are diagnoses that are very simple, and there are diagnoses that are complex or that leave no doubt. Hopefully, all diagnoses had patterns as easy to recognize as this one. Someone who saw shingles, well, they won't forget it. No more questions. And if they had chickenpox at some point in their life, the lesions arrive, they are in dermatomes, they are vesicular, painful lesions, and well, you realize it's shingles. But sometimes diagnoses are difficult because the presentations can be variable or subtle. Several diagnoses can present similarly. Let's take this simple example: a baseball player who comes to you because his chest hurts, one side of his chest hurts. So, we could think of several diagnoses. And if we wanted to order those diagnoses from 0% probability to 100% probability, each of those diagnoses would be arranged differently. So, we could think, this is a young baseball player, whom we might know, and apparently, he has no risk factors. He doesn't smoke, doesn't use drugs, doesn't use steroids, because we know him, whatever you want. So, if his chest hurts, thinking about a heart attack, well, we have to keep it on the differential diagnosis list. But perhaps because we know this, we would put it low on the probability list, and perhaps we could think that a rib fracture from a slide, a fall, a blow, could be more, more probable, or perhaps a simple contusion could be more probable. So, in the same way that we are organizing these diagnoses by probabilities, we start to think about thresholds without realizing it. In what sense? On one hand, we have a threshold on this side, below which, if we consider the diagnosis to be below that threshold, we consider it so unlikely that we say, "Well, I don't think it's worth doing cardiac enzymes on this patient because I don't think it's a heart attack. I'd rather look for other diagnoses first because I consider it very unlikely." On the other hand, if any of these diagnoses were up here, above this second threshold, perhaps we would say, "Well, I'm already sure he has it, I'm going to start treating him. I don't need to do any other type of test to confirm or rule out the diagnosis." So, in reality, in this little space between these two thresholds is where diagnostic tests are useful to us. The problem is that we are accustomed to thinking of diagnostic test results as black and white: positive, he has it; negative, he doesn't have it. But let's take examples like this: if a child has a fever and cough, and has a blood count with 20,000 white blood cells and cough, does this mean he has pneumonia? Are we sure he has pneumonia? If an adolescent has hepatosplenomegaly and a negative mono test, does that mean we have ruled out tuberculosis? No, he doesn't have it. If a child has a general urine exam with 30 white blood cells, is that equal to a urinary tract infection? The reality is that it's as if they never explain it to us so openly, and it's kept more or less a secret that diagnostic tests are not perfect, even though we would like to think of them so simply: if it comes out positive, yes; and if it comes out negative, no. But then, if they are not perfect, we have to see how reliable the tests are. And the indispensable condition in any research study that wants to see how accurate a diagnostic test is, is that that diagnostic test must be compared against something else. Because otherwise, how are we going to see if the positive result of the diagnostic test is real, or if the negative result of that diagnostic test is real? Therefore, all diagnostic tests that we want to evaluate must necessarily be compared against the best diagnostic method that exists. We will call this the gold standard or reference test. And sometimes it is very simple to identify a gold standard, which is what does the diagnosis best. For example, we are talking about a general urine exam for the diagnosis of urinary tract infection. Well, we can think that the best method to confirm this diagnosis is a urine culture. There is our gold standard. So, we can then compare: the white blood cells in the urine came out positive, and see if the urine culture came out positive, to see if it is a true positive. Let's think about the gold standard for a urine pregnancy test, those sold in pharmacies. What would be the gold standard? Someone might say, "Well, perhaps the gold standard would be a blood test." But rarely, but there could be a false positive in blood due to some medication they are taking, due to some tumor, something rare, but there could be false negatives. What would be more infallible than that? Well, someone might say, "Well, see the baby on the ultrasound." And if we don't have an ultrasound, what would be more infallible? Something that is definitely infallible? Well, perhaps we could wait. We wait nine months, and if a baby is born, it means the result of that pharmacy test was correct. If not, the result was wrong. It sounds a bit funny, but well, this is the result of the gold standard. You don't need to have it at the moment you do the diagnostic test you are evaluating. What you want is for that gold standard to confirm or not confirm the result. So, waiting, observing the clinical evolution, could also be a gold standard. Likewise, ultrasound for diagnosing appendicitis, because you have a patient with abdominal pain, and the ultrasound comes out positive, or negative. How will you know if the positive or negative result is real? Well, one way is if those who operate check the pathology, the slides in pathology, and there you could confirm the diagnosis of appendicitis. And if you don't operate, there could be a second gold standard. Well, if you don't operate, well, follow up day by day, for enough time to rule out appendicitis. So, there must always be a comparison. In the case of diagnostic tests, we need to compare that diagnostic test we are evaluating with a gold standard, the best possible. Even then, gold standards are sometimes not perfect, but well, we have to compare it against the best, the best diagnostic test we have. Having that comparison, then we can get values, values that speak to the accuracy, to see how accurate the test is. So, we need two things to correctly interpret the result of a diagnostic test. First, those accuracy values, to know how accurate the test is, as through a study that has compared that diagnostic test with a gold standard. And second, we need to understand the patient's context, the probability the patient has of having that disease before doing any test. And with these two things, I could estimate the probability of that particular patient having the disease according to the result of that diagnostic test. To illustrate this, I will give you an absurd example, and it's an example I've used for many years. Don't think I came up with it because of what's happening right now. Again, take it in good spirits. It's something, although it's fictional, it will help us make the analogy from this fictional, absurd, funny thing to real life. So, I will give you the example of a clinical case. I saw this clinical case in my practice at some point, and it's a patient who was bitten by a zombie. What happens if a zombie bites you? If a zombie bites you, does it cause the disease? Will you turn into a zombie? If you are not familiar with all of this, please study. There are the clinical practice guidelines for the zombie apocalypse from the CDC. It's not a joke, and you can consult them on the CDC website. You can also read this review article, a good narrative review from the BMJ on zombie infections. But well, for those of you who are familiar with all of this, you already know that if a zombie bites you, you will turn into a zombie. You become a zombie. And what is the treatment? The treatment is something radical because you have to sacrifice the patient who is going to turn into a zombie. Or if it bit you, well, you cut off the limb and take two paracetamols so it doesn't hurt, sorry, so it doesn't hurt. As I say, you have to sacrifice. So, here we have a clinical case in which we have a patient who is at risk of the disease. We want to know if he has the disease or not. Since he was bitten, we want to know how probable it is that he will turn into a zombie, because I'm unsure. I've seen movies where someone was bitten and ultimately didn't turn into a zombie for some reason. Upon examining a patient, we realize that the wound is just a superficial scratch, not even a proper bite, a very superficial scratch. So, that makes you more doubtful. "Oh my, will he really turn into a zombie? Will this patient truly be sick?" So, we need an immediate, accurate diagnosis. Why? Because on one hand, we have a radical treatment, a treatment with many side effects. We are going to sacrifice the patient. On the other hand, if we make a mistake and say he wasn't going to turn, and let him go home, and he turns, we have the possibility of a pandemic, an apocalypse. So, it turns out that in the hospital laboratory, we have a diagnostic test. These rapid tests that give you results in seconds, like a test strip that you touch, you touch it with the test strip on the wound, and it changes color in 5 seconds. According to a study, this test has a sensitivity of 60% and a specificity of 95%. We apply it to the patient, and the test comes out negative. What is the probability that this patient will turn into a zombie if the test is negative, and we have these values? What is the probability that the patient is sick despite the test coming out negative? Obviously, it cannot be 0% because the test is not perfect. We are seeing it here. So, how much? Some of you might say, "Well, the probability of being sick with that result is 40%." Or some of you might say, "It's 5%." And that's a very natural response from most clinicians because we see these values and try to translate them directly to the patient. We see this, and we say, 60%, well, if it came out negative, it must be the remainder, 40%. Or we see this, it came out negative, so it must be 5%. Let's take a break. Let's put aside this test and this patient, and let's think with our brain, with the right hemisphere of the brain. Let's put faces. This is another disease, another diagnostic test, another population. Here we have 100 patients, of whom these 70 above are healthy, and these dark ones below are sick. They are 70 and 30, okay? We are going to do a laboratory diagnostic test on all these patients, whatever you want, to see if that diagnostic test is useful for diagnosing the sick. We do the test on them, and it comes out like this. All those that remained shaded red, all of these, the test came out positive. All those that came out shaded green, the test came out negative. So, if you notice, it's a more or less good test. Of all those who are sick, see how it diagnosed most of them, but it made a mistake on these ones here. Of those who were healthy, it diagnosed most of them, the test came out negative, but it made a mistake in this little part. So, let's calculate the sensitivity and specificity. What is sensitivity? Sensitivity is the probability that the test comes out positive in patients who are sick. So, if it's the probability that the test comes out positive in the sick, which patients are we going to focus on? Only the sick ones. And of these sick patients, how many had a positive test? This little part here, which is 24, had a positive test. 24 out of 30. If we calculate this proportion, it's 80%. This test had a sensitivity of 80%. How do we calculate specificity? The reverse, with the other part of the population, with the healthy ones. Specificity is the probability that the test comes out negative in those who are healthy. Who are the healthy ones? All these who were in clear. Of these, this little part had a negative test. 56 out of the 70 who were healthy. If we calculate this proportion, it's also 0.8, 80%. Therefore, that test had a sensitivity of 80% and a specificity of 80%. Now, returning to the zombie case, I asked you, "Okay, what is the probability that my patient has the disease if the test came out negative?" And you wanted to answer, "Well, not you, but perhaps some of you wanted to answer based on sensitivity." And we calculated sensitivity from this part down here. Someone might have wanted to get the result from specificity, and we calculated specificity from this part up here. And in reality, our patient, if the test came out negative, would be one of these shaded green ones. So, it's neither the sensitivity part nor the specificity part. If the patient had been one of these, they would be one of these green ones. So, if my patient's test came out negative, and the test came out negative in these 62 patients, and I realize that of these 62 patients, these above are truly healthy, which are 56 out of these 62. The proportion is 0.9, it's 90%. So, if my patient were identical to the patients in that study, and that were the test I did on them, and the test came out negative, the probability of turning into a zombie would be 10%. And we get this from the negative predictive value. That which we calculated from that shaded green part is the negative predictive value. If we had calculated, if the patient's test had come out positive, we would have taken all of this from the red side, and it would have been the positive predictive value. So, here we obtained these three values from this image: sensitivity 80%, specificity 80%, and negative predictive value. But we have a problem right now. Perhaps it's starting to get confusing. The values I need for my patients are not sensitivity and specificity, they are the predictive values. But we have a problem with the predictive values. Look here, what is the prevalence of the disease in this population? The prevalence of the disease is the proportion of sick people in this total. If there are 100 patients and 30 were sick, the prevalence of the disease is 30%. Now look at this population here. There is a higher prevalence of the disease, and let's see how here the prevalence is 60%. 1, 2, 3, 4, 5, 6. This population has a 60% prevalence of the disease. And notice, if we calculate the sensitivity of the same test we are using at the beginning, if we calculate the proportion that came out positive among the sick, it's still 80%. If we calculate the specificity from this upper part, it's still 80%. But if we calculate the negative predictive value from this green part, look how it changed, did you see? Here it was 90% with 30% prevalence, and here it changes to 72. Did you see how the predictive values change? They change depending on the prevalence. So, if your patient, you are seeing them in your office, and your patient does not belong to a population with a prevalence similar to those in the original study where these values were obtained, the predictive values are not useful to you because predictive values vary with prevalence. So, if sensitivity and specificity speak about the test and not the patient, and if prevalence modifies the predictive values, then what do we need to use it for the patient? Because neither of them serves me, because one speaks about the test, and the other varies with prevalence. So, we necessarily need to convert the values that come from these studies into something that could be useful for our individual patient. That useful thing is what is called the LR, or likelihood ratio, or probability ratio, or likelihood ratio. And look at the formula for each of them. The positive LR has a division with a subtraction. The negative LR also has a division. And when they put divisions in the formula, it gets very complicated. But well, don't worry about calculating this. We already have internet pages that do it automatically. We have calculators on our phones that can calculate it automatically. But to do this example, we will calculate it manually. And I told you that this test strip is for 5 seconds and had these sensitivity and specificity values. Let's transform them into LRs with this formula. So, all I am doing is putting the sensitivity and specificity of this test. I am putting 0.6 because the test had 60% sensitivity. I am putting 0.95 because the test had 95% specificity. Be careful when doing these formulas, don't put 60 here, put 0.6. Don't confuse integers with percentages. Just be careful with that. So, doing this operation with a calculator gives me this number for the positive LR of 12. We will see later what it's for. Let's calculate the other one, the negative LR, with the other little formula, and it came out 0.42. So, we have the first step I told you: knowing how accurate the test is, sensitivity, specificity, and LR. These will be useful for the individual patient. The next step is to know the patient's context. What would be the patient's context? Well, what is the probability, based on the patient's symptoms and signs, of turning into a zombie before doing the test? How could I have these values? One would be if I had a prevalence study at hand for my population, my individual context. But sometimes it's difficult. So, we have to rely on clinical information or scores or something that helps us assign a probability to the patient before having any test. So, I could say, "Okay, I already have a lot of experience seeing zombies, and of the patients in the last few days, how many have turned into zombies?" Well, since I have a lot of experience with zombies, I can tell you that with the wound this patient had, I've seen many, and half have turned into zombies. Okay, there I have a probability before doing the test. With this probability, I can already start thinking about what I will do with my patient. If the test comes out positive, if the test comes out negative, thinking about the consequences of assuming he has it or assuming he doesn't have it. This is how I will set the thresholds. If I have a treatment like the one we were discussing, which is to sacrifice the patient, a treatment with many side effects, I want a lot of certainty that if the test comes out positive, that 50% will come all the way up here above this threshold to treat, because the treatment is radical. So, I want a lot of certainty that he has the disease. On the other hand, if the test comes out negative, I want this 50% to come down here enough so as not to risk an apocalypse if I make a mistake and he goes home. So, we had said, this patient's test came out negative. With these values, I can now use these values to interpret the result according to this individual patient. If the test comes out negative, the number that will be useful is this negative one. If the test had come out positive, the number that will be useful is this one up here. And how will I do it? How will I interpret it? With a diagram that is sometimes found almost always in internal medicine and pediatric emergency books, which is there but we never use because we don't see its utility. But this is useful. It's called the Fagan nomogram. And in this first column, I will put the patient before doing the test, the probability of having the disease before doing the test, the prevalence, or the patient's context. And I had told you, well, just like these patients, half turn into zombies. So, we start from here. The test came out negative. So, we will consider the negative LR. We will look for the LR in this second column, at 0.42. In this column, it's here, 0.5. So, 0.42 is more or less here. What I will do is draw a line from the patient through 0.42, and it comes out like this. So, with that result, my patient's probability of turning into a zombie, which was 50%, decreased to thirty-something percent. So, I had the patient here, and it decreased to 31%. What was the use of that test? Practically nothing. It left me with the same uncertainty. I wanted the negative result to move the patient all the way over here so that he could calmly go home. If the test had come out positive, and from 50% it had gone up to 60% or 70%, it would have left me in the same situation. I wouldn't have been able to treat him because I don't have enough certainty. So, this test with those values didn't help me. The negative test translates to this patient having a 31% probability of turning into a zombie. So, here we say the exercise with this absurd example. Now let's do it in real life. Here we have a diagnostic test, those serological tests that are being used for the new coronavirus. I was very struck, it seemed almost incredible to me, that at the beginning of the pandemic, they wanted to sell me a test. A friend sent me a message saying, "I have a supplier of rapid coronavirus tests. I don't know if you're interested in buying them for the hospital or your practice." On March 6th, they sent me this message without any concern about how accurate that test might be, what its accuracy was, what the consequences of using it would be. No, it doesn't matter, let's sell them. Well, now there are already several serological tests, including tests that are considered rapid antibody tests, approved by Cofepris. But in no way can we assume that these tests are perfect. In fact, that's why they put these legends on the results: "probable infection," "probable." But how probable? The first thing that is obvious is that they are not useful for diagnosing at the moment of the illness, which is when we want to do it most. Why? Because IgG and IgM start to increase late in the disease. So, they are not useful at the moment we want them. But then, perhaps they would be useful in other contexts, for other objectives. Why would we want to do an antibody test? Well, perhaps a businessman would want to use it for labor decisions with his employees, meaning, depending on the result, I will tell them whether to return to work or not. Another person might say, "Well, I haven't been in contact with anyone, I've been staying home, but I'm curious to know if I had it, if I caught it from somewhere for some reason. That's why I want to get it." And another context could be that one of us had a patient with cough, fever, and symptoms suggestive of this disease, and the PCR came out negative, and I want to confirm if they really had it or not, because, well, I have to rely on another diagnostic test. These are three different contexts. So, let's do the exercise. First, we have to know the accuracy of these tests. So, I will give you an example of a real test. If we have a real test, I have to obtain the sensitivity and specificity. And where do I get the LRs from? One way to get them could be to just look at the leaflet inside the box of this test, the manufacturer's leaflet. And well, in those informational brochures, they generally describe wonderful sensitivities and specificities, almost 100%, from what the manufacturer puts. So, let's use those. Perhaps it's worth looking for a study where that test was used in another context, in another population, so that it's not the manufacturer's, and perhaps in that other population, that sensitivity and specificity won't be so wonderful. So, it turns out that for this, we found a validation, a validation that I believe exists in Canada, and we found the values for IgG, the sensitivity, and the specificity. Okay, we have these values to get this first step: how accurate is that test? Very well. Now, the patient's context. In the patient's context, this would be transforming those two values from the test we found into LRs. Ignore these references, they are from another talk where I put them. But well, in this validation, we saw that it had a sensitivity of 69% for this IgG test, and a specificity of 92%. So, what did I do? I transformed those into LRs. With that, I have the first step: the accuracy of this particular test, this IgG test, sensitivity of 69%, specificity of 92%, and I calculated the LRs. Now let's look at the three contexts I gave you: the labor decision, the curiosity, or the suspicious patient. First, I do this IgG test on my employees to see if they can return to work. And this is a question that exists, that many employers, many companies were asking themselves. In fact, we had barely been in the pandemic for two or three months when they were trying to make decisions with these tests that were on the street. At that point in the pandemic, the first question we have to ask ourselves after that accuracy is the context. So, at that point in the pandemic, how much disease do we have in the population? What is the prevalence of the disease in the population to make that translation of the results of this test? So, it turns out that if we take values from other places at these points in the pandemic, in May, in June, studies said that even in heavily affected areas, the seroprevalence was around 5%. So, we would have to start from there to interpret the results of that test and see what would happen with our employees and with the decision-making with those employees. So, let's round these values to 5%. If we were to use that test with that prevalence of 5%, and it's a more or less large company of 1000 employees, and I want to spend money on testing those 1000 employees to see if I tell them to return to work or not, or how much they can or cannot. I could make decisions with that test. But look what would happen. Out of 1000 people, I would estimate that at that point in the pandemic, in the epidemic, only 5% would have the disease. So, out of 1000 patients, 50 have the disease and 950 do not. Already from there, I see that out of 1000 patients, I am only considering that 50 will come out positive. But the test is not perfect, it has these values. So, of these 50 who do have the disease, only 69% will come out positive. So, of 50, this test will detect 34, and of these 50, it will tell 16 that they didn't have it when in reality they did. Look on the other side. On the other side, of the healthy ones, let's take the specificity value. And look, look at the problem. Despite a high specificity, of these 950 who don't have it, 76 will come out positive. And of these who didn't have it, it came out as a true negative with 76. So, look, to these who didn't have it, it came out as a false positive. What will you tell them? You will tell them to stay home when in reality they didn't even have it. And then to these, you will tell them, "No, they didn't have it. They can't return. Stay home," when in reality they did have it. And to these, you will tell them to return when in reality you don't even know how long immunity might last, and perhaps you will tell them not to even take precautions because the test came out positive. So, these are the consequences you could have for not considering the prevalence at the moment you are doing it, and for not considering the accuracy of that test. These are the errors you could make. And that would only be a snapshot at a static time. What does that mean? Well, these values at the moment you tested them, when you drew blood for that test, these values are no longer valid tomorrow or in a week, because you could be in contact with other people. So, those would be the consequences in that work environment. Look at the second context. I no longer know if I had it, I get tested. If this specific test, we had done it in the context we were talking about earlier, with a general prevalence of 5%. From there, I have to start. If I am a person who is not attending patients, who is staying at home, and is not going out to see patients, well, perhaps that is the general prevalence of that population at that moment, and we would have to start from 5%, just as we did here. So, assuming that person's test comes out positive, we would take the positive LR result. The test came out positive. So, let's cross from 5%, let's cross it through 8, which is around here. So, from 5%, it goes up to 31%. This is the certainty you have that this patient had contact or had the disease or got infected. So, look, even though that person has a positive test, it's more likely to be a false positive than that they truly had contact with the disease. Look, on the other hand, what would happen in a patient without symptoms, but already in a population with a higher prevalence? Perhaps the pandemic has progressed, and we are at 10% seroprevalence. Even then, the result is 45% certainty, it leaves you with many doubts. So, getting tested out of curiosity when you haven't had contact with patients, or haven't gone to clinics or hospitals, you are not a health professional, and it leaves you with many doubts even if the test is positive. This particular test. Look now at the utility in a suspicious patient. A patient whom we, as doctors, see, and who did have symptoms: cough, fever, headache. Well, that patient had those symptoms 20 days ago. But that patient had at least a 50% probability of having the disease given the context we live in. So, if we start from there, from 50%, and the test comes out positive for that patient who did have symptoms, look at the difference. That patient's probability of having had it is 90%. So, you see, we cannot interpret diagnostic tests in themselves as black and white. In this particular case, the interpretation depends on different contexts and on a particular test with a validation, with a validation study in particular. But interpreting all these tests will depend on the accuracy of the specific test you are trying to evaluate. There are better tests than the example I gave, and there are much worse tests. There were even tests with an 18% sensitivity, very bad. But there are better tests than the one I gave. So, it depends on the test. It depends on the day you do it. We know now that before 14 days, they are useless. But they are more or less useful between the second and third week, and up to 35 days. After that, we don't know how good the tests are after 35 days. Then it also depends on the context, how many sick people there are in the population, or the individual context. If they had symptoms or not. If you had contact or not. It depends on who you are and what symptoms you have had. That's why all these tests must be guided by a doctor, because the interpretation depends on the symptoms, on the prevalence. So, today we saw two parts of this critical reading, which begins with the validity of the results and their applicability. In terms of validity, we saw the indispensable condition for diagnostic test studies, which is to compare that diagnostic test with a reference test or a gold standard, against what best diagnoses it. And in terms of results, we saw how to interpret, well, we saw what sensitivity, specificity, predictive values, and LRs meant. Because all of this helps us to see if we can confirm or rule out a disease. And the truth is that this numerical exercise that I presented, we don't have to do it with all patients, obviously. But I think it's worth doing this exercise just once with the diagnostic tests we use most frequently to get an idea of how much we can rule out, how much we can confirm the disease in some of our patients.