Transcription
欢迎回来。呃,欢迎各位听众和热情的参与者。我们已经进入了医疗保健领域人工智能的第六个模块。今天由安德拉·加古马教授博士为我们带来本次课程。他是阿肖卡大学蒂迪学院生物科学与研究系的院长。在此之前,他曾是IGIB的所长,并且在2014年获得了最年轻的尚提·沙鲁普·帕特纳加奖。他是一位受过全球培训的学术爱好者,从As到Baylor学校,我们很荣幸您能来到这里,先生,欢迎您。主持本次会议的是钱德拉·S·林古谢蒂博士,他曾是美国Core Care Clinics的首席执行官。他在印度完成了医学学士学位,随后在英国和美国接受培训,并在哈佛大学接受了医院管理培训。现在将时间交给您,先生。还有一个小小的公告,这将是我们所有人的一次盛会,因为我们正在尝试创造一项世界纪录。欢迎各位观众,现在将时间交给您,先生。 >> 非常感谢Suresh。我假设你们都能听到我的声音,我很荣幸能代表国家委员会向大家发言,并且越来越荣幸地看到我们正在努力打破一些记录。因此,我将以一位对人工智能感兴趣的医生身份,谈谈为什么机器学习在医疗保健领域如此重要,但更重要的是,何时不应使用它。我认为,医生们未来最大的挑战将不是如何学会使用人工智能。人工智能已经融入我们的生活。我们每天都在学习。但学会何时使用它,何时推广它,以及何时不使用它,是我们的患者将寄希望于我们的地方。我们今天的目标相对直接。我们希望建立对变革潜力的认识和理解。因此,我们完全同意印度政府刚刚于今年2月17日由卫生部发布的“印度人工智能与医疗保健战略”,该战略认为人工智能有潜力改变印度的医疗保健,但同时也承认其局限性和道德界限,这使得医生成为人工智能使用方式的合作伙伴,甚至领导者变得非常重要。人工智能和健康最根本的原则是,它们不是平等的两半。不是“为了人工智能而健康”,而是“为了健康而人工智能”。健康是更重要的一半。最终,特别是从战略角度来看,目标是改善医疗保健的可及性、可靠性、质量、可负担性、公平性和正义性。从公共卫生的角度来看,这才是最重要的。我们关于如何使用它的许多决定都将基于此。但随后,我们需要每个人都成为使用者和信任使用者。无论您是患者、提供者还是支付者,保险公司也需要信任人工智能,我们需要有一种团结感。当引入一项新技术时,不可能百分之百确定不会犯错。错误是会发生的。只要这些错误是在追求上述目标的过程中发生的,我认为我们都可以有一种团结感。我们是为了一个更伟大的事业而努力。我将直接陈述两个有些矛盾的原则。我们没有人完全理解人工智能,而人工智能最大的危险之一可能就是我们过早地认为我们理解了它。另一方面,有疑虑是可以的,但利用疑虑来阻止人工智能的采用或拒绝使用人工智能,并不比仅仅穿上白大褂假装成医生更理性。我必须说,人工智能并不新鲜。早在1970年,《新英格兰医学杂志》上就出现了这样一句话:计算科学将通过增强,有时甚至在很大程度上取代医生的智力功能来发挥其主要作用。这是《医学与计算机:变革的承诺与问题》。年份是1970年。即使在那时,您也有这样的用例:一个程序可以帮助医生解决一个复杂的基于资产的问题。您有一个血气分析(ABG)。您想了解患者混合性问题的根源。早在1970年,您就可以进行这样的完整对话。现在会发生什么呢?如果您看到医生有一个非常狭窄的回答。如果医生用通常的人类语言说话,程序就会崩溃,因为计算机实际上无法与我们交流。它们只能执行我们为它们提供的逻辑框架中的某些计算。那时的人工智能和现在的人工智能是完全不同的东西。不花太多时间,您就必须认识到,在20世纪50年代、60年代、70年代,我们正处于人工智能时代的开端,其理念是将大量输入信息输入到某个东西中,然后产生输出。您可以学会分类事物。您可以对复杂的事情获得一些决策支持,而这些事物中的每一个看起来都像一个神经元。多个输入像树突一样进入,一个输出像轴突一样出去。然后您会意识到,对于像异或(XOR)这样的更复杂的问题,一个神经元是不够的。您需要许多神经元分层排列。然后它就开始看起来像大脑,因为大脑有多层相互连接的神经元,层与层之间只有很少的连接。这就变成了神经网络。大约在1980年,当Jaw Finton想到一个主意时,为什么我们不将网络输出的误差反馈到系统中,并告诉系统“你自己想办法最小化我的误差”。这成为了机器学习的真正开端,您所要做的就是提供正确答案和大量数据。计算机将计算每次迭代的误差,您将学会最小化误差,基本上通过机器学习来解决问题。我将单独处理这之后的部分,但非常基本地说,这基本上就变成了一个问题:您将什么信息输入计算机系统?例如,如果您想识别这是一只狗,您是否会输入整只狗?这没有必要。在座的各位听我说话,只需看到一张狗的图片就能认出它。我甚至可以模糊掉大部分狗,只留下鼻子,您可能仍然能认出它。如果您看手写体,看英文字母,我可以遮住一半,您也能毫无问题地阅读。所以关键是,我们从来没有真正看到图像,我们看到的是图像中的特征。我屏幕上现在显示的内容是,您希望提取特征,以便机器变得更好。例如,这里有一个C。C包含两个直角。我放了一个卷积核。这个卷积核会找到一个直角。 wherever it finds a right angle, the number becomes three. Where it doesn't and finds a straight line. The answer might be two. Very soon this image doesn't look like a C anymore. It is simply broken up into a map of features. These features can be a variety of types and given enough examples and that's where the internet was important. Millions of examples of images and that's shown on the right hand side. Very soon this thing called commuted neural networks where you use neural networks but you provide them features that is not uh the actual image but parts of the image you they became very very powerful with classifying things. What difference does it make? If you look at a 1993 paper on the use of neural networks to read a mamogram what you see is human beings put in 43 features. Computer scientists build a three-layer network. Then they train them on 133 cases. The computer learns to make the classification with only 14 features. But each feature has to be fed in individually by radiologist looking at the image. After they are fed in and this curve for those of you recognize an AU, you will see the neural network became better than the attending radiologists. But it was useless because if the attending radiologist is going to feed in the features and the computer is going to make the diagnosis, the amount of time wasted makes it useless for a doctor in clinical practice. And the 14 features that the neural network would use to come up with the final diagnosis are listed over here. All those of you who read mammograms will see these are very important. But all of you recognize that if you had to say yes, no, yes, no to all of these, it'll take you more time to simply do this than to just read the mamogram. But with CNN's you, humans did not have to feed features. Features were already extracted. Now the machine pretty much builds a model by itself. And very soon in the world of CNN you can not only find the abnormality you can get a map showing that this is the part that we think is abnormal and you can get a classification with a high accuracy as per the computer that this is a potentially malignant nodule. And this is the transition that we have gone over the last many years driven by better compute, better data and digitalization of the entire world such that the amount of data floating around is massive. But a huge problem still remained which was that if you have to label every image a lot of doctors are required. So for example when Google did the diabetic retnopathy they needed 30,000 diabetic retnopathy images annotated by clinicians that type of effort is simply not possible especially in India where you will never find a good doctor with plenty of available time to sit down and do annotations this is where something very cool came in AI this is probably one of the biggest beginning was of the current golden age it's called generative adversarial networks All of you, if I were to show you a new diagnosis on HS X-ray, after seeing 15 cases, you'll understand it. Why? Cuz you know what a chest X-ray is. You don't need to learn it every time. What do I mean when I say you know an X-ray? You can draw it. I can just give you a blank sheet of paper and you can draw a chest X-ray. And that kind of shows you understand what a chest X-ray is. So what if I had two AIS? One jet generates noise and creates an X-ray. That's called generator. Other one is a discriminator. It doesn't have any annotated X-rays, any labeled X-rays. Just knows these are all just X-rays. All it the discriminator is supposed to do is look at what the generator is sending and say is it a real X-ray or not? In this specific case, is it a real echo cardiogram or not? And then they go round and round playing with each other until the time the generator can create an X-ray on echo cardiogram that looks just like a real echo cardiogram and the discriminator can no longer tell. And you will see how in every round epoch 1 2 3 4 5 the echo cardiogram starts looking more and more like a real echo cardiogram. Although there is no patient, no nothing. It's all synthetic made out of mice. At the point the two are unable to discriminate, you could say the system has an understanding and using something called transfer learning. Now you can give real cases and today it is entirely possible only using a few hundred good quality cases to train machine learning systems that have already been undergone transfer learning and this has been a big power recently. But this alone was not enough. These are good for narrow applications. But if you want computers to have the power of human knowledge, they need to read our knowledge. And reading is a huge challenge for a computer because it we don't always write things in a way that is intuitive. Now some people say that Sanskrit is a perfect language but unfortunately most of modern medical knowledge is not in Sanskrit. So let's take a sentence. The dog did not cross the road because it was tired. What is it? I could have a second sentence. The dog did not cross the road because it was white. All I've done is I've changed tire to white. And here it is the dog. Here it is the road. It is just because of the way we use English. You can see that GPT4 understands what we are saying. But no previous AI before GPTs would have been able to easily do something like what I have done here. The reason is the concept is that when you see it and you see the word dog and the road and you see tired, your attention immediately goes to dog because dog and tired go together. In Hindi, a road can be tak. But in English, you will never say tired for a road. Similarly, a dog can be white. There's nothing logically saying a dog cannot be white. But a road is much more likely to be white. This particular thing called transformer architectures completely changed machine learning because after this attention where do you go to when a new piece of information comes became obvious to computers and after that they could digest massive amounts of knowledge and build what are called foundation models trained on the sum total of human knowledge at least the one that is written and documented. Where are we today? This is articulate medical intelligence explorer built by Google deep mind by Vive Natrajam's group. Here, Amy the patient, had conversations with Amy the doctor, and Amy the critic criticized the conversations between Amy the patient and Amy the doctor until the point that Amy the doctor became good at doing a full conversation for health purposes with patients. To test how good it became, they took 100 oskies. All of you know what Oskis are. Patient actors were used because it would be unethical otherwise. And they tested it in India, Canada and UK using a set of clinical problems from each of these countries. Now Amy or a real human physician shown here as PCP, primary care physician would interact with the patient actor, not a real patient. At the end of the interaction, it would make a decision, a diagnosis, a judgment, escalate, manage, and then a specialist physician would look at the transcript. This is all text by the way, although there is a new Amy that can also handle images. And rate them on accuracy, management, and escalation. The patient actor would rate them on confidence and care, perception of openness and honesty, and perception of empathy. This is called a radar plot. The closer you are to the boundary, the better you are. And I think all of you will see it clearly that at least in this specific thing which is a patient actor probably describing the typical presentation of a disease may not be the rare presentation may not be unusual languages outperforms a human primary care physician. This was last year. In fact, it was year before last. The paper was published last year but the preprint from Google Duke Mind was out in 2024 and we are well beyond this already in terms of where we are in practical terms today. So let me ask you well I'm going to answer it. Is there a need for AI in healthcare? Clearly healthcare is changing in ways that our health workforce cannot keep up with. Sometimes the volume and variety of information is just so much that we can't keep up and sometimes the content for example if today people can get genome sequenced they can have variables providing days worth of data at millisecond resolution it is simply too large or too complex for the human mind to manage we cannot keep up with the volume the variety or the complexity which means AI will be required whether we like it or not number two there aren't enough of us. We both need ways that we can train people to be like us and do this faster and better. Yet our medical colleges are lying empty with faculty positions not filled. Same is true for other healthcare workers. And then the number of people on the ground need to figure out a way to work faster, more efficiently, more correctly. And we also need to provide access to health care to people who live in areas where doctors and other healthare workers are simply not there. There are plenty of such places in India. As you all know what people need from AI depends on who you are. A physician who is using AI that's physician facing AI may simply want to screen for potential emergencies amongst all the routine cases they are seeing. They may want a camera that is looking at a baby that is not moving much in the arms of its mother in the people sitting outside their clinic and say this patient is sick. I need to see them first. It might be you know a scan with a bleed amongst many scans. We may want AI that simply automates our routine tasks. Does measurements and radiology generate prescriptions from the previous one. Generate a discharge summary for us without us having to say anything. assist us in complex tasks. We may want some help in the final diagnosis. We may wanted to know the protocol, write complicated chemotherapy prescriptions and cancer. A frontline healthcare worker may have completely different needs. They may simply want referral to further care. They may simply want automation of the filling of their forms that the government requires them to fill. And they might want simple assistance in following the guidelines set out by the Ministry of Health. Patients may need none of this but may simply want how to manage my illness given a plan that my physician has given to me. They may want information on all the drugs that have been prescribed to them but nobody bothered to explain to them what these drugs were for. And importantly and dangerously they may want the AI to help them know is it time to go see a doctor and if so which type. Now this is a relatively dangerous one but in no world that I can imagine for the future will patients not use AI for such purposes. We will have to control how they interpret the information that comes. But just like they ask each other, they will ask AI. So what are the clearly right ways of using AI? Wherever AI increases quality, wherever AI increases quality, there are no safety concerns, AI should be used in areas where there is no alternative. The Tina effect we talk about in India. Genomics, public health surveillance, randomized control trials are done or simply nobody else can do it. That's a very good place to use AI. The other area to use AI is where it increases economic or functional efficiency with low risk. A frontline healthcare worker using an AI to screen diabetic retinopathy. There might be some safety concerns of false positives, false negatives. They don't necessarily have the ability to judge uh whether it's being right or wrong. But still, it is just a referral, right? There is low risk in referring someone to be checked. And obviously we can put a safeguard that everybody who's symptomatic with a bad vision even if the healthare worker screens and finds the answer to be no it's a normal retina they still need a referral simply based on symptoms alone not difficult to set up the safeguard on the other hand if you're doing something irreversible giving a drug giving a dangerous drug you cannot use AI to enable a frontline healthcare worker you will never see a situation where a thrombolytic can be safely given based on the call of an AI in the hands of a person who does not know how to use these things. Education, training and testing. This is an area where I think we will see increasing amount of AI. In fact, it is ethical that we learn on virtual cases before we learn on real cases. If we have surgical simulators, it'll be good for us to learn suturing and many other things in virtual simulations as opposed to real human beings. Is there a risk? Yes, there are some risk. But again, safeguards are easy to create. So when you are asked should you be using AI for a task? This is a good way to look at it. First, is it a narrow and well- definfined task? most narrow tasks, well-defined tasks, you can train an AI to do reasonably well. Is it primarily about pattern recognition, for example, in radiology rather than a clinical judgment that could be more complicated? If these two are correct, it's an area where AI can often outperform humans. It's an appropriate type of task for the AI to do. But the second thing you must ask is is the decision that will be taken based on the AI high stakes or irreversible. If the answer is yes, AI can only assist never decide. That's the example thrombolytics I was giving you. And the other thing you have to ask is what is the worst case harm of a wrong result. Sometimes reassurance can be more dangerous than a false alarm. Sometimes a false alarm can be more dangerous than a reassurance. It all depends then what we are talking about. Then are the technical parts. If physician should ask was the model trained on data similar to my patients? It has been shown again and again when AI is used in settings different from where it was trained, it fails. Humans are much more flexible in the way we learn things. AI simply does not have that level of flexibility yet. Number two, you have to ask yourself what may be underrepresented. An AI train in America is very unlikely to have ever seen ferasis. It may have never seen traoma. You have to understand that bias heights in small subopuls and these are the kind of things you need to understand can be used or not. Then you have to look at the quality of evidence. Has it been tested prospectively? AI is brilliant at retropestic data because the modern AI's foundation models have pretty much seen all the data out there. It's only in prospective testing in real world settings on patients similar to your patients that their true performance comes. And then of course are the performance metrics clinically meaningful? Don't just use accuracy. You need to know your PPVs, your NPVS. You need to know the rate at which you'll generate false positives, false negatives. What will you do with them, the false positives, false negatives, what will be the outcomes of going wrong? All these things you need to look at from the point of view of evidence. This brings us to safety and validation of AI and healthcare. It has technological aspects and application aspects. On the technological side, you basically are usually thinking in terms of accuracy. But I have to be very careful in this. Validity is the more important thing. Varity is not the same thing. Validity is everything in a health technology assessment. It takes into account safety. It takes into account quality. It takes into account performance on the field. It even extends almost up to effectiveness. If your purpose for using an AI for screening diabetic retinopathy in the field was to reduce preventable blindness then you have to ask the question is my tool proven to be valid in reducing preventable diabetic generated blindness in the population. It may be there is no data but you still have to ask the question. Similarly, if you're looking at an X-ray tuberculosis solution, you have to ask in the district where this AI solution for tuberculosis on a chest X-ray was diagnosed was deployed, did we find more patients with TB being diagnosed and treated? Did we see a reduction in tuberculosis cases in the years after? Validity is many things. Confusing accuracy with either safety or validity is a mistake many of us make. We read a paper that says very high sensitivity, very high specificity and we decide that we've already found our answers to safety and validation. And safety is definitely not accuracy. In fact, safety may be higher by compromising on accuracy measures. If the model is not certain, even if it is 75% likely to be correct, compromise accuracy by telling it not to answer. If the type of thing that happens after the answer requires a 95% confidence, you can shift a decision threshold. If you are in doubt regarding a bleed in a person who has undergone trauma has been stratified, change the threshold that you only say nothing normal if it's really definitely normal. You might sacrifice sensitivity to have high specificity. In a different setting, you could do exactly the opposite. You normally in screening for referrals would go for high sensitivity and that's approach for intraanial bleed. But when you're going for a final diagnosis and you're going for a what you would call an irreversible decision, you might go for very high specificity while you observe the patient uh in different settings. Again, I think doctors are best equipped to do this. Also understand there are two different types of AI, narrow AI and broad AI. Broad AI is the generative AI of transformer architecture that I talked to you about. Narrow AI is very good for specific tasks using well understood approaches. Errors of narrow AI are typically along expected directions. With generative AI, we have a relatively limited understanding of the internal process. It's extremely powerful, but hallucinations continue to be a problem even today. Here is a case. This is a case where there is a head injury. The AI has diagnosed a subdural hematoma. I know you can most of you cannot see it. I also cannot see it incidentally but my wife who's a radiologist can see it and it turns out that in this case the AI is correct and the doctor confirmed the presence of a subtle subdural hematoma. These are the type of challenges that all of us will face. AI might mark something. I might be an emergency physician not a radiologist. I might not see the SDH even after it is marked on this film for me. By the way, it's around here. And because of it, I might be tempted to dismiss the AI. But if I knew that this is a high-risisk setting, I knew the performance characteristics of this particular AI, I know that it's very accurate with high specificity, and I know that there's something I need to get an opinion on. I will behave the right way when the AI gives me something I don't understand. And sometimes it'll go the opposite direction. This is a patient with shortness of breath. The AI has diagnosed pulmonary emolism. Why pulary emolism? You will see that some of the vessels are opaque. Contrast is not seen. A doctor who fully understands and has seen multiple pulmonary angiograms will understand the problem is not enough contrast has been given. The ventricles are not sufficiently bright. and there isn't enough contrast or the timing is wrong then the vessels will not get opicified and therefore it'll look like a PE even though it's not a PE in this specific case the doctor rejected the air diagnosis calling it inadequate opacification there is as of now no substitute to a domain expert for critical decisions with AI AI is very often right But frequently enough wrong. The doctors need to be in the loop. When you look at genative AI in particular, the biggest problem is that architecture is for coherence, not correctness. Coherence is best defined as what fits well, what word comes next. This is true for vision transformers. This is true for everything. And in absence of something called grounding, external reference, it is impossible for genative AI to not hallucinate at all. The more you can provide external reference and grounding, the less it'll hallucinate, but there is always a risk of hallucination. You could input a chest X-ray with a clear plural eusion. It could suddenly tell you the X-ray is absolutely fine. You could even tell it, look at the costic angles. It'll even reply to you that yes costtophrenic angles are a good place to look to find whether there fusion. I have looked at the costtophrenic angles and there is no blunting and it can say that when there is clear no costtophrenic angle visible. It is just being coherent on its first statement the next statement and this is something that all of you will face if you use generative AI but still as I showed you with the Amy example very often can be very good. It is just not very good with vision right now but is particularly good with text interactions symptoms and things like this but still it goes wrong. How can we reduce this problem? A is today we can create agents. A broad AI agent can have a narrow AI tool. The narrow AI tool can look at just X-rays and plural eusions after the broad X-ray knows what it is supposed to be looking for. You can have iterative reasoning. You can go back and forth between multiple agents. And last, I'll not bother to explain this much. Something called context persistence in which you can have a large amount of clinical history accompany the entire thing. You could have in fact the entire previous medical records with very long context models. The net result of all this is you get much fewer hallucinations and error rates today are not very different from a human physician. But when they go wrong, it is very hard to justify to anybody else. How could we possibly have missed this? When you think about AI for India, you have to look at AI that will work here and you have to look at AI that will work here. You have to take advantage of the digital transformation at multiple levels. ABDM, National Health Stack, UPI, UI. But you also have to remember there are places like this where there is digital but there is no use for digital in places like this. Something as simple as voice to digital data where you no longer have to enter forms. You speak the AI is already taking care of things for you might be the most important type of AI in settings like this. the ICU. You could have complicated algorithms that are continuously monitoring a baby or an adult for any sudden decline with immediate interventions through alarms. It all depends on your use case. Now I come to contentious parts. Who should be using AI in healthcare? The safe view is that anybody who can do the task without the AI is best qualified to do the same task with AI. If you can't do it without AI, don't do it with AI. If that is followed, safety and validation issues are modest because a person who is fully competent is the user. That's true for most doctors using most AIs. But scalability of health services are enhanced only to the extent of efficiency gains of their individual physicians. And if you don't have enough physicians, then clearly that does not solve too much of your problems. Public health value therefore with this view is medium. The strategy for AI for health in India is saying innovation over safety but not to compromise on safety either but innovate. If safety and variety can be reasonably assured we can have users doing task they could not do without AI and there's nothing fundamentally wrong in it. When doctors use AI to interpret a genomics report, they are already doing things they would not be able to do without AI. It is just a question of thoughtfully percolating this idea further down. What, how, where, when, how it should be overseen are important questions we all need to deliberate. There are no answers that I can certainly provide. But as we operationalize the national strategy for AI and health, these are the things we'll have to think about. There are lots of problems remaining. We have a shortage of data. We need representative data for India. We need extensive testing in a desired use case. In fact, we will never be able to do that many randomized control trials. Chest X-ray for tuberculosis that worked in Africa, failed in India, worked in North India, is failing in Northeast India. AI fails unpredictably and using it at scale is critical. Then there's an expertise gap in everyone from people who are providers like you including people who are supposed to be regulators and governance people. Everybody has an expertise gap and that's what this series by NBME is really about. Then governance needs to be risk based and benefit based. If something is high risk and low benefit, just don't do it. If something is high risk but very high benefit, do it carefully. And of course, we don't need to discuss low risk and high benefit. That's obvious. And last, we need to educate the public. Unlike previous medical technologies, this will probably bypass health systems. We need to let people know how to use AI. Well, we cannot leave it to them. And we need to participate in educating the public on the correct use of AI. So how do I see the way forward? What we are doing is we are working with locally tuned free models. We are trying to build platforms to create places that doctors can engage with show their preference, see what works for what specialtity in which region and use their feedback to give evidence back to the nation. And of course we want to use voice as a primary mechanism of health communication and regional language related layers is V2DD is an important part. All these things need to be regionally scalable and working with the government is critical in this going forward. Last I'm going to come back again to the weaknesses of AI. The fact they are very creative now the genative AI is their weakness. Respect for autonomy, privacy and ethics are not algorithm based. Unbiased collection of biased data does not eliminate bias. Many of us who are researchers will say we collected the data in an unbiased way. Doesn't matter if the data is fundamentally biased. You are still going to be biased. And last, those of you radiologists will understand this. Activation maps will tell you whether the AI is looking but doesn't tell you what it is actually seeing. Big difference. It's not really understandability. Let me show you an example of unbiased collection of biased data. America bid algorithms to look at who is going to get sicker and the high score meant you're likely to become sick so you get extra care. The dark line are black people, the yellow line are white people. The score is on the x-axis. Their medical problems are on the y-axis. What you will see is to get the same score a black person had to be sicker with more problems. Why? AI is not racist, but black people are poor. Poor people tend to become sicker before taking off from jobs to go seek healthare. Because they cannot pay, they will often be sicker at the time somebody decides to admit them. AI does not understand human society. It will simply normalize the difference and simply assume that people are equally sick, all initially admitted people. And these are fundamental problems that you can just imagine what will happen in India if we trained AI on the type of data that will be between public hospitals and our private hospitals all mixing together and letting the AI learn. It might decide that our poor people are so strong they don't need hospitalization at all. Then sometimes we don't know what we want. This is an AI algorithm for allocating livers. The thought probably was to give the liver to the person who has the highest. What occurred very soon is that young people stopped getting livers while old people got the livers. Old people were probably the ones with burnt out liver from alcohol and cerosis and things like this while young people probably had more genetic disease. What is right, what is wrong is very hard to say. But we know that humans who are allocating livers follow a very different set of principles than simply the risk of dying. A very old person is clearly at higher risk of death. But at the same time, they're going to remain at a very high risk of death even after they get the liver. There are variety of ways to address the problem. I'm not going to dive deeply into this. But we need to know what we want the AI to do. It cannot read our minds. And the more we offset decision-m to it, the more the problem will be. My final slide in a sense the ten commandments. AI is only for decision support, not decision replacement. Do not deploy AI without local validation. Accuracy is not safety. Rare events will break AI. If you know it's rare, don't use AI. High-risk decisions require you to be actively in the loop. Bias in data will become bias in care. understand where the data is biased. Automation bias, so you become a rubber stamp to what uh AI says is a real clinical risk. Explainability matters most when stakes are highest. You don't want a non-explanable blackbox to decide on whether to give thrombolytics to someone. Governance of AI is as important as everything else. And the very best AI for society will improve health equity and foster trust. These are my key messages to you today. I'm going to stop sharing and hopefully we'll have time for questions and discussion after this. Thank you. This is wonderful Dr. Andra. I think I haven't had a better lesson of late in the machine learning. You made it so easy starting from how to read the dog space to how a can hallucinate and everything else. uh lot of questions I can see other than some of the background noise and uh thank you messages in this uh Chundra uh over to you and then we will wrap up with few questions to the public whether they understood or not. Dr. Chandra over to you. >> Absolutely. Thank you so much. Um what a wonderful uh what a wonderful talk Dr. ag um that was a very thorough and very thoughtful overview of the entire AI landscape starting from the basic level machine learning deep learning neural network you name it you covered everything including the agentic system as a clinician our task as we all understand right now is not to understand every technical aspect of the AI but um just understanding where the responsibility lies across all platform terms of AI as you all know the most important distinction is is the AI advising us or is it deciding for us and you know the answer now the interesting part is u you know um interesting part of this session is to look at the questions um let me just share my screen real quick to bring your attention to a couple of um important slides here so what you're seeing here is a decision support matrix per AI tools. As you review this slide, you may see uh across all forms of AI, the question is not the capability. The question is how much authority we allow the AI to uh take over. As you can see in the matrix, this matrix essentially will help us out to decide where we should use AI and where we shouldn't. This is just a mental um sort of cue mental anchor if you like. So now um moving on to the next slide. I have segregated most of the questions into these four important buckets and these important buckets will really help us to stay grounded to what are the categories and what are the u broad categories categories of questions or categories of concepts that we need to talk about in the next 10 minutes or so. So with that I will just stop sharing here and um there are several questions as as you might have expected. Uh they all fall into the first bucket where it helps and you have already covered it elaboratively. Um and if I made a one example I may add one just one example to it. In one of the innovations I led, we use machine learning models applied to self-reported health and lifestyle surveys to estimate population level risk for the chronic diseases, cancer, and lifestyle diseases. The value wasn't the diagnosis. It was enabling payers and providers to prioritize the screening, support early detection, early intervention, and managing those cases. This is a good example of AI supporting prevention and resource allocation while clinical judgment still lies with the clinicians. So in your opinion, I would be interested to know your perspective on where you have seen AI adding most value in prevention early identification at a population health level. Dr. I would say the best use case I have seen so far probably remains a chest X-ray screening for tuberculosis that is automated and I'll give you reasons why this particular one worked while others didn't AI already now is better than human at detecting an abnormal X-ray and a normal X-ray has reasonable predictive power for at least excluding significant tuberculosis and This is something that a doctor does not need to do if half the cases 60% of the cases that come for an X-ray at volume screening can be disposed of directly by the AI. In fact, in a company called Oxipit, they got C certification in Europe. If they said the X-ray is normal, no human radiologist needs to see them. It also fit well into a national TB program and a national high health priority. And even there only stratified screening was done. So either you had to have symptoms or you needed to be a contact of a person with tuberculosis. So I have to also give credit to the national program for thinking it through. You don't go house to house and doing an X-ray on everyone. >> Absolutely. >> And because of it, it became a success in districts where it was deployed. The number of case finding of tuberculosis rose. And while we don't have follow-up data, it's my expectation that in time to come we will see actually a drop in tuberculosis. Now, that's a great example of something that worked. And in fact, ICMR did a health technology assessment. It actually showed that deploying AI saved money. So, it benefited people. It made doctors happier. It did not create extra work for them. Many X-rays are disposed directly and it saved money to the government. What can be better? On other on other hand there are areas where these things don't work well as well and where the incentive structures are wrong or you know doctors get fed up my wife's a radiologist she has a practice in which um you know scans get labeled with um triage sometimes there is AI fatigue every third scan gets listed with something abnormal so sometimes it doesn't work but this is very specific use case and I think the more specific you make them the better it is >> that is a wonderful example and I Okay, >> Janju, I saw a very interesting question here. >> Mhm. >> When to use decision tree versus machine learning. >> Ah, okay. So, see in the original definitions of AI, even a decision tree, good old good oldfashioned AI, goofy was also AI, but it was not machine learning in the sense that you provided all the instructions, the machine didn't learn on its own. Almost anything to do in images is almost always ML. In fact, apart from implementing national guidelines via rags, but the top-end chatbot will still be AI. So, you can combine decision support of national guidelines along with the AI type interface on the top. The example I showed you from the 1970s was a pure decision support while interpreting ABG. Why did physicians not like it? If you entered even a letter wrong, the system would suddenly throw an error. So I don't think pure decision support alone will ever exist in the future. At the very least, the decision support will be in form of a guideline. The guideline will go into a document and context for an LLM and there will be a heavy amount of rag to follow that document and it'll be almost be like decision support but still it'll be AI and machine learning. Wonderful. That is a that's absolutely fascinating and one of the thing I probably may be able to bring to the table is as you rightly pointed out machine learning. We use machine learning techniques um such as um random forest and also um dition tree definitely and maybe maybe a little bit deeper as well um neural network to look into stratification of uh patients in insurance company when we you know when we sent um surveys and when we also had nurses calling the patients and getting their acute and chronic morbidities and their hospitalization history, use of ER and whatnot. And then we analyzed the models and stratified these patients based on their future uh risk of hospitalization and their future risk of get getting any catastrophic diagnosis or catastrophic conditions. Um that really helped us as well. So I think uh one of the questions that is a nice segue into one of the questions who was uh one of the doctor is asking what is the difference between AI and what is the difference between machine learning and the deep learning and AI and as Dr. wallal um astutely mentioned it is um ML is a part of artificial intelligence it's like a syndrome artificial intelligence encompasses all these technologies starting from all the way from a simple um machine learning and all the way to um neural networks and deep fakes and deep learning and whatnot um very wonderful and I think it's it's uh interesting how uh we are able to um use AI in population health um larger broader interventions in terms of public health and again in the chat box I see an important question related to uh will AI be useful for preventing and also predicting the pandemics will AI be useful to um plan for the pandemic management and whatn not I think that is very very important and uh my answer is yes as long as the clinical risk is low and you are using AI to try you are using AI to summarize and pattern highlightation highlightation and risk stratification and workflow support. So um very wonderful nice segue into >> I can add to that. So genomic surveillance of sewage is an area that is now thought to be critical for pandemic readiness and in fact the international um society the WHO basically IPSN pathogen surveillance network international
病原体监测网络已优先在污水监测的连续基因组数据中引入人工智能驱动的分析,用于检测新的核酸进入系统,这将提供早期预警时间。我的意思是,噪音和假阳性的数量还有待观察,但事实上,三月份在英国有一个关于这个主题的会议。太好了。
第二个重要的领域,这里确实引起了很多关注和兴趣,那就是人工智能变得危险的地方。我想我可能会说,人工智能可以重塑医学,或者成为病人的噩梦。从您的角度来看,在临床工作流程中,您认为代理人工智能最大的风险是什么?是技术本身,还是临床医生过度信任其生成结果的倾向?您有什么看法,先生?
我最大的恐惧是自动化和技能丧失将对人工智能产生什么影响。我的意思是,你已经可以看到它了。我的意思是,让我们以前面的例子为例。现在,除了我妻子的电话号码,我记不住任何电话号码了,所以基本上如果我的手机坏了,我唯一能记住的电话号码就是她的。几年前我能记住很多。医生们随着越来越多地使用人工智能,他们将开始失去技能,而这种技能的丧失将是棘手的。
另一个问题将是人工智能在关键情况下的幻觉,而没有建立足够的保障措施和约束。迟早会有人习惯使用人工智能。他们会认为自己完全胜任,我可以举一个全科医生的例子,他们可以访问人工智能工具,他们会说:“我为什么要咨询放射科医生?让我自己来做。”我将从美国举一个实际的例子,这个例子对每个人都有效,但它可能每次都失败。
当病人出现疑似中风时,扫描结果会自动由人工智能读取。如果人工智能认为结果是阳性,神经科医生会自动收到通知,放射科医生则完全不在其中。神经科医生过来,做出判断并决定。现在,这假设神经科医生在解读特定的 CT 或 MRI 方面与放射科医生一样好,这在不同情况下可能是真的,也可能不是真的。
现在,它之所以有效,是因为每个人都能得到报酬。你可以为人工智能的解读收费。神经科医生即使过来并说人工智能是假阳性,他们仍然会向病人收费。病人会将这笔费用转嫁给健康保险系统。所以,每个人都很高兴。
但是看看医疗费用的增加。是的。所以最终我们都买单。因此,低效的系统将出现,每个人都会暂时满意。但长远来看,它将导致医疗费用不断升级,我们都将被卷入其中,而且可以肯定的是,那些大公司将能够找到如何将人工智能货币化用于健康领域的方法。
先生,你说得太对了。强烈要求再次展示“十诫”,因为至少我看到至少有30个人在询问。
让我来做吧。
绝对。
好的。
如果没有其他问题,我只想感谢 Agarwal 医生和 Ted Shandra 医生主持了这次会议,并以最好的方式进行了引导。那些是的,这里我们感谢 Onrag 医生为此所做的一切。这就是“十诫”。我希望我至少让几百人高兴,仅仅是通过让你这样做。
所以,Andagi 医生将会在 PPT 中展示一些问题。您可以随时回答。这不是反馈表,我们将以此结束今天的会议。Andagi 医生,再次交给您。
所以,有几个快速的问题要问您。什么最能定义医疗人工智能的基础模型?我想我在我的演讲中已经给了您答案。这是值得记住的。再次,您应该自己找出答案。我将为您解答,但我会给您几秒钟时间让您看看这四个选项。特定任务系统,大规模模型,基于规则的算法。有人问到决策支持和仅限于研究环境的实验模型。哎呀,我本来有答案的。没关系。
问题是,人工智能模型在回应临床查询时提供了详细但完全虚构的引文。哪种类型的人工智能会造成这种麻烦?我想,如果您什么都不记得,请记住这一点。有一种人工智能,当你和它交谈时,它是最聪明的。它也是最有可能不时误导你的。而且我们有很多这样的案例。我想,因为我知道现在我点击它,答案也会出来。但没关系。让我们快速进行。这可能也很重要。
如果工具增加了,我不知道。什么是避免假阳性诊断的弃权率?这是一个策略吗?是提高特异性吗?是减少幻觉吗?是通过改变阈值以牺牲准确性来提高安全性吗?是提高训练速度,还是仅仅取代人类?
说到这里,我将停止了,您可以按照自己的节奏进行。
非常感谢 Andra 医生。这些问题非常有趣。感谢 Chandra 医生加入。我知道现在是美国的午夜,您非常忙,尽管如此。感谢您抽出宝贵时间加入我们。NBMS,您有什么我从聊天框中遗漏的评论吗?有什么我遗漏的吗?或者我们可以结束会议了。
我认为先生,一切都已涵盖,时间也到了。
非常感谢您。
谢谢大家。
谢谢您,先生。
谢谢大家。
再见。
感谢 Mak 医生带来了如此精彩的会议。感谢 S 医生主持了会议。感谢 S 医生共同主持了会议。感谢大家。非常感谢。再见。
谢谢。