Transcription
Thanks everybody for being here. Um, I'm really glad to be back at Stanford where I did my PhD and really talking about a body of work that has been spanning at least 10 years but has really kind of taken off in a new direction that I'm really excited to cover throughout the course of the talk today. And it's really going to be going over this idea of trying to unlock the concept of autonomous surgical robots.
I want to start with this idea of what are robots actually, you know, leveraged for and I think one perspective that you can take is that it is really to meet this growing need for skilled labor, right? And this need is happening across all kinds of industries of which healthcare is one that you know, is both significantly impacted but also industry that we absolutely are intrinsically tied to every one of us. If we look at the shortage of healthcare workers in this country alone, we're looking at currently numbers of tens of thousands of surgeons that are not trained or not not sufficiently trained in terms of meeting the needs of the patient population. Simultaneously, we're seeing shortages of hundreds of thousands of nurses necessary to meet the short the the demands of the healthcare population. And so, you know, with this observation, robots are actually a very very strong opportunity to kind of come in and provide significant value.
Um, they have some very nice properties such as uh they don't get tired. I mean, they need charging, of course, uh but they don't get tired. Um they're very available 24/7. They have uh precision that often can go beyond human capability. We talk about, you know, CAD and CAM precisions uh from the machinist kind of point of view in sub-millimeter ranges. And we could in certain platforms get that type of precision as well in robotics. The very interesting thing about robotics as well as a platform technology for scaling skilled labor is that in the healthcare practice today, we have this notion of one, you know, physician or or one nurse training side by side with another physician or nurse, right? And this is a very linear kind of increase in um uh in in um making that, you know, skilled labor uh uh group grow. However, for robotics, these can be conducted kind of on a fleet-wide update basis just like autonomous vehicles. I'm not sure how many people, you know, uh drive Teslas in this uh room, but of course, you can have an example today where overnight you'll get a new uh software update and the the system will have acquired some sort of new autonomous skill. Finally, this idea of programmable expertise is really interesting. And how many people have gone into healthcare settings and wondering, well, what is the quality of care I'm getting in this setting versus another setting or another individual. In this case, you can program, you know, uniformity of expertise across these platforms.
I I to give an idea of where current you know, paradigms of medical robotics are today. This is a da Vinci surgical system. It is a teleoperated surgical robot that get the moves inside is placed has tools that go inside your body and the surgeon operates these tools using a joystick console as shown here and the tools inside your body move basically directly proportional to the movement of the surgeon. This robot has been around for well over over 25 if not 30 years at this point. Um and if we look at where the landscape of surgical robotics today, it's really exploded, right? And this is just a snapshot of all the surgical robotics out there. There are a significant uh Lee higher number of robots that I couldn't fit on the page. But the point I want to make is that all of these robots are teleoperated today. They lack inherent uh they have zero inherent skill. The skills coming from the surgeon and it does make surgeons more precise and less tired, but what it doesn't do is it doesn't help us address the population issue, population crisis issue. It does not help us do more surgery and in fact, when you're doing robotic surgery often times the surgical team that's required to perform the procedure requires more personnel than if you were to do it without a robot.
Okay. Now, if you look on the other side where there is autonomous robots being deployed is manufacturing, right? And in manufacturing, it's all rigid pre-programmed motions. They operate in a very structured environment. They require extensive setup and human oversight and they but they do excel at speed, precision and repeatability. This is autonomy but in an unskilled way. So, what if we try to merge the concepts of autonomy and uh surgical robotics? This concept is actually not new at all. Um, if you can believe it, this is, uh, you know, concepts that have been uh, dating back many decades, um, both from things like Robodoc that was used for, um, orthopedic surgery and doing, uh, small alignments for implants to Aesop, which is a voice-controlled, um, robot that holds a surgical camera for laparoscopic surgery to, um, methods from university labs that have looked at using what is very popular today, uh, concepts of imitation learning or learning from demonstration to show things like autonomous suturing.
Okay. So, all of this has been kind of a progression of new techniques, uh, back in the day that have been kind of applied over and over to try to address this issue of, uh, surgical autonomy. And in certain cases, we've been able to show some successes. Um, but if we look at where are we recently today in terms of general robot autonomy, we're looking at these ideas of large foundation models, uh, or vision-language-action models, which usually require large data sets, um, which is massive data collection at scale, uh, or developing what's called world models, like simulations or generative models of the world, um, that are trying to just scrape how describe how the world evolves over time. And, uh, what we see in this space right now is a large variability between accuracy and success. And so, there's a lot of tasks that are kind of shown that can be very successful at decent accuracies, um, in the like mid-to-high 90% range, but, you know, uh, uh, there's also many that are failing, uh, dramatically. Um, and all of these, uh, these foundation models have been, uh, you know, developed in what I would say is relatively clean and controlled settings where stakes are low, the environment can be easily reset, and for the tasks in which the demonstrators are abundant. So, people that can sit down at a console and just uh show a robot how to like fold a piece of laundry. Almost anybody can do that. But, surgical robotics has that issue of none of those properties really exist for surgery. So, that you have scarce data, uh the setting cannot be controlled, right? The setting cannot be easily reset. You know, when you're doing a surgery, you can't kind of roll back time normally. Um safety is absolutely critical, right? So, few demonstrators, right? There's few We already talked about the lack of surgeons, right? So, there's even fewer demonstrators, and there's very low incentives for them to be sitting and doing demonstrations for a robot uh learner. Um data collection in these situations is basically at the lowest priority, right? You're usually trying to address a patient and not to address a machine learning algorithm. Um and finally, world models right now have no concrete path to maturity, given that there is a limited data, that data is protected behind and and reasonably so, behind privacy uh laws, right? And so, how else can we approach autonomy?
And so, we can kind of look back at these what I would say is like kind of the foundational pillars of autonomous robots. And I would say that there's four that I would categorize as one is perception, um seeing kind of and and perceiving the world, and feeling the world. Second one is modeling and simulation. So, understanding like uh I uh rough models of how uh you know, things move, the the general physics of the world um at a small scale. Uh planning, understanding how to make decisions given that you know how the world evolves, and then control, actually executing those actions uh in a precise way. And you know, all together, you know, in in various types of uh um methods, we can bring these all together to get what we call embodied intelligence and autonomy.
So, at my lab at, uh, UC San Diego, we have been playing around with these four pillars quite a a lot and trying to build up the technology to make surgical autonomy uh, possible. And we've looked at, uh, a large range of different areas of autonomous suturing to, um, assistance to deformable modeling and probing to, uh, very specific robot, um, applications like, um, biopsy and cancer. And I think what we've learned, uh, and what I feel like I've learned over the time is that at least one very viable path to unlocking, you know, autonomous intelligent robots in surgery is a structured path of building context awareness, uh, moving towards in- incorporating multiple, you know, uh, knowledge, um, cues to kind of build on, uh, on this knowledge to be a lifelong learner, and then finding common common embodiments so that we can deploy it at a larger scale. So, I want to kind of cover these ideas, um, during this talk.
So, uh, for achieving context awareness, what I'm really talking about is perception, okay? So, robot autonomy, in my opinion, can only be good as good as its perception and its ability to recognize and understand the world. And only recently have we been able to get techniques very much, you know, on the backbone of, uh, uh, large data, uh, collection and, um, foundational models to segment the world, identify, localize, and track targets. These are just two examples of the state of the art, uh, and I'm sure that these will change within the next 6 months to something, you know, even more powerful, um, as the field progresses in computer vision. Uh one thing I will say though is that surgical robotics is one of the hardest environments to work in uh for these perception tasks because often times when you're trying to uh understand a surgical environment, you have typically a very narrow field of view. You don't have context understanding globally. You have bad or no depth measurements. Usually you have uh either a single camera um to understand in 2D what the space is like, but you have to infer depth from things like shadows and and um even understanding of like the anatomy that is being shown to you. So this is a lot of prior knowledge you need to have built up over time. Uh or in the other case, you have cameras that are you know, stereo cameras. So you have you know, two cameras that can do perception um stereo perception, but they're the the baseline between the two eyes of the camera are so tight that you have very poor depth perception. Okay. You also like I said have specular uh reflections from uh from fluids, blood, and there's smoke, you know, when you're cutting through tissue. You're always kind of cutting through the scene with large instruments um and so it's always occluding your view. Everything is deformable. Sometimes it's hard to find an anchor point in the scene to know where things are static. Um and like I said, there's insufficient data to train modern neural models. And all of this is wrapped around the fact that whatever we learn, often times we want millimeter precision. Okay.
So how can we kind of break this down? Because this is a very challenging uh problem and I will say we took uh spent all of the last 10 to 11 years I've been at UCSD just continuously working on this problem. And it can be broken down into several categories. Here here is like the kind of four main ones, which is precise robot proprioception. That is the idea that as a robot it has to know exactly where its hands are or its instruments are in space. Precise object localization. It has to know exactly where objects are in space relative to those instruments, right? 3D scene reconstruction. Understanding the background of what it's interacting with and especially in surgery the background actually it is a continuous kind of gradient from foreground to background, right? And so understanding 3D scene reconstruction and then scene fusion so that you even despite the point of despite the fact that you have a very narrow point of view over time as the camera moves around you get a better sense of where you are in space. So we have done a lot of work in this area across surgical tool tracking, things like needle tracking, even suture thread tracking for the explicit purpose of being able to very precisely get our robot instruments to grab objects in space and we have also done you know the third and fourth task of 3D scene reconstruction and fusion which is really about reconstruction reconstructing an entirely deformable environment.
I wanted to spend a little bit of time on this because reconstructing a deformable scene is really important not only from the perspective being able to track where you know anatomy is as you manipulate it and move it around and that's obviously important when you're doing something like excising a tumor and you want to be able to make sure you're cutting with you you have enough margin around the tumor as you cut and release it so that you don't miss any lesion. But it's also important from the perspective of physics understanding. And so if we start with this technology that we had developed over many years to get deformable reconstruction from cameras of the scene, how can we kind of move that into a useful a useful set of information to inform how do we do cutting and dissection? Okay. So if we go from vision and understanding of the scene to go to then instruct a robot how to cut, we need something in the between and that is our model of physics. Okay. So what is absolutely necessary in this scenario is this understanding of the physics and mechanics of deformable deforming tissue. So this idea is generally, you know, coined as digital twinning. And there's many techniques to do digital twinning. One technique that we were very excited about um when we started was this method that was developed uh in computer graphics um called position based position based dynamics. And the idea of position based dynamics was that you would simulate a deformable scene using particles and you would run the particle interactions using a set of constraints that define how those particles would like either bounce off each other or were constrained to each other. And the most important aspect of this was number one, it could run faster than real time. And number two, it could satisfy position constraints exactly. So the first one I'll explain. Um the fact you can if you have a simulator that can run faster than real time, um then that means that you can simulate many possibilities. Like uh and kind evaluate all possibilities of interaction with the tissue without, you know, actually interacting with it uh and waiting for a long time to figure out what the consequences are. So, at the end of the day, you can close the loop on that and create a model-based controller that is very standard in classical robot control theory. So, that is a very, very useful and I actually it's like a a property that's um absolutely necessary to apply those techniques. The second uh property of position-based uh dynamics was that we could take what we observed from the cameras, which generated this like a topological um surface map, and we could force in one shot the simulator to match that. And that means that we don't have to do many, many iterations to get the matching to occur, and therefore it keeps that um uh faster than real-time simulation um uh from slowing down.
So, one of the things that, you know, you we immediately found out was that if you go into a uh surgical type of environment and you recreate the environment from scratch uh and then you immediately try to simulate it, because you're recreating the um digital twin from just a surface image, you have to guess what is happening under the surface. And so, right when the model tries to um uh uh move the the tissue um in simulation, the simulation won't match real life. And so, what we can do is to correct for this behavior, we can use something called differentiable rendering. In a differen- differentiable rendering, the idea is that we can both correct for uh the we can we can take observations that are streaming in from a camera and we can look at how those observations match up to the simulation. We can get a loss between the observations of the simulation and the the camera image and then we can back propagate that loss to find how the simulation simulator should change to make it match the real life and in this case what needs to be changed in the simulator is the mechanics properties of those tissues as well as things like if the camera moved. Now what we can see if we do that is that we can get the simulator to continuously match the observations coming from streaming video and you can actually get a significant degradation decrease in the amount of prediction error from your model before you do this iterative correction to after and the difference is from several millimeters 5 millimeters down to sub 2 millimeter error accuracy. So this is for topological surfaces sorry surfaces that are you know kind of mesh like and topological in that nature but we can also do it for things like rope like objects so this is taking an image of a rope and having a initialization of a simulator of that rope pulling on the rope and understanding and learning the physics properties of that rope especially it's elasticity and viscoelasticity. You can imagine that this is applied in surgery to things like manipulating vessels and understanding the stiffness of those vessels which is of course important for safety. We can do it for fluids as well fluids is just the extreme end of a soft body right and this is just an example of us looking at a video of chocolate milk being poured and we're trying to estimate the viscosity of that chocolate milk. Now of course, how does that relate to surgery? Well, uh we were very interested in hemorrhage control. And so, when you see a hemorrhaging happening, there's going to be a flow of blood and it's going to start pooling. And what we were very interested in was to use a very quick understanding of the um the properties of that fluid to use closed-loop control to suction up the fluid as fast as possible. And so, if we define the problem um like that, it becomes an optimal control problem. And so, we basically did uh model predictive uh control uh on um this, where the model itself was uh this uh physics simulation.
And then another kind of a thing that uh was very important to us was to uncover where connections are uh occurring. Uh one of the one of the, you know, um um hallmarks of surgery is that, you know, you're going to be dissecting through tissue, and you have to have to find where the tissue is connected. And one of the ways that we found that tissues uh could displ- uh uh the One of the ways we found that tissues could show that they were connected was when we pulled on it, we could see that the texture from the side we're pulling on at some point would have a sharp discontinuity uh in terms of translation. And so, when we see that sharp discontinuity, that shows us that there is a connection in those locations. That's shown on the left-hand side uh in the purple. What we can do as well is to wrap this up in a Bayesian inferencing framework, where not only do we have a guess of where the connections are, but we can attribute uncertainty as to where we think those connections are. And once you have a measure of uncertainty, you can incorporate that into a safety-aware control approach. And this is an example, at least in simulation, of how a robot autonomously would you uh wrap a safety-aware controller around tugging on tissue to find out where things are connected before cutting. >> [snorts] >> So, in this case, the simulation of moving around this connected tissue is where the robot uh effector doesn't know that it's kind of connected on the red regions. And as it's moving around, it reveals those red regions are connected, and it will try to maximize the amount of information gained about other regions uh while minim- making sure it's below some energy threshold, which is basically the tearing energy. Um and so, given all of this, you know, you can imagine that you've found a place where it's connected. Now, it's time to cut. And so, here's the the example of uh us doing this type of procedure. All I would should have mentioned that all of these videos are of butcher shop meats. Um yeah, uh probably should mention that ahead of time, but we you know, this is all um uh on the bench a type of work. So, so here's this is an example of our robot learning how to basically peel back tissue uh to expose the region of cutting for a secondary cutting arm to come in, and this is all autonomous uh based on model predictive uh control methods. >> [snorts]
Now, one of the things that is um what is called like an emergent property in modern um foundation models for robotics is uh this idea of recovery. And so, you know, if you query, you know, a foundation model today and it fails at a certain task, often times people will show you that the the policy will go back in and try again, right? Which is really great as an emergent property, but it's also completely uncontrolled um in the sense that there is no actually guarantee that the robot will come in and try again. It's relatively random chance, at least at least like currently, right? But if you have an explicit model about where where things um uh if you have an explicit model about what's happening in the environment and you've tied it to physics, um you can query uh in that model whether it's failed or not and then you can go in and you can be targeted and basically have surgical precision, for lack of better word, of where to attack next and and uh perform the procedure um so that you complete the task. So, one of these cases is where you're trying to cut through a piece of tissue and you have failed the cut. And if you look at using a model like a this real-to-sim digital twinning framework and uh applying um this uh continuous iterative like a matching routine to it, you can see that we can model where our cuts were effective and where they were not and we can then go in and specifically cut those regions out all in an automatic fashion. Um and so we can here's this is an example of uh one of the you know, one of our um autonomous routines cutting through this chicken. Yeah.
Now, uh all of these are um techniques that we've built up from, you know, the vision and the uh modeling and and uh planning and control side. Uh, basically this toolbox, right? That we can uh, put together uh, for for developing autonomous surgery. And, you know, by advancing these tools, we can improve the capabilities of these tools and these tools individually, and we can strengthen their reliability over time. Um, and by assembling these tools in different configurations, we can broaden the capabilities of what we can automate while explicitly retaining explainability, which is extremely important in automating robotic surgery, where you have to be able to say, like, what is my robot doing? Is it going to be safe? And making sure that, um, you you as, you know, the the supervisory, uh, um, individual is in control at all times, and it's not really the robot. Um, so so that's great, but everything I showed you really has been, you know, piece by piece engineered together. Um, and so every single task you have to hand engineer. Um, and so would there be a better way to scale this over time? And so, given a set of behaviors or policies, we want to, um, find a good way to sequence them and combine them in a way that can be learned autonomously, and also in a way that can be learned uh, in which new skills can be added over time. And so, this idea is uh, called lifelong learning, and it's still a very, you know, uh, uh, uh, early and, uh, uh, uh, early field of research, and there is still, um, a great deal to, kind of, um, uh, explore here.
One approach that my lab has been looking at is this idea of combining all of these kind of learned or or engineered behaviors into a neural network that kind of glues them together. And so we call this um um knowledge grounded reinforcement learning. And in this case, we have basically knowledge modules, as you can see in the green. Um and those knowledge modules contain the actual behavior of a small set of like a a small set of behaviors. Uh and those behaviors could be very uh very different. So some behaviors could be um a robot camera that's like moving around to scan something. Or it could be a robot tool coming in and doing a cut, picking up objects, transferring objects from hands, uh many different kinds of behaviors to to uh uh that you can incorporate. And then the idea here is that using um a very sparse kind of neural network um architecture where there is a little bit of waiting that's used to glue these know- uh knowledge modules together, you can create a framework where through reinforcement learning it learns by itself how these knowledge module modules should be combined so that you can get um a longer sequence of a task done. So, for example, you have a knowledge set that you uh learned. In this case, we learned it using reinforcement learning in a simple simulated environment. We specifically uh looked at a task that was part of the what's called the fundamentals of laparoscopic surgery, um which is like an exam basically that uh laparoscopic surgeons have to take to um prove that they are able to do the uh laparoscopic surgery effectively. And so these are things like um uh picking up and and grasping suture needles, uh moving blocks from peg boards to from peg to peg, transferring objects between hands, and so on and so forth. So, we can learn that on in using reinforcement learning. We can deploy them in real life and show that they are effective, you know, going from the real for them simulator to the real world. And then using the same kind of approach, we can show that it can learn to sequence the set of steps for doing multi-throw needle multi-throw suturing. >> [snorts]
So, with this kind of ability it enables um you know, I think a very interesting framework for building autonomy. But, one thing that I've recognized we struggle with is that everything we've shown has been on this da Vinci platform. Okay, and there's all these kinds of robots out there and does this mean that we have to do the same thing for every single robot? You know, is there a best way to scale um the usefulness of, you know, robot autonomy for health care when the robots are so diverse? And so, of course, recently as everybody knows, there has been a very significant interest in humanoid platforms. Okay, and so, back in about a year and a half ago, all of these humanoid platforms were coming out and they were doing back flips and kung fu, a lot of things that didn't really speak to me about like how's this forwarding our society in a meaningful way? And you know, of course, there was interest in their use in home care, although a lot of it was also in the use of uh, their use in manufacturing, but not a single video or talk instance of anybody talking about its use in healthcare system. So, what we did was that we went out and we deployed one of these humanoid robots through mostly teleoperation, and I'll get to exactly how how much teleoperation later, but through mostly teleoperation in as many hospital settings as we could. And so, this is what that project led to. Um, this robot we named Surgi. Uh, Surgi is a humanoid robot with humanoid five-fingered hands. Um, we specifically looked at common procedures that you would see and especially in emergency scenarios where we could justify the the use having a robot present to do a procedure when there wasn't a physician present had significant value and had a lot of it was easy to justify. Some examples of tasks that are actually challenging, you know, for non-skilled individuals to perform. This is an example, you know, of intubation. You have obviously only a few minutes to get intubated before you suffocate and die. And so, if you're in a situation where you can either trust a robot to try to do it autonomously or teleoperated or just, um, you know, not have that, then I'm sure you would take the robot option. >> [gasps]
So, uh, here are just examples of what we were able to do, um, through through this Surgi platform. And I think what this is really getting towards is that there is a real value proposition for humanoid robots in a healthcare setting because if we think about the surgical robot I showed, the intuitive system, right now each system is over 3 million dollars. Okay? Um and you think there's very few hospitals even in the United States that can afford more than one or two da Vinci systems. And to think about the remote communities that have limited access and capital and resources. Um and that is only for laparoscopic surgery. Right? You can also think about scenarios where you have, you know, surgical teams where a surgeon is working side by side with another surgeon and this is actually more common than just having one surgeon operating in a operating room. You're usually working in pairs. If you're in an environment where you do not have enough surgeons, then you are actually taking somebody that's skilled that could be treating somebody else um and putting them in a scenario where they're doing very like relatively limited skill assistive tasks like holding a camera or retracting tissue and holding it back so that the primary surgeon can get a good view. Right? And you can think about circulating nurses, you can think about scrub nurses handing off instruments, you can think about ultrasound techs. Uh all of these applications there's really not that much justification for having another robot and another robot and another robot. Right? For for um that a hospital would have to purchase and get approvals for um to run through the the entire um uh to to run through the entire um uh pipeline so that uh they have um sufficient coverage of all these application areas. But of course, humanoid robots with enough skill can do this. So we look at ideas of remote surgeon avatars, autonomous surgical assistants, and autonomous nursing and technicians.
Okay. In the remote surgery aspect, we have been playing around with what it looks like to do teleoperation of a humanoid robot with laparoscopic tools. And this is just an evaluation of what where are we in terms of this versus a da Vinci robot versus somebody manually controlling the tools. >> [snorts] >> And what we observe is that um if you manually control the tools, that does require a lot of skill because when you're doing laparoscopic surgery inside a body with manual tools, all of your motions are flipped inside out and like uh the axes are all inverted. Um and you're kind of working at a very different scale. But with robotics, you can change the scale, you can invert the axes just in software. So, by doing this, even on the surgery platform, we can show it is better than uh a lot of people had easier time doing it with humanoid than with manual tools. And we, of course, could not expect to and we did not reach the level of a da Vinci system, but that's okay because a da Vinci system has been engineered over 25 to 30 years, right, to be perfect for laparoscopic surgery. But I think the point here is that the humanoid robot is, at least in this case, an interesting platform to keep looking into for those type of tasks.
But uh beyond those tasks, um one of the things that the video showed was all these diverse instruments that were being held and interacted with. And this is where the autonomy comes in because where we what we found was that it was nearly impossible to teleoperate a robot hand. Uh and what I mean by teleoperate a robot hand would be like we tried using cameras to look at our hands and then we tried to have the robot hand match the like hands of the um that was seen in the camera and try to pick up the objects and and instruments and try to interact with it. They kept falling out of the hand and we try often squeeze too tight and and things like that would everything, you know, you could think of would happen. And so uh and the reason because the reason for this is because even though they look the same like robot hands and human hands, they might look the same but they're not the same in many ways. Number one, robot hands have a reduced number of degrees of freedom. You have size mismatches. Most of these robot hands are much bigger than you'd think and much bigger than human hands. I mean, like basketball player size hands. Uh and what's really really important and is going to be a really long road to development is that they have limited sensing. And so on the right-hand side uh of this figure is one of the hands we use and the red portions are where there's tactile sensors. Okay, so in these red regions is where you can get a sense of what the tactile forces are on your fingertips and like in the middle phalangeal regions, but anywhere you touch outside of that you have no tactile sensing. And you can imagine how difficult it would be to manipulate something like, you know, a forcep like a scissors when you are lacking tactile sensing in those regions.
So, instead uh one of the things that we would do we could do is allow it to just figure out how to use these instruments on their own no matter if it is similar to humans or not. And so we do that using reinforcement learning again and so we massively parallelize uh robot arms learning to kind of hold and grasp onto uh existing instruments. Here are some uh examples of the behaviors of instruments that we were able to uh learn to to use. What is especially difficult in all these cases and why I'm showing these ones to you is because these are what we call articulated tools. and so in one hand, you both have to have a very stable grasp of the object, but you also have to have the flexibility to move the hand to open and close or activate that object. Right? And so you have tongs, you have pliers, forceps, and also a laparoscopic tool that has a trigger to it. And this is all used done using reinforcement learning. Um of course, you know, we were still kind of on this uh desire to use learning from demonstration if we could because if we could do that, we could save us a lot of computation time. Um cuz reinforcement learning a node or is uh is uh definitely um very computationally heavy. Um and what we quickly realized was that haptics, you know, the feeling of touch was actually very important, and it was not only the uh being able to recognize touch at the robot side, but able to deliver the feeling of touch back to the hand. And uh especially in the hospital setting, there's so many pieces of equipment that you're interacting with with all kinds of, you know, buttons and and knobs and um and uh plugs and wires that you need very um uh expressive types of haptics. And what we found was something as simple as pushing a button forward, like pushing it forward to hit a button, nothing that we found online really could uh found commercially uh could really do that. So you had gloves that you could put on that would constrain your fingers from closing, right? But if you push a button straight on, that has nothing to do with constraining your fingers from clo- closing, right? And so that type of haptic feedback was just not available. Um so how could we, you know, provide that back to the user? And so we went on to whole uh effort on building our own haptic glove. And so this is the first haptic glove that we know of that can do directional haptics on every finger. And so this idea is that you have linkage systems with multiple actuators on your hand that will push and you know, pull your finger in many directions. So that things like pushing or you know, sliding across the surface, you can actually get that tactile feedback or sorry, that force feedback back to your fingers. It is a pretty chunky system, of course. But the one of the reasons why it's chunky, especially at the hand area, is because one of the design principles we wanted to ensure is that all of the system was actually at the hand and there was not mechanical tethers to like your wrist and cuz that could cause ergonomic strain. So here's an example of you know, my students operating the haptic glove. You can obviously use it also for simulations as they're shown there for doing virtual interactions where you can feel the forces. And this is just an example of the kinds of how the multi-directional forces are being generated from the linkage system. Um and so yeah, we we had a you know, strapped this to our humanoid robot. We could show we could push buttons and things like that and really deliver a data collection approach as well as a teleoperation approach that was was much more effective.
So I talked about the use of these robots in hospital settings, but beyond hospitals, one area that is that I found was very under developed was this idea of um uh managing pain and especially like chronic wounds. So, chronic wounds is a very a very significant problem um not only in North America, but around the world in terms of wounds that are healing slowly or will um maybe never heal and they constantly need, you know, bandages uh removed and changed. Uh 2% of the population in North America have chronic wounds. Uh and so this especially for those who have mobility impairments, they need uh folks uh caregivers um to help them change out bandages and and change out these um uh change out these dressings uh very frequently, often on a daily basis. And this is a massive toll to the patients, caregivers, and families. You can imagine that I already said that there's not enough nurses in the US. All of this usually rests, you know, fall like the majority of this falls on the family members, right? And so you can imagine that in these cases over time the dignity of the patients are affected and they absolutely are looking for solutions where they can uh independently live and have ro- autonomous robots assist in these uh complex uh cases. So, [snorts] we started on this project um and uh we started bringing in all of our tools from deformable modeling to doing optimization based control. This is a case where, you know, we were teaching a robot how to peel off bandages uh with a notion that you would peel in a direction so it would minimize stretching of the skin and reduce pain. And then um we further on went to looking at, you know, how do you manipulate tape and then tape back on new bandages. So, the idea of dressing removal, preparing the tape, placing the tape down. And uh here are just some examples of um See, some of these Oh. Well, the tape the tape video didn't work of a peeling the tape. Uh but but here are just uh some examples, recent examples, of the robot trying to autonomously place the tape down um with all of the vision uh haptic models that we did develop. And this is still early work. You can still see a little bit of shaky hand happening there, but I'm very uh hopeful that as we move this forward that we're going to get um these robots to, you know, solve a problem that is really um impactful to our society. Um another kind of fun thing is that none of these robot hands had uh fingernails. So, we actually had to create our own 3D printed fingernail for our robot to peel the tape.
All right. So, I mean, where do we go from here, right? I think I I feel like I've in the you know, years that we've been working on this project, uh these types of projects, there's been a lot of validation as to the approach of building vision systems, models that describe key areas of um physical interaction with humans and their bodies, um to kind of trying to scale in um in controlled and um uh specific ways. And I think that we should keep pushing along this path. I think there's value in large foundation models and interacting with them and companies that are trying to develop push that forward. But as researchers in this area, I also feel like the more that we can build mathematical models, understand the physics behind the techniques that we want to deploy in robots and more generalizable and adaptable they can become. And so there's still a lot of work to do. So all of this all these projects, every one of them has been motivated by and has been developed in close interaction with our clinical team shown here. Many students have worked on these projects. One of the things I'm very proud of is also the student graduates who have now become faculty members at their own institutions who are now developing medical robotics and moving them forward as well. Thank you for your attention. I am very happy to answer any questions. >> [applause] >> Thank you very much.
To your last sentence what where you see the value of foundation models and where not and then the foundation model can help with like right now the performance whether you see me foundation whether they will have specific Yeah, so the question of like where do foundation models go from here and how does that maybe integrate with what we're doing? Well I at least that's how I interpret it. So because I do think that there is a integration a stage here. We are using foundation models for vision, for segmentation, we're using foundation models for tracking. You know, as these models get developed and more generalizable, we very much are happy to take what is out there if it's good enough and integrate it in as like either the main method for one of those pillars that I mentioned or you know, a supplementary method and this mixture of approaches. I think leads to robustness. So you don't necessarily have to rely on like either black box architectures fully one way or the other in terms of where foundation models will head and whether they're all these companies are going to eventually consolidate to rely on one model versus another. I don't know to be completely honest, but I will say that we have so many problems in robotics and they have so such diversity that it would be hard to think that there's going to be a one robot model that will work for everything. I think if you take surgery as an example, that is such a specific domain with very specific needs and very specific limitations of data and and safety that it would be a very very long time before I can imagine a model that could you know, clean up your kitchen is the same model that could perform surgery on you.
Yes. The very interesting work here. So I'm going back to the early part of your talk when you mentioned visual difficulty of knowing what's going on underneath the surface and I can understand why you focus on vision and you come from teleoperated robots and those that deliver a camera view to the surgeon. But as part of your work, I'm just wondering why you would what would be useful thing of other modalities so that they could like acoustics or ultrasound to look underneath the tissue in addition to the the vision and somehow integrate that. When you think of autonomous robots, don't necessarily need to be the same as you know, what you would use if you're teleoperated. Yeah, absolutely. So so the use of things like ultrasound are ones that we've actually played with in the past and I didn't show here but uh yeah I it is absolutely important to use whatever imaging modality you have and like you said a robot can not only interpret the images as humans generally could especially if you look at the development in radiology today and how AI has kind of transformed that field but it can accumulate you know and juggle all that knowledge in some ways much better than than humans could. So spatial awareness as you're scanning an ultrasound probe and and and incorporating that into understanding of the tissue geometry underneath. Yeah, ultrasound yes I think is a very very exciting area. Just as a context one way you would deploy ultrasound here is that there's ultrasound pickup probes. So what you do is that you drop a probe in that's ultrasound scanner into the body and one of those surgical tools will actually pick up the probe and now you can scan it and so so yeah so we've actually you know done
work like that before. Yeah, really interesting. Thank you. Yeah, thanks.
So when it comes to training data, is that I'm assuming this is one of the primary bottlenecks right now? And what would you say are like some bottlenecks in general in like when it comes to training data? Is it like the raw data? Is it labeled?
Yeah, I mean, so, so of course, yeah, if you have, if you're getting data from basically the wild, right, then then there's a challenge of like what is useful data and what's not useful data. I mean, there this is a really big question of what are the challenges, right? I would say like more tied to what I talked about, one area that I think people are starting to clue into, um, being something that really needs a lot of work is, uh, capturing force data and tactile data because so much of our interaction with the world depends on feeling the world. And so the challenge, and why probably this is, um, not the first thing that everybody goes to, is not only do we have limited interfaces to feel the world, as I showed, like we have to build our own, but we also have like, uh, material science problem here because most tactile sensors, uh, today are very, um, there's there's very limit, there's a lot of limitations in tactile sensing today, um, that they, they can, if they're sensitive, they can break easily or they are not robust, and if they're robust, then they are very sparse in nature and they, they're not very sensitive. So this kind of trade-off is a really difficult material science problem, and, uh, for that reason, um, at least I think it's going to be quite a long road ahead, um, to to solve it. But at the same time, I think we still need to forge ahead to incorporate whatever force and tactile data we do have and see how far we can get with what we have today.
Yeah. Uh, yes. So speaking of the, uh, touching, I can mention on the start of the presentation, the tearing energy that you have. Um, I'm very curious to see if that has anything to do with the foundation models of physics, uh, in those foundation models that you talked about. Um, is it, is it sort of like the importance of these foundation models, the fact that they can give the wrong amount of energy on this tearing energy?
Yeah, so in terms of, uh, tissue tearing and using that in closed loop, that was done using the local physics simulator in position based dynamics. So, so this PBD simulator, you can get an energy metric out. And so it was not using any foundation model at all.
Yeah, and how do we, how do we associate that physics model with like a boundary energy metric? Is from vision because we can see how things are moving and we can associate it the viscoelasticity that way.
Yeah. Yeah.
Um, I'm very curious about, um, feel like you have to, it's not like a robot, it's not like a manufacturing robot that it's doing. It's very repetitive work. So that's why it sounds like you have to create a very unique time and space for each task. And then so, so this is the kind of, you don't hear about the scalability, you know, machines. So I also see this one is a combination of hardware, software. So just have, just want to get a general sense about how much is about hardware and how much is about software, you know. I know it's, it has to be embodied and it's really challenging.
Yeah. I mean, okay, so you, you're probably asking somebody that's pretty biased in this opinion of how much is hardware and how much is software. I come, I started in mechanical engineering. So you're probably going to guess my answer. Um, but it, I think obviously, like, uh, you know, software is absolutely critical in here, but, you know, I think having the right hardware is, or sorry, the having an effective piece of hardware. I mean, right, is I don't know if there's like always a right, you know, solution to everything, but something that's effective is important because I don't think any amount of software can make a crummy robot, you know, achieve what it needs to achieve. So a lot of the robots we use in surgical robotics are designed very carefully to be inherently safe. They're back-drivable. They have like very little like backlash and so it doesn't jitter and it's very smooth. Um, so yeah, it's very, very, um, well engineered from a mechanical side. Yeah, hardware side. Design is very important here.
Yeah. Yeah. We're observing this tissue from tearing. It's a question that comes up a lot. Sort of about the next surgery. Have you looked at trying to set boundary conditions to avoid tearing tissue when your instruments are pushing on the side?
Yes. Yes. So, this is, um, this is absolutely, yeah, something we've been really interested in as well. So, we look at it from a retraction standpoint. One of the things that I showed was a video of the, uh, tool that was pushing like kind of revealing the next location to cut. That actually was, uh, incorporated in that was this metric, um, of how far, how much pull is too much, right? So, that I'm stretching it too far. And so, the same metric can be applied, um, uh, uh, both to be, you know, to inform, you know, what is, uh, safe and unsafe to tear.
Yeah. Yeah. You please. Is that we don't understand the degree to which we can manipulate tissue before you produce an a bad outcome. Is that, are you aware of, are you interested in work like that?
That's That's very interesting. Yeah. I mean, I'd love to talk after about this after this, uh, about it because I, I think, you know, I, I suppose you can kind of come at it from two perspectives, right? Like as as a neurosurgeon yourself, I'm, I'm guessing, so, so you have built up prior knowledge about what visually you think is going to be like within safe regions, right? Like manipulation. And I think we would have, like, I, I would probably approach it from the same angle. Like how do we teach a robot to visually understand what is within safe regions? And then every time we enter into a new, like patient, uh, then we apply that prior knowledge and, you know, uh, we, we maybe adjust it for, um, uh, the, the, the patient, like, uh, the patient, like age or or or gender or whatever, um, properties, um, that might, you know, affect affect that. But, but I think we can look at it from a human training angle. Uh, I think that would be the way to go.
The question about laparoscopy is, so you showed that there was all of the differences in performance. And and there was that sort of that trend in the that toward cognitive to have two people operating surgery in DaVinci. Correct me if I'm wrong, but you using, so the teleoperation interface, you were still using the hands of the DaVinci for both controlling the UI and the DaVinci and so on. And this is what I was, there any difference in the view that the teleoperators were getting?
No, so, yeah, this is a great question. Yeah, so for that study, we're using the same user interface and we're using the same camera. Um, so the difference here, which is, it's a very interesting point because this is actually a big issue. So, and it's actually a big issue for DaVinci, but, uh, they solved it, which is that, um, in laparoscopic surgery, you often are in like a, you actually have to work in a pretty large workspace because you're at a fulcrum, you're entering in a fulcrum, and your hands have to kind of move around pretty far. And the problem was actually that we were constantly hitting the range of motion for our humanoid robot because the G1 robot is not actually that big. Like it, it had that large of a wingspan, basically. And so, so we kept hitting these joint limits and so people were, um, taking longer for that reason. Yeah. And the, the intuitive system has kind of been designed to minimize that as much as possible. But with the G1, you still have to program for the like the remote center of motion?
Yep. Yep. Everything like that was programmed in already.
Great. Thank you.
>> [applause]