📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Stanford CS25: Transformers United V6 I Advancing Science and Medicine with Collaborative AI Agents

Stanford Online1:06:33

Transcription

Let me um introduce our speaker for today. We are delighted to have Vive, who is a research scientist at Google DeepMind, leading research at the intersection of AI, science, and medicine. He is the lead researcher behind MedPM and MedPM 2, the first AI systems to obtain passing and expert-level scores on the US medical license exam, uh, respectively. Vive also co-leads Project Amy, a research program aiming to build and democratize medical superintelligence. And prior to Google, he uh, he worked at uh, he worked on multimodal assistant systems at Facebook AI Research, and he is also part of the faculty for executive education at Harvard T.H. Chan School of Public Health. Uh, let's welcome him.

Um, thank you, Karan. Uh, thank you to the uh course organizers and uh, thank you everyone for coming in and uh, sorry for not starting exactly on time. Uh, but hopefully, this will be like a useful session and uh, yeah, today I want to be telling you uh, about some of the work that we've been doing over the past few years uh, at the intersection of uh, agentic AI systems and um, science and medicine.

Um, uh, I mean, the underlying theme of this work uh, is to kind of like build out uh, generally capable and general-purpose AI systems uh, that can act as like uh, collaborative partners for experts in the scientific and medical domain, that is, namely uh, scientists and doctors. Um, and so the first project I want to be telling you about is this effort called Co-scientist. Um, and so, so the system that we're trying to build out over here, uh, the the mission for this work is to essentially kind of like build out like a collaborative partner for uh, scientists. Um, and the hope is that, you know, with this system, like when we can really like accelerate the clock speed of scientific and biomedical discoveries and really like give like scientists uh, like superpowers to do the work that they are doing.

Um, and um, it's it's funny because like this work actually has its genesis uh, in a talk that I gave at like Stanford, like a few years back. Um, so I think as Karan mentioned, like um, back in 2023, like we had this paper uh, published in Nature uh, where we introduced this uh, large language model called uh, MedPM. Um, and so MedPM was one, like, like one of the first examples of like a medically tuned uh, large language model. Um, and so we kind of like built out the system to, you know, do well on tasks like medical question answering and summarization. Um, and back then, it was like, you know, one of the first systems to um, do well on this benchmark, which was like reflective of uh, US medical license exam style questions. Um, and so it was like one of the first exam, like, first models to reach like passing level, and then later on, like a future version of the system, like reached like expert-level scores on this benchmark. Um, and so that like made like headline results. And um, so I think it was like October 2023 when I was at Stanford, and I was actually uh, giving a talk on that system. Um, and at the end of the talk, like there was this uh, one professor uh, over here, like Dr. Gary Pel, um, and he kind of walked over to like me and my team at Tao, and he essentially like sent us like a follow-up email later, and he explained that, you know, like we are using the system for uh, medical question answering and summarization. Um, but he also thought that, you know, you could use large language models for like another kind of application. Um, and so, so his idea was essentially that, you know, uh, you know, because like LLMs are like, you know, trained on like a lot of like scientific and medical tokens. Um, you could actually like, you know, prompt it to kind of like do hypothesis generation. Uh, for example, he wanted to actually use it to do like, identify like causative gene factors uh, for like rare diseases and so on and so forth.

Um, and so it was like an interesting idea, but um, you know, like back then, we were still like, you know, dealing with our PaLM class of like models over here. Um, and so for like, you know, those of you who don't know, like PaLM is like the precursor to Gemini. Um, and I hope that some of you are at least using Gemini. Like, you know, trust me, it's like good. I mean, maybe not as good as ChatGPT, or I'm kidding, but like, you know, you should you guys should like, you know, try it. Um, but um, you know, so it was like the precursor to Gemini, and Gemini itself is now like, you know, three generations uh, old. Um, and um, so, so at that point of time, it was like, you know, still like struggling to get like, you know, like reliable answers from the system. Like there would be like hallucinations all over the place and things like that. Um, and so the idea that, you know, you could use like an LLM-based system to do something like hypothesis generation, um, it kind of like fell wild because, you know, hypothesis generation and, you know, coming up with like novel, creative scientific ideas is squarely like in the realm of like expert scientists. And to think that you could build out like an AI system to kind of like do that, uh, it felt like, you know, many, many years away.

Um, and so like I actually went back to my team and I asked like, okay, like, should we do this project right now? Um, and, you know, the overwhelming kind of like feedback was that, uh, we're just not like ready for it yet. Um, but, you know, sometimes like, you know, the best projects, like, and I think you probably have all have like experienced this. Uh, but for me, certainly, I think like uh, some of the best projects that I've like worked on are essentially like examples where we actually don't know how to reach the end goal, but we still like, you know, decide to like just go ahead and do this. Um, and so this was like one of those kind of like projects where like, you know, we did not know exactly how to exam, like, you know, how to clearly like approach this problem of like, what, what, what was going to be like a solution uh, that would lead us to this end goal. Um, it's almost like, you know, kind of like jumping off like a cliff, and you kind of have to like, you know, figure out like building out like a flying machine or like an airplane on the way down. Um, and so this project is kind of like an example of that, like where we just decide to go ahead and do this, and we just try to like, you know, okay, let's construct some evals, let's like make some progress, let's just see like how far we can push. Um, and um, and that's what like we continue to do. And so like what we initially did was just using our PaLM-based system, uh, we put together like an agentic scaffold. And remember, this is all like 2023, like the word "agent" was not being used. Um, but we still like, you know, put together this scaffold which had access to some databases and some tools, and it could, you know, go and like retrieve that information. It could read that literature up, and then we would ask it to kind of like, you know, predict like, okay, like what are maybe like the most uh, likely genes to cause this rare disease variant and things like that. Um, and we started doing that, and we actually even like saw some like promising results. And so there was like this one particular hypothesis uh, that this uh, system came up with, uh, that actually Gary went on like uh, to like, you know, validate in his lab using these CRISPR knock-in experiments on mice, um, and it turned out to like actually work, and that ended up being like a peer-reviewed paper that was published in Advanced Science like a few months back. Um, so that was like promising.

Um, but in the process of like doing that, like we also like realized, okay, like what are the limitations of like using, you know, LLMs in like a naive manner for this hypothesis generation task. Um, and if I were to like, you know, summarize it or like put it crisply, uh, I would say that like, you know, when you're like traditionally using LLMs, like say Gemini or ChatGPT or Claude, like I would argue that these models are kind of like, you know, primarily doing uh, like System 1 style thinking. So they're like, you know, coming up with like very fast responses, and so they're kind of like, you know, akin to like, you know, coming up with like very uh, intuitive responses, like without like, you know, deep thinking. So it's mostly like, you know, responses with like, you know, surface-level correlations and pattern matching and things like that, right? Um, so, so that's helpful, but at the same time, like when you think about like, you know, the, the kind of like thinking that is like needed for science and like scientific discovery, um, it's actually like a very different kind of thinking. So you would kind of like notice that it's actually uh, a lot more like slower. Uh, it's actually like, you know, more deliberate. Um, it's like rigorous. Um, and I think you all might have like experienced that, like, but when you talk to like, you know, some of the best scientists, like they'll tell you that they've had like their best ideas. Um, when they've been like, you know, thinking about like a problem for like, you know, weeks or months, sometimes even like years, right? Um, and so they'll tell you that, I mean, I've always had this thing going on in like the back of my head. I've been thinking about this, and suddenly I had this like, you know, like nice creative "aha!" moment over here. Um, but essentially, the point is that that's actually a very different kind of thinking. So it's more like System 2 style thinking, and I think that thinking is really like the hallmark of like, you know, science and scientific discovery over here. Um, and so like the key research question for us was like, okay, like, how do we build out these AI systems that can actually perform this structured, rigorous scientific thinking, um, that is really like the hallmark of like science and scientific discovery?

Um, and so that's like, you know, one framing of the problem. Um, another way to kind of like think about this is like, okay, like if you want to say, build out like, you know, a truly scientific, like, you know, superintelligent system, like, how do you like progress over there? Like, how do you quantify like progress, right? Um, and, and so for that, like uh, I have this graph over here. Um, and on the y-axis, you'll see like, you know, different scientific tasks, ordered by complexity. Um, and so, so they range from things like, you know, literature review and hypothesis generation to kind of like more complex tasks like, you know, writing like a research paper, uh, or like, you know, coming up with like a PhD thesis, to something that's like, you know, at the highest levels, which is like, you know, coming up with like, you know, paradigm-shifting theories and breakthroughs, like, you know, say, like the general theory of relativity. Um, and then on the x-axis, um, I have like the time it might like take like an average human scientist to like, you know, perform these tasks. And so it can range from, you know, something like, you know, minutes and hours to like, you know, days, weeks, months, years, and sometimes even like decades and lifetimes to like, you know, really do like, you know, paradigm-shifting like research and theories.

Um, and so, and then if you were to like, you know, start asking this question, okay, like, where do some of like the popular AI models and systems that we have around us, like, okay, where do they actually, you know, fall onto this graph or this plot? Um, and so, like a very popular example is AlphaFold. Um, and uh, so if you were to think, okay, like, where does AlphaFold fall on this graph? Um, so you'll see that I have it over here, which is uh, at the extreme end of like the x-axis. Um, and so if you ask me like, okay, like, why do I have it over there? Um, I would argue that, you know, um, AlphaFold has, in some ways, like, it's already done the work uh, that's equivalent to like, you know, millions of like human scientists' years' worth of work. Um, and then, I mean, if you come and like ask me why, like, you just have to think back to like the status quo before AlphaFold existed. So, like, you know, before AlphaFold existed, like to get to like the structure of a single protein, uh, that was often like, you know, like an entire PhD's worth of like effort over here. So it would almost take like three to five years of like dedicated PhD work to get to like the structure of like a single protein. Um, and then literally like, you know, once uh, AlphaFold, like the model was developed and like the system came online, um, like literally overnight, like we have like, you know, predictions for uh, like millions of proteins, like literally overnight, right? And so, so if you were to like, think about that, it's kind of like done the work equivalent to like, you know, millions of years of like human scientists' worth of like uh, time and effort over here. Um, so that's very cool, obviously, and it was recognized with like, you know, the Nobel Prize for that reason.

Um, but at the same time, uh, if you look at this graph, uh, you'll actually notice that I actually have AlphaFold outside of this graph. Um, and the reason for that is like, at least when I think about like a scientific superintelligent system, um, I actually think about generality being like a very important property over here. Um, and so, and again, if you were to ask me, like, I would say that like the one thing that AlphaFold doesn't have is this property of generality, right? So it's like a very specialized model, and so like the only inputs that it recognizes are like, you know, protein sequences. Um, and the only outputs that it can like produce are like, you know, protein structures over here. Um, and so if you were to like, you know, give it like any other kind of like scientific problem, uh, then it won't be able to like, you know, make sense of it, right? Um, and so in that sense, it's like a very, very specialized system, and it's, and it's obviously like very, very useful, and it's solved like a, you know, a grand challenge that, you know, dated back like decades, and so that's why uh, the system was like, you know, recognized with like a Nobel Prize over here. Uh, but at the same time, it's like an example of like a highly capable but like a very specialized system over here. Um, but if you want a truly superintelligent system over here, you would want this property of generality. Um, and what do I mean by generality? Uh, it is just the ability to kind of like, you know, take like literally like any problem and, you know, figure out like a way to like, you know, make progress on it. So you may not necessarily like solve it, but what you would want is like your algorithm to be able to like, you know, make sense of it, like, you know, understand it, break it down into like steps, and make like reasonable attempts towards like, you know, solving them over here. So that is what I mean by generality.

Um, and so again, like I don't know like um, how many of you have like seen this "Thinking Machine" documentary, uh, which kind of like chronicles like the history of like DeepMind, starting from its founding to like AlphaGo to like AlphaFold. Um, so if you haven't seen it, I would definitely like, you know, recommend uh, seeing that documentary over there. Um, and um, in that documentary, there was like, you know, there's one uh, clip where they actually show like IBM's like Deep Blue machine uh, taking on like Garry Kasparov. Um, and so this was all back in uh, 1999. Um, and um, yeah, and so obviously like, you know, Deep Blue was able to like, you know, beat Garry Kasparov over there. Um, but actually Demis Hassabis, who's like the CEO of DeepMind, like he actually uh, makes this comment, which I thought was like quite insightful. And so what Demis says is that like, he was actually like, you know, more impressed by uh, Garry Kasparov, uh, rather than by Deep Blue in that sequence. Um, and the reason for that is just simply that, you know, Garry Kasparov's brain was like remarkably general, right? So it was able to like, you know, not only compete with this giant black box of a machine in this very complex game like chess, but it was also able to do everything else that like humans can do. So, like, you know, Gary could like, you know, talk in like multiple languages, he could like appreciate art, he could like, you know, uh, think about like, you know, physics and philosophy and poetry and whatnot, right? And so like, I mean, the human brain has this remarkable property of generality, and that intel, that in turn kind of like allows us, allows it to kind of like, you know, do things like science, come up with like creative hypotheses for like complex scientific problems, and things like that. Um, and so I mean, in many ways, you could even argue and state that like, um, I mean, so far, like in the entire universe, like the only existence proof that we have for a machine that is capable of this kind of like generality and this kind of like, you know, hypothesis generation capability, uh, that's actually just the human brain over here. Um, and so like the question for us is like, okay, like, can we uh, you know, build out like an AI system that can like, you know, mimic this capability? Um, and so, so that's like the core problem.

Um, and then so, so if you had to ask, okay, like, okay, like, I mean, we're talking about generality over here, like, what is the key building block? Um, and again, I would argue that the key building block over here is actually like language, and natural language. Um, because like, if you have natural language as like an input-output interface, uh, much like we humans do, like we might be able to like, you know, understand like different concepts, and we might be able to like, you know, make sense of it, and we might be able to like, you know, solve like different problems over here. Um, and so, so that's like a great starting point, because like, you know, um, the large language model systems that we have around us, they're all built on this paradigm of like natural, natural language input and output, right? Um, and so if you were to think about like, you know, Gemini or like Claude or like GPT, I'm sure like you're all like talking to it about like a whole bunch of topics. I'm sure like many of you are using it to, you know, understand your lectures, your homeworks, but also, you know, things like, you know, physics and like philosophy and poetry and whatnot, right? I mean, like a whole array of topics that you're like talking to this, and and these models obviously like, you know, are able to like understand it and interpret it. Um, and so that's like a good property because they have this generality notion, and so like, you know, you can pose to it like, you know, questions about like any scientific problem, and hopefully that they will be able to like understand and make sense of it. But at the same time, like um, maybe even at least until like say, like a few months back, um, I think the the amount of evidence that we have of these, you know, LLM-based systems, like, okay, like the kind of scientific tasks that they can do, um, I would argue that it's still like very, very early. Like I'm sure like many of you have been using like these deep research modes of these uh, systems, like to do things like, you know, uh, evidence synthesis and like, you know, uh, literature uh, summarization and things like that. Uh, but in terms of like, you know, more complex scientific tasks, like say, hypothesis generation or like, you know, writing like uh, research papers, like we've not seen like any evidence of that, right? And so like uh, in that sense, there's like a lot of like, you know, dark space or like, you know, blank space in this plot, and so there's like a lot of room for us to like, you know, make progress uh, up and to the right over here.

Um, and so like the goal for us with the Co-scientist effort was simply just that, like, okay, like, can we build out AI systems that are probably making progress on this plot, like up and to the right over here? So essentially perform like, you know, more complex scientific tasks over like longer time horizons. Um, and so hopefully that framing is helpful in terms of like what we're trying to uh, accomplish. Um, and so in order to do this, right, right, in order to like, you know, build out this uh, say, structured scientific thinking engine, um, the the approach that we took, um, actually like, you know, borrows from like the uh, rich lineage that we have at like Google DeepMind over the years. Um, and so if you go back to AlphaGo, like, and so this is like the 10-year anniversary of like AlphaGo over here. Um, so you'll see that like uh, you know, AlphaGo, like the way it was developed, like the key principle was this idea of like self-play, right? Um, and to like, you know, very simp, uh, state it in like a very simplified manner, um, the key idea was that, you know, you had like uh, agents playing against each other in this environment, um, and um, like depending on like the outcome of like the game, which was obviously this uh, uh, game, game of like Go over here, uh, the agents would like receive these uh, reward signals. Um, and then, you know, based on like the reward, like the winning moves or like the uh, the decisions that led the, or the strategies that led to the winning moves, that they would get like, you know, reinforced in like the weights of the network. Um, and then the ones that like say, like led to um, uh, you know, the defeats would like, you know, get like de-prioritized in some sense, right? Um, so that was like the key idea, right? And so like it's like a combination of like self-play with like, you know, reinforcement learning and search. Um, and uh, essentially the, the, the beauty of AlphaGo, like essentially, or like actually it was more like AlphaZero, was the fact that, you know, you could create this environment, and in this environment, you could just have like, you know, agents like go against each other, you know, using this self-play setup. Um, and this setup actually like, you know, scales like remarkably with compute. So there's like no like, you know, human interference over here. Uh, you could just like start the agents from scratch. You could set up this environment, and if you give like the right kind of like rewards at the right time, you what you essentially see is that, you know, these agents like keep on improving over like a period of time. Um, and then at the end of like, say, three months, like that was actually how how long like AlphaZero was trained for, um, the system became like superhuman over here, right? Um, and so, so that's like a, a pretty remarkable property to have, like right, like essentially throw compute at a problem, and your algorithm is just so good that you don't have to do much. You just let that system run for a period of time, and then it becomes like superhuman at this task. Isn't that awesome? Like, isn't that wonderful?

Um, so like the question for us, like, okay, like, now how do we take that approach, like that idea of like self-play, and apply it to kind of like more complex domains over here? So I mean, like we started with Go, and then like in 2019, in 2020, like people showed that, okay, like you could uh, have these systems like with a very similar approach, like work in like more complex environments, like AlphaStar, for example. Um, but like, I mean, games are useful, I mean, they're like useful testbeds for us to like, you know, develop learning strategies and algorithms, but you want to like, you know, apply these uh, agents and these AI systems, uh, and develop them for like, you know, more complex real-world tasks, such as like scientific discovery or like, you know, in medicine and things like that. Um, and so pretty much that's what we've been doing. I mean, like uh, over the last few years, like the, the, the areas that we where we've been, you know, applying these systems, like the complexity of like the board games and the environments, that's just become like more and more, uh, I would say like, uh, sophisticated over here, essentially. Um, and so they've kind of like evolved to become like uh, the realms of like, you know, medicine and scientific discovery and so on and so forth. Um, and so, so the way we generalize the idea of like, you know, self-play to uh, scientific reasoning and scientific discovery over here, uh, is through this notion of like scientific debates and like self-debates. Um, and so what we have within our agentic setup, and I'll talk about that in like more detail in the next slide, is essentially like uh, a team of agents. Um, and we have these agents like continuously engage in like debates with each other over here. Um, and so they kind of like, you know, continuously generate like scientific hypotheses. Uh, but they also like, you know, refine and review and critique each other. Um, and we let this uh, process go on for like a period of time. Uh, and we continuously uh, introduce like net new knowledge and like reward signals. Um, and using that reward signals and using this net new information, uh, the system is able to kind of like, you know, self-improve itself towards like, you know, better quality like scientific hypotheses uh, over here. So that's kind of like the key idea. So we just let this process of like self-debates or like scientific debates uh, happen for like a period of time. Um, and then the composite AI system that emerges from this multi-agent setup is actually capable of like uh, scientific hypothesis generation. Um, and the ability to kind of like, you know, solve like complex scientific tasks over uh, long time horizons over here. Um, so hopefully that's uh, helpful uh, framing.

Um, and um, in this slide, I have like a deep dive uh, of like the, the architecture of this multi-agent setup over here. Um, and so Co-scientist, I would say, is like one of the first examples of uh, a general-purpose multi-agent system for uh, scientific discovery. Um, and uh, again, there's like a lot of text over here, but trust me, it's actually quite simple. Um, and um, and as I mentioned before, like the, the mission was to kind of like build out like a collaborative AI system or like a partner for like scientists over here. So we imagine that like the human scientist uh, would always be at like the, the uh, the driver's seat over here. So they would be the one who would be like, you know, guiding the system. Um, and so, so that's how like the input-output interface is set up. Um, and so it's all like set up in like natural language. And so what that means is that, you know, like you can literally like, you know, post to it like any kind of like scientific problem over here. Um, and so, you know, the way that happens is like, you know, the scientists can come in and they can like, you know, specify this uh, research goal. Um, and so this can be quite short, uh, or it can be quite long, and so you can like add in all the details that you want about like the scientific problem that you want to solve. Um, and you, and the idea is to kind of like give as many details as you need to the system so that it can like do a good job of like, you know, coming up with like the right kind of hypothesis or like solutions that we want over here. Um, so in addition to like the research goal, uh, the scientist can specify, for example, like, uh, what are like the some of the constraints, um, that the system must like satisfy in the solutions that it generates. Um, like what might be like, say, some, um, uh, axes or like rubrics to like, say, rank like different comparable hypotheses over here. Um, they can also like articulate like, you know, some of like the uh, the preferences that they have. Um, but not only that, like, you know, much like, you know, say like uh, your PhD supervisor, uh, might like, you know, tell you like, okay, like, go and explore this initial direction over here, like the scientist can say, okay, like, you know, like this is the problem, but I think like the solution might lie over here, so go and like figure this out, essentially. So they can give the system like initial, like, you know, seed directions to explore over here. Um, and uh, again, like they can also give like additional forms of like multimodal data over here. So it can be like, you know, uh, PDFs of like, you know, related publications and things like that. Um, it can also be like, you know, additional uh, experimental data over here that they might have generated in the lab. Uh, so essentially, that all forms like the context that is needed for the system to uh, perform its um, uh, hypothesis generation task. And so that's all like structured and set up as this research goal. And that's kind of like the input to the system.

Um, and then what the system does is, it kind of like does like a bunch of computation, and as I said, this is like dynamic, and so it can like go on for like as much time as you want. So it can like go on for like, say, a few minutes to like, you know, a few hours to sometimes even like days and weeks, depending on like the complexity of the problem here. Um, and at the end of that, what the system produces is essentially, like I would say, like a research report with like uh, a bunch of like hypotheses or like solutions towards like solving the problem that the scientists had articulated over here. Um, so again, like it's all like the, the interface is like very simple. It's like, you know, research goal in, um, and like, you know, research proposal or like research summary out with like a bunch of hypotheses or like solutions to solve uh, the problem. Um, and then within the system, like as I said, it's like a multi-agent system. Um, and the way these agents are configured are essentially to kind of like continuously uh, generate like different kinds of scientific hypotheses, but also like, you know, review, critique, rank, and prioritize them over like a period of time. Um, and uh, these agents are all essentially like our latest and greatest like Gemini models over here. Um, and we kind of like set them up with like, you know, different kinds of uh, strategies and system prompts. And so we kind of like force them to do like some very specific, specialized uh, uh, like tasks within this multi-agent setup. Um, and so again, like there's a bunch of text. Um, but I would say that the simplest way to think about this system is to essentially think about it like a computer program with like a while loop. Um, and so there's like, you know, like one function or one method which corresponds to like, you know, continuously like generating different scientific ideas and hypotheses. Um, there's like another function to kind of like, you know, review and like critique those hypotheses over here. Um, there's like a third function to kind of like, you know, rank and prioritize these hypotheses. Um, and then there's like a fourth function over here, which is to kind of like, improve and evolve those hypotheses. So that's it. It's like a while loop with these four functions that are kind of like running asynchronously and continually. Uh, and all these functions are implemented using like uh, agentic LLM approaches over here. Um, and uh, yeah, and so like each, as I mentioned, like each of these uh, agents have um, um, yeah, so, so each of these uh, agents have their own uh, system prompt set up over here, and so we call this like a library of strategies. And so we kind of like give the each specific agent like, okay, like a set of like strategies that can like select from to perform like the tasks that it has been given. And so for example, to like, you know, generate like scientific ideas, like one strategy is to like, you know, go read all related research papers to that given goal, and then like use those research papers to like, you know, build up like a foundation of knowledge, and then use that to like propose like some idea over here. Um, another strategy to like, you know, come up with scientific ideas is to kind of like simulate this debate, right? And so like uh, I think some of you might have also like experienced this, where like sometimes you get your best ideas by like talking to a friend, right? So you like, you both like start off with some initial ideas, but then you have like a back-and-forth conversation, and then you like, what emerges out of that is like an idea that is like much better than what you had like both initially to begin with, right? Um, and so, so we kind of like try to simulate that within like an agentic setup over here. So we ask the agent to like, you know, simulate like an interaction between like two experts, let it go on for like a few turns of conversation, and then at the end of that, like come up with like a scientific idea or hypothesis over here. So essentially, there's like um, hundreds of different like strategies over here, like within each agent, including like the generation agent. Um, and in fact, like um, like one of the favorite uh, like pastimes of our like our team members, uh, is essentially that they would like, you know, spend time like listening to these podcasts from like say, Terence Tao or like David Deutsch and others, and they would see, okay, like, how are these people like, you know, uh, thinking like in an abstract manner, like how are they coming up with like new scientific ideas, and then they would kind of like try to like abstract and summarize them and implement them as like a prompt within this uh, library of strategies that we have. Um, and it's kind of like a continual, like evolving body of knowledge over here. And so like you can keep on adding it, and then at test time, you just have to like, you know, sample like one strategy and then use that to kind of like generate a scientific hypothesis over here. Um, and so, so yeah, so similarly, we have like a bunch of these strategies for like the, the, the reflection agent, which is uh, essentially doing like this review. Um, and then the, the agent that I want to maybe spend like more time on over here, uh, is actually this uh, ranking agent. Um, and so what this agent does, uh, is it kind of like takes the hypotheses and the ideas that have been generated so far. Um, and then then it kind of like tries to like, you know, rank and prioritize them. Um, and the way it does that is using this debate mechanism that I uh, mentioned before. Um, and so what it does is it kind of like, you know, orchestrates these uh, debate matches between these ideas. Um, and then it kind of like tries to like, you know, simulate like a few turns of conversation, and then essentially the goal is to kind of like, you know, rank these ideas. Um, and say like, okay, like which one is better based on like the uh, the criteria that the scientists had like, you know, articulated like initially in their research goal. Um, and so that's kind of like the setup. And so it, it, it does the scientific debates. And once you have like a few of these debates like set up and run, uh, what you can then start doing, um, is you can actually start like computing these ELO scores. So you know, much like in a Go or like in a chess tournament, like you can use the ELO ratings to kind of like sort and like, you know, rank players. Uh, similarly, you can take the same approach over here. Uh, and use that to like, you know, rank and like prioritize the different scientific hypotheses over here. Um, and um, what purpose it serves is uh, essentially this, right? So, uh, if you were to like, simply have like only like a generation or like a review agent from the system, then I think what that would end up being uh, is like a system that's kind of like generating like many, many like, you know, good ideas or like decent ideas over here. Um, but I would argue that that's actually not enough to like, you know, move the needle in terms of like uh, scientific discovery over here, right? Uh, because like when you're like, you know, talking to scientists, like you would see that they're actually like not short on like ideas over here. Like, like the expert scientists have like, you know, more ideas than they have like, you know, time and resources for. Um, and so if you're like truly building out like a collaborative partner for these scientists, then you have to kind of like respect their time and their attention and only surface like ideas that are like really worth their time and their attention over here. So you have to do like a very good job of like prioritizing the outputs from the system, like really like summarizing them, and also like have this notion of like epistemic humility within the system, right? Um, and when I say epistemic humility, like what I mean is that like the system should be able to like, you know, really convey uh, when it actually doesn't know about something. Um, so it should be able to clearly articulate like, what is its confidence about like a given hypothesis, and also say like, okay, like, what are the key uncertainties that need like resolving over here? Um, and so that is what this ranking agent kind of like tries to do. So it essentially tries to like, you know, prioritize these ideas. It tries to like say, okay, like which one is better, and which ones like, you know, spend, uh, is like worth spending like time and attention on. Um, and it also like tries to like, compute, okay, like, what is the uh, uncertainty associated with each idea over here, and like propagate that to the scientist. So that's like one purpose that it serves. Um, and then the second purpose that it serves is essentially because these uh, debates are all like happening in like natural language, um, what you can do is you can like, you know, create like summaries of them. Um, and then you can like put all of those summaries back into like the overall like memory of the system. Um, and then what happens is like, you know, the next time like any of these agents are doing their uh, work, what they can do is they can like use it, I mean, read this summary from like the memory, and then use that to kind of like, you know, say, come up with like better hypotheses or do a better job at like generating reviews and so on and so forth. Um, and what that helps is it kind of like creates this nice feedback loop in the system that kind of like leads to like a self-improving cycle over here. Um, and so, so that's kind of like the crux of the computation. So as I said, it's like quite simple. It's like a while loop with like four methods over here. Um, and if you were to like, kind of like zoom out and think about it, um, it's almost like I would say like uh, an agentic uh, in-silo implementation uh, of like the thought process that happens in a scientist's head, uh, when they're like trying to like come up with like say, like new scientific ideas and things like that. Um, and then once you let this computation go on for like a set period of time, and that depends on like, you know, when you reach your end state, um, and again, like there's like a few mechanisms in which you can like reach this end state. One uh, way could be like, okay, you say like, keep on going until you have like X number of ideas. Uh, or like another method could be like, you could just tell the system like, okay, like, um, maybe just like, uh, stop when you can no longer solve the problem or like, you're kind of like stuck over here. Um, and so once the system reaches that end state, like what it does is it kind of like, you know, visually clusters all the ideas that it has kind of like explored so far, um, and then like puts all of them together into this nice summary talk, which in turn is then like, you know, fed back to the scientist over here. Um, and so that's it. So that's like the crux of the system, and as I said, like again, all of these are like uh, agents implemented using our latest and greatest like Gemini models, and so they have all the long context capabilities, the multimodal capabilities, the agentic tool use capabilities, and everything that we need over here. Um, and so, yeah, so hopefully this was helpful, but maybe I'll stop over here right now and maybe take like a few questions because I thought that that was quite me. Anyone?

>> Yes.

>> Um, how, how is the like reverse systems like, how would you evaluate whether a hypothesis is good or not? Like

>> Um, I mean, even for scientists, like different people have different views and how this agent system could evaluate um, hypothesis.

>> Yeah. Uh, so that's the next half of the talk, but uh, yeah.

>> In the 25 Co-scientist system,

>> In the, in the 25 frequent, the Co-scientist Agora system, didn't really see much performance saturation or plateauing with test time compute. Have you explored those scaling of further symptoms?

>> Yeah, um, so, yeah, I mean, we can uh, sorry, we can discuss that um, in more detail. So I think it, it comes down to like the, the class of problems that you're working with. Um, and so there's like a class of problems where like the search space is like so big that if you keep on like, you know, throwing like more compute at the problem and let the problem like intelligently like explore the search space, then there's very likely to like, you know, come up with like better solutions. Um, and so, so for those classes of problems, um, essentially there's like, you know, in some ways, like no limits, right? So there these class of problems where you're trying to optimize something, and so you'll all, I mean, if you have like a well-defined like fitness function, then you can keep on like, in some ways, like optimizing it and making it better and better as long as you have like more compute. Um, and so for those class of problems, we essentially do not see much saturation over here, and so you can like keep on pushing the system. Um, but then there are going to be like, say, like, like two classes of problems at like both extremes, right? And so there's like one class where the the problem is like so easy that there's like no point in like, you know, spending so much compute over here, right? I mean, like why do you want to spend like, say, like millions of tokens when like it's like just such a trivially easy problem, right? So it could be like some retrieval problem or like, you know, something that's extremely like simple. Um, and then at the other extreme are like problems that like, I mean, we just don't have like the knowledge to do to like, you know, solve. Like for example, if you ask the system to do like, say, like, you know, build out like a time machine, it's not going to be able to do this. And so like no matter like, you know, how much uh, compute you throw at the system, it's not going to be able to do this. Um, but the good thing is there's like a large class of like problems in the middle, and so over there, like you keep on like throwing more compute, you keep on like introducing like new information, and essentially

You see this like very nice scaling property where the system is able to like, you know, come up with like, you know, better and better hypotheses over like a period of time.

Yes. Actual types of science fall in these different categories, not like...

Yeah, um, so that's the rest of it. I'll show you like a few validation examples, uh, uh, where we like applied the system.

It seems like a very linear way of scientific progress, like falsified. But modern action science is like, are you like in our keyboard? So, but you have like a system where you get like a bunch of PhD pieces automatically, but you never actually get to the next level of having a breakthrough. There's a bunch of like the side research. So, how do you still keep the human in the loop to be informed enough to actually make the breakthrough, if that's like a, you can achieve, or do you not believe in that?

Yeah. Um, I think that's a very good question over here. And so like, I think, uh, one of the goals that we have is to like, uh, generate like such compelling ideas and hypotheses that the human who's in the loop, like the expert who's in the loop, uh, they would be kind of like willing to drop all their ideas that they have and work on this one. Um, and so if you're able to kind of like do that, like enough number of times, uh, then that means that your system is working. And the hope is that that kind of like leads to like, you know, more discoveries and like hopefully breakthroughs. Um, and so, so yeah, so that's kind of like the goal in some ways, like to just develop things that, you know, scientists did not think about. Um, and yeah, I mean, like maybe all of them don't pan out, like they don't work out because like that's the nature of like science. And, uh, but like, yeah, the idea is to like, you know, keep pushing and like generating like such compelling ideas, um, that, yeah, the scientists are incentivized to like go and then try them. Uh, but I mean, you, you are spot on. Right? I think like we, we'll soon have like a, like we'll soon enter this era, or maybe we're already there, but like there's going to be like so many like, you know, compelling hypotheses that actually the bottleneck will be in like the verification and the validation over here. Um, and so, so yeah, I mean, that's kind of like the next step, okay, like how do we, uh, prioritize them? How do we like allocate the right amount of resources and pick like the best ones for like validation?

Yes. Sorry, specialization.

Yeah. So there's no specialization for science per se. Um, they are essentially all our Gemini models, but some of them might use, say, like a flash variants versus like our Pro variants. And so it depends a little bit on like the nature of the task. And so if you're doing like a review or like a verification task, then you might want to use like the Pro variants, which are maybe a little better at like, you know, reasoning and thinking over here. Whereas if like the task is like a little bit more simpler, then you might want to use like a flash model just because of like, okay, you want to save some tokens and like reduce the cost. Uh, any other questions?

Okay, great. Um, so over the next few slides, like, uh, what I will do is like, I want to walk you through some examples of like validations of the system, right? So if I'm saying that, you know, system is like capable of like, you know, scientific discovery, uh, I mean, the real proof is to kind of like, um, validate it in the lab and show that the hypotheses are actually working out. Um, and so, so the first example that I want to highlight over here, um, is, uh, it's actually more of like a recapitulation over here. Um, and so what happened is like, I mean, when we were developing the system, and this was back in, uh, 2024, uh, during the Thanksgiving break. Um, and so we knew certain researchers at Imperial who've been like, uh, working on this concept of like, uh, antimicrobial resistance. Um, and so we kind of like reached out to them and we told them that, hey, like, you know, we're building this system called like the co-scientist, uh, do you have like any use cases or, uh, applications for it? Um, and then like, essentially what they told us was that, like, that actually made like a very important and like compelling discovery, uh, related to like this very novel like, uh, horizontal gene transfer mechanism in, uh, bacteria, which in turn is like, you know, responsible for like antimicrobial resistance. Uh, but the, the crucial aspect was that they're actually not yet like published it anywhere. So they had the data and they were kind of like in the process of like, you know, writing up the paper. Um, but like the rest of the world did not know about this discovery or this breakthrough over here. Um, and so, so when we reached out to them, they said, okay, like, we, we think like a good test of your system might be to kind of like, see what it produces if we were to kind of like, give you the same research goal that we've been like, you know, working on for like the last, uh, 8, 10 years. And so essentially like their research journey or their timeline, uh, is captured on like the left-hand side over here. Um, and then on the right-hand side, uh, is our experiment. Um, and so, yeah, so they gave us this research goal. Um, and we kind of like, you know, let the system run for like a couple of days. Um, and then we shared the results, uh, back with Jose and Thiago, uh, who are like the key PIs, uh, associated with this, uh, research over here. Um, and, uh, maybe you can just hear from them like what they thought about these research, uh, results. I don't have internet access, but it's fine, but yeah, maybe I can just tell you what happened. Um, and so as I said, this was during like the Thanksgiving break over here. Um, and, uh, when, when we shared the results, like literally like, you know, 30 minutes later, uh, I get this, uh, email from like Jose and he says, like, uh, I need to talk to you right now. Um, and I was like, I mean, Jose, hold on. Like, you know, this is like the Thanksgiving break. Like, I'll, I'll call you. And then, uh, I end up like calling him. And then he was like, literally like, you know, like, no, hi, hello, nothing like that. Um, and, um, like he was straight up like, do you have access to my email? Um, and the reason he was asking that was because he thought that we were cheating. He thought that like we were like, you know, reading his email and like, you know, uh, using that to like guess the hypothesis over here. Um, and then like, you know, we had to like, you know, tell him that, you know, I mean, Jose, like, you know, we do many things at Google, but, uh, we don't read your email. Um, and then he was like, immediately like, okay, like, do you have access to my ChatGPT? Because like he was using ChatGPT to kind of like, you know, help him write his paper over here, right? Um, and then we again had to like, you know, tell him, like, you know, Jose, like, I mean, come on, like, we're not friends with OpenAI just yet. Um, but, uh, anyways, like, you know, when, so Jose is like not someone who gets like easily excited by things. He's like a very, you know, seasoned, uh, researcher, right? Um, and so when he had that, you know, visceral reaction, um, I think that was like the first moment when we felt that, okay, like, we were on to something with the system, uh, because at the beginning of the talk, like, as I was telling you, like, we were like, at all, I, we were not, not at all sure that, okay, like, this is like even possible at this point of time. Um, but like, you know, when, like someone like Jose had that reaction, um, I think that is when we thought, okay, like, okay, maybe we are on to something with the system. Um, and so since then, like, what we've been doing is, uh, we've been kind of like opening up access to the system to like, uh, more and more like scientists and researchers around the world. Um, and, uh, and I think what we're starting to see is, uh, this emergence of like a new form of like, uh, AI human scientist collaboration, collaboration over here, uh, that's actually like leading to like new insights and discoveries and breakthroughs. Um, and so for example, um, there were like, you know, certain physician scientists at Houston Methodist Hospital in Texas. Um, and so they were able to use the system to, uh, like identify like new repurposing drugs, uh, and like combination therapies for like, you know, very complex forms of cancers, like, uh, acute myeloid leukemia. Um, similarly, like, you know, Dr. Hey, Gary Pel over here at Stanford. Uh, he was able to use the system to first like identify like, you know, new, uh, epigenomic targets for like liver fibrosis. Um, but not only that, he then went back to the system and he said, okay, like, okay, give me like an experimental protocol to validate this in my organoid setup that I have in my lab, and, and also like, you know, suggest like some new drugs, uh, that can help me like validate this target over here. Um, and so again, I'll, I'll show some results in the next slide, but that also like worked very well. Um, and then finally, what we've been doing with the system is actually doing like, you know, more and more complex tasks, like discovering and engineering like, you know, de novo proteins with like different kinds of activities, uh, that we want for like a bunch of different applications. And over here, the system is like using tools like AlphaFold and like, you know, binder design models, all in the loop, and performing like these complex scientific effect tasks over like, you know, longer time horizons. Um, and so, so yeah, so this is like the, the, like, just a snapshot of like the results from like the, the drug repurposing experiment that we had. Um, and so like on the right-hand side, like essentially what you see is like a bunch of these, uh, IC50 curves. Um, and so like these are essentially like validations of like the predictions that the, uh, the co-scientist had for like certain specific drugs. Uh, and they were like tested on like a bunch of these different cell lines. Um, and essentially what these curves are trying to see is like, okay, like at what drug concentration, uh, are you seeing this, like tumor inhibition activity over here, like, okay, like when are you like actually like, you know, uh, killing the cancer cells, uh, and the lower the concentration, the better it is. Um, and essentially what we're seeing is like, across these cell lines, uh, like the, the, the drugs are actually working at like reasonable concentrations over here. Um, and so that's like exciting. Um, and then the second thing is like, again, like I just copy-pasted like a snippet of like the, the novelty review that the system had like generated for its own hypothesis over here. Um, and so what the system had predicted was, okay, like, let's try this drug called Kira 64 AML over here. Um, but it was actually like quite calibrated in its like own sense of like novelty over here. So it was very clear and said that, okay, like, uh, like this is not like completely like a breakthrough thing because like there are other drugs that work on the same pathway or like the same target. Um, and they've been tried in the context of like AML before. Um, and so, so that gives me confidence that like Kira 6 will also work because it works on the same target. Uh, but essentially like Kira 6 has not been like studied extensively and there's been like no clinical trials over here, and so that's why it makes sense for you to like, uh, try it. So it's kind of like calibrated over here. It's like grounding its recommendations on like evidence and like, you know, past, um, literature and so on and so forth. Um, and so that's kind of like helpful, and that's the reason why these physician scientists like picked up and ran these experiments. Um, and so this one has like the results from the, the, the liver fibrosis experiment that I mentioned. Um, and so over here, like, uh, Dr. Propels, like he took like four suggested drugs from, uh, the system, um, and he tested it in his hepatic human organoid setup. So these are like liver organoids. Um, and essentially like, you know, all the four drugs over here, um, like all of them show like very promising, uh, anti-fibrotic activity. Um, and in fact, this one, like this suggested one over here, this is actually like a well-known and safe and FDA-approved like anti-cancer drug over here. It's called Voras. Um, and the reason this is, uh, interesting is because it kind of like shows the complementarity of like the human intelligence and like the AI over here, right? So, so if you're like a liver fibrosis expert, you may not be aware of like what's happening in like the cancer space, right? So, you may not be thinking about like all the therapies, all the progress, all the pathways that like people are researching. Uh, whereas this AI is kind of like able to like, you know, go broad, uh, and see what's happening over there and make these unexpected connections and present that diverse human, uh, viewpoint back to like the human scientist who can then use their own judgment and their deep expertise to figure out, okay, like, does it make sense or not, right? So in some ways, like humans have this deep expertise, whereas this AI can like go broad, and if you're able to like intelligently bring these two kinds of like intelligence together, then hopefully what you would get, the sum of the parts, hopefully that's like much greater than like the, the individual systems alone. So this is like a good example of that. Um, and so this is like another like more interesting, like a recent one over here. Um, and over here, like, uh, what happened is like, uh, there were a few researchers, uh, based out of like the Science Lab, uh, in the UK. Um, and so they actually had like a bunch of these, uh, alpha structures over here. Um, and, uh, they were trying to like figure out, okay, like, how do I actually analyze this, like, how do I discover like, uh, anomalies in this data, like which ones, which structures are the ones that are like, you know, like worth paying attention to over here. Um, and so what they asked is like, they used the co-scientist to like, you know, build up this, uh, structural novelty index over here. So it's essentially like a bunch of these features. Um, and then they use these features to kind of like, reanalyze the full structural predictions in the data. Um, and what they were able to uncover, uh, was actually like a, a completely novel and like previously unknown, like a massive, uh, potato-like immunoprotein over here. Um, and so usually, like, you know, what plant biologists assumed was that like, you know, most of these, uh, plant-like immunoproteins, I mean, they're called like resistance genes, they'll have this hexagonal shape, whereas like what this new, uh, insight of this new analysis, like what it uncovered was like, they're actually like much more novel and massive, like immunoproteins, like, so it's called Levomer over here, essentially, it's like that, that's like quite, uh, wild in some sense. And so this was again, like a preprint that came out like a few, uh, weeks back, and I think this has like quite massive implications in terms of like, uh, decoding plant immunity and how that might like impact like, you know, agriculture and like global food security and so on and so forth.

This is like another example. And so this is where, um, we have like co-scientists using AlphaFold in the loop. Um, and what we have it doing over here is to like, you know, design these better Oct proteins. And so Oct4 is one of those classic like Yamanaka factors. And so these were like the set of, uh, transcription factors for which like, uh, Professor Yamanaka was awarded like the Nobel Prize back in, I think, 2010 or 2011 over here. But the thing with these, uh, transcription factors is like, you know, while they rejuvenate the cell, I mean, they can also cause like, you know, uh, like tumors and things like that. So you want to like, you know, have like better versions of these proteins that have like the same activity but are kind of like, you know, more safe over here. Um, and so that's what we, you know, asked the system to do. So we asked it to like, you know, predict and design like better sequences, uh, that have like the right activity and like the right, uh, delivery properties and so on and so forth. Um, and so we kind of like asked it to like use AlphaFold in a loop, and that's what it does. It kind of like, so it makes like a prediction of like a sequence. It uses AlphaFold to like look at like the structural predict, uh, stability, and then uses that to kind of like refine its hypothesis, and it kind of like does this in silico, like a bunch of times. And then what we observe is like the final sequence that emerges, uh, that actually has like the right kind of activity that we're looking for over here, and the right properties that we are looking for over here. Um, and so we're doing a lot of this right now. We're taking and testing all these de novo protein designs in the lab, and they're all like looking like extremely, uh, promising right now.

So this is like another example. So this one's like, uh, unpublished yet. Um, but over here, like, what we asked the system to do was to kind of like, you know, go read all the literature, look at like all the, uh, data, and then like, kind of like, predict like these new, say, genetic factors or these secreted proteins, uh, that might actually take like a cell, uh, and rejuvenate it or like reduce like the number of like senescent cells over here. So senescence essentially means like your cell essentially becomes like so inflamed, uh, that it doesn't divide anymore and causes disease essentially. So you want to like reduce the number of such cells in your body, and then hopefully like you can like, you know, reverse them back and like, you know, make them like younger and be like, you know, more reflective of like, uh, younger individuals over here. So that was what the task that we gave. And so we told the assistant, like, okay, like, go find these novel or like, you know, unstudied, uh, genetic factors and things like that, and just like, give us like a few. And so that's what it did. So it ran for like a few days. It gave us like a list. Um, and so we tested like a few of them and we compared it with like, there's positive control, like Clot, which is like very well-studied. Um, and essentially we see that like the AI-nominated factors, um, they actually have like the same, like fold reduction in terms of like, uh, the, the number of percentage of senescent cells over here compared to this positive, uh, control. Um, and so essentially, if this like plays out, I mean, we have to do like, you know, more experiments over here, uh, we might have some like net new, like rejuvenation factors, and that would again be like extremely cool, if this like, you know, works out. Um, yeah, and then maybe like another example that I would want to like, you know, very quickly walk you through, uh, is, is a case study related to, uh, Alzheimer's disease. Um, and so this is kind of like similar to like the, uh, antimicrobial resistance example that I showed before, which was also like a recapitulation. Um, but over here, it's a bit more complex in the sense that, like, there were certain researchers at Mass General Hospital. Um, and what they were trying to do was they were trying to like, uh, study this linkage between, um, like ACE inhibitors, uh, which are like again, like very popular, commonly used to, uh, treat like blood pressure and like heart diseases and things like that. Um, and so essentially they found this very interesting thing that like, you know, uh, people who are like, you know, taking these ACE inhibitors, they actually have like heightened risk for like Alzheimer's over here. Um, and so what they wanted to do was they wanted to kind of like understand like, why is that reason? Um, and so, so they did like, you know, years of like experimental work, and they were able to like, you know, come up with like a mechanism and a prediction, and so it had like a very complex, like nine-step, like, uh, cascade over here. Um, and so what we asked the AI to do was to kind of like, okay, we gave it the same problem, and we asked it, okay, like, what do you think is the reason why ACE inhibitors like lead to like, you know, increased risk of hypo, like Alzheimer's? Um, and essentially what it did was, it kind of like recapitulated their entire, like nine-step, like mechanistic cascade over here. But the, the cool bit was like, it was not only that, like it also like predicted like one of the key steps that the scientists had actually missed. Um, and so again, to like simplify this a lot, like essentially the thing is like, um, so if you take ACE inhibitors, they kind of like modulate this, uh, chemical signal called bradykinin, and that in turn like leads to like, um, there's a, there's a surface of, there's a receptor on like the surface of your brain cells called like this B2R receptor, and so it kind of like triggers, uh, this B2R receptor, and that in turn like leads to like, you know, neurodegeneration over here. Um, and so essentially like, what the AI said was that, uh, the, the missing link was essentially this link between, uh, bradykinin and this B2R receptor. Um, and so, so what these researchers then did was they went back to the lab and they did this, uh, protein stability assay, which is called a chase assay over here, and essentially that, like missing step was also kind of like, again, like validated. So the AI was not only able to kind of like, you know, recapitulate what they had exactly done, but it was also able to like fill in the finer details over here and like complete their hypothesis. Um, and this is actually like a good, um, example of like how this agentic scaffold is actually better than like base LLMs. Uh, because what we did was we also gave like the same problem to Claude and GPT-5. Um, and essentially, while both those systems were able to get like this first step and like the high-level hypothesis right, um, they were actually not able to like get the full details over here. And, and so you really needed this agentic hypo, uh, like harness to be able to like arrive at like the key details and the specifics over here. Um, and so that's again like, it's telling you like how this agentic harness is actually like helpful, like the time that you're spending on like verification, like, you know, coming up with like the right kind of hypothesis, that is how like you get to like, you know, much better results and like, you know, genuinely novel scientific insights, and discoveries. And, uh, yeah, the last one I want to again like, very quickly highlight, uh, again, like one of the things that I do besides like trying to help, uh, improve the system is actually just ask it like interesting questions. Um, and so like one of the things that I've been like interested in is like this relationship between, uh, neurodegenerative diseases and cancer. It's actually kind of like well-known that like, you know, if you are an Alzheimer's patient, you actually have like much lower risk of like cancer. Um, and so that's kind of like, you know, not super well-studied. Um, and so what I asked the system was to just go and like read up the literature. Um, and then tell me, okay, like, what are maybe some, you know, understudied pathways, uh, that are, you know, very common in like neuro diseases, like neuro genes, essentially, uh, that might actually be implicated in cancer as well. And, and so just tell me like, okay, these genes and these specific forms of cancer that we should like, you know, go and look at. Um, and so when I gave this prompt, and this is, this is the inverse comorbidity problem, right? So you have like one disease, and that just reduces the chance of you having like another disease over here. Um, and essentially, if you were to like, you know, identify enough of these pathways, they could also be like, you know, targets to kind of like, uh, build like drugs for those specific indications. Um, and so I gave the system this prompt, um, and, um, like it came up with like a pair of like genes over here, which are like common, like neuro genes, and so there was these two genes called DHX9 and SRRM4. Um, and not only that, it also told me, okay, like, you know, if you want to validate this gene, you should reach out to this professor, Dr. Filippo Bellajia, based in Germany. Um, and so that's what I ended up doing. So I just took the hypothesis and I wrote like a cold email to this professor. Um, and what he did was, he had like a bunch of like, uh, whole genome, like CRISPR screen data, and he went back and looked at it, and essentially what he saw was that in the context of like, small cell lung cancer, which is what the system predicted, like these two genes were like way overexpressed compared to like other forms of cancer over here. So essentially the hypothesis was right. Um, and so again, it's like another cool example. It's like unvalidated, but it's essentially like the system like making a connection and like surfacing it. Uh, and that's like, you know, leading to like new, new insights and potentially like discoveries and breakthroughs. Um, so yeah, so I'll maybe stop, uh, on this note, and so you can see like, uh, some of the, uh, the results from the system. So this is like a snapshot of like the hypothesis that I generated. Again, these, these reports are like, you know, super detailed, like 100 plus pages, but you can see that it's all like nicely explained with like images and things like that. Um, and again, this is like one of the research contacts that I mentioned, like it tells you like if you want to like, you know, validate this hypothesis, you may want to like reach out to these people, and it also tells you like, okay, like what are some of the unexpected things that it discovered over here. Um, so yeah, so that's like the co-scientist part, and maybe I can stop over here and take like a couple of questions if you have any.

Have you tried playing around with knowledge cutoffs at all? So like, uh, Hockey has been, uh, been getting pressed lately, 13 billion parameter model, but only on pre-1931 text. Can you predict stuff that happened after the cutoff?

Yeah. That's a great question. So I think the Imperial example that I showed was exactly like one of them, and even the Alzheimer's one as well. Was that like over? Yeah. Uh, I think it's a bit difficult to like isolate, uh, corpuses like that because like there's like so much leakage that's happening. Um, uh, but one thing that we've been kind of like setting up and doing is like, we're using the system to like, you know, predict like future events that might happen. So for example, like there's like a bunch of like, you know, clinical trials that are happening, like predict the outcomes of them. And it has like reasonable success so far. Uh, but again, it doesn't have all the information because like a lot of this is like proprietary in nature. Um, so yeah, I mean, my, my intuition is that if you have like the right kind of like knowledge, but like not like fully revealing the details, then these could be like extremely powerful in like predicting clinical trials and things like that, and that that could be very helpful. Uh, but yeah, I mean, like doing the experiment that you're talking about, like, you know, put like, yeah, like that's something.

Because we have to have more and more papers, but peer review is only takes too long, and papers are only going to be read by other agents. So isn't there a better way to kind of have a peer review system and like another space for other agents to work on this other than waiting like nine months before it's finally passed like science?

Um, it's, it's a great question, and honestly, I don't have like a good answer to it. Um, I mean, we obviously starting to see like a lot of like AI usage in like papers, and uh, uh, I mean, I think it was just this morning where arXiv came up with this new policy that, okay, like, if you have like hallucinated, uh, references in your papers, then they would block you from using arXiv for like a year. But I don't think that's like very productive, uh, in my opinion. I think we have to have like better mechanisms. Um, and again, like people are using AI in peer reviews, and that can be helpful, uh, but like the risk is again that, uh, if you don't use it like judiciously, then you're going to like favor only like certain kind of papers that kind of like pass through this AI filter or funnel, and then you leave like the other topics, uh, unexplored. Um, so yeah, I think so that, that is going to be like an important challenge over here. Um, again, I don't, I don't have like a good answer. I think there's also like things that we can do to enable like agents to share like the discoveries and information. Maybe there's need for like a science context protocol over here. There's need for like, you know, formats that can help you like, you know, standardize like, uh, scientific data and like how they are generated, how they're like kind of like shared and like, you know, decentralized. Um, so yeah, I mean, there's there could be a lot of like cool engineering work that we could do, but again, like there are going to be like some thorny issues also that we have to like address with like the flood of papers that's coming, because like humans are not going to be able to like, you know, keep up with it.

Uh, any other questions?

Yes. So for the reports that you generate, you get around 100 pages.

Yeah. I just wonder how, I guess humans going through, how, you know, it's quite...

Um, so like one thing we try to do is that again, we want to have all the details in the reports, but we also tell the scientists, okay, like where to spend time and attention on. So we tell, okay, like, uh, maybe this is like the most compelling or interesting hypothesis. Um, and so you may want to read them first. Uh, but again, like it's also possible that the system is kind of like too conservative. Um, and so even if it thinks that certain hypotheses are like not viable, uh, they may be still having like some interesting clues or like some interesting insights that might help the scientist, right? And so, uh, that's why we still like keep them in the report, but we clearly tell the scientist, okay, like, you should focus your energies on this particular section or this particular path because that seems like the most promising one. Um, and on the flip side, there might be like, you know, certain explorations where the system just comes back and says, like, you know, none of the ideas that I have, like so far are going to work, and so you may want to like, you know, reframe the problem or like, you know, make it like, uh, you know, maybe focus your attention on like a sub-problem and things like that. And so that happens quite a bit in like mathematical problems over here, where if you like make it solve like the full, like come up with like the full proof, then it kind of like breaks down, but if you like break it down into like the right steps and the right chunks and have like the human work like iteratively and like, you know, interact with it, then it's able to like, you know, make progress over here. And so again, like we have all those mechanisms set up in the system to be able to do that.

For the case studies that you just talked about, you mentioned you run it overnight or over a couple of nights. How many 3.1 Pro or whatever tokens, order of magnitude? How many agents there are in parallel?

Uh, maybe I'll tell you once this is not recorded.

Any other? Yes. Uh, do you guys, how do you guys think about in terms of bad actors that might have?

Yeah, it's a great question. And so, yeah, we have like different layers of like, uh, safety checks over here. Uh, so there's one that happens right when the scientist is like specifying the prompt or the goal, and so over there, we check and make sure that, okay, like, it's not something, um, nefarious. Um, but often times like, you know, you can get around that, you can phrase it in like different ways. Um, and so, so what we also do is, um, we also like, you know, monitor like the, the ideas that it's like continuously producing and like the exploration that's happening, and again, if we see like, say, a percentage of the ideas are getting into like this unsafe territory, and that's right now, I think set to like 10% of the ideas, uh, again, like we halt the computation and we tell the scientists that, okay, there's like an unsafe research goal, essentially. Um, and then finally, like, again, like we, because these agents are all based on like Gemini, and that itself has like CBR and text, uh, testing and things like that, that has like happened. So we are able to like, uh, get some of those properties from within the base model itself. Uh, but yeah, I mean, like when you have like a multi-agent setup, like the, the surface of, uh, where things could go wrong and like where things could be used, like, uh, in an improper manner, that also like expands quite a bit. And so we have to do like, we have to have this multi-layered approach to like safety, essentially.

All right. Unfortunately, I think that's all the time we have for today, but a round of applause for our amazing speaker.