Transcription
Activating the lab of the future through an open, modular, and cloud-native approach is what we're going to be talking about. Uh, now, uh, my name is Ravine Chararma. I have about 25 years of experience serving this space. I've worked at uh multiple small startup companies, and I'm currently at Deloitte where I lead the Lab of the Future practice.
Um, with me, I'm very privileged to have Yurus Vanam, who's making his way up. Uh, Yurus is a data and digital healthcare entrepreneur with 20 years of experience. Uh, he spent the majority of that time in pharma R&D and diagnostics, liberating data from the labs, clinical research, and clinical practice to accelerate development of new drugs and therapeutics to amplify their impact on public health outcomes. Earlier this year, Yurus started a new venture to democratize evidence-based care with the help of AI. So Yurus is going to help talk about kind of real-world applications and and kind of the battle scars that he's had with working with data and trying to get data liberated.
Right? So, what we're going to talk about is one, you can't ignore your labs. I'll get into, you know, why we think that's been the case. We're seeing incredible increasing investments in this space. There's still a lot of challenges and barriers. You know, it hasn't been solved. It's not a solved problem. We're going to talk about closing the loop between the wet lab and the dry lab. What do you need to achieve that? Um, what are we building with AWS? And, you know, what does the roadmap look like? And then it's all about making your scientists happy, productive scientists, or happy scientists.
So, you can't ignore your labs. So, we've got three kind of vignettes, three snippets on this slide. Uh, the first is, you know, the the story of the of the productivity in R&D, as measured by rate of return, has been declining pretty much for the past 12 to 15 years. There's been a slight uptick in the last two years, and we're sitting at about 5.9% return on investment, and two points of that is directly attributed to GLP1. So, you can say for the vast majority of pharma companies, they're sitting at about 3% return on investment, which is really, really low.
The other dynamic that we're seeing is uh loss of exclusivity. There's about 190 drugs that are going to come off patent in the next five years, and that represents a loss of revenue of about $250 billion. So, a quarter of a trillion, right? So, how do you, how do you accommodate that? Right? And so, you have to have a strategy of sustainable pipeline development. So, you know, for the short-term chronic problem, if you haven't planned for it, you're going to be buying compounds, right? And the price of those those compounds are are going up, right? So, really, as we've kind of interacted and interviewed um a large number of heads of research, heads of R&D in large pharma, basically they've come to the conclusion or they they've been been advocating that you need a balanced approach. You can buy compounds, obviously, but you're not going to buy yourself out of this problem. Number two is you have to invest in your own innovation engine, right? And so, you really need that balanced kind of approach. And this is why we're seeing a ton of interest, a ton of investment, and a re-look at your own labs, right? And this stat here kind of, you know, alludes to it. 69% of pharma leaders believe they will lose competitive advantage if they fail to connect and automate their labs. Right?
So, that kind of first-mile problem, you know, of getting the data off of your instruments, out of your labs, your on-prem, if you will, getting that efficiently moved to the cloud where you can start doing compute to that data and, you know, power all the ML um technologies and ML products that you've been hearing about today. That's that's a big that's a big problem still, right? And what we're seeing is again, there's investments being made in this space. So, from the large pharma all the way to the the AI-native or cloud-native pharma, the smaller uh innovative ones, if you will. But, you know, AstraZeneca, this is on their website, right? So, they're very proud of the fact that they fully automated a given lab. In this case, it's medicinal chemistry, right? Um, in silico. They're very proud of the fact that they've got robots generating data, and it's automatically moving into this PanOmics data environment, right? And Recursion is very proud of the fact that they've been assembling very rapidly one of the largest datasets in this area. Right? So, this is all great, but these are all pockets of innovation, pockets of investment, right? As a pharma company, you've got way more than a medicinal chemistry lab or a high-throughput screening operation, right? So, how do you get the rest of your labs kind of, you know, brought up to par, right? So, there needs to be a way to scale. You need to scale this.
The barriers that we see um, I'll go through it in a second, but what this represents is there's an organization uh called the Pistoia Alliance that a lot of pharma companies are members of. It's a pre-competitive kind of consortium where a lot of information, good information is shared. And this is uh two two surveys, one done in 2023, one done in 2024. And they interviewed about or they surveyed about 200 people, executives, lab managers, you know, senior research scientists, um, R&D IT folks who support the lab operations. And they asked them, what are your top investments as you're as you're looking forward, either, you know, both in 2023 and also in 2024? So, 62% responded that they're going to be investing in AI/ML. Okay. So, yes, then the next kind of biggest category was 30% responded that they're going to be investing heavily, their top investment are cloud computing platforms. Okay, that that kind of makes sense. And then the next two, the number three and number four, were were kind of interesting, and it helped us kind of understand really what's happening, you know, through this survey, what's happening in the labs. And one is that ELNs are being invested in heavily, right? And LIMS systems are being invested in heavily, right? So, you can start kind of maybe thinking, you're hypothesizing why that might be, right? And then, you know, the next kind of line in the slide, what are the barriers, right? When these folks were surveyed. So, you know, they're pretty much the same in both years. You know, data silos, unstructured data, it's hard to process, hard to get your hands on it. Um, you know, in 2023, there was resistance to data sharing and collaboration, which is called the data mining problem. You know, that data is mine, this data is mine, I'm not sharing it. It's an oldie but a goodie. Um, so the way that got fixed is through governance and through a top-down mandate saying, if you work for this company, if you work for this organization, you are going to share data. You are going to input the metadata fields for that experiment you just designed and had executed, and you're going to share it with the rest of the community within this organization.
So, you know, we see a shift in 2024, a year later, where, you know, you still have data silos, unstructured data, but what comes up is the lack of metadata standardization. And then that ties back into why people are investing in ELNs and LIMS systems is you're fixing your source data. You're investing in fixing your source data so you have that entire metadata record. So somebody besides yourself who may have run that experiment can actually understand what you did. Right? So that's very important before you feed it into an algorithm and have an agent do its thing on it. Right? So, really, what's holding us back is this concept of unfair data. I don't know where it is. I don't know what the context is about it. So, what unfortunately, we actually see happen more often than people care to admit is because they can't figure out how that experiment was run, what the samples exactly were, what the annotations are. I'm going to run that experiment again just to make sure, right? So that that's what's happening. It's it's unfair.
So, with that, Yurus.
Yep. Thank you, Ravine. Great uh um great segue. We've seen some fantastic examples earlier in this session and in the keynotes this morning of of benefits of implementing AI um in the labs, right? And that's what we tried to illustrate here. U at an abstract level, if you do more upfront work in your dry lab um, you can make better use, you can optimize the use of your scarce resources in a wet lab, right? Avoid running tests again, avoid running tests that are going nowhere. And and how do you predict that? Well, you need to know what you've done in the lab before. And when I joined Pharma R&D 20 odd years ago, the only way to find out what's been done in the lab before, you needed three things. You needed the actual data from the um from the experiment that was in a lab notebook. You needed a description of the experiment that was typically in a in a document store, probably somewhere on E-OT, totally dating myself. And you needed a person who worked on the experiment. With AI, you can replace that third thing. Think of AI as a great tool to rummage through your data. AI is not a magic black box. You can ask Perplexity today, what test should I run in my wet lab next? It's going to tell you something, but it's it's going to be rubbish. The real power of AI is feeding all your prior tests to the AI model and then asking that same question again. That's where the real power comes to bear.
So, how do you unlock that power? By making your data rummageable. The the the real term is FAIR, right? Findable, accessible, integrable, reusable. But that's what we're talking about here. And again, we've seen some great benefits once you make your data rummageable and and provide it to AI of um benefits of running the experiments in the lab itself. But that's not the only question that keeps our our leadership and our and our boards and our shareholders uh awake at night. They also look at organizational questions. I mean, over the past 20 years, how many iterations have we had from, oh no, the most productive R&D organizations are individual autonomous lablets, and then the new head of R&D came in and said, no, we have to have one central organization, and then a new head came in and said, we have to outsource everything to you to academia, and then the new head came in and said, the the point is not that um, if you organize your data well, it works better for one model or the other. The point is, if you organize your data well, you give yourself the organizational flexibility to do that. If you democratize the data that comes out of your lab, as I think the folks from Genentech called it, it doesn't matter how you organize, and you can try, and you can change it. Gives you much more flexibility from an organizational perspective to organize your uh lab experiments in whatever best way fits your your culture or um um or the time of day.
Um, that's particularly true in in diagnostic companies where labs is what they do, every day, right? It's it's their uh it's it's the product that they bring to market. Typically, the way lab companies start is a couple of brilliant scientists in a small lab somewhere, and they go to market as an LD LT, right? Right. Laboratory Developed Test, and you are allowed to do that as long as you run the test within that lab only. And lo and behold, the test becomes successful, and you want to market an IVD. Today, that's a rebuild because none of the scientists properly engineered uh the software pipelines and so on and so forth. So, investing in your um in your data infrastructure um gives you much more flexibility from an organizational perspective, from a scale-up perspective, but also from a portfolio perspective. The most tedious, the most daunting, and the least rewarding project you can ever give to your IT organization is selling one of your compounds. A close second is buying a new compound. But then at least you can blame the biotech. Selling one of your compounds means that you again have to rummage through your data. And if you haven't properly organized it, it takes days. Sorry, it takes months. Once you've organized it, you can reduce that to days. But without that, literally, you go through lab notebooks, and this one stays, and that one goes to the other company, and so on and so forth. Carve-out, as we used to call it, of a um of a molecule from your from your portfolio is a really, really daunting, expensive, lengthy task that again, you can uh reduce the complexity, the cost, and the timelines of that from uh from months to days. Because, and I realize I'm probably uh preaching to the converted here, the product that we all make in R&D is data. Data is our product. When you out-license a molecule, you don't out-license a protein structure, you out-license the data that you have about that structure. When you go to the FDA, you don't send them an envelope full of molecules. You send them data about the safety, toxicity, efficacy of your uh of your product. Even when you go to commercial, you tell your commercial colleagues about the efficacy, where, who's responding, who isn't. So, the product of our enterprise is data. And that data is going to live all the way along the product lifecycle, long after the molecule leaves our labs, right? Uh, after the molecule progresses from your labs, hopefully colleagues in clinical development will be writing investigative brochures, will be submitting INDs, and a few years later, you will be submitting to the FDA. And again, there are beautiful AI modules today that actually help your colleagues write those submissions if you make the data available, if those new AI products can again rummage through your lab data to help uh draft uh that IND application or that regulatory submission. Right?
Last example that I'll share is that the benefits even extend well beyond after you launch uh your molecule uh to the market, your your new drug. It will happen every so many years or so that your blockbuster will fail, and that's not anyone's fault. I mean, it's just by law of numbers, right? Every so many years, you will think it will be a blockbuster. Your leadership will think it will be a blockbuster, your board members, your shareholders will think it will be a blockbuster, but a signal emerges. And again, there's a lot of numbers because you only test your new drug in so many samples and so many patients before you go to market, and then hopefully it becomes successful, and within a few months, tens of thousands, if not hundreds of thousands of patients um are using your drug in the wild, and a signal emerges, a safety signal or or an efficacy signal. And the only saving grace you have then is your data. The only uh uh the only option you have to uh to rescue your drug at the time is to, you know, it rummage through your data. I can give you a real example of that. When I was working at a a diagnostics laboratory, the way you uh do R&D in diagnostics for three to five years, you collect samples of positives and negatives. You bang them, and then after three to five years, your uh assay is ready. You run the assay against all the positives and negatives, and then you have your sensitivity and specificity, and you launch. Right? Our specificity was not good. We had too many false positives. But because our data was rummageable, what we were able to do is link um uh the uh the samples that we collected at the time to the current health records. And what we were able to to see is that patients who were supposedly negative actually within six to twelve months after they contributed the sample turned positive. So, we did not have a specificity problem. Our sensitivity was better than the current screening levels. Now, that's a much better message obviously to launch your diagnostic with, right? So, that was really a saving grace because we had made our data FAIR, or rummageable, however you want to call it.
Last point I will add to that is, so I spent the last uh three years verifying our data at a diagnostics company. We produced about 10,000 lab samples per day. So, it's quite a big volume, and we made a big clinical genomic database of of 10 million, 10 million rummageable records. Technology was the easy part. The technology is done. The technology is way ahead of our organizational capability. The the challenges we ran into was, you know, that proverbial, that guy somewhere out in a lab from a small company that we bought, saying, you can't have my data because it's mine, and I need approval from the SVP, from the SVP, before I can share my data with you. Like, dude, we bought your company. And you get over yourself. And we screen-scraped his data over the weekend without him knowing it. But that's a different story. The point is, if you want to make your investment in the technology yield, you got to properly set it up. It's not just about the technology, about the process and the people as well. You will be running into a data privacy officer who's going to ask, is this PHI? And maybe a protein structure is not a PHI, but according to nine states in the US, a genomic structure is. So, these are all the questions that you'll have to uh work through. And that's where the real heavy lift is going to be. And well, I think you can benefit from some expertise from, that's a segue to Eurovvine. Where are you? Right there. Back to yours.
Okay. So, those are a lot of like real-world, real real-lived experiences. And it kind of ties back to this where, okay, if you have a data foundation where you have the data collected and it's FAIR, you can do those use cases that Yurus was talking about. Knowledge management, so knowing what you have. Um, how does that data get in there? It's connected instruments. It's having lab ops supporting your scientists so they don't have to wait around for reagents for for other consumables. It's wrapping a digital culture around this, and then having the right governance, operating model, automation, and orchestration to make this as efficient as possible. So, the message here is that it's not just a technology problem. All of this has to come together to move the needle in productivity, uh, particularly in your labs. But being here at this conference, AWS and Deloitte, we're going to show you, we're going to talk about the tech enablement journey here. How do we accelerate that with you for you?
So, what we're building is going from left to right is a systematic, scaled, orchestrated, governed way to move data from your instruments or, you know, on-prem servers, your Isilon servers, moving that in an organized way using AWS native services, the the primitives, which I'll go into in a minute, and getting that moved to the cloud, having a control tower around that movement. Okay? So, you can start thinking about, well, I've got 10 instruments in this room in one lab. You start multiplying that out. You know, I think I've seen stats where, you know, a large, large top 10 pharma companies have on the order of five, six, 7,000 instruments, right? All over the world geographically. You've got labs that are external that are sending data. So, having a control tower that helps you see what's happening, you know, what's failed, what's moved, okay, is super important. There's also a ton of detail. You think, hey, I'm just moving a file from point A to point B. But if you're living that experience, there's a ton of detail in there that we've gotten into. So, for instance, you know, I've got a panel of experiments. I want to wait until all those experiments are done before I move the data because it's no use moving like a piece by piece. Let's move it all at once. Right? So, all these business rules can be incorporated into this into this uh data movement capability. And then once you get it into the cloud, okay, you got to, you know, we talked about making the the data FAIR, you got to have that complete metadata record about the samples, about the experimental design. So, having integrations to where that metadata lives, your metadata repositories, your ELNs, in some cases your LIMS systems, your lab execution systems, all of those need to be integrated, right? So, as the data flows, the metadata comes with it, and it's stored in a data foundation, and, you know, those data foundations, they form data products. They're foundational data products that uh we can explain to you. They are data products that are built for supporting AI and ML. And then what we do is we apply standards where it makes sense. So, the Allotrope Foundation has a lot of good standards. They're not always adopted, but it's a place to start, right? Then you've got the consumption on the back end. So, how do we do this? So, as I mentioned, we use AWS DataSync, the Greengrass IoT framework for messaging. That's to get the the data movement. Um, S3, obviously, DataZone for governance, and then HealthMix, Bedrock, and SageMaker Unified Studio for some of the bioinformatics and AI. But this is all modular. So, this isn't obviously the entire universe. It plugs into many different types of applications. Right?
So, what do we actually do? We take these, you can think of these as um, you know, the small Lego blocks in a Lego block set. This will be the one with the one nub or the two nubs, right? And, you know, if you love that kind of stuff, you'll spend days or weeks building up your, you know, Millennium Falcon Star Wars thing. So, what we do is we pre-package that, right? So, we make them into these big Lego blocks that make it easier to deploy. A lot of the tinkering has been taken out of the equation, and it's been parameterized, so it's much simpler. So, the key is it's still working off of that tried and true, scalable, cloud-native componentry, right? That's what we're talking about here. And so, what we're doing is we're building this and have built it with client feedback in mind. So, the core tenets are open, modular, cloud-native, scalable, and standards-based where it makes sense. So, we're not interested in holding your data hostage where you might have to go to somebody else's cloud and download it into yours. Um, you know, we're interested in making it as modular as possible. So, if you have a fantastic knowledge graph, if you have all your data verified, maybe your problem is possibly you need to get the data into that environment as fast and as efficiently as possible. So, it's modular. You pick, you pick what makes sense for you, right? So, what we've been doing is working with AWS. We developed the core framework around this, around the data movement. Um, we are at this point right now where we are working with clients and looking for more clients to try this out within their environment, to deploy it and configure it and customize it within their environment at, you know, one, two lab scale, or at an entire site, or, you know, labs across across the world to bring that information together. As we're doing that, we're still, you know, we're working in many things in parallel. So, another aspect of it is working to define data products and how they should come together, working with partners to build integrations to ELN and LIMS systems. And then, you know, what we also have in the pipeline is later on in the year is the ability to perform multimodal analysis. So, we'll have a flavor of it. You know, there's there's other vendors, there's other partners that will have a flavor of it, but the idea is you've got all this data coming in. You've got the complete metadata record. You've got external data and publications and at uh at collaborators. How do you pull that together across the different modalities, across the different experiment types where they weren't really intended to be integrated or work together? Right? That's where the remit of using specific LLMs for that domain to extract the entities, extract the knowledge, and populate one methodology is to populate a knowledge graph. So, you can look for those correlations. So, that's coming out later.
So, really, you know, you can think of this as the left to right is your flow of your data. You got to get it, move it efficiently to the cloud. You got to organize it, integrate it, and then you got to use it for benefit, right? So, what we're looking for are early access clients. We've already got some. We'd like some more. Really, it's to be involved. Give us your feedback. Try this out. um help us with your priorities. What are things that you are really you really want to crack, right? What problems do you want to solve now rather than later? So, we can tweak our roadmap. Then, of course, there's incentives and investments that go along with this. So, it's a collaborative journey, and really, it's to it's to essentially get this get this pipeline moving uh much more efficiently. Um, impact, we've we've talked about this, right? It's really you want to create better, highly differentiated therapy therapeutics or drug candidates faster, right? With less hassle, with less friction, more automation. Um, the idea is you want to give time back to your scientists to spend on science. Not because you want to reduce your workforce, that would be foolish, but because you want those scientists to be able to work and progress a larger pipeline, right? So, you want them to be productive. You want them to be happy. And that's what we're about. So, thank you.