Transcription
Yes. Okay. Great. No. So, welcome all of you. I think we have a few people uh coming in dribbling in, but we're going to get started.
So, this is a joint um initiative from Smart Lab Network under Tim Cer Consortium and Oxford Global uh who are organizing this conference. So, this is the first of of three events. Today, we've got you got the foot side. Yeah. So, we're going to focus on digitalization. So, we've got three great speakers for you. uh Burkart Schaefer is going to kick off and then we're going to hand over to uh Lea from Sheekch from Bayer. They're all going to give you different perspectives on digitalization and the challenge and then we're going to go to uh a question and answer session with a panel. So if you've got any questions as you hear the speakers uh please hold on to them until we've heard from all of them and then we'll go we'll open up the microphone to everybody and uh and go from there.
So I'm going to hand over first of all to to Burkart Schaefer who's a a long-term member of the sealer community, also leads the animal project um and also has his own company Splash Lake. So Burkhart, I think I'm going to hand over to you now.
Great. Thanks Patrick. Um and uh now going to try to get the presentation up on the screen here. We do have a little issue with uh with Chrome here. I think I'll have to quickly rejoin. Okay. It worked fine when we did the rehearsal. Honest. Try one more time. You you will we all know this. But uh did Yeah, we've tried different platforms, but we finally ended up in on Chrome, which used to work quite reliably. I think the issue is that Tim, you're still sharing. Would you mind unsharing? Sorry, I thought you could just kick me out. No, maybe because you're the boss. Maybe that's why. Okay. Oh, no. It let me. Thank you so much. Sorry. My bad. That did the trick. So, I think we should be doing fine now, right? Can you Can you see this? Yes, good. Okay. Wonderful. Great.
So uh so thanks again for the uh for the invitation to contribute today and I just wanted to frame this this topic of from data discovery and look at digitization opportunities um from a standardization perspective on the one hand from the perspective of yeah and animal as organizations that help contribute with standards but also from a perspective of a vendor that builds solutions why does it make sense to bet on standards and uh and how does that help us all um create a digital ecosystem um in uh in our labs.
And so I think a good way to to to start this conversation is to think about yeah the life cycle of uh of of of data in a lab environment. And the traditional play was always oh yeah let's let's do some experimentation right. So we do the data generation um and then we capture all of that data. We represent it appropriately. We store it somewhere and then eventually we can we can talk about how we apply that data, how we use it, how we exploit it. And uh I mean as we all know and everyone loves talking about this part because this is the interesting bit right. So this is where you can drop your AI and MLB buzz words and you can where you can really see the things. Yeah it's visual. You can take action on the results. Um and that's beautiful but in order to do that we know we've got some work to do. So we really need to start here. We need to think about how do we generate the data appropriately? How do we manage it? and um uh and that uh that will then allow us it'll it'll form the digital baseline for being able to do that. And in today's day and age, we think even further because whatever we get out in terms of data interpretation, we might feed into what comes next, right? So at that point we are closing the loop and we then use different technologies such as machine learning or foundational models including large language models and others to define future process parameters and then to inform what we want to do next. Yeah. So uh at that point again the data generation and the data management piece are a very important part and obviously this is also where standards can play a role where CLA 2 as a communication standard can greatly facilitate the experiment execution the instrument control. um so that we don't have to build bespoke interfaces here and then um once we've had the big bang of the of the of the generation of the data the execution of the experiment we deal with the fallout and that's the data yeah so at that point being able to create data packages that span multiple types of instrumentation I think that's really really interesting and if we are looking at this type of a closed loop approach or even even if we only have these three steps without the closed loop uh we notice that we have different types of orchestration that needs to happen um for this to work because here in the data generation we are talking about the traditional automation orchestration bit and when it comes to the data management and assembling a reusable data package that is data orchestration and those are different types of skills and different um and different types of tools that I need in order to to do that.
So let's drill down on these um and and really think about um how can we complement this well-known automation orchestration bit with a data orchestration approach that complements that nicely because we've we kind of have an existing landscape here where we have um kind of a multi-layered system. So at the lower level we have the local automation orchestrator or a scheduleuler right which which drives the equipment um and which is either using standards or proprietary drivers control the equipment. um and during the run it produces files and there are a couple of examples on the market for systems that do that. um and then we've got some things on top. So we might either have a workflow engine on top or even at the highest level things like ELN, limbs, design of experiment tools, method design tools, uh things like that and that's the existing landscape. What they do is they make the model train run, right? So at that point we can get the instruments to move. We can get the experiments conducted. So what does data orchestration then bring to the table here? Data orchestration basically puts a couple of little vacuum cleaners next to each instrument that are participating in such an automated flow and it collects the instrument data during the run and it assembles a run data package from the individual data sets that that come from the instrument ideally already in a standard format maybe has to be converted on the fly so that we can normalize that content and really do something with it. It will also integrate with other consumers for that data package whether that's limbs ELN or some type of a data store or analytics package and um and it will then also allow for visualization of that integrated run data. um and this gives us a well- definfined and reproducible path to things like data analytics and closed loop. And some aspects of it will be generic in nature and we can apply patterns that make that scalable and sometimes they will be specific to the measurement technique, the process technique or even the concrete workflow. So at that point um we're talking about data orchestration that nicely complements the automation that we are all already familiar with. So we see that as a big opportunity and what that allows us to do is it gives us rich instrument data at our fingertips because now we no longer just get a couple of Excel files and text files coming out of a run but depending on the type of instrumentation that is participating um we getting um a rich data package um that is then also in open formats and that can be visualized and that can be repurposed in a number of ways and um that means different types of data whether it's chromatography whether whether it's time series data and then we exploit that content um in a number of ways.
What we can then do is we can think about how how can we evolve this to a common blueprint for sustainable lab data and you have seen I assume a circle like this or a very similar one where you start with the overall life cycle. Yeah. So you've got your experiment planning that you do in leading systems. You have the execution that is the work that's being performed in the lab. You have the data capture from the instrument. You normalize it into a common format. In our case, that would be animal. Obviously, um we enrich it with metadata so that we get context. And then we have data visualization and analytics which can then lead to planning of additional experimentation. And if we're talking about automation um and potentially having multiple pieces of equipment involved even if it's on a automated work cell we find that this this might actually especially that execution bit might actually decompose into a number of um of separate life cycles of its own that that operate locally um on on specific instrumentation. And so at that point we see these individual runs take place on different pieces of equipment and we need to assemble that to a more complete data package. And so that assembly of that data package is actually quite interesting because it allows us to um to do things that that actually track the full um the full workflow in terms of a material flow and a data flow through the automation platform. And what we've done here is we've we've taken an example of uh of a formulation recipe that includes a couple of different steps in order to make a particular batch. And here we can see the different operations in in yellow and the different samples or ingredients that get used. We see those in white. All of these get added to a particular batch that have formulated and then eventually that goes into measurements and then those measurements could be part of a chromatography sequence. So those samples are then injected. We have detection on it. um We have results coming on it and so on. So we no longer just look at individual results but we see them connected. So we see the big picture and this is obviously the baseline if we want to train machine learning models um and if we want to do any kind of uh predictive uh work and the nice thing is about these data packages and what we're looking at here is just animal files right they can then be um propagated on to other areas. So in this case we have a limbs and we can actually bring the data that comes from the individual experiment into the limbs UI and um and here we need to look at um how how how we then bring the data to the point of use rather than the user coming to the place where the data lives. So that's a that's a different way of thinking and part of that is also how do we deal with these different types of automation. Yeah. But also how do we make that work in the analytical work world. So we just looked at it from from the automation and high throughput perspective. Let's look at the traditional analytical lab where we might be doing chromatography spectroscopy and so on where we have chromatography data systems where we want to generate sequences and so on. So CEDA is also doing work there together with animal where you might have um a leading system such as a limbs and ELN or SAP um and a chromatography data system on the other side and you want a proper way of um of them to talk to each other. So what we're doing here is we're exposing the capabilities of the chromatography data system as a sealer service. So you can discover your instruments, you can see what methods are available, you you can submit samples as sequences, you can amend them, you can get the status of instruments and so on, but also generate a traceable result data package. And the way CILA works is CIA has this notion of feature definitions. So you can experience you can discover the capabilities that an instrument um offers over the network. So here I can see aha this is instrument discovery. I can see what instruments are connected which methods are there. um I can read the sequence I can submit new sequence elements and so on and in the end I can get results back. um and now we have that implemented for a couple of CDSS. I want to show you some examples here. um where for example in chromium um you you started at the limbs you said these are the things that I want to measure it comes up with the sequence and then it it does the run and in the end you've got your chromatoggram and that is then exported again as a data package and now we have the equivalent of that data package in animal format which can now be looked at in any web browser which is really cool right so you can really close the loop and this is the same data format that we just used for the high throughput um automation and um And that obviously makes it nicely scalable or if we're not in chromion. Let's say here's another example from openlap cds2 um from agyant. And so this is how we can uh get the sequence built and uh and here's a chromatogram on the agulant side. And then this is what that also looks like on the animal side also with peaks and with baselines and peak demarcations and so on. And that can then obviously also again taken further into uh downstream consumption. let's say again by a limb system. So in this case we've taken the same data we've put it into uh into Levantage we've put it into labware with the embedded visualization. So here's kind of the the the idea that if we put that all together we can really assemble a nice story around a sustainable digital lab because we have a common way to talk to our instruments with CLA 2 a common way to represent our instrument data with animal. We have seen how we can track material and process flow and we also get a nice view on our resource consumption and utilization because that's all passing by in front of us and that it also gives us new ways of leveraging our data again and again.
So the the big question is how do we get there because it would be nice for all of those standards just to be universally applicable and available and and that's not the case yet. We need to work uh with the vendor community to to generate more and more adoption. And for that to work, we said, well, how can I get my instrument data into an animal format, how can I get a sealer driver for my instrument? And here what we've created, and that's obviously now with a vendor hat on, um we've created a pool um of of converters and and adapters that um is maintained by a dedicated development organization. And so interested users can join the pool. they get access to the entire pool um at um um at at no extra cost and the development is then co-unded. Yeah. So if anyone who needs a particular driver, they can get together with others. We then spread the the load of doing that. um The resulting adapter goes into the pool that everyone else has access to. Yeah. So that that is then maintained as a product and is fully supported also as a software as a service with a service level agreement and also yes somebody is there to fill in your cyber security questionnaire and the nice thing is that such a shared cost model makes standardsbased integrations less expensive than the proprietary custom integrations that we need to do today. Uh so the idea is to get that scalability not only from the capability perspective of what the standards allow us to do but also from a commercial perspective because that the total cost of ownership of such a solution um is actually greatly reduced and we can bring those standards to fruition uh in our lab lab environment and uh and I think that is that's sort of a quick perspective of um how CA and animal can can complement our digitization efforts um in in drug discovery, drug development. And uh that's from that's it from my side. And so I'll hand it back to Patrick and we can do Q&A um toward the end of the webinar. Thank you.
Okay, thank you, Burkart. That's uh that's very nice. We've only given you a few minutes, but uh we'll it sets the scene. So I'm going to hand over now to Abishek. Abishek, are you with you? Are we you with us? Yep. Can see you right. Do you uh I've got some your slides. Do you want to show them or do you want me to show them? Uh, you can share them. Okay. Okay. So, Abishek is one of the speakers at the at the event in June in Basel, the Oxford Global Discovery and Development. Uh, we've only given giving Abishek a few minutes here to give us all a taster of what he's talking. So, I don't want to steal all his thunder. You've got to go to the event in order to hear his whole story. But, in the few slides he's got, if I can just find them, uh, he'll tell you a little bit about what he's doing. And you'll see that that links back a little bit to what uh what Burkhard was telling us earlier. So hang on now. I've got to navigate this thing. Maybe one quick side note. So if there are questions, please keep the questions for the end. So we'll have a question answer session at the end to make sure we stay in time. Okay, we got this. Can you see the slide? No, not it. What am I doing? H. Okay. And now Yep. Sure. Okay. I'm going to go full screen and then you're you're on. Whoops. Let me go to the beginning. Okay. Abishek, over to you.
Thank you. Hi all. I'm Abishek Chri. I'm working as a principal engineer in Bayer Pharmaceuticals in RWD global commercial department. Before that I used to be part of evidence generation. Technically I'm still doing evidence generation but now more in the commercial side. Uh my objective here is to not replacing anything but automate as much as we can with the modern technology. So uh if you go to the next slide, what we are doing is we are revamping the complete uh real world evidence of inside and we are bringing it to more uh with say software ccentric approach where all the data is accessible all our tools are integrated and that can be used and for example if in any internal tool is being built we want it to be used but a user need not to know about it. So they will just ask a question like get me a patient population with this this this criteria and behind the scene most of the work we want to automate that's the core objective uh to do that we are you of course leveraging AI with generative AI lot of engineering and pipeline and automations and what we are doing is we are adding quite a lot of governance on top so based on your login uh we go to a governance engine and which decide this particular user user even if it asks a data which is not being accessible by the user authorized by user it should not show that. So that's the main control. So uh building on top of that with a very high regulation and governance is the most essential part and that is the first layer on our platform we are building to make things easier. So we work with researchers statistians epidemologist and many different and even from when we collaborate with the different team companies universities we have a different kind of usage genre. So for that we are uh creating a next layer which is called avatars. If you go to the next slides it's a glimpse of what it is. Uh you can see so what we call it is a RWD workbench where you get many things. So for example this is one of the agent chats. Uh these are aars. So a data analyst aar a writer epidemologist critical thinker research assistant. So what it does is all this aas currently being triggered with a dynamic prompting. So it understand behaves. So we I talk to researchers I try to understand how they work and you can imagine I mimic an author out of it by creating prompts. But what happens when the researchers are testing or if somebody's using and they are giving a negative feedback that this answer doesn't seems all right or it is nothing to do with what I'm asking. It goes back to the researchers for validation. We have a layer called validations which validate and try to understand what went wrong, what they desired and what didn't work out and then we change we optimize that layer. So this entire protocol you can select a different kind of uh agents here like researchers and you start asking questions. Right now you can imagine it's an alpha stage where it gives you uh you can go 20 30% of the steps but over the time based on feedback and learning and learning we anticipate that we can go close to 60% of the task which can be already answered by it. This engine has connection to everything. Our data warehouse, our internal library, our protocol, our studies and based on the person who logged in it understand what and what cannot be done. What kind of answer they this particular person can ask. Even this agents are some of the agent are selected based on the particular user. um high profile user will get all the agents. Some won't even get it. uh researchers won't get the agent of engineering because they don't want to directly interact engineering and find the code which is it is not logical. So that's one of the agent we have. Another agent uh which is not getting in the screenshot is data insights where we are not technically solving your entire question. We are giving you insights of the data uh again in the glimpse. So we know text to SQL. We know LLMs are good at generating SQL but in our era it is just not the SQL we work with. We have our internal tools which need to go through it. Then we have more additional layer of optimization on top of those. What we are doing is inside this engine everything like as Bard explained we have an orchestration framework inside which based on the question understand which tool to trigger which library to add into this engine which virtual environment it goes and then it gives you a viable answer what you are actually looking for and again it's a feedback model we will still go wrong because the hallucinating model so it does the less hallucination what we are doing is we are creating a valid validation engine layer on top of it. So just to be uh clarify we are not targeting here to give you response in a seconds we are targeting here to give you response maybe in a 5 minute window but those things would have taken maybe weeks for you now so it's 5 minutes 10 minutes is still awesomely good but what we do is all this process validation is very visible to the user they can see what is happening why it takes time where the status is and then they get all the steps and all the responses that this This is what we came and this is our journey. That's how we found this answer. This give a trace back model. So you can say a flow workflow model where they can say this happened this particular data set happened this library they used and that's how they so same like odyssey and any other data set they want CT and what kind of data they are using those optimization is done here and the and that way we are integrating more and more tools to it. So you can um uh the target is to find out how much we can automate not to replace anything. So we don't believe that uh we can replace uh we cannot do automate the entire RW cycle but we can already do many things which are repetitive. So we find repetitive patterns and we are trying to engineer that we are trying to integrate that and we we have this single window login where you find every information. So if any new tools come I need not to tell you you can just go to the tools and you can find what is available and then you can add as an integration and start chatting. So it is not a collection of one tool. So it is a collection of many different tools. It's just that you need not to do the work of integration. The same thing with R studio. We have a hosted environment. So it's in the browser. It runs on the browser. We are using uh web assembly was some kind of layer which actually open in your browser the UI. So you want an analytics. We use some open source dashboard. But you need not to know it. You just say get me the uh count of number of user doing blah blah blah. It just gives you the chart to you and then you can do interaction and you can say no I want the hosted one. So that gives you a link. So it those are the works we are trying to automate with the entire cycle of RWDA and this is what I'm going to talk about in the thank you very much.
Okay, thank you Abashek. I'm going to take off there. Right, so as I said, hold on to questions. Uh, we'll hear first of all from our our last P uh speaker, Rea from Astroenica. So Rea, I think you're on the first day. You're giving one of the opening keynotes and sharing a the session. So uh I know a little bit about what you're going to talk about just in a few minutes. So I hand over to you. Do you want me to bring up your slide? Yeah, that's that's fine. Okay, let me go find it. So, um, so while you're, uh, doing it, uh, hello everyone and, uh, thank you so much for the opportunity and I'm quite excited uh, to be part of this and also participating in the conference. Um, looking forward to it. Uh, so just a little bit about me. Um I head uh the globe I'm the global head of data office uh with Astroenica and uh prior to that I worked in um different sectors um performing uh similar roles uh in with respect to IT and uh uh uh data uh spanning over about uh um 19 years of experience. Uh so what I'm going to talk about in terms of from the uh the conference is in the keynote speaking is how we could leverage um AI and data uh to support the innovation and uh the benchtop patient right not just the only from the drug discovery but uh expanding it beyond and that's what um I'm currently uh working in Astroenica because we do have we have various aspects right I think it's not Astroenica are the only ones in that like in any global organization they have to deal with uh multiple uh challenges whether respect to commercialization uh respect to the discovery and uh tech transfers and also leveraging like you know reaching to the market in a in a quicker manner and the traditional process is no longer viable I will not say it's not valid but the traditional process is not going to be viable and I know in last three weeks many of you experienced in terms of this uh the tariffs episodes and almost all the companies uh have to scramble through in terms of how we could build these resilience in terms of the supply chain in spite of all these challenges. So we definitely we need to look for uh more digitalization opportunity but at the same time we also have to streamline the process. So we are looking at uh both aspects right you know where we can uh seize the opportunity to digitalize you know we are uh seizing that I mean we are looking at that uh uh um you know working through leveraging the AI uh but we are also looking at in terms of uh streamlining the process because if you are not going to do it it is also going to be a huge uh hassle especially from the global organization and I als also would like to cover the other aspect um as well is uh not just with the geopolitical risks which we are dealing with but many countries right now there is a regulatory criteria has been enforced which we need to be factored in with respect to the data like how the data can be shared across regions because that needs to be seamless when you have to work on in terms of the end to end but we also need to factor in in terms of how to accommodate uh the regulatory criteria has been enforced by different countries uh which also enforces different set of challenges. So I I I think it is we are in the interesting era. I have to say that and uh not only the data needs to be trusted but the data also need to be AI ready so that uh you know like Burkhard and Abishek talked about in terms of how we can build this agentic AIS right with the uh different uh personification and how can we get the access to the data sets much more quickly and we could trust that information so That's the effort which we are currently working in uh in Astroenica. I wanted to just give a kind of a flavor in terms of uh uh in my keynote in terms of how we could come up with uh uh amicable proposal right or amicable solution in terms of addressing these challenges but also exploiting the opportunities or the challenges right you know converting these challenges into an opportunity where we could excel in terms of this product launches the diverments or switches you know covering the entire breadth of the uh product supply chain uh or the product life cycle just from the discovery uh to the patient value chain. So that's the the glimpse which I'll be covering uh in the conference. With that uh back to you Patrick. Thank you so much. Thank you. Thank you Ricka.
Okay so you've heard there are three different perspectives, three different talks at different levels of granularity. I'm going to open the floor up to questions. Uh, who have we got? We got lots of thumbs up, which is delightful to see. Do we have any actual questions? Yeah, I mean I would have one of course especially it's actually to all of them. So first to book it. So if I wanted to get started with with animal, I mean this is really like of course what we all dream of that you have animal. You put all your data in animal container and then you can feed it directly into the AI and you've got all the metadata, all the raw data all together and the AI can start to train or predict stuff. So is there a pipeline? Is there a anything prepared one could use to not start from scratch every time? Yeah, I think there are a couple of things um that you need to look at. Obviously what you're describing is an entire life cycle, right? You so you start somewhere at the data generation, you run your experiment. At that point you need a way of ingesting the data and and doing the normalization the animal conversion right and then uh you need a place for that data to live and then you need the tooling to to then um feed that into let's say your your machine learning model or expose it to a large language model describe it annotate it appropriately so there are different bits and pieces and uh and I think we're in pretty good shape when it comes to the the first part of that uh of that value chain. um where which is pretty well understood um and I think now the interesting bit is now that we are at a level where that data is easily accessible that's where that's that's where the innovation is now is now moving right so so how do you then put the ML on top of it how do you put AI on top of it and that's where we are actively looking for for um for folks that want to work with us to drive that forward yeah so so I think the the first part is is is fairly straightforward um and now we can innovate on top of that but to get started hands on like like we in academia we are using for example you mentioned the limb systems which support um animal we use an open we use rsp space and and freeto use um eln are there any freeto use which already exist to export um animal files or which can then be reused or what is a good starting point if I wanted to use animal yeah I mean as as I said we really have to look at that entire chain, right? um because the ELN does that have all the information you need in the end. It's usually not that hard if you have if you're starting from a structured data source. Yeah. um and that's the question, is your ELN already structured? um or is it more of a Yeah. And that leads to the question of data quality, right? Does your ELN is that just a a fancy spreadsheet or is it a is it is it more document oriented or do you really have a um a strong data model that uh that that is there. So it's always a question of the data quality. um tooling I think is not the problem and there's there's quite a bit of stuff and obviously in in this this uh public forum I don't want to mention particular tools and things but um but it's more um it's more of a question of what's your source of data? what's the quality of the data and then to kind of get into that pipeline.