Transcription
Welcome to Data and Biotech, a podcast from Cordai, where we explore how companies leverage data to drive innovation in life sciences. Every two weeks, we sit down with an expert from the world of biotechnology to understand how they're using data science to solve technical challenges, streamline operations, and further innovation in their business.
In this episode, we sit down with Yesper Ria, Director of Computational Biology at Merck, to explore the evolving intersection of neuroscience, immunology, and AI-powered discovery in the pharmaceutical industry. He discusses how his team uses tools like knowledge graphs, single-cell data, and dashboards to support target discovery and pipeline stage projects. Yesper emphasizes Merck's lean, partner-first approach, validating external platforms via pilots and applying generative AI for tasks like literature mining and patent analysis. Finally, he highlights the importance of interdisciplinary collaboration and being mindful of data's limitations. Let's get into it.
>> Yes, Berea, welcome to the Data and Biotech podcast.
>> Thank you. Thanks for having me.
>> Awesome. Well, just to kick us off, would you mind giving us an introduction to your background and what brought you into the field and what brought you here today?
>> Yeah, sure. I originally started as a biophysicist back in the day, trained in Copenhagen at the Niels Bohr Institute. I'm originally from Denmark, before there was something called bioinformatics. So, um, I was very interested in both topics, and at the time, I started also developing an interest in genetics and neuroscience. Um, so I was actually, the first exposure to two experiments was actually electrophysiology, you know, measuring electrical signals for neurons and trying to model that in a simple way. So that was the kind of the physics at the time, and I think it's still valid. Try to build simple models that enable understanding of whatever you're exploring, in this case, biology. And I, I've really always appreciated that, even though models are becoming more and more complex, I think we're all struggling with how to interpret them and understand, you know, what is actually happening and how does it help us, for instance, understand biology and diseases. So I, I'm trying still to find a balance, but that was my starting point. I, I continued with the neuroscience. I thought it was extremely exciting and interesting. Uh, so I did a neuroscience PhD in Karolinska Institute. Um, I was there for several years working on the spinal cord, since it's actually a quite good model system for understanding how neural networks work. You can record electrical activity in vitro while it's oscillating in a similar way as if it was sending signal out to the to the muscles when you're moving or walking. And so, you know, you could actually access and and and record from the cells, uh, in vitro in this system. And it was a good learning experience as well. And we were trying to connect that at the time with the microwave technology that was emerging, you know, molecular aspects like looking at different cell types that we could label, for instance, with person markers, uh, and then collect several of them, not single cells, but, you know, populations of cells, um, and then look at the molecular profile and try to relate that to electrical signals we were seeing. It was very challenging, um, and not as simple as we had thought. Uh, we also looked, uh, at a disease model, specificity, which was in the tail model of the mouse. So it was actually more gentle, that, you know, the classic models that were used at the time, because it's just the tail that is affected and develops this kind of phenotype. And then we were looking at motor specificity that projects out to the mod muscles. And we were hoping by doing microwaves that we would understand the molecular mechanisms and we could, you know, had a good hypothesis that there were some calcium channels that was driving this, um, and again, it turned out that biology is much more complicated, and we did not find a simple explanation, and it was probably the whole network that is changing. Uh, the balance between inhibition and excitation was probably, you know, shifting as well, but it kind of triggered a lot of, um, appetite for trying to bridge different types of data to understand diseases. So I moved on, uh, to do a postdoc in EPFL in Switzerland as part of the Blue Brain Project that was actually building in silico models for the cortical column of rats, but based on data from the lab. And I, I was fortunate enough to have a collaboration with one of the pioneers in single-cell transcriptomics, Karolinska Institute, because I, I kind of bumped into there before I left. And we were looking, uh, then at at mouse in the cortex, but also in in the dopaminergic system, and started working a little bit on neuromodulation. And that's kind of led into the exploration of Parkinson's disease and mechanisms there. But just characterizing all these different cell types were emerging and is still, you know, moving forward. But we, I think all this single-cell data has really given us a much better understanding of the molecular profiles of different cell types. And and the big challenge in neuroscience, like it's just how do we make sense of the billions of cells? Like, are there subtypes that are very characteristic and we can kind of think of them as interacting in as groups of cells, or are they very individual? And, you know, you can base it on morphology, like where do they project? And then you have the molecular profiles as well, that is really helping in in driving this understanding. And I think when it comes to disease, maybe also look at the dynamics. ICS. So that was the next step, right? To see what can we have reliable disease models where we can look across time, because the big challenge in humans is that for brains, we don't have access to, you know, tissue samples. You, you just get them postmortem, and that, I think, makes the brain quite unique in terms of other tissues and other diseases. So we, we still rely heavily on good models in the mouse and translating them to the human. So I, I, I was working, uh, at the time on that, but I, I also kind of felt like I wanted to be closer to the translational aspect. So I transitioned into the industry, first in in a company that was working on clinical diagnostics on sequencing. It was, it was quite distant, but it was a good learning experience to to come from academia where you have a very pragmatic approach to scripting, you know, and and programming, and it's, it can be a bit messy, but whatever works, it's fine to a production system where it really has to be neat and, you know, you have to have version control and and everything has to be structured and organized in a completely different way. Uh, and and and that I have kind of carried with me since then. And and, you know, that was a good learning experience, but I missed a little bit the scientific parts. There was a very, kind of engineering approach. So I, I, I continued into the pharma industry where I'm still, uh, located. And, you know, that was really like in early discovery and translation of biomarkers at the time. I was fortunate to join a medium-sized company, and that can also make a difference, right? Big companies tend to maybe be a little bit siloed. You get expertise in one topic, but you, your your colleagues might stay far away, and you might not get insights to all the activities in the company. In a medium-sized company, you kind of see all the different departments, even commercial, legal, and and you get a kind of better understanding of how all these things play together and how things are prioritized in in a company. So you, you might be interested and passionate about a certain disease, and then you bring that to the management, and they just say, for commercial reasons, this is just not interesting for us. Right. So, so you start learning like, okay, there there are different aspects that are important in in prioritizing what direction you go and in terms of, you know, the the diseases and targets that you will continue move forward with and explore further. Um, and now I find myself, uh, in Germany, in in Merck Germany. So maybe I should also clarify the distinction here that that, you know, and and and a lot of confusion for myself as well when I joined. The original Merck is here in Germany, was established more than 300 years ago, and one of the family members created a branch in the US more than 100 years ago, and for historical reasons, that branched off into a separate entity, but with the same name. They, they insisted on keeping the same. And now the complication is, we're entitled to Chorus Merck everywhere apart from US and Canada. And in US and Canada, if somebody refers to Merck and Company, it's is the US part, and there we're known as EMD Serono. So, so that will be the reference. Uh, and I just, I always need to clarify this aspect. So when I say Merck Germany, or even if I say Merck, I, I'm referring to the to the German part. And I was tasked to to basically set up a computational biology team within the research units of neuroscience and immunology. And I've been doing that for one and a half year now, and it's super exciting. There's so much happening, uh, within Merck, uh, and also outside of Merck in in biotech companies, developing like data-driven AI tools for target discovery. Um, yeah, it's, it's really amazing how fast the field is moving forward.
>> For sure. Um, so, you know, can you give us a little bit of an overview of some of the, I mean, you mentioned you mentioned research science and immunology, you know, what are some of the projects that your computational team works on?
>> Yes. So we're tasked with basically supporting the pipeline projects until they reach the clinical stage with whatever questions they have. It's typically mode of action, understanding better, like the how is the drug actually exerting its effect in in vitro and in model systems, but also translation. We, you know, needs to translate the mouse models and in vitro systems to human. And and I would say that we have become incredibly good in in curing and treating animals, and mice in particular, and and not so much always humans, because the models don't translate very well. So that's also something that that we are challenged with and and are focused on. And then I think the big chunk of what we are currently doing is actually early discovery. So implementing different data-driven methods, uh, to really improve the discovery. Because the challenge is that until a few years ago, it was actually an immunological research group, right? They were focused on on autoimmune diseases, and then they realized that there's a big opportunity in neuroinflammation and neurodegenerative diseases by bringing in that knowledge, but moving into the CNS space. And that basically means that we have to kind of build, you know, both on the computational and on the lab side, and infrastructure. And that has been the challenge for the last two years. So supporting the lab part by characterizing the initial system with OMIX. So a lot of it is OMIX based, right? And the other part is maybe on on the machine learning, NLP on on the literature, and trying to bring in the relevant data sets, analyze them, and bring insights for target discovery, target validation. And then also supporting the pipeline projects as they move forward. And also making these accessible to bench scientists, right? So, so, you know, they they don't necessarily program in Python and R. So we need to have like dashboards and other types of interfaces where they can go in and and answer maybe more simple questions.
>> Yeah, I, I think so. Um, there's so much that I want to that I want to unpack there. I want to get, I want to make sure that we get to, uh, translating mouse models into in vitro systems into humans and the early dis, uh, the concept generation and the target discovery. And then I also want to make sure that we spend time on both the, uh, multiomics and multi and multimodal data, data sets, and what you're doing with NLP on the, uh, on on the literature search. So, I, I don't know, let's take, let's take concept generation and target discovery to start. Can you just explain in your world what concept generation means in practice and how you, how you as a computational team support that?
>> So for us, concept generation is really proposing new targets for treatment of a disease. And traditionally, it was done by bench scientist experts looking at the literature and say, hey, there's a phenomenal new paper that came out, there's a clear causal collision between this gene and this disease, we should pick up on this, and then they start, you know, validating those experimentally and moving forward with that. We still do that to some extent, but it's clear that scientific literature has just exploded, and nobody can keep up with all of that. So we really need like computational tools to basically leverage on the existing knowledge. Uh, so that's one challenge, right? So one of the key things of concept generation is, if they still bring this proposal forward, this target is relevant for this disease, how can we capture all the existing knowledge in an easy way that doesn't require this scientist to go and read like 500 papers? So one thing that we've been doing is looking a lot of knowledge graphs and seeing if we can capture all this information in a knowledge graph for, you know, this gene is associated with this disease, and this drug is targeting this this protein, and and bridging all that together and then pulling out that information. So we say, for this target, this is kind of the knowledge base around it, and this is the diseases that are associated with this. So, so that is one attempt to kind of making it easier and to harmonize the data retrieval to support and validate these proposals. But I think the other aspect is also to to have computational data-driven proposals, right? So we, we use the data or mix or also the knowledge graph, because there we can maybe do link prediction and say, are the patterns in the topology of the graph that suggest that, you know, maybe not not everything has been explored, right? These are incomplete, where we're still learning, we're still adding information to this knowledge graph in a way. So there might be certain genes that are important for a disease, but that link has not been established yet scientifically, but the graph might help you because they're connected in a certain way, and then they say, there's a very strong likelihood that this gene is really involved in this disease, and then we can also maybe experimentally try to validate that. So that's one, uh, method that we're exploring when it comes to concept generation, where you kind of build on the existing knowledge.
>> That makes a lot of sense. So it, can you provide a little bit of insight into how you integrate all of the new literature that's that's coming out all of the time into into the knowledge graph itself, especially given the reproducibility issues that exist out there? There's sort of like a, there's a data quality issue. You could be adding a lot of links to your knowledge graph and and, you know, and introducing a lot of noise, but also, you know, expert time is really expensive. And so I'm just interested in, you know, what sort of processes have you have you developed for constructing the knowledge graph so that like all of these sort of downstream tasks are kind of, uh, achievable?
>> But actually, we're cheating a little bit because we're very pragmatic. Uh, we're a small team, and and we need to really be quite lean, and and we can't develop everything in-house. And I think that's that's that's an insight and a learning I think the pharma industry has done in the last few years, that it doesn't make sense to develop everything in-house. There's a lot of companies out there that has a lot of knowledge and expertise that create some of these infrastructures for us. So I think that became clear first of all from the OMIX data that my team was spending too much time on bringing in the data sets, curating them, making them ready for analysis rather than just analyzing them and and generating insights. So we started looking for like external providers that could allow us to to just access that and and externalize that in a way. So now we're using Kaiogen's Omicssoft product for that. And then in that process, I also started exploring their knowledge graph. So they've been building up that resource for many years. Maybe some people are familiar with this IPA Ingenuity Pathway Engine, and and in principle, everything that's under the hood is curated information they've been building over the years. At some point, they decided, why don't we just license that, you know, enable that data scientist that don't care about our our user interface, that just want access to the knowledge graph to do with it whatever they want. And that's kind of what we've been playing around with for the time being. So, so they are in a way providing it to us and building it for us. Um, and then we, we created some pilots around that as proof of concepts, also because for us, it would be a huge effort to first build a graph without knowing that there's actually value in this type of analysis. So what we often do is to create pilots and see if we can find these resources elsewhere. And then once we've shown that there is value, we can then think about, is it good enough that we can use it as it is, or will we make an internal effort to improve on that now that we know that there's value in going forward along this path. So currently, you know, that our our knowledge graph is is, you know, not, we're not building it.
>> Yeah. No, that's great. And, uh, we had an episode, uh, on, uh, Omicssoft, uh, earlier that we can that that we can sort of link to in the episode in the episode notes on this, if people are interested in learning more about that. So you've got this knowledge graph, and you mentioned, you mentioned sort of link prediction and also selecting which links in that knowledge graph you want to sort of experimentally validate. How do you, you know, either in the context of a specific problem you're trying to tackle, disease area, uh, you know, concept that you're in the process of generating, how do you decide which of these aspects of the knowledge graph that you want to experimentally validate and sort of, you know, layer your own internal knowledge on top of the knowledge graph that you have at your core?
>> Yeah. So that's that's a great question. I mean, basically from the link prediction and and the existing knowledge in the graph, we can have different types of focus. So we, we tried one, we say, let's try and detect the most novel relationships between genes and a particular disease. And then you, with all these type of analysis, typically end up with at least 50 or 100, you know, genes that you need to rank and like you say, prioritize, which one do we follow up experimentally? Which ones do we think are valuable? And I think that's that's the other aspect that I think is very important for the pharma industry, also when we interact with external partners that is also doing something like this or academia, and they come with this phenomenal proposal, this gene is really really valuable for this disease. And yeah, but is it druggable? Is it something that's worth exploring for us? Right? So there's there's additional features that you need to consider. Right? So there we say, okay, but what is known about this gene in addition or set of genes, right? Is it druggable? Do we have the structure? Is it amenable for small molecule design? Transcription factors, for instance, have been very tricky. Right? So maybe we will down prioritize those safety issues. Is it expressed specifically in the cell types we're interested in, or is it expressed everywhere? Maybe it's not a a roadblock, but we have to think of mitigation strategies because we have to then targetly deliver a drug to those cell types if it's expressed everywhere, not to get like, you know, off-target effects. And there's a lot of types of information that we can then include to, you know, make a ranking and prioritize the gene targets. And then we can, you know, with the biologists, often dive into a handful of them, uh, and then decide on which ones looks most promising and then start validating experimentally. It's a long process, uh, and and even that part, we're also looking at at leveraging knowledge graphs, but knowledge that might be coming from databases. Open Targets is is a great resource. So again, we are not reinventing the wheel. We're trying to grab as much as we can from either the public space or seeing if we can license things that can really move us forward fast. And then for for the gaps that are left, and then then we, we, we, we fill in those as much as we can.
>> Yeah, that makes a lot of sense. And so, you know, you you mentioned with the biologists, you sort of dive into a handful of them. I'm imagining that like the sort at that point when you have a handful and they're in the hands of biologists, there's sort of like custom experimental regimes that need to be constructed around around a given concept in order to in order to collect the data that either like validates or invalidates the the, uh, the target that you're trying to get to. Am I thinking about that right? Or is there like more of a standardized process of just, you know, running running all of the targets through like, you know, uh, a lab automated process with with, you know, uh, with certain assays that you're doing over and over again, certain, you know, multiomics panels that you're doing over and over again, and then that data alone is sort of enough to get you a lot closer to the answer that you're trying to get to?
>> But we can certainly as much as that data exists, try to leverage on existing omics data, right? So if, if we're looking at a specific disease, and there are like transcriptomic data, suell or bulk, that we can look, is that gene upregulated, or is that pathway upregulated? And that would at least give additional support that that's worth pursuing. I think on the in vitro side, the the big challenge is to have the right models, right? Is it expressed in a neuron? Is it expressed in a microglial? If it's in immune cells, but their effects on on the CNS, it can be very challenging to combine all of these things into one in vitro system. So that can be a selection criteria as well. And we say, you know, we're, we're just focused on microglial partially because they seem like important players, and partly because it's feasible experimentally to validate targets that are in these. We can, we can create aspects of the disease in vitro, right? Um, and then we can see if we knock it down with, you know, CRISPR or silencing, does that have the the effect that we're looking for? And that are those are the type of things that that you need to consider as well, right? When it comes to feasibility, how easy is it to follow up on this? Is it, is it, do we have the the right in vitro models? You can also have broader experimental systems, right? The perturb is is a new technology that come out that we're looking into, right? Which actually single-cell CRISPR. So, so in one disease model, if you have like a set of 50 genes, you can actually screen all of the 50 genes at the same time. And then you can actually from the data itself figure out like which cell had which gene knocked out, for instance, from the CRISPR panel. Um, and then see if that brings the cell state back to normal. And then that also gives you a good idea. And then you can in one experiment get a a good understanding of which of these are good candidates and which are less good candidates. So, so there's a lot of new powerful techniques that are emerging as well. And I think this has even been bridged. So you can do it in in animal models. So we have spatial resolution with spatial transcriptomics. Um, and and this is very exciting, you know, that that we can accelerate the discovery process with these technologies.
>> Yeah, that makes a lot of sense. And, and I, I actually, you know, you mentioned earlier that you had the opportunity to work with one of the early pioneers of of single-cell transcriptomics, and then, you know, now you're talking about spatial transcriptomics data. You know, my, my understanding is that these, is that these approaches are sort of much more computationally heavy, but then also richer in terms of the insights that that they can give you. And so I'm curious, what have you seen in terms of your ability to, you know, generate concepts or, you know, validate, validate targets with these new, uh, approaches? How has it changed your workflow?
>> For now, it hasn't, but we're very curious about exploring these technologies. I, I think it might also be depending on the disease, but for the brain that is so complex, and we're looking at immune cells that are innovating this organ in disease, like it it will be really crucial to have good quality spatial transcriptomics, so you can see, okay, what immune cells are innovating the brain, and what is the the microenvironment around these, which cells are is it interacting with, and what is also happening in these neighboring cells? That would be extremely insightful, I think, in terms of understanding disease mechanisms in neurodegenerative disorders, for instance. I think the platforms have evolved rapidly, right? But the spotted ones where you had like overlap of multiple cells that you had to deconvolve was a bit challenging to use. I don't know how many insights it has generated for us. But now that you get image-based spatial transcriptomics platforms coming out where you have true single-cell resolution, but okay, you might be limited by panels, but the panels can be up to several thousand genes. So you, you can maybe leverage on single-cell data to make sure that you get the genes you're interested in captured in the spatial transcriptomics. And we are actually exploring like collaborations that we can generate data, right? Because there's very little out there. So fortunately, the tradition in academia has been that if you publish an OMIX study, you share the data. But since these technologies are so new, there's not a lot out there. So we, we, we, we have to generate this data ourselves. And I think the big promise was that there's a lot of biobanks with, for instance, postmortem brain samples. And if we can, you know, use this technology on these samples, I think that could generate a lot of insights in terms of disease mechanisms. And we are, we're of course very, very interested in that. But it is early days, uh, and we, we will have to see how many insights, you know, we, we can generate with this. But I'm, I'm very excited and I'm, I'm very optimistic about what what we can learn from this.
>> Yeah. You know, when you look at when you look at partnerships to either, you know, uh, acquire data directly from someone who's already who's already generated for for you, or partnerships to sort of generate data on your behalf. I'm curious, you know, how do you, you mentioned the pilots earlier, the pilots earlier that you do when you're onboarding external tools like Obsoft and seeing whether the value is there or how you whether you need to create that value, whether you can create that value more effectively in-house, or whether, you know, onboarding the external tool is there. I'm just interested in how you think through the process of onboarding these new partners or of bringing in these new data sets, um, and how those pilots tend to work.
>> Yeah. No, that's also a great question. And I, I think we had transitioned a little bit from a situation where we would focus extensively on a on a small handful of diseases, and and we would just scout everything that was out there in terms of OMIX data and just ingest it, and it was manageable. I think now we are, we are pivoting quite fast between different diseases. I think that's a typical aspect of of immunological assets, right? That you, you target a B cell or T cell or mechanism, and can be relevant for multiple diseases. And the competitive landscape changes over the years. So what was relevant today might not be relevant five years from now. And then when it goes into clinical development, there's a push to say, you might think you wanted to go into this disease, but we are not sure that's the right, you know, commercial potential. You need to go into maybe XYZ, and then you need to ingest that data to to, you know, see, okay, but is is that supporting basically the mechanism of action from this? And and that is also reflected early on, right? That we, we're exploring more and more diseases, and we need to bring that data in quite fast and quite dynamically. And so we, we, we have had some efforts in scouting and identifying, you know, the data sets. Um, but then, you know, we've looked a little bit also of course at which companies can then provide that service for us. And that is actually not as easy as it sounds, because, you know, there's a lot of companies out there, and then you Google, you try to look at the home. But if you don't know, if you don't have your network in place, or you haven't been exposed to that. Fortunately, in a big company, that we actually have the the the business development that that supports these activities. So I'm, I'm fortunate to have a colleague that actually helps me scouting, you know, so if I define, I say, we have this challenge, we, we clearly have gaps, we have looked for the data sets, they're not there in this indication, we want to generate them, can you help me find the right partner for us, you know, so they will actually come with proposals, they will look for for the companies, and then we will, we will reach out to them and start a conversation to see if they can help us with that. And the same for the curation, right? Like I mentioned earlier, who's the right company for for that? And there's also a lot of different players, and then you have to be very clear in what your expectations are. What is it that you need from them? And then you can, you can try and engage in these partnerships in in this way. But, but it's, it can be a challenge to to find the right partner. Let's put it like that.
>> Yeah. And I, I can imagine that, you know, when you talk about expanding the scope of the data that you're, uh, that you're gathering, and the, and the diseases that you are targeting or might be targeting in the future, and sort of like building the data platform for all the different use cases computationally across the organization, and then you're, you know, thinking about onboarding a new partner, you've got to sort of pick the right, um, problem and the right, like narrow data set with, like, you know, to initially trust that partner to bring the data in, and then you get to then you like get your hands on the data in inside of a specific use case so that you could understand, you know, how would this apply to the rest of the scope of the work that you're doing. Am I, am I thinking about that right, or how would you update that?
>> Yeah, I mean, sometimes you can get sample data sets in another disease, but it's the same method, and you can, you can kind of validate that, yeah, this is good quality. You know, it can also be a matter of curation and metadata, like, and again, examples of of data sets that they generated or curated can be valuable. So that's, that's one starting point. And then, like you also mentioned, right, we can do a pilots just to, to say, like, do they deliver what they promise, right? And then you can stage the engagement as well, right? And say, there's certain milestones, and we, we basically do an initial, kind of, let's say, validation, and then when you have delivered that, and we have shown that, you know, it satisfies our criteria, then we move on to the next stage. And you can, you can build that in as well.
>> Yeah, that makes a lot of sense. Um, so I, I want to, uh, zoom out a little bit and talk more about, like, the computational workflow as a whole. You know, um, how do you think about, like, the role of a computational team in in biotech research, you know, and maybe it's worth talking about, like, how your computational team sort of complements the rest of the research that's happening in the organization, and, you know, what is the interplay like between your, between your team and, um, people who are in the wet lab or other stakeholder groups that sort of touch your work?
>> I think there's two aspects there. There's supporting questions coming from the different teams, using data, right? Typically OMIX data, or like I said, text mining from the scientific literature, and having a more supportive role, um, and simple questions sometimes can require quite a lot of work to to onboard the right data or do the right analysis. And the other aspect is actually driving projects from the computational team to generate insights, and especially in the target space, right? Think about new ways of integrating data sets, like I mentioned, the the knowledge graph link prediction, right? Like something that would not have come out of maybe of a request, uh, from a lab person. So we kind of do both things. And, you know, that's, that's, you know, the challenge a little bit to then build up the data assets that can support these things, right? And like, like you said, we build up data assets, OMIX in different indications, but that also enables us to look more globally at some point, right? And say, but what is the commonalities between diseases? Can we actually in a smart way find targets that are maybe related to an immunological process that is shared across several diseases? So I think there's also like a benefit in that sense from the computational kind of point of view to to look more globally and more holistically at, you know, these processes, and not just answer individual questions that are coming out of these research teams. So these, this is kind of the two aspects that we're dealing with.
>> Yeah, that makes a lot of sense. So you've got sort of the reactive, the reactive question answering aspect of the of what you all are doing, and, you know, you've got you've got people who are studying very deeply, you know, a certain disease category or certain therapeutic approaches, and so you're, you're answering questions that support them. But then you also, like, another thing I heard was that you're, you're building the platform to make answering those questions, you know, relatively relatively easy, and making it so that you have the have the data in hand to answer the most important questions, um, that, you know, scientific questions across the organization. But then once you've built the platform for answering these sort of reactive requests across the organization, you also have the opportunity to do that proactive exploration of the data sets that you've developed to, you know, establish these commonalities across different diseases. You know, for example, the immunological process that you just shared. Am I, am I thinking about that right? And if so, like, how do you, like, how do you prioritize between the different aspects that your team is doing or allocate resources across those different aspects?
>> It's not so difficult, really. I mean, we, we have these two accountabilities in a way, right? That's pretty well defined, but of course, they can change over time. But we have to support the teams in in answering these questions. And I think the other aspect is also just to kind of promote the say, the digital culture, uh, within the organization. Some bench scientists are very excited about, you know, all this computational stuff, and are really happy if they, for instance, get access through a dashboard. They might not start, you know, programming, but they're very eager to dive in if they have tools that are a little bit more intuitive to use. So, we also focus together with other teams in the organization to try and build these self-service, as we call them, uh, platforms and dashboards. And then you have the the the people that are more slow to to adapt these things. Um, but then as they see examples in other projects where we are generating insights and answering questions that are similar to they to the ones they're faced with with a computational pro or an OMIX data that was out there, they also maybe get inspired, then reach out to say, you know what, I have a similar question, can you help me with that? Um, so, so that's a way also to, you know, improve a little bit, like the, the impact we have on the on the different projects. And then the discovery part is the other mandate, right? That we try to really drive new computational strategies, or in a way, I, I see it a little bit as two sides of the same coin, right? The experimental side and the computational side. In a way, we're just generating data sets that are so big that you can't just do a t-test like you used to do to see what is significant or is there an effect, right? You, you need to just have a bigger toolbox of of computational tools to deal with those data sets. But in principle, it's still an experiment that is designed to address a certain question. So there is the the experimental biology aspects, and then there is just the computational aspects, but it's still biology in a way. Uh, we are not doing in silico modeling or something really purely theoretical in that sense. And I think that's the mindset. I think that's also changing a little bit. And I think maybe younger generations will not think so much about, because they probably grow up with some sort of skills and programming and will not find that so challenging, even if they're in a lab setting. And so a lot of the work is really like, but do we generate different types of phenotypic screens, right? It, it's an experiment that, you know, you have a disease model, um, and then you screen for, for instance, the drugs that bring that model back to normal. And you can do that for thousands, thousands of compounds, so it just becomes a very massive data set. So, so there's really a close relationship between, let's say, the computational and the experimental side. I don't see it as so separate. And and then we're just supporting, you know, these kind of experiments, because it can seem intimidating for some biologists, uh, to engage in these things and say, I don't have the skill set. But once we have a team that I know is dedicated to this, this, you know, all of a sudden becomes a very exciting collaboration in a way, right? And then you bring people to the same table that have these skills that can basically push these for projects forward. And and I think that's super exciting. So we have a few of these those also, like starting up, um, in the T.
>> Yeah, that, uh, that's really interesting, and it makes a lot of sense that as the assays are generating more and more data, that, you know, even if the experiment was designed in order to generate the the data, and that experimental context is sort of critical, and also, you know, follows a logic path that came out of the scient, that came out of the scientist's brain, that, sort of, the computational processing of that data is still going to require, like, somebody with a computational mind or computational skill set in order to get the the resulting insights, insights out of there. And so, like, that partnership seems to, like, these, these two sides are being driven together, being driven together, or attracted together, as much as they're as much as they're being forced to work together. One of the things that you mentioned earlier that I'm glad you brought up again is the is the idea of self-service platforms and dashboards for, uh, for, you know, scientific users. And I, I think you mentioned, you know, maybe this is related to, like, the knowledge graph and the link prediction and stuff like that. But we just be really interested in what are some of the, if there are any success stories you have with, like, self-service platforms or dashboards, or things you have in the wild right now that are really useful to your end users.
>> I think a lot of them are conceptually rather simple, right? But if you have a lot of OMIX data, it, it might not be easily accessible to non-computational people. And you just have a dashboard where they can look up, is my new target of interest expressed in this disease? Is it upregulated? I found an experiment in in a publication, very recent one, and the data looks interesting, but they don't look at my, you know, favorite gene. Uh, can we bring that data in? And then we can put it basically in these self-service platforms, right? And make it available for them. And then they can at least do some of this, let's say, simple analysis on their own. And and I think that's, that's the biggest success stories, right? Where where they can go in and and independently explore and answer questions. Uh, and then the more complicated stuff, you know, we, we might not build into a, a self-service platform. We, we just do that together. I, I think also, we, we, we have a team that is building up an internal, uh, knowledge graph. And I think that's also becoming more and more powerful because you can really in a very nice way integrate different data resources. It can be like a licensed, uh, platform for, like, competitive intelligence, and then it can be like some open, like Open Targets and other open resources. But since you can integrate it all together, you can, for instance, now we're, we're working on a dashboard that just summarizes in a, in a more harmonized way, all the information that's relevant for evaluating a target, right? So, you know, where is it expressed at the kind of single cellular resolution? What is the competitive landscape? Is there a lot of things happening there already? Are there safety concerns? Like, what is the effect when you knock it out in an animal? Are there CRISPR screens that have looked at this? A lot of different aspects, but, you know, not everybody was aware of all the platforms that are licensed internally or exist for free externally. So people were looking in different places, providing different types of information, or not simply being aware of some of them. And and now having one access point that integrates everything is is becoming increasingly valuable, I think, and and across the whole organization, not just for us.
>> Yeah, that makes a lot of sense. And are there any, like, applications? I mean, you mentioned, you know, summarizing in a more harmonized way. So obviously, that triggers the idea of, you know, sort of like generative AI or bringing LLM into the mix. Are there any interesting ways that you're using AI or LLMs internally at this moment?
>> Well, that's just, let's say, the standard use, right? I mean, I, I, we have our own internal platform for, you know, uh, it's called like, like typical version, right? And we're building agents around that. And I think that's, that's quite interesting. So it can be really simple stuff, like just, it can be related related to non-scientific aspects, HR, or or other things. But it can also be like, we were looking into patents, right? There, that these patent applications are hard to read and extract information from. So if you can just get an LLM to to digest it and spit out the results, that that will be very valuable. And so this is not really fully developed, but that, that's a use case where where we could see the value, right? And and then there's just the day-to-day stuff where people are just asking, instead of going to PubMed, they're using, you know, ChatGPT instead, and saying, what's known about this disease? Is this target relevant? But you have a, a conversation with it, right? To kind of challenge it a little bit. It's not giving you a final answer. And I think also the other learning is that, it's, it's not accurate all the time, right? Sometimes it, it hallucinates, and and you have to challenge it a little bit to see. And and you might only become aware of that if, if you're an expert in the field you are asking it about, and you see the mistakes. But then if you go into another disease area where you're less knowledgeable, it might not be as transparent. And so that's also, let's say, an insight and a cultural change, but people are embracing it across the organization. And and I really see that we have projects where you then
Try to build agents that can interact with each other because it, it can't just magically solve every question. But you can build an agent that's a specialist in like Omix data, then you can build one that's a specialist in something else, and then they can interact with each other and cross-check each other and then finally produce an answer.
And of course, it's extremely powerful for normal non-tech users as well as the data science community to use these tools to just integrate all these diverse data sets to bring you insights. And then also for programming, I think we're getting more efficient in a way, right? But there's also caveats. But you can ask it to review certain parts of your script and improve it, or like make it more memory efficient because you did something like quick and dirty and now you have bigger data sets and everything is crashing, and then you say, "Okay, can you optimize it?" I'm not really sure how to do it, and it also works quite well for these things, right? So, I think on the programming side, I also see like a big benefit to using these tools.
>> Yeah, that makes a lot of sense. Um, so as we come toward the end of our conversation, I just want to zoom out a little further and ask you some questions about like the future and what things look like. So, you know, are there any emerging technologies that you're really excited about, um, kind of transforming your work in computational biology? And so, you know, that might be on the assay side or the lab automation side, or on the, or, you know, the data processing side, or, yeah, AI, any of the things that are out there.
>> I think we've touched on them, but I mean, spatial transcriptomics for me is super exciting and I think will generate a lot of insights into disease mechanisms and multiomics aspects, right? It is just crazy how much data you can generate. I think the big challenge is that, like you said earlier, you generate a huge amount of data from a very, very tiny area of the human brain that is quite big. If you want to cover a bigger part of the brain, it's still challenging. And in pharma, typically you want to have many patients, right? You want to look at a hundred patients and a hundred controls, and that's a big challenge with these type of technologies. So, I think this is still a long road ahead, but I'm very excited about, you know, these new developments. And the other thing is generative AI, right? I mean, it's just again, crazy how fast that is going, and I clearly see that as a powerful tool in the future to bring different assay data types together and generate new insights and support, you know, a lot of the activities we have on the experimental side and on the computational side as well.
>> Yeah, that makes a lot of sense. And, you know, data integration, data integration and harmonization is one of the places where generative AI has already demonstrated the capacity to do what we need it to do. So it makes a lot of sense that that just has the capacity to make it easier to get that full picture of all of the data that you've collected and how the pieces sort of fit together.
Are there any questions that you think people in pharma R&D should be asking about their data, but, you know, probably aren't right now?
>> No, I would maybe just caution that you have to be maybe aware of the limits of your data. You know, what can it not answer and what is the quality of your data? There's a lot of hype, and for good reasons, about single cell, right? But there are technical limitations, and sometimes a computational scientist might not be aware of those because they don't have hands-on experience. But there might be certain cell types that just don't get a signal for various reasons. I mean, we had an example in psoriasis where in bulk samples, there was a clear signal from neutrophils, but we were not seeing them in the single cell data because, basically, the way the platform works is that they get captured in a droplet, they get lysed, and then because they have these internal lysosomal organelles that are, you know, phagocytic, they were basically digesting themselves and the RNA, and they don't get any signal from them. So you just don't see those cells appearing in single cell, and then you make conclusions about disease mechanisms and all this kind of stuff, and you're missing an important player. And I think that you need to build sanity checks as much as you can. You know, generate bulk and single cell and see if, you know, maybe make pseudo-bulk from the single cell and does it correlate with the bulk that you're seeing? Are the signals that are not getting appearing in a single cell? It's just one example, right? But there's many of these where you kind of like try to build in some sanity checks and be aware of the limitations. There's also a lot of interpretations on single cell in terms of proportions of cells, and they might also have biases, and you have to be aware of these because you often say, "This cell population was increased, and this was decreased." And I think it's also been shown, right, that they might be just more fragile and they might fluctuate for various reasons, but they can be confounding factors that are hard to decompose in these data sets. But it's good to be aware of them. And every technology, I guess, has these kinds of problems, and you, it's important that the data scientist is aware of these when they make conclusions, so they don't end up going down the wrong path.
>> Yeah, I think that makes a lot of sense. And I guess, you know, it's difficult to be an expert in absolutely everything. So I'm wondering like how, you know, how you or your team think about sort of building up the healthy skepticism and like the internal knowledge about the boundaries, like the inferential boundaries of these different data sets and, you know, what some of the assumptions are and how they can be used, and what the sanity checks look like. Is it just sort of a collective conversation, or how do you approach that kind of problem?
>> It depends a little bit on the expertise you have in your team. I think for our team, a lot of people emerge from labs. They have lab experience. They trained in a lab, and at some point, they had an interest in computational aspects or they were exposed to Omix data that they were generating and wanted to analyze themselves. So they have a good understanding of the experimental aspects as well. That might not always be the case. I think it's best is interdisciplinary teams. I think that's by far. If you don't have that in one person, bring these people together at the same table and have a discussion around, like, what are the pitfalls experimentally? What can you do computationally? And then try to understand, you know, what you can and cannot do, right? What is the question ultimately that you want to answer, and how do we get to that? And what are the limitations in whatever you're proposing? And is it ultimately going to give you that answer, or do we need to do it in a different way? And what people need to be part of the conversation, you know, to make sure that we do it the right way.
>> Awesome. Well, yes, it's been a real, it's been a real pleasure talking to you. Are there any final thoughts you'd like to share before we let you go?
>> No, it's been a really nice conversation. I really appreciate being here.
>> If people want to get in touch with you or follow your work, what's the best way for them to do that?
>> Well, they can always find me on LinkedIn. I find that that's the one thing that stays, you know, no matter where you move. So, I would just suggest that people find me there. I don't reject any connections, and you can message me. So, yeah, just look me up there.
>> Well, Esper, thank you so much for joining today. It's been a real pleasure talking to you and look forward to connecting down the line.
>> Thanks a lot for having me.
>> And that's it for this episode of Data in Biotech. If you enjoyed the episode, please subscribe, rate, or leave a review in your podcast platform of choice. See you next time.