Transcription
Just introduced this webinar from Bioforum IT Digital and Data, which is all about managing data as a product for digital transformation in the pharmaceutical industry. And, uh, this is produced by a team called Data Enablement for AI. Uh, the teams from right across the pharmaceutical industry, and is one of about 120 teams in Bioforum. And this is just to illustrate the groupings that we have called forums that are all working to make things better in our industry.
Just to get you started thinking about our topic today, which is all about managing data as a product in the pharmaceutical industry, we're asking people to think about why are we not getting the full expected value out of our data? And there are many answers, um, to that. So you've got some really interesting thoughts here: faulty assumptions about the data, poor requirements being specified for the data, responsibility confusion, continuity of data ownership, the fact data is sometimes not contextualized, maybe a lack of data integrity, sometimes data format, bad collection or inconsistency, a good number of different reasons coming through and why we're not getting the full expected value out of the data. Someone's got no strategy, it's not real-time yet, there are data silos, it's not adopted. Oh, that must be getting, if you put all the effort into building something and nobody's using it. Um, sometimes there's bad models, digital literacy, lack of connected systems. Whoa, the, uh, cloud is going wild as people are adding their thoughts to this question.
So over the past year or so, 40 members of this team have been meeting every two weeks online and actually for three days face-to-face in the Netherlands. And they share their case studies and strategies for how they are managing data as a product. And all have discovered it's quite challenging. So, uh, today we're not presenting a new theory. It's what we have collectively learned as practitioners putting theory into practice. We have six speakers, um, and they are also authors of a publication. Some of them are live today, some of them are pre-recorded. And it goes something like this: Sandra first talks about the current situation, why data is hard, and why pipelines frequently break. And then Maru steps in with the elements and characteristics of what, uh, does it even mean to manage data as a product? Wilfred then talks about, um, the evolution to data as a product because he's tried it, and it's quite tricky to get there. De reminds us that there are a few different types of data product, and that sets the scene for Kate to show what we've learned about the practical realities of organizing teams to managing data to manage data as a product. And then Bob comes back to that problem of things breaking and the need to manage the whole lifecycle of a data as a product. And again, it's a combined view of cycles and phases that we actually use as practitioners.
So, um, first of all, Sandra, which is going to build on the things that you guys have been adding to the word cloud. I've asked her to introduce herself and say why this whole topic is relevant to her and her work. She is here today, but she's chosen to pre-record this segment. And the questions for Sandra are things like, what's the current situation? Why is data hard? Why do things frequently break? And as she talks, please feel free to jump down in the chat points that resonate with you. We might pick up on some of those afterwards.
Okay, hi. My name is Sandra. I work at Novo Nordisk as a data architect in a central architecture team. And if you can move to the first slide, I'm here to talk about the current situation. So, we, I guess, why is data so hard? We all want to make more and better use of the data that we have. So why is it so difficult to actually create value? I guess this is a common question that multiple enterprise-sized companies have actually asked themselves. But I do think there's some unique challenges within the pharma industry. We do have a lot of OT data from many different sources, making, for example, interoperability very difficult because we have to stitch data together from a lot of sources to create a unified view of the world. We do have key business units that are non-digital native with lots of legacy systems. We have, we are in a heavily regulated industry. Uh, we do have a lot of external partner organizations, and there's a lot of siloed data domains.
So let's start by understanding why creating data is so hard. I guess what we're actually looking for, um, is not necessarily only the data as such. It's also the, the insight, the information that we can get out of the data, the knowledge, and maybe even the foresight. Um, so if you look at the drawing or the illustration, then I guess it's the top of the iceberg that really provides the value. It's these assets that the business rely on to understand what is happening and to support informed decision-making. It's the dashboards, it's the visualization, and the reports. And I think maybe even moving more into models these days, it's the wish to make sense of the data in the vast quantity that we are collecting it in and converting it them into actionable insights.
So how do we do that? We've been talking a lot about that. Um, so if you look at this, then you can see that that is the top of the iceberg. But we need to create the foundation to provide these meaningful insights. And I think that is all of the sort of the "list fun" work, um, the work that we don't see, that is less obvious. It's everything that lies underneath the top of the iceberg, um, that is where potentially a lot of the true complexity comes into the picture. It's foundational work, maybe it's up to 80 or 90% that we need to do in this sphere to actually make data of high quality, that is consistent, that is findable, and interoperable, and reusable. And I think also because we talk about large, uh, amount of data suddenly, then it's also about being effective and flexible, um, in this new schema.
Then we can ask ourselves, what kind of disciplines do we need to in order to do this? And if you look at the right side, you can see I talk about data strategy, data governance, cataloging, master data management, um, data ingestion, transformation, and even essentially data platforms, also buzzwords, and artificial intelligence platforms. And I think that maybe all these capabilities and tools and disciplines are new to some of you, at least they're new to, to the pharma industry in general. So it therefore requires effort and financial resources and dedication to master, um, and probably also new ways of working. And I think an additional complexity is also that there isn't one golden path on how to master, master this. But there's a common set of, of, what would you say, challenges that we've been identifying.
So if we jump to the next slide, then we can see some of those. So I think some of the common denominators is that we're not getting the expected value out of the data that we do have. Um, so why are we not getting the expected value? And I think one of the sort of predominant reasons is that the expectations have also risen. It's become a lot more a competitive edge to actually master this. We have this idea that data will be the key way of increasing our productivity. We want to make our sites digital, predictable, and adoptable, like you know from the digital maturity model. And then suddenly there's a huge, uh, focus on large language models, ML, AI, and all of that requires data, and it requires more data than we have been collecting before.
So I think that also leads us into the next one, uh, skill and complexity. Because to do all of that, we need to collect more data, and it's still from unharmonized sources. Um, and furthermore, it's not necessarily only sensor data that we're talking about. Uh, so it's not only OT data. Suddenly it's also image data and large, for example, gene sequencing files. And we also have requirements of of real-time data potentially. So it's a large scale increase, and suddenly also a complexity increase. Then there's a general consensus that some of the, the, the, out there, they don't really match, uh, the practicalities of the pharma industry. So there's a theory and practice mismatch. I think also, I alluded to before, that the capabilities, uh, within the pharma, pharma companies are, I, I think they, of course, they're specializing in developing, um, and manufacturing pharmaceutical ingredients. We haven't necessarily taken the steps to become, uh, data, um, native companies. So it requires an upscaling in capabilities within the different disciplines that I mentioned before, and also, of course, the tooling components. Suddenly we have to have data catalog, we have to have lineage tools, quality tools, and different ways of working.
Then, of course, concepts like we're going to introduce data products, they're not well understood. And if you don't understand something, even as a business, uh, manager or as a PO, it's really difficult to actually implement. So I think that the requirement to achieve such good data products, they're not understood, and therefore very difficult to handle. Then we have the whole thing about ownership. We talk about ownership a lot. It's a central concept in both Fabric and Mesh and other theories. Um, and there's this kind of, I would think, a duality in it because there's both a benefit from owners if they know about the engineering and understand the engineering behind the product, but also if, so a technical person, but also if you actually know the data, so a business person. So what, like, how should you actually do this? Then the whole way that we are organized around it's, I think it's, um, it's very organized around sites. Um, and when we suddenly create data products where we potentially aggregate data across both sites, but also across different systems, then there's a mismatch between how we organize and how we actually want to handle data. Um, so when we create data centralized data solutions or global data solutions, how do we fit that into a site sort of organization where we want to potentially have a matrix organization? That is also one of the sort of, um, the reasons why it's difficult.
Then there's an immature governance component. Um, there's, if you Google data governance, you get a lot of different frameworks. They often have like stuff like quality, they have lineage, potentially access security, quality, I said that already, of course, ownership. Uh, there's also, let's say, people, processes, and technology as a component of it. And I think actually, we're immature in all three parts, at least I can say that that there's work to be done in Novo. Then the lifecycle part of it is not really managed that well. There's a tendency that that a technical team comes and they build something, and then they hand it over to business people to run it. But I think in more mature product, um, development cycles, you build it, you run it, but we're very far away from that one. So I think in general, if we want to get more out of our data, it needs time and priority. But let's look at the current situation. Now, we've talked about some of the obstacles. If you jump to the next slide.
Okay, I talked a lot about different obstacles, but I would also say there's a lot of good work being done out there. Um, so it's also very promising. I think that we can all agree that the data exploration part of it, there is a lot of success to be had. We have a lot of use cases, there's lots of potential, a lot of business or data-capable business people go and make great solutions. So there's like stuff worth doing out there. It's simple architecture and requires less scaling. But what then happens also is that when we sort of over time, and when, or when we want to move something to production, uh, it's difficult to scale, it's difficult to make global solutions, it's difficult to professionalize the code with version control, error handling, all of that. Data often becomes untrustworthy, pipelines they fail. We don't treat, for example, the cloud platforms that are part of sort of the data landscape. We don't treat them in the same capacity as we would do with other production systems. So we don't talk about change control or manage all of that. Um, so pipelines fail, and it doesn't scale. And when we talk about the impact to the business that actually relies on these solutions, then they don't trust our solutions. They're time-consuming to maintain, and they require a lot of, a lot of resources. It's also the general consensus, if something breaks, then we have to fix it quickly. Um, and I think we are in a situation where we're building up technical debt unless we do something in a different way. So I think to sum up, a practical strategy for structuring, uh, the responsibility of working around data is needed. And I think that is some of the stuff we'll cover later on in the presentation. So that was my part.
Fantastic to, uh, hear what Sandra had to say there. I haven't seen any comments coming in the chat, but please do add those if you'd like to. Next up, uh, it's Maru, and he'll introduce himself in a moment. But what do you, oh, lots of applause coming in. Good stuff. Sandra, you can put your, um, camera on and take a bow there if you liked. While I'm talking, but as you're thinking, what do you think is essential about managing data as a product? Um, Maru has got a really clear way of explaining what this means, and for me, his explanation was a light bulb moment. Um, for me, data products are not just a new name for what I did 20 years ago when I was doing some data reports and so on. And what Maru has to say will also help you answer the next question on Mentimeter afterwards. So let's hear from Maru, also on pre-record.
Thank you. Uh, and I, uh, so my name's Maru Kagel. I, uh, lead the Data Foundation team within AstraZeneca, part of a wider Data Analytics and AI, uh, competency center. And I look after deliveries on data platforms like metadata, data quality, master data, um, and one of the key things that we're doing is around data product delivery and data product, uh, excellence. Yeah. So I guess it's great to see in the chat there's, you know, some of the things around accessing data, finding data, accessing data from multiple sources, you know, and this is the most common use case for a data product, right? Is for people that are struggling to access and find data, and that's what it is. So, you know, we try to come up with the definition. So, you know, the definition of what a data product varies, um, quite a lot, I think when you're talking to different people, talking to vendors, talking to consultancies out there. But what we were trying to do was to create a common definition that everyone agrees with. So a data product collects data from the relevant sources, which could be multiple processes, and contains, uh, enriched metadata to make it findable, accessible, and understandable to anyone who needs it to meet a specific need. The specific need is what gives it its value. So I think, you know, the debate around there is, is around, you know, the data product container, is it a data asset? Is it much more? Is it a report? But ultimately, everyone has agreed that, you know, ultimately the heart of a data product, there is data, and that data should be reusable. Um, you know, we shouldn't be creating things for one-off use cases, but something that's foundational that people can go out and find when they need access to data. We have metadata that enriches it, um, and this is probably where the most debate happens, is well, what metadata should we be storing around the data? Um, and I'll, I'll, I'll get to some of the definitions around, um, what we believe the metadata should be and the attributes of the data product. And then lastly, but not least, is an owner, right? So if we see problems with the data, we have questions of the data, you know, who, you know, just like any product, even outside of a data product, a product runs for the life, right? So who owns it? Who helps drive the vision of that product? Who helps prioritize what we do with that product around? Yeah. So it's really important that every product has a name.
There, move on to the next slide, please. Before we do, Eric is interested in the last sentence where you say you've got to do a fair score, and he wants to know how do you do that? Yeah, that's a really good question. Um, so I mean, for us, we have a number of different measures that make it, uh, findable and accessible. It's not an exact science. Some of it is around, you know, the metadata that's being stored and cataloged. So, you know, we take different personas, uh, so what would it take? What needs to be in place for it to be findable? You know, so, I.E., you know, it's, it's got a description, it's in a marketplace or a catalog that people know how to reach, you know, for accessibility, it's in the database that I can connect to, consume. Interoperability is obviously around the master data and the controlled vocabulary and the reusability. So we do have criteria. I know at AstraZeneca, it's about 16 points, and that's what we use to score it. But I would say it's only an indication. Yeah, it's quite a difficult thing to score. Thank you. Good. Carry on.
I think, you know, going back to what David said, you know, back in the 90s, right? Or when we're creating data assets, we were doing data warehousing. It's, uh, you know, basically we had some code, generated an ETL job, and it produced some data. And that data could have been for a single use circumstances, or certainly started as a single use content, you know, for maybe a report. Then we came up with the concepts of, of, you know, more enterprise data warehousing and came up with data models that were much more flexible and reusable. But still, the main focus was on data. Um, and so, you know, certainly one of the first things that when we first started out on our data product journey, a lot of people just took their data, their data models, and said, "Oh, I've got data products now." But actually, a data product is much, much more. So if we move on to the next slide.
So these are the different elements, you know, that make up a data product. So yes, you know, at the heart of it is still data. You know, it's still data, right? Have to have data. Um, but we also, you know, and now we start talking about what's the metadata that we collect around that data. So absolutely, the code is part of that, right? But we look after of the code, versioning, the code, it's part of the product, it's part of the lifecycle, we treat it that way. Um, I'm going to go down a little bit to reusability. Right? I've spoken about this a bit more. You know, we shouldn't be building data products for single use. You know, as you can see here, there is effort in building a data product. It's much more than the slide before and creating a data set. So we have to pick and choose when we do this, right? We do this when it's valuable and reusable. Yeah. For single use, you might just want to create your data set as the one above. So, you know, within AstraZeneca, we have enterprise data modeling standards, conceptual models, so we, you know, that we use to link the data together so that we're creating data products that are sort of categorized or packaged into certain areas that can then be interlinked going forward. Um, okay, they need to be connected. Um, and, you know, because we could have data products, let's say for sales or for our manufacturing, uh, process orders, or it could be for financial budgeting. And quite often our processes go, you know, just like our processes go across systems today, and we need to interconnect the data. Sometimes our processes will go across the data products, right? We need to be able to combine them and bring them together. You know, an analogy is brought up within AstraZeneca for certainly is around Lego bricks, right? You think of a data product as Lego bricks. You take the ones that you want and you build whatever you need for the insight, right? And you drag and drop them. Um, so, you know, having that important connectivity is key, right? So if you don't have good master data, you will struggle to connect data, especially if they're in disparate systems. And also, you know, around controlled vocabulary. So, you know, if one system is using a different terminology to another, we need to have a way of aligning them and have a common language throughout our data products.
Going to go to data quality. So if people are going to use, you know, these data assets, one of the things they're going to want to know is what's the quality of the data. So there isn't, you know, in terms of our framework and methodology, it's not necessarily around fixing the data. There's lots of other methodologies around that. But it's really important that we're able to show the quality of the data. And that could certainly feed into a backlog, right? For the data product for improving it, right? We never build the perfect data product, but certainly we want to highlight the quality. Is it good? Is it bad? Is it trusted? Uh, and so we define metrics. Some of those metrics are around observability. So common ones, you know, what's the latency? I'm looking at the data, how old is it? Um, you know, how complete is it? But then we also, uh, create metrics and definitions that are much more tailored to the data product around, you know, maybe sales, you know, that might be an important field in sales, like order number, and we could measure the completeness of that. We could also associate policies around certain standards, you know, maybe lengths, and highlight and put monitoring in place on top of that.
Going to go to the top and talk about catalog. That's really important that we put metadata into, um, a catalog. Yeah, you know, whether it's, um, a BI solution, you know, there's many out there on the marketplace or something that's homegrown. It's really important that there's a central place for people to go in and search for data products and understand what's there. You know, and it's really important that we augment the products, you know, with descriptions and tagging, you know, so that, um, yeah, it makes it as intuitive as possible. You know, what we find is the less metadata you put there, the more results come out. So if you went into our catalog, for instance, and typed in "batch," you would get thousands of results. But once we start augmenting that metadata with descriptions and tagging, we're able to hone down onto the specific data products that we want. You know, part of the trust is also, um, lineage. It's really important that we understand, I guess, two things: one, where's our data coming from? Where's this, how's this data product been built? You know, is it taking its data from an SAP system, a financial system? But it's also important to understand usage, right? Which is, who is using the report, right? And so the lineage will take us through and show us maybe the reports that are used, and then we combine that with the number of people that are looking at those reports. We can then start assessing, you know, that's one of the metrics that we can use to start measuring the value. The number of people that use it, um, you know, if one person's using it or 500 people, it's not a definitive indicator on value because one person could be using it for something very, very important. But it's an indicator as to the value of the data products and whether we should continue enhancing that product or whether we should even think about retiring it if we're not getting usage. And then the last one on here is around semantics. For us, this is quite a hot topic at the moment. It's, um, getting a lot of, uh, visibility. But it's really around, you know, data products connect, and we know they connect, and we know data in manufacturing can affect our supply chain. And so, you know, when we're looking at data or interrogating data, obviously, you know, as a human that understands those relationships, we can go in there and sort of create a report by pulling the data in ourselves. But what we want to do to make it more machine actionability, going back to FAIR, is that we want those relationships to be stored somewhere, right? So we create either ontologies, some of those stored in reference systems, some of those are brought into knowledge graphs. And what this is sort of helping to drive for us is our solutions around generative AI, right? It's a very hot topic at the moment. And obviously, if those relationships are in there, we can start asking natural language questions that sort of transcend the data products, right? Be very, very difficult unless we went in and manually started creating very specific use case views on it.
Let me interject. Emma made a comment before you showed this slide, actually. She said, "Just because it can be found doesn't necessarily mean it's usable. How do you define reusability? Is it that it can be cataloged? Is that enough? What about on top of ology and taxonomies? Do you define the minimum metadata, or is it up to each data product PO person or something?" Yeah, product owner. Um, so we have a standard that goes across the board, um, around some of this, um, but some of it, you know, certainly the additional metadata could be up to the data product, uh, owner, basically. Um, you know, so for instance, for us, ontologies is not a mandatory metadata, uh, today. We want it to be, but we're also on a maturity curve as well. You know, so we don't go from, uh, you know, when we build a data product, it doesn't have all of this, just to be clear. You know, what we do is, you know, this is the standard, this is where we want to get to. Our gold data product, we may start with some code and data, even, and then start some of the cataloging, then we start building to this. But we do have a minimum, uh, data set in terms of reusability. It's more around, you know, sometimes it's a very specific AI or analytics use case, and that data will only ever be required for that use case. Then we wouldn't build a data product, right? Because for us, it's not reusable. There isn't any other use cases. If it's a single use case, right, doesn't make it reusable. If the data product can fulfill multiple use cases, then we define that as reusable.
I just really some of the questions, but I'll take. There's one around data products, GMP validated. We do, uh, we qualify our, uh, the platforms where our data products are, um, and, you know, validate any processes to be with the GXP compliance. Okay. In terms of data product characteristics, how do I know? You know, we've set those standards, we've sort of raised some of those elements. But ultimately, what's the acid test? How do we know when we build a data product, right? And it has to meet these characteristics. Has to meet all of these characteristics ideally, and if it doesn't, we don't, uh, we don't think of this as a data product, right? So has to be discoverable, right? And and that might mean different things to different people. You know, an IT person could have a very technical interface, they might even go directly into a database, right, and find it through schemas. But someone in the business that is maybe using, uh, something else, they would need a much simpler interface, right? And so it depends on who you are. But ultimately, you need to understand who the consumers are of these data products and ensure that they are able to go in and find the data products. It needs to be addressable, right? So once we found the data product, well, that's great. We can't just find because in essence, what we're finding is the metadata. We don't want to find the metadata, we actually want to be able to get access to the data. So we need to understand where it lives, right? And what are the access protocols, you know, especially if we're looking at strictly confidential data, there's probably an approval process that we need to go through. All of that needs to be captured. It needs to be trustworthy, right? And again, you know, how we define that, some of it is made up of those elements, you know, understanding where it comes from, you know, what the quality of it like, who's using it, maybe I can go and talk to them and ask them what decisions they're driving. Um, and it's also based on the use case. Does your data need to be 100% accurate? Well, it depends on what decision you're making. Sometimes it does, sometimes 90 something percent is good enough to make the decision. It needs to be self-describing, you know, if I go in and look at a data product and I look at a title or a description, and I don't understand it, and I have to go out and reach out to people, it's going to cause delays, right? So we need to be able to understand that it needs to have enough metadata in there so that whoever is trying to access these data products can understand. Um, interoperability, right? So if we're bringing multiple data products, we need to be able to connect them, right? And from a secure perspective, I mean, yes, we can allow anyone to go in and see the metadata and see the data products, but we have to ensure the right controls are in place so that not, you know, so not everyone can access all of the data. Um, so, you know, the right data to the right person, where is needed. And ultimately, we go back to, you know, one of the third things I said, which is it needs to be valuable, right? We shouldn't be building this unless there's some value attached to the data product. We shouldn't be building data just to build data and hope that someone comes in and uses it, right? So we do that based on use cases as how we develop our roadmap for data products, and we start small. Sometimes there could be five or six fields, and we expand as people want to use it, right? And that ensures that there's value out of what we're building.
So Maru, thank you so much. But I do have a question for you. This was inspired by a fidget spinner. Whose was the original fidget spinner that inspired you? It was my son's. Fantastic for about a year of his life, wouldn't put it down. So Maru was describing the wonderful world of proper data products, but it's quite challenging to get there. And actually, all of the member companies in our team have started in slightly different ways on their journey to managing the data as a product or data mesh or data fabric, or whatever. And I wondered, I'll just post it into the chat, a question for you: One point you think is important to get right about data products early in your evolution. And that might be from your experience, you're well on the way, or the thing that you're thinking is the most important to do next. If you're not on your way to data products, it might be some of the things that Maru said in his presentation, or it might be something else entirely. So please could you, 185 people on the call, let's get some views going on what is important to get right about data products early in their evolution?
Um, so the next segment is again pre-recorded, and this one is from Wilfred. And Wilfred is going to talk about his experiences of that evolution. So here we go.
Thanks, David. Um, so my name is Wilfred Mascare. I'm with the Eli Lilly company. I'm in a central role in our manufacturing and quality data and AI organization, and my primary responsibilities are, um, innovation and, uh, strategy in the space of data and AI. So as David said, you know, we all agreed that the data product is the right approach to building your, um, data maturity and making data available to end users and consumers, right? But there is an evolution in terms of getting, uh, data products built. And we had various discussions in terms of how do you go about doing that, both considering organizational capabilities, organizational maturity, the way, uh, companies are organized, so on and so forth. So we talked about, uh, data mesh approach, which is basically a federated approach where you have, you know, owners across their companies, everyone owns their data, everyone builds their data and makes that data available. Then you have the data fabric approach, which is more of a centralized approach where everything is centralized, um, and, you know, the data is governed centrally, is built centrally, is monitored centrally, so on and so forth. And then you have, you know, you could have data contracts approach, which could be building data and then having just a contract saying, um, you know, this is what I'm giving you, and this is what I will promise you. And then you have the data warehouse approach, which, you know, DAMA has a lot of work around that. Uh, and then of course, we talked about the data products. So as you think about data products, I think Maru, um, you know, showed this diagram, which I think is a money slide. I really like this slide, which explains in a very concise way what makes a good data product. But doing this is a lot of work. And what we said is, hey, organizations should not think unless you do all of this, you cannot start the journey, or unless you have all the tools to do everything, you know, you can be successful. Um, so what we, what we talked about is, you know, you don't need to start your journey with thinking about all those eight elements for defining a good data product. You could start small, and depending on the data product, whether it is GMP, non-GMP, which is high risk, low risk, right? You can take different approaches. Even at the start of the journey, you could have some data products that have six elements out of the eight, some could have three out of the eight, so on and so forth. Even at Lilly, you know, we said this is the minimum you've got to do for GMP. Maybe you want to do some more, or there may be a higher minimum standard for GMP assets and so forth. This could be the minimum, uh, viable product, or definition, if you will, of a data product. Creating the data set is the most, you know, without a data set, you don't have a data product. So obviously, you have to have a data set. Then on top of that, you could build a data catalog, and a data catalog could be as simple as a SharePoint list, could be, you know, as, uh, complex as buying some product out there. Additionally, using AI to automate a lot of the work around cataloging. And then we also talked about a marketplace, which is, you know, you may have multiple catalogs that feed into a common marketplace, which is where people go to not only find things, but also provision, provision their data products, get access to data products automatically, and get everything they need to start using the data product. Lineage is a very important aspect. We all agree that at the minimum, we, you know, we need to know where the data is coming from and how it is being consumed all the way to the reports. And last, but not least, you need some usage metrics. And then data quality, without, uh, data quality, you don't have a data product in my mind, especially in our world, from a compliance standpoint and a GMP standpoint. So even data quality can be as complex as automated data quality, a lot of AI-based rules, etc., but it could be as simple as, you know, having the basic sets of data consistency rules, data integrity rules, etc. We think this is the minimum. But at the same time, I want to point out that even though this is the minimum, you can also start small, in terms of what you catalog, how much you catalog, how much you track in terms of a usage standpoint, and how much, how you mature from a data quality standpoint.
What we realized is, companies, in the manufacturing space, or organized based on manufacturing sites. A lot of companies don't have centralized domain teams. Right? Within Lilly, we do have a quality team, we have a team for MES, for manufacturing execution system, for lab systems, etc. But large organizations may have that, but smaller organizations may not have that same thing. In large organizations, there's a lot of distance between, you know, the producers of the data product and the consumers of the data product. So you have to be aware of that. How do you bridge that gap? How do you make sure that what you produce is actually consumable at the for the end user? So involving your consumers is very critical in terms of defining the data products and defining the, you know, what constitutes an actual consumable data product. We also talked about, to, uh, basically mitigate that challenge between consumers and producers, you know, why don't we give the ownership of the data products in the hands of the operations team? Right? And we very quickly realized that that's not possible because they are really focused on getting manufacturing operations and getting product out the door. They don't have the time to sit down and figure out what is a good data product for them. So there is a lot of, um, tension there in terms of spending the time to build data products and define data products. So, um, then data literacy, we thought was a great, another challenge. The space has evolved so fast and so quickly. How do you, you know, bring the entire organization's data literacy up so that they understand the concepts of data cataloging, data governance, data ownership, master data, reference data, so on and so forth? How do you, how do you have at least some basic understanding across the entire organization? Code and data is not always, uh, have the same owners in practice. So many times, you know, system owners may not be the ones owning the data. They own the data in the sense they own the data from that is generated from the system, but they may not be thinking about downstream analytics use of that data. So our data becomes more as an afterthought. We also discussed, you know, that, um, not everything needs to be centralized. For example, if you go the data mesh approach where everything is federated, you can't have 10, uh, catalogs, right? You cannot have 10 marketplaces, you cannot have 10 databases, each one managing their own database. It would be a huge cost overhead to have a completely decentralized approach. So no matter which approach you use, we all agree that there is some amount of things that need to be centralized, and you need to think about very carefully in terms of what needs to be centralized to optimize your costs, etc. At Lilly, we've used a hybrid approach. So we are doing the data fabric approach, bringing all the data into a centralized data store. We are fully GMP validating that data store, and we have centralized organizations in quality, you know, labs, each function has their own centralized organization. So we're working very closely with those central organizations who, in turn, you know, work with all of our manufacturing sites to get the right requirements from those sites, which then gets articulated into a centralized data team that is building the, uh, the data data products. We do domain analysis to make sure the data product is appropriately defined, has all the, you know, has all the metadata, has all the master data, so on and so forth. So we do with the domain analysis more at the function level, but the build of the data, the platform qualification, all of that done is done centrally, within a central organization. And we've seen, and when, and from a data mesh standpoint, you know, we are using that approach more from a governance standpoint. So the governance is federated, but the build and the ownership and the support of the data is all centralized. So that's kind of the approach we've taken, and it seems like it's working very well in terms of, you know, how we are organized, as well as our maturity in the data space.
Fantastic. So, um, I'm going to give you guys a chance to do another little Mentimeter. There were six topics that Wilfred came up with. And as you guys start to answer that question, we should start to see a bar chart or these sort of, oh, wow, they're reordering themselves. But I also noticed lots of interesting comments about where we start. I noticed Martin De Boer from Waters, I believe, asking, "Should the vendors create source of land data products?" And Dej responded to say, "Yeah, it's a great idea." And I know others have been talking about that within our team. Yeah, people are talking about the security model, getting that right. And Thomas says, "Getting the goal is critical." These are things, you know, getting the right use case with the actual valuable purpose, that's important. Ownership from Karen as a key issue. And let's see what else there is. Traceability and auditability, that sort of lineage, but also I imagine from a GXP perspective. So we've only seemed to have had five people responding. No, maybe I need to click. We've had 20 people responding, and they're putting first so far, the data literacy requirement is high, is one of the challenges of federating data ownership, and then ownership of operations and data products don't mix well. Some things need to be centralized, others decentralized. It is fascinating to see which order things are turning up in from your, you guys, and your experience.
Now, Dej, I need to see your face and know that you're here, ready to do the next segment live. So if you could do that, that's good. Just to introduce, there are actually different types of data products, and in the case studies, there are at least six different ways of calling or classifying data products in each organization. The concepts are similar. So, could you just introduce yourself to tell us about the different types of data product and why that matters for how we manage them?
Okay, so my name is Dej Kyle, um, from Lilly, similar to same as Wilfred. Um, my role has been in leading the building of the first data products at Lilly, and I've also now transitioned into a role around introducing the, um, data governance to manage them. So this whole topic and working on this whole publication has been really helpful for me, um, and for us at Lilly for progressing our data strategy. So as David mentioned, there were many different definitions of data products that we all discussed through this publication, but these three concepts were where we very much aligned that there is three different types of data products. And if you look at this slide from the data, from the bottom up, we're talking about the analytical plane when we're talking about data products. So we have all our systems at the operational plane where you're all running our our business, and that data is brought up to the analytical plane through ingestion to the cloud or whatever other method where we build these source-aligned data products, which is the first type of data product we talk about, and I'll be giving some examples of those in the next slide, but they're linking to the source, just making the transactional data more, uh, consumable or useful for analytics, as much of the data in the operational systems isn't required for analytics or isn't in the right format. So that's the purpose of those data products. Then we have the aggregate data products, where we combine data from different sources, according to a specific use case. So this is where, as other speakers have talked about, the use case is important. And then when you have an important use case, you look to build aggregate data products where you bring the data together in the right context and relationships to make it valuable for answering questions about that use case. That's the second type. And then the third type of data product is the consumer data products. And we've talked here, kind of very generally about two types of consumer data products, the formal ones where you go all.
The way to building an Absol, a solution on top of the data product. So you're are building, um, maybe a validated report that can be used to submit something for regulatory, for example, or a dashboard that's very much meeting a use case and can be filterable and used in that way.
But the big difference about what we want to do with data products is we also want to make them consumable for ad hoc data analytics. So as we have many more types of users with many different needs and capabilities, with the whole explosion of AI and all the different modeling, we also want to make the data products available for these use cases, uh, to enable those types of users who have the capability and the and the questions to ask and and to be able to give them those insights.
So the next slide will give, I'll give a few examples of these different, um, data products. But before I give an example of a data product, an important thing to think about, which Maro mentioned earlier, is that in the operational plane, this part at the bottom of the slide, it's important to think about the fact that there is always a need for application-level data solutions. Not everything needs a data product because, as you can see, you know, it's an investment and sometimes if you're looking at a single source, it's more relevant to look at something that's given to you out of the box by the application. Or if you need something in near real time, something that's more focused on operations, sometimes using something at the operational player level is more appropriate.
But then to explain a little bit more about from a data product perspective, you can see here in the diagram that we've outlined some data sources, um, and their data is being copied up here to the source-aligned data products. And this is the question from the chat, you know, is it appropriate for vendors to help build some of these source-aligned data products? It definitely is. And and my next slide goes on to a little bit to, you know, who the different, um, owner is or, uh, who's most appropriate to to manage these data products. But these source-aligned data products are very much relying on people who are experts in the source system data.
So one example that we have at Lily that we're building at the moment, uh, around source-aligned data products is for our LIMS system. So our Laboratory Information Management System. We're moving from a from one, uh, system to another across all our manufacturing sites. But users will want to be able to look at their samples, their tests, their results, regardless of what transactional LIM system that data was entered into. So for us, an important source-aligned data product will be to create one data product that for LIMS data that will extract the sample results, tests, data in a very consumable way for analytics, um, across those two systems. So that users don't have to go figuring out, uh, which one they they need to extract from.
And then going to the aggregate data products, an example from our experience in Lily, and the first data product that we went live with, um, a few months back, was the asset data product. Um, and this was based on the use case of asset qualification monitoring. Um, and what we did there is we brought together our maintenance, um, data with our deviation and change management data. So and built that data together from those source systems in in that context with those relationships to be able to ask, answer asset qualification monitoring type questions. And this is version one of that data product. There are other data sources that would make that more powerful to be able to further, um, fill in or further answer the asset qualification monitoring questions, but also to go broader and give us more data about assets for more insights that we want to get there.
Uh, and on top of that asset data product, we then built our first consumer data product, which was an AQM asset qualification monitoring dashboard. So this, uh, has all of the asset-related data from all of our different sites across, um, our manufacturing networks. But their dashboard is, um, computer system or validated for GMP purposes. And it's allowing users to filter it for the monitoring of their own site or their own data in the specific area of a site, but also maybe to look across multiple sites where there's similar, similar type of assets or equipment.
As well as creating that formal consumer data product on top of that asset aggregated data product, we have also, uh, made a GNA, uh, solution on top of that where we have natural language processing asking questions of the asset data product with the same data set. Um, very much in the exploration phase of that one, but already we're finding it really useful. And it also shows the power of the same data product being able to be used for for multiple purposes.
So the final slide for me then is just talking a little bit more about the different rules in managing these data products. And here again, just to what I wanted to really highlight about the source-aligned data products is having the, um, input from the SMEs of the source system data, the the source system data experts is key here. Where is in aggregate data products, what's more important is that you have a business product owner, so somebody who understands the use case and the data that will meet the needs of that use case and how to relate that data to to the different sources. And then from a consumer data product, if you've built your adequate data products well, then, you know, there should be multiple options. The the possibilities here should be significant. And then it's a case of meeting many different needs for different types of consumers. And that and then the the insights and the questions that they want to ask will influence future versions of the aggregate data product to make sure that they continue to evolve to meet their needs. And I know Kate's going to go into more detail about some of the roles around these, uh, managing of data products in her section coming up next.
Right, your mentee should now have moved on and allow you to answer this question. What balance of data engineering and data domain knowledge do you think is needed for these three types of product? And a difference in points of view, actually. The graph has many bumps on it. So the balance of data engineering and data domain knowledge is a tricky question, isn't it? Because it's hard to find deep data engineering skills and deep data domain knowledge in the same person. And so one of our conclusions, you might need a team most of the time.
And de, just come back on the line a sec. We've got a question here. So why don't you have a go at that one? And shall I read it out for you, just to give you a moment to think? Hi, dear. Any reference system architecture? How different data, different products in the three layers connect to each other? We don't have a specific reference architecture on the connection. Um, what we like from from a architecture perspective, our source system data comes into raw and refined, and we are building these data products at the conformant layer in the in the cloud architecture. So all of them are at conformant.
Mao mentioned it, and we've also used the analogy of Lego bricks. So the idea of the source system, source-aligned data products being sort of the bottom layer of the Lego brick, uh, you know, if you think of a wall, and then you have your aggregate data products built on top of them. But even then, with the aggregate data products, when it comes to the batch data product that we're working at, some of the batch data is foundational and would work for all use cases. But then we'd have different modalities with different types of genealogies, for example, that would almost be another building block on top of that batch. So it's not a reference architecture. We do have reference architectures around what tools and things like that we're using in the cloud, but that's kind of the best, uh, answer I can give for the data product slant. It's not strictly, you can only have three layers and three products. There might be more, there might be extra ones, etc. That's your s. Thank you, Prad, for your question.
Good. Do you want to introduce yourself and tell us a bit about how do we organize ourselves, um, so in teams? Um, I'm Kate Leelan. I head up data standards and governance for CMC in GSK. Um, we're looking at how we, um, get to FAIR data for CMC and bridge the data governance gap with our, well, initially with our manufacturing organization, but broadly across the enterprise. So just in terms of how we operate, we are a big pharma. There's no one model that works for all, but a really good reference guide is this Team Topologies book. So the note here of the quote saying, "This is the secret source for successful data products." It is a multi-team effort. We're looking at the business, tech, data access, potentially a lot going on, high cognitive load. And Team Topologies recommends that you have small, stable teams with clear boundaries that work together to deliver against a common value.
So how do we apply this to data products? So just reflecting on the flow around data products, starting with those operational systems on the left, moving through source-aligned data products, taking that into the analytical plane and generating aggregate data products, bringing many sources together, and then consumer data products that are the business-facing data, the data that matters to drive those business decisions. So what are the roles along this, um, spectrum? So we're starting maybe on the left with some of the operational systems. If you move the slide on, David, we start talking about earlier, we mentioned that we may have some business system owners taking accountability for source-aligned data products. This is a relatively rare way of working. There are always challenges if you're giving people with operational roles a challenge of then looking after data, particularly if it's data that is downstream of their particular source.
You could also look at having, um, engineering teams accountable for data products. You may have a specialist engineering team who are accountable for ingestion pipelines, creating those source-aligned data products. This is quite a common model. Your engineers are probably quite data literate. They may well understand the source data and how to then aggregate that onwards into consumer data products. However, engineers are a scarce commodity, and we may find that again, having a persistent data management role when there are further operational needs for these teams can be a real challenge. So, I mean, you could also then look, if you move the slide on, David, at the domain data product team, so where you have specialist groups of analysts or engineers, data managers in the domain, generally multidisciplinary, many skills coming together to deliver the consumer data product. However, rare to see these types of teams taking on accountability for source system data products. However, largely they may well take on accountability for lifecycle management of consumer data products, especially those that continue to deliver continual data to to projects, assets in the pipeline.
So many models, bringing it all together. There's a couple of more complicated slides coming up that show how all of these roles could come together, adding in the data governance body, so covering the flow of data from source to use. Federated data governance teams who may get involved in one or more areas of these teams. Similarly, underpinning all of this, the data and analytics platform teams who have accountability for the tech products that deliver not just the data products and the platform, but the data management products as well. We wouldn't normally have domain teams, especially in in our side, setting up tools and infrastructure. That is the accountability of the tech teams, and we look to have common platforms for all.
So it gets even more complex. Moving the slide on, you might have many different models coming together. So operational teams on the left, multiple versions of potentially LIMS or MES from different sites, flowing through potentially a data engineering team who move through many sources, putting together aggregated data products at source, potentially, you know, accountable for then the consumer data products. Or you may have business domain teams who are accountable for the entire data flow from source to user. And we generally see a mix of models. You may have data engineering teams who are SAFE Agile, they make data products for anybody, um, in terms of the operational systems. They own code, but not the domain data sets. And you may have data analysts in the domain teams alongside data engineers.
So what's some common threads here? It's complicated. There's no one size fits all. However, just to wrap on the final slide, there are some common elements around enablers of some of these product teams. Really crisp communication channels between teams with firm boundaries, regardless of your operating model. Clear roles and responsibilities, even if you're flexing, uh, between different operating models. We mentioned at the start, this is very high data literacy work, and a strong data culture helps us get that piece of work done. Agile is helpful, especially if you're looking at multiple teams who need to come together to deliver value. And underpinning that, of course, these bottom three: transparent partnerships, change management, and psychological safety for teams to challenge where, um, work may may need to be adjusted.
So thank you for the opportunity today, David. Back to you. And you've actually been trying to establish some of these teams, haven't you, in your world? So this is all from your experience. Fantastic. Please do add some comments. We have one, um, actually, you could tackle it, Kate. It's hard to find aggregate data product SMEs. How do you solve that? It is a particular challenge. I think it came up in the, um, the last poll we did around aggregate data products. Um, often join up multiple sources. And to understand how the data joins, you often need to be very data literate, obviously understand data models, the code to wrangle the data, but then also understand why the data matters. So what is important to a scientist to have in a particular data set? So it's that real kind of, almost a unicorn, a domain. What we're generally seeing is these people come from the scientific domain and pick up data skills because they understand the science and how to develop an asset first. And then they come together with the data teams and upskill around code, wrangling, data platforms to deliver the data products. Thank you, Kate.
So, um, I have just, oh, did you want to join in, Bob? Go ahead. Yeah, I just gonna add a comment to that. Um, it's jumping a little bit ahead to the next topic about lifecycle, but when you put this together in a lifecycle of development, and as was said earlier, if you, you don't do quote data products for everything, you do data products for things that have a very strong business need. So when you do that and you start going through a lifecycle that will hit on here in a second, that's where you start pulling the teams together. So this idea of making sure you do recognize, oh, if I'm trying to get to a particular point, I'll give an example here in a minute of something on tech transfer. You know, the idea is, okay, who are the key people involved in tech transfer? Let's get that team pulled together. And then let's go through, you know, a definition of what's required in order for a tech transfer to work successfully. And then go through a scale up. So you're never going to find the unicorn, but you will across the different groups find people that understand what's needed and the technologies that are that are available, where the where the data is coming from to start bringing it together. And that's how you put it together from what we've seen, anyway. Um, thank you so much, Kate. And actually, everybody else who I forgot to thank, and everybody who's contributed on the chat. I want to move this on and just, uh, return to that point that we made at the very beginning, that sometimes things break if you don't take care of them. There's something to be said about managing, uh, data products through their lifecycle. But what is that lifecycle? As we've been, um, working on this together, we've realized that we all had different drawings of our lifecycles. And, um, I want, um, for, uh, Bob, just to take us through a little bit of this now, if you could.
Yes, my name is Bob Lenich. I'm the business director for Emerson's, uh, Process Knowledge Management application. And I'm working with a number of customers to do digital tech transfers across, uh, the development lifecycle for products. And so obviously, when you're in that kind of a space, the data around what's being captured in individual phases of development, how that's packaged up, and then moved on to the next phases are important parts of the overall process. So the key thought here, uh, and like I'm trying to to capture down at the bottom, is, you know, historically, everyone understands products, and everyone has historically understood the paperwork. But as you start moving from a paper-based approach to things, which is the way the industry still is, but we're we're really trying to change to a data-driven approach where you get rid of paper, batch records, and you want to get away from anything that's paper-based, and now you're working just with data and data in applications, the lifecycle of both the application and the data itself become a very important part of what you're managing because you have to think of the data in the same way you thought of paper, but it's data as opposed to paper. So you need to be, you know, kind of changing your mindset around how to approach things.
So what we recommend doing, and and what we're starting to see with a number of groups on on the next slide, is to start really looking at the data and the data applications that you're, um, trying to drive here as you convert from a paper-based process to a digital approach to things, is to look at a lifecycle of the data products that you're involved in. So again, like we said earlier, you don't do this for everything. I mean, otherwise, this would be way too much overhead and and nobody would ever really do it. But for key things, and like I said, I'm in the the digital tech transfer space. Digital tech transfer is clearly a very important business opportunity out there. So looking at that type of an application across the lifecycle and the data associated with that application as a data product across a lifecycle is a really important thought. And so these these three main bullets, you know, experiment, you know, basically do kind of the proof of concept, industrialize it, or, you know, kind of get a program going that says, okay, great, we proved that it worked. Now let's make sure that it works on a on a large scale across the application and the enterprise that we're interested in doing it in. And then, uh, we use the word maintain. I also like to append the word maintain with proliferate because hopefully, if you're doing things well, you know, you're looking to expand it, not just do the first one and then maintain it. But you do have to maintain them, you know, throughout, you know, the course of all these things. So there's a little more detail behind that on the next slides.
And so what the the thought here is, you want to start, and again, when you think about managing data products, and we talked about the metadata, experimenting as a part of this initial proof to say, great, what kind of information am I going to need? What are the key metadata elements that I need to attach to that data and then manage as you're going through that? And then a real key thought is the, I'll say, the the traceability of that. Um, we we brought it up a little bit earlier, but the idea of once you start having data and you You've got a source of data, the traceability from where that source is coming from to where it's used all the way through as going through transitions is a really key part of the experimentation phase that we've seen at least in the the the tech transfer space. So the people can say, yes, I can show a chain of identity for the data, you know, going from where it's sourced to where it is being used. So that as you're moving from, you know, a clinical phase to a commercial phase, you can show that yes, everything actually worked. And when you get into your validation process, you've got that full chain. So going through and doing that type of experimentation to say, great, in order to do that, we need to make sure we've got tokens or some something attached to the metadata so that you can follow that kind of flow is a really good thing. And that's that kind of happens in this first P, this first section when you're doing this experimenting and as a part of initiation.
Once you've kind of proven that things work, then you want to get into the program and to really industrialize it. So the idea is, and again, I'm referring back to the kind of things we're working with customers on, taking a product as an example and saying, great, we're going to we're going to start working on the digital transfer of a product from site one to site two and going through all the the work around that. So we've proven that we can do it. Now let's take a product, let's go through the product, get all the data associated with that product, you know, captured together in one place, and then go through the the activities of moving it from site one to site two, doing the gap analysis, doing all the change management that's needed, any all of the scale-ups, things like that. And once you again, once you've proven that now that the concept worked, now let's do it for a product to show that the product works, we can industrialize things, and we can validate that overall piece. Great. You've proven that you go live, you have a release, and then hopefully you're able to then say, okay, we've done it for one product, you know, one or two sites. Now let's start extending that across other products, other sites. And that's the the maintain proliferate where you're taking what you've already done and you're going through that exercise, that that exercise of, um, expanding it so the people are using it.
And then the the last part, and again, when you think about things in a paper standpoint, and I have to admit, I haven't hit this one yet, but at some point in the future, you will get to the point where you're going to retire that information. So now there you do have to look at all of the record retention, you know, requirements that are out there because data has the same type of record retention that that paper does. So you will need to make sure you're managing that. But at a minimum, you want to be able to then start, um, you know, exporting it out of a system and at least putting it in some kind of a backup or an archive to where it's still available, but it's not necessarily being used because you're on version 10 of something and you really you need to keep version one because you need the history, but you might not actually want to use version one. But the key thought here is that as you're looking across, when you go to a digital process for things, you are now starting to look at how to use a lifecycle process for the data itself in addition to the product you might have, and keep those things kind of in parallel so that as in our case, or in my case, as we're doing, you know, quote product tech transfers as a part of product development, the data associated with those products are also being managed and migrated through an overall product lifecycle. Thanks.
Fantastic. And thank you, Bob, for that really interesting example, um, of using data products for tech transfer. And what I want to also point out is that there's another webinar coming up, not from this team, but from a parallel team, thinking particularly about the challenge of data from other organizations. So it is definitely relevant. Um, and this team has been thinking about how do we work with partners to in a playbook to establish digital connectivity with another organization. And I think you'll find that useful if you've enjoyed today's seminar. Again, it's, I think it's seven speakers that time, so another epic BioFor webinar. But, you know, it's the industry coming together, so we need to give everybody their voice. I've put a link to that in the chat if you want to register. Um, we meet fortnightly in this team, Data Enablement for AI. We share case studies, we're scoping some new topics, we deliver some outputs. We have an annual face-to-face. The next one is in September 2025. And if you're members of, uh, your company is a member of BioFor, get in touch if this is relevant to you. Get in touch with me, actually. If you're not members, you can get in touch as well, and we'll work out what needs to happen for you to participate if that's what you really want to do.