Transcription
Um, yeah, so thank you. Uh, I know I introduced myself earlier, but just, uh, just as a reminder. So I focus on, um, a set of our AWS Services, uh, that are, are, are about, um, processing data and making use of data, uh, and governing data, uh, in, you know, kind of all different types of applications. Uh, one of the things we're going to focus on today, though, is really how generative AI is transforming, uh, the business and, and how data is such a big part of, of, of that journey. And what do we need to be thinking about when, when we're using data for, for generative AI applications? There we go.
Okay, so, uh, the one, the, the kind of first thing I wanted to start with is, uh, what we talked to customers about a lot, which is, um, uh, data is the center of, of the universe, uh, today, uh, in, in every modern business. And, and it is a shift in thinking about what is, where is the most valuable assets that, that a customer has, uh, a business owner has in their business. A lot of times we think about customers as our most valuable asset or our people, and those are absolutely true. But data, the data that we have, uh, available in, in modern business, is really one of the key drivers for innovation and thinking about how you make use of that data and how you treat data as a, as an enabler and a, something that can empower innovation, is, is really a very important part of, kind of thinking about how do you drive, uh, your business in a, in a different way.
And so, uh, the, the, the changes that have come about are really driven by, uh, these new innovations in Gen, Gen. And so I showed the, uh, earlier, uh, I showed the, the, uh, kind of transformation of, of Gen, uh, over the number of years. And really, Gen has reached a tipping point, uh, that helps, uh, drive forward new innovations. And it's really based on a combination of things. It's based on a combination of new, uh, compute capabilities. It's based on a combination of the ability to process more and more data. And it's empowering new, uh, experiences for customers such that are different than what we had before. So where before you might have had, uh, somebody who was, uh, like a, somebody who was working in a support environment, uh, very knowledgeable, knew where to go get all the data, uh, to help a customer, that, uh, that person now can use a knowledge assistant. And that trained model knows all the different, uh, data about a customer to present to that, uh, support, uh, person on a support call, uh, before even starting the conversation. And so that absolutely transforms the experience of, of customers when you have more and more data made available, but made available in a way that's very, very personalized and pertinent to an individual's customer needs. And so that's the expectation now. And that's what we all have to be thinking about as, as, um, leaders and, and business owners, is our customers have a bigger, more kind of, uh, higher-level expectation about how their data is used and, and made available, uh, to those, uh, to, to, you know, as part of those experiences. And so they'll expect personalized experiences. They'll expect that, uh, when you make a call to a support center, that the, uh, support person on the other line already knows all about you. And so how do we make that happen?
And so, uh, just to kind of, uh, kind of highlight some of the other, uh, things we talked about. Obviously, earlier today, you heard about Martin, uh, from Martin at Trellix, talking about how cybersecurity is transformed with Gen. Um, you know, we'll hear a lot more later from, from Mark, uh, around enhancing customer experience, uh, but it, it's transformative across many different, uh, parts of the industry. And, you know, like, uh, you know, I spent a long part of my career as a software developer, uh, writing Java applications and Python applications. And looking at what I did in my career in the early days, and what now they're teaching and talking to computer science students about, uh, is, is light years different now. It's all about how do you use, you know, trained coding assistants to help you write your code, instead of knowing all you need to know about, like, the, all the object P, um, object-oriented programming and, and patterns for, uh, for building out applications. It's really how well you make use of, of coding assistance. That's going to be in all industries, in all the different walks of life. It's going to be, how are you as an individual empowered by AI to do your job better?
And so, um, one, one more thing on the, the, kind of growth of generative AI is, I spent, uh, I've been in Amazon 16 years. Um, I spent seven years, uh, in S3. And really, the key thing here in driving this, this growth in, in Gen has been the, the massive proliferation of data. And so when I was in S3 for seven years, it started out where we were mostly focused on archive and backup, and then it shifted very heavily towards data usage. And that, that trend continued. And really, then combined with the power of, of, you know, modern scalable computing, it's really changed the way in which these models operate.
Okay, so let's look at a little bit about how that transformation and change has, has happened across, uh, the, the, kind of AI and machine learning space. So we started out originally, or we still have models that are trained on a specific set of inputs and, uh, are targeted at a specific task. Uh, and that might be something like a translation, uh, process, or, or, uh, even, you know, some sort of very, very simple, uh, kind of text generation. And the idea here is a simple amount of data inputs and a simple output. Then we progressed to deep learning, which allows us to have a lot more variety of different types of tasks, uh, that the, the training model can, uh, can accomplish. But it's really still focused on a specific outcome. And it might be something like computer vision, recognizing objects. That's another example of where deep learning really, kind of, helped us innovate. And so then we come to foundational models and Gen. And the big difference there is that you have a wide variety of very general-purpose data. You don't really constrain the type of data that goes into these foundational models. And so over that wide variety of data, the model is developing a way to reason about that data. And that reasoning then allows for very, very complex outputs.
And so all of this comes together as an AI strategy. So you shouldn't think about, how is my data powering my Gen applications? It really should be, how, uh, data is powering all of my AI applications. And so you, you will continue to use, uh, machine learning models. You'll continue to use deep learning for specific tasks. And you'll also more and more be using these foundational models to really expand the scope of what, what you can impact, uh, with AI. And so the, and so here, data, building a data foundation that is core to enabling something like, uh, the outputs that you get from, uh, use of Gen. And so really, it's about connecting to a wide variety of data sources, providing all of these different types of, of, uh, unlabeled data, feeding into your foundational model. That foundational model is figuring out how to make use of that data. And it then adapts and works towards developing, uh, tasks and, and capabilities to do things like text summarization, info extraction, question and answer with a human. All of these things come out of using a, a massive amount of, of data. And so really, the key here is, you need to connect to your data. And so, and, and as I called out earlier, it's really not just about connecting to just a subset of data, it's connecting to as much of your data as possible. And so you have to think about data and, and kind of blocking and, and unblocking data access is the, is, kind of the core part of the strategy here, and making the best use of Gen.
Okay, so there's three themes that I want, want to highlight, uh, and then hand it over to, to Mark to talk in more detail about specifically how, how these things are applied at Workhuman. So, first thing is, uh, availability of data. And the way I think about that is, is you really need to make as much of your data available as possible. And so what does that mean? That means connecting to data when it's going to be used, uh, as fast as possible. Uh, there's going to be different use cases for your data, uh, in AI. It might be used in training models. It might be used in fine-tuning models. And it also might be used when the model is actually processing data, uh, through something like retrieval augmented generation. And so really, the thing there is, make data available. And so in, uh, AWS, we have a, a wide variety of data sources that you can pull, uh, data from. That could be data lakes and data warehouses, where a lot of, uh, folks throw the vast majority of your data. But there's also data across all different types of, uh, databases, from relational to document, also analytical sources, uh, that could be query engines like Amazon Athena. Um, it could also be, uh, you know, even, uh, you know, log analytics tools, all different types of, of analytics tools can be a source of data for your, for your Gen application, as well as then, uh, the most recent innovations with vector databases and, and feature stores, which help you, uh, at the actual model evaluation, uh, time to be able to more efficiently, uh, map data to, to the outcome that you want. And so really, it's using your data to customize the behavior of the model. And really, that means that you're going to be, as I mentioned earlier, using your data in different parts of, uh, the model's lifecycle. And so you have the ability to, uh, kind of change the C and customize the behavior of the model. But these are going to be different types of optimizations. With fine-tuning, you're really changing the model behavior for some, uh, from, for some sort of persistent behavioral change you want in the model. Uh, for, for something like retrieval augmented generation, that's going to be able to, to take the latest and up-to-date information and inject it into the model's reasoning, so that it can think about something like, this customer made a purchase last week. It needs to know information about that purchase last week. And so using something like RAG to be able to inject that into the model evaluation is going to be essential. And then there's also going to be a need to continuously refine the model and train the model. And so in certain cases, you may be, you may be doing continuous pre-training on the model. And again, those are for situations where you want some sort of long-lasting behavior to be, to be relevant and impacting the model. But in all of these cases, it's really about getting as much of your data as possible into the model to make it more effective.
Okay, so what does this look, you know, see some of these patterns, uh, in, in later sessions, and so we'll go deep dive into RAG and all these other, uh, use cases. And so you'll see these different patterns, uh, in more detail. But this is kind of a high-level view of what's happening behind the scenes when you think about the Gen application as the outcome that you want for your customers. And so really, it's going to be a combination of a bunch of different data sources. And then you're going to have different ways that that data is going to flow through the application. You have situations where you need the most up-to-date data. That's going to be, uh, the, the application itself accessing data from something like a NoSQL store of conversational, uh, history. That might also come from a stream of latest and up-to-date data. So you might have a streaming data source that is also being accessed by the model or injecting into some sort of up-to-date, um, data store. You're also going to have situations where you have batch ingestion, uh, that's going to go into either data lake or data warehouse. Might have some stream processing or ETL that's involved with it. And that's going to then be also be used by the model for different purposes. And that might be used in certain cases, uh, when the model is evaluating during, during something like RAG, but it could also be when you want to retrain the model. So again, there's these different proc, different kind of, uh, patterns that you want to apply and using your data. But it all starts with that data source, uh, set of, of, of capabilities. You really need to make sure that in all of these cases, you're connecting to as many data sources as possible.
Okay, so then in AWS, we have a wide variety of services to help. Um, you know, I called this out earlier, our focus is on that comprehensive set of services. Uh, and so for, you know, a wide variety of, uh, analytics stores, analytics processing engines, we also have, um, multiple ways to work with your models. Uh, that could be through SageMaker, uh, when you're doing say, training of your models, building your own, uh, models. It could be with Bedrock, where you're utilizing foundational models that, that already out there. Most of our customers really are building on top of existing models. And so Bedrock helps you manage that, that and go through the process of, of, kind of augmenting the model with your data as well. And then also the end goal, what's the end goal of, of using, uh, Gen is the experience that your customers are going to have. And so we've plugged Gen into things like QuickSight. We've also got things like Q to help you build an experience that can take your customer's data and then make it available to your customer in an interactive, uh, experience. And so all of this needs to be, has a baseline on, uh, on governance. And I'll talk about that in just a little bit.
Okay, so what's the next theme that we need to be thinking about? And that's quality of your data. And so the, the key thing here is that when you have quality data in your model, it's going to protect you against things like bias and drift and, and challenges that you might have in, how the model behaves. So you need to be thinking about quality throughout the entire process. And so if you look at, kind of, what that qual, that flow looks like and, and where do you apply quality controls, uh, as data is flowing through your system, you start with your variety of data sources. And then you will be, uh, pulling that data in, uh, either through something like, uh, initial process, uh, run by, uh, Amazon AWS Glue, or could be something like, uh, uh, Data Wrangler and, uh, SageMaker, where you're trying to prep the data. Regardless of where the data comes from, you want to have that mechanism to be able to, to enforce quality or guarantee quality. And so using tools that are built into the AWS services can really help there. There's also situations where you need to think about how do you get humans in the loop on data quality. And again, that could be through something like SageMaker Data Wrangler to be able to, to take a look at the data, reason about the data, and make sure that you have quality data that then flows downstream into the use inside the model. And so, and in thinking about quality, it's really important again to connect to, uh, you want to be able to have a solution that can connect to a wide variety of data sources, build consistent pipelines, reuse those pipelines over and over again, make sure those pipelines are checked in, source controlled, so that you have quality built in, and also consistency built in, uh, when you're thinking about, uh, how you, kind of, keep that quality measure up to date. You also want to be able to have access and understand all of the data you have. And so you want to be able to include governance, discovering and cataloging into that solution as well, through, through, uh, tools like Glue Data Catalog or DataZone. Um, and then also, it needs to scale. And so you want to focus on serverless capabilities because you don't want these, these data quality checks or other, uh, mechanisms in your data pipeline to be a blocker for getting the latest and up-to-date data. And so you want to use those, the scalable, uh, solutions that allow you to do that at a cost-effective way. And one of the ways, I mentioned this earlier, one of the ways we're helping customers is through making ETL simpler. And part of making ETL simpler is really doing the least amount of, of ETL yourself as possible. And that's really where we've, uh, been innovating, uh, significantly in what we call Zero ETL Integrations. And the idea there is, ETL needs to happen, data movement needs to happen, but it's important that that data movement happens in the simplest way possible. And so being able to do a one-click, uh, uh, movement of data from Aurora to Redshift, and then we're also focused on many other different data sources. Um, you can see here, it's not just relational databases, it's all different types of data sources, including things like DynamoDB, and other, uh, other, uh, sources that we're going to be iterating on to provide these capabilities. And Zero ETL is also about connecting to data. And so using connectors, federated queries, these are all strategies that you should be thinking about and making use of all of your data across all of the different, uh, sources.
Okay, so the last thing I want to talk about is the protection of your data. And the idea there is, protection isn't meant to stifle the use of data. It's not meant to, to lock down the data. It's really meant to make sure that at each stage, you have governance and protection over how the data is used, because you still want the data to be used. You just want to make sure it's secure, it maintains security, and it maintains its protection and, and, uh, throughout the, the entire data pipeline.
Okay, so in AWS, we have a wide variety of different tools to help you with, uh, in, in, kind of governing your data throughout the, throughout the pipeline. We have Glue, uh, as I mentioned before, to ensure data quality and cataloging. Really, what that means also is that you, you have consistency in the tools that are processing the data and applying those permissions and, and governance controls throughout the pipeline. You have Lake Formation and DataZone. Um, we've, you know, recently seen a significant amount of, of growth in customers, uh, you know, building out data meshes and, and trying to figure out how to collaborate and share data across the enterprise. And that's really where DataZone and Lake Formation help is that you need to know what data you have, but you also need to have a process by which you, you collaborate on data. And so DataZone can help with that as well. Then, and on the, you know, if you think about the other side of, of data usage, uh, beyond just the data sources, you want to make sure that you have governance and how the data is used in the model and in the processing of, of the data, uh, in the Gen application. And so really, SageMaker governance, uh, is another set of tools to allow you to do that, as well as governance provided by Bedrock. And so really, think about it all the way from the data sources, the data pipeline, all the way to the application itself, that you want to be able to use all the tools and capabilities to apply governance. And then finally, part of governance is also understanding where your data is being used, auditing your data. And so you really need to be able to, uh, to, to have traces, logs, all of these different data sources of how your data is flowing through the system, how your data is being used by, uh, your AI application. So making sure that you're focused on enabling these events, these CloudWatch logs, um, all of the different traces that are available from the services that are making use of your data, such that you can then, um, utilize the tools that we have like CloudTrail and, um, CloudWatch to be able to do alerts, to be able to, to, kind of investigate, um, data usage. And so really, it's both a combination of the data sources for, for logs, as well as then the, the solutions to be able to process and reason about your, your, uh, the usage of, of your data in an AI application.
And then finally, there's an overall, kind of encompassing area of being responsible in, in driving your AI applications. And so there's different ways to frame, uh, responsible AI, but it's really about a wide variety of different services applied at different stages. And so you can, kind of, think about all of this coming together in, be your, your AI being responsible with the data that it's, that it's using, being responsible about how it interacts with users. And so that's going to be things like, um, you know, the, the, and the ability to understand how data is being used. Uh, it's also the guardrails for, for Bedrock that I mentioned earlier, that can, that can constrain what comes into the prompt or what comes out of the response from the model. Uh, it's also the auditing, uh, services that I mentioned before. So you really need to think about all of these things as part of, of, kind of developing your responsible AI strategy. And so just, just to, kind of, uh, close it out, it's really about thinking about all of these different themes as part of your, your data strategy. And then also thinking about what do I have available so that these things can be built in, uh, from, from the start. And so think about this as part of, uh, you know, using our, our all set of AWS services that I mentioned earlier, our comprehensive set of storage services, query services, where these things are built in. Use the pipeline capabilities we have through Glue, uh, to ensure data quality. And then also use the governance capabilities at all levels, uh, to ensure protection of your data.
Great. So, uh, I'd like to hand it over to, to, to Mark, uh, as, as next, to talk more in detail about how Workhuman has applied all these, uh, all these, uh, different aspects into their, uh, data foundation and their data journey. Thanks very much, Rick. Uh, hopefully everyone can hear me. I'm Mark. Um, so, uh, we're Workhuman. We're an Irish company. Um, we're the global leaders in social recognition. And we were started by Eric in 1999, so we've been around a long time. And this is not going to be a sales talk. We're in Kilkenny, which is Ireland. And it's fantastic to have a conference like this in Ireland. We're an Irish company. We're very, very proud of that. And it's one of the reasons why, um, I love working at Workhuman. And the second reason is that everything we do and everything we're about is about making people's work lives better. And we've a lot of big, big customers. Um, and what we, we do three things basically. We help employees thank each other. We help employees talk to each other. And we help employees celebrate with each other as well. Um, so it's a really important thing to be thanked, actually, for work you do. And we've, we have lots of insights that show how that really improves people's work lives. Um, and we thank, we help 7 million or more than 7 million people around the world thank each other in 180 countries. What does that mean for us in this room? It means enormous amounts of data, gargantuan amounts of data, in fact. And so we have our Workhuman Live conference next week in Austin, where it's a pity that the weeks landed that they did, because we've a lot of exciting AI product launches next week, but we can't talk about them. We can talk about them next week. Um, but I'm going to talk about lots of things that, that we have done. And Brené Brown is a regular speaker at Workhuman Live. And one of the things she talks about is that, um, maybe stories are just data with a soul. And if that's the case, and we want to tell stories, we need more data. And we find that our customers need more data, our teammates need more data, our product needs more data. And actually, when I started at Workhuman 4 and a half or a little bit more than that years ago, we were fully on-premise. We didn't have elastic capacity. We didn't have a cloud data strategy either. So, so we knew we needed to transform, and pretty radically at that. And today, our working with cloud is built upon data. Data is an asset. It's part of our moat. And it's a kind of a fundamental underpinning of everything we do. And don't mind the room next door with all the fancy AI stuff. If you don't get the data right, that's all built on sand, right? You got to get the data right. Everything else comes after that, right? So it's really, really important.
Okay, so we embarked on, uh, about two years ago, a three-year journey. So we're not there yet, but we've made, kind of, remarkable progress on that. And Kamal is going to talk a little bit about some of the specific architectural things we did. But one of the things, one of our main strategies was, we had a huge amount of applications, a huge amount of technology with huge amounts of data. Um, and when data is lost, it's lost forever. So we wanted to take all of that data and preserve that data because we knew that we needed the data and we couldn't predict how we would need to use it in the future. And I think a lot of companies who've gotten into generative AI in the last 12 months have seen that, you know, it's really, really important. You cannot predict how data is needed. So how do we make such a radical transformation? Um, well, I'm going to talk about eight primary lessons that we learned. And it, kind of, doesn't matter where you sit in an organization, whether you're an intern, whether you're an engineer, what are you, the CEO reporting to the board? There's a lot of things I think that people can, can maybe learn about our journey and how we were able to make this transformation, how we're able to begin.
So the first one is, tell a story. When I started in my career, um, I used a lot of facts and I used a lot of data to make points to people, and no one ever remembered anything I said. And maybe I just wasn't very persuasive. But one of the things that I realized was that it's stories that really inspire people to to make changes and and to move with things. Okay? And so to get people to come on a journey, it's really important to tell a story. And the story has to be believable, and it has to help maybe solve a problem or show a brighter future. So in our case, when I joined, I went around the business and I talked to as many people as I could, and I tried to get an understanding of what their understanding was in terms of technology, in terms of data, in terms of how they used it. And then I started to tell people stories. And I told different stories for different people. I told different stories to people in technology, to people outside of technology, different stories to senior people, different stories to engineers, for example. You know, and I use the word "imagine" a lot. "Imagine what you can do, Sandy, if you had this data. Imagine the insights. Imagine the decisions you could make in your business if you had these type of insights." And that was really the beginning of helping people to imagine a world in which a data strategy could really affect their lives.
The second thing then is to take action, to do something, do anything. It, it almost doesn't matter. What's really important is you have to create some momentum. You have to create some movement. Um, and to do that, you're going to need to find some teammates to help you deliver, because nobody can do everything, uh, themselves. And there's always people in an organization who know that they could do more, who have more to give, who have, who are willing to come along on a journey, who are willing to take a risk with you. So you have to find those people. And then you have to ask for forgiveness, not for permission. So in our case, I found a number of people, and some of them are in, in the audience, I think, here today, um, who were willing to come on that journey. Um, so we created a little bit of a coalition, and we worked away in the background. And it's a very busy company. You know, we have massive Fortune 500 companies, and we have a lot of demands on our time. But people were willing to do that because they knew that they could do something themselves to really impact things, right? And actually, even the CEO, he, he was a little bit, um, he knew we were doing something. He wasn't quite sure what it was. He didn't stop us. We didn't ask for permission either. Um, and he became much happier, uh, with us after the next lesson, which is to demonstrate value.
So it's important at some stage to create something useful. I think that's probably obvious, right? If somebody has an issue, if they have a challenge that they're facing, um, try and make their life better. There's, there's an old saying that if you make other people's lives better, they will help make your life better as well. And so that's what we, we tried to do. And in our case, strangely enough, and we're here at the Amazon AWS conference, who the number one e-commerce store, uh, on the internet. Our e-commerce store is actually one of the 30 biggest e-commerce stores on the internet, and it's private. It's only available to our customers, so it's not publicly available. So it's a billion-dollar, uh, store, right? So it's, um, there's a lot of significance to that. And when we were on-prem, we were somewhat limited in terms of how we were able to understand every single thing that was going on with users. But one of the things that we were able to do, in fact, it was two things we were able to do. The first thing was, as we went on our journey in the cloud, we were all of a sudden able to give our data scientists almost unlimited capacity to scale their their training and to scale, so that they could generate much better insights. And we were also able to join disparate sets of data with unexpected insights, unexpected results that we couldn't have seen before because we were limited by our capacity. Now, we were only limited by the checks we signed to our friends on Amazon. Um, but, but that was a big, that was a big step forward for us because all of a sudden we were able to show, demonstrate real value to the company, to make a real difference. And so that really helped cement this journey as something that was important to the company.
So the next thing is to expand person by person. So people are actually always up for an adventure, I find, for the most part. And everyone wants to improve their own lives, and and most people want to improve the lives of their customers, I hope. Um, so it's important to find a tribe of like-minded people and kind of root them out. And that's what you have to do, kind of, almost person by person. That's what we did in, in the organization. We had a lot of people who were interested in data, we had a lot of people even outside of technology who were interested in data and how to use it, but they didn't really know how to get started. So what we were able to do is bring those people together, create that sense of community. And once you start that wheel spinning, it will spin on its own. People will then be the carriers of that message through the organization. Um, and everyone expects technologists to talk about technology, but when business people start to talk about technology and how that's affecting them, that's really when the whole business starts to sit up and take notice.
The next one is get executive support and go as high as possible with this. So once the groundswell starts to begin to grow, you have to make it official at some stage. And the reason you have to make it official is because you're going to need funding, and you're going to need funding to pay for infrastructure and to pay for support. And without that, you're not going to be able to support users. And all of the goodwill you build up initially will just dissipate because you will begin to disappoint people more and more and more. So that's really, really important point. If you bootstrap too long, that's a recipe for a bad result. So at some stage, it's really important to get that executive support. Try to go forward instead of the back.
Lesson six, bring in experts to help. Um, so we, we brought in experts. Um, we had great support from, from AWS, of course, and they came in and helped with a lot of consultancy and advice. They did some labs with us, and they were always great to sanity check and to offer support and help. Um, I often say as well that when you're embarking on a kind of a new technical journey, it's important, a beautiful way to to set up a team is to have internal experts in the business and internal experts in the technology that's inside the company, and to bring in external experts who are experts. I'm using "expert" a lot, who are who are, uh, very experienced in the new technology you're trying to work with. And in our case, we brought in two companies, a company called Spark HQ, Irish company, and a company called SoftServe. And they worked together with us, and we were able to learn from them, and we were able to then pick up a lot of the techniques and a lot of the lessons that we needed to be able to really bring forward what we, what we needed to do.
Lesson seven, do good work and tell people you do good work. So we had a Data Day last May, uh, in Workhuman, and it was a full day dedicated to data. And it was set up by, by Mick Rice, who's the director of data, uh, platform at Workhuman. And he didn't ask for permission either. He just did it, with like David McAteer for, for this conference here, as we heard earlier on. Um, and he ended up with so many speakers that we had to run two streams in parallel all day. Now, from where we say, 18 months before that, where we didn't really have a data strategy, to have a Data Day for the whole company to turn up in and to participate and to learn more and more about data, that's an enormous step forward. It was a big marketing tool, but it also showed the level of interest that had been built up by some of the excitement that we'd that we'd shown and some of the work we'd done. Um, data is actually now part of our strategy. I said it's a differentiator. It's, it's actually reported to the board now. So it's, it's part of our company strategy board. So it's absolutely vital, vitally, uh, is vitally important, and it's part of our competitive differentiator.
And lesson eight, there's always more to do. There's always more to do. Um, technology changes, the business changes. And just, you know, generative AI is going to change business. Now, business is going to adapt to that. Then some more technology will come along and change, and business will need to adapt to that again. Um, you know, and I mentioned early on with AI, you know, getting the data strategy is, is really, really, really important. Um, the other thing that's really important is enablement, and safer data, particularly is data literacy. So to push that enablement through the organization as well. If you create a great asset like a really strong data asset, it's important that people are literate, and they need to be able to use it.
So this is called Workhuman IQ. So, um, and this is, I, I think a really great example of how we use data to give companies a full health check of their organization. And at first start, you might look at, you might think, well, that's just a UI on top of some data. But, but that's not what it is. This is value. These are unbelievable insights. No company in the world, I think, has access to data like this. And if companies really believe that people are the most important asset a company has, well, then this should be one of the most important things that executive leadership and leadership and people at every level in an organization should be looking at. How is my culture? You know, what's my? Are there? What's the linkages between between teams? You know, are these people talking to each other? Are people getting shut out? You know, how do I have schisms? Do I have divides in there? And this is, you know, the culmination, I think we have been in, um, we've been working in AI actually for 10 years. And, and one of the things Rick brought up earlier on, which is really, really interesting idea, is AI is not just generative AI. It's deep learning. It's machine learning. It's all of the tools that are available through the AI stack. And so we've been doing this for a decade. So we have a lot of expertise in, in AI. We've built a, a lot of things on this. And we can use these insights to really help drive change and really change people's work lives for the better.
This is an, uh, this is called Inclusion Advisor. And this is something that companies are always blown away by because we have since 19, so 25 years expertise at looking at how people communicate to each other, both in private and in public, around recognition, around job performance, around these type of things. We have built, uh, a huge amount of, of models and huge amount of, uh, intellectual property around how people communicate with each other. And what this is, is if you say you want to thank somebody like Johanna here for, uh, some work that, that, that was done, um, and you use language like, you know, "coming in on weekends or on holidays," this looks at microbias. Now, at first glance, it looks like, okay, it just corrects the, you know, so that you're using language as more appropriate. But actually, what this is doing, it's training the people to work in concert with the values of the company. So it, it trains people over time as to what's acceptable and what's not acceptable, right? Um, and there's another example here, which, you know, my secret weapon, right? So, you know, like a personal favor. So this is really, really important. But these are models that we've built up over years. And companies are, you know, massively, massively, uh, interested in this. You often hear about biased data where, um, AI can be a problem. But here's a great example of where AI powered by data can be really be a force for good.
Okay, so, so there we go. That's basically, tell a story. Um, make sure you're demonstrating some value. Um, put something up there that, that people can, uh, can use. Get some budget, get some money. Always money is important. Um, do good work, tell people you're doing good work. And there's always more to do. Um, and I'll probably put a write-up, uh, on my blog about this next week with some, uh, suit bad jokes, is what people usually tell me about that. And, uh, and next, I'm going to hand over to my, my very talented colleague, Kamal S. Kumar. And Kamal runs, uh, the data architecture group in Workhuman, and he's going to talk a little bit more about some of the architecture specific things we've done. Thank you.
[Applause]
Thank you, Mark, for talking about the data strategy. How did we start, uh, in Workhuman? I'm going to talk a little bit about how did we transform a data strategy into a technical strategy and an architecture. Today, I'm, I'm going to go through a high-level overview of our data platform and architecture that we put in place. I will also go through how did this architecture currently support our current needs as well as for our future needs. I will touch a couple of solutions that we recently delivered to our customers, which I will cover in the following slide. One is Data API, and the other one is self-service reporting.
So before I start talking about the layers within the architecture, I suppose it's important to understand when we started this journey, we didn't have anything. Uh, we started from ground up. So everything was built in AWS. That was one of the reasons. But it's important as a group for us to understand foundational components that we need to identify at the beginning stages. We also wanted to identify what layers should we build in our platform, how the data should be flowing from source to customer. These questions when we asked helped us to create and formulate what layers we need in our platform. So we ended up with four distinct layers in our platform. They are extraction layer, curation layer, modeling layer, and consumption layer. I will go through each one in detail and give you a high-level overview what it does and how it leverages purpose-built AWS services.
Extraction layer. This primary responsibility of the extraction layer is to extract the data from various sources, both internal and external. As you can see from the diagram, we deal with many data sources, and some, and they're broadly categorized into three areas: one, product-related data; two, back-office data; and three, operational data. So for the, for, for these three data types, we have to leverage AWS services such as DMS to bring relational data store, EMR for bringing API-based solutions like third-party systems, and we used a combination of Glue. We, we also built a generic ingestion framework which complements the ingestion service by interacting with AWS ingestion service through configuration. The benefit of doing that is it helps us to bring the data at speed, which is critical, as Mark mentioned earlier, turning the volume, turning bringing the data and delivering the business value. So it's important for us to, to, to design a framework that allows us to bring the data at speed.
I'm going to talk a little bit about the curation layer, the next part. This layer acts as a data refinery. This is where the data gets cataloged, the data gets validated, it gets cleansed. It then consolidates the data into a prescribed format that is ready for analytics and reporting. So we transform all the ingested data into a consolidated format that is par, that we consider as a standard format.
And moving on, modeling layer. This is the, this is where the actual magic happens. The data transforms from source-confirmed data into business-confirmed data. So we use a combination of DVT, leveraging Redshift and Spark to model the transformed logic. What this helps us to do is, is basically to deliver the business value from the source data to insights that just Mark talked about earlier, showing the Workhuman IQ application, for example, that requires quite a lot of detailed transformation that we needed to do. So how do we do it at scale? So we need to establish some standards within our modeling layer. Every analyst and engineerings engineers should speak the same language. They can express that in SQL, and they can run it depending upon the workload, either in Redshift or in Spark. We also have a structure. All source data are mapped into a base model, and then they are grouped into a domain model based on the product feature areas. Then they are transformed into a core model, which is then surfaced to the consumable layer.
The last layer of our platform is consumption layer. The consumption layer is the layer through which we deliver data to the end users. There's many elements in that. I'm going to touch a couple of them. One is the Data API, which allows our customers to consume the data dynamically. And then we have QuickSight, which allows the users to consume through a business intelligence tools. And we also have exciting work that is going on behind the scenes, and it's going to be announced next week, which is Gen, generative AI. We are building a generative AI Workhuman assistant that is also treated as a consumption tool because we have all the data accumulated over the years in our data platform, and we are leveraging that to drive the generative AI strategy.
So moving on, how did we realize the value of the data platform? We have invested for three years. It took us to build this robust architecture with our platform.
Data is consumed in three areas: internal users, product, and customers. I'm going to talk a little bit about two solutions that we have recently launched. One is Data API and a self-service reporting. Both of them are targeted towards our customers, and they deliver value and have a profound impact on our customers.
Let me dig a little bit deeper into a Data API solution. So, before I go into the solution and walk through at a high level, I want to give you guys a context on what was the challenge we were trying to solve. Customers have their own data platform and needed their data to be integrated into their analytics platform in an automated way. They wanted to interact with an API similar to a SQL interface. For example, they want to provide how the filtering should happen, what data they need, what file format they need to download. So, we took this as an opportunity to add a new consumption capability to our platform, and we called it as a Data API. This allowed us to move from our traditional file-based legacy applications that were delivering static files to SFTP servers. And this also created an opportunity from the business point of view to have a B2B solution delivering a bulk data API to deliver transform data models from our data models to our customers.
Let me go through in detail how the overall solution operates. The overall solution is an asynchronous operation. Data feed generation is done asynchronously. It has three distinct endpoints: one is create, next one is to poll, next one is to retrieve data sets. So, let me go a little bit on those functional how does the function operate for the three endpoints. Create one: it is responsible for creating a data set. Once the client application is authenticated, then the request is taken into our platform, and then a functional check is carried over to verify whether the caller's identity is valid to execute a data generation process. And then it needs to map to the underlying data model. What is the data set request and how does it need to be translated? The processing also translates the API request into equivalent unload SQL statements, which then delegated to Redshift or Athena API to unload the data back into S3.
The second function, sorry, I jumped on the slide. The second function of the API is to pull. Because it is an asynchronous job, the client application needs to pull to understand whether the job is completed or not. So, it's just a simple polling status and gives you whether the job is completed or not. And the third one is to retrieve the data set. The retrieve data set is through which the client application can download through S3 pre-signed URLs. So, at a high level, this gives us a generic way of delivering multiple data feeds to our platform. And the real benefit to our customer now is basically they can now get their data in real time, in an instant manner, and in a dynamic way.
The next solution I'm going to talk a little bit is recently launched a new self-service authoring and reporting capability to our customers. As like I want to mention about the challenge before we talk about the solution. So, our product offers capabilities to our customers through reporting capabilities as static reports, primarily through pre-defined reports. And the major drawback of a pre-defined report is the data is static, it's not dynamic. Customers can't customize it. They still see the static interface. If they need to make a change, they have to put in a request to us, and we have to develop a software and deliver it. This is whole cycle is taking longer to reflect. So, we looked at our data platform and we looked at how we are consuming data internally. We are consuming transformed data models through a QuickSight interface, and we are sharing that with internal business users. The internal business users have an ability to further customize and share it with a subset of their users. We looked at that and we looked at that as an opportunity to see how we can customize it and reuse the solution.
So, to reuse the solution, we thought about using QuickSight's capabilities. One of the things that we wanted to use was how do we do it at scale, how do we do it securely, and how do we do it at a rate that we can turn it on for 500-plus customers. We focused on four areas. Area one was security, and security was handled by QuickSight's namespace. We were able to isolate users and groups and provide a multi-tenancy solution using QuickSight's namespace feature. Customizable analysis: we used QuickSight templating API mechanism to create customer custom copies for each customer, and we also allowed embedded QuickSight featuring and authoring capability to deliver an in-product experience.
As a result of the overall solution, we have a seamless experience that we are delivering to our user through our product, a complete authoring experience and a viewer experience. And this solution is going to be turned on for all of our customers by end of May. Mark mentioned the volume of users that we have, 7 million users on the platform, and we are going to, out of the 7 million, there's going to be a million users who are eligible to work on reporting and authoring capability. So, the impact is going to be broad. So, the whole solution is going to be scaled up to a variety of customers.
So, I'm going to wrap it up with a few key principles that followed through our journey. We established layers in our data platform architecture, and that allowed us to clearly define the purpose and the responsibility of each layer. We applied standardization. We standardized data models by applying modeling practices, and everybody speaks the same language. We adopted a lakehouse design pattern and ELT approach that allows us to flex so that the future use needs of the data use cases that we can address and flexible. Our architecture was flexible to adapt any AWS ingestion services. As you can see, we have used many of the capabilities that AWS has provided. And with that, I will wrap it up. Thank you.
Really wanted to close it out and say thanks to Mark and Kamal for sharing their journey with us today. A lot of the stuff you saw here will you'll see in other presentations through the rest of the day. So, definitely take what you learned here and kind of figure out how to apply it, but also learn from all the other sessions that we have coming up today. Thanks.