📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Community Call: Can you make AI 100% deterministic? | March | 50th edition

Hasura1:02:31

Transcription

Topic today for episode 50, which is incredibly hard to believe, uh, we are going to be talking about making your AI 100% deterministic.

As usual, we have a pretty packed agenda. We're going to start off by talking about these predefined AI agents and if they're really the answer, uh, with how quickly things are evolving in the AI ecosystem. We're going to look at ways that we can actually build agents on the fly and have them be specific to whatever problem it is that we're trying to solve.

After that, we're going to take a deeper dive at looking at a couple of areas we've made some product improvements in for PromQL. One is going to be with Pervine, where we talk about chatting with our files. So if you're want to talk with your data, that's what we're going to be doing today. And then finally, uh, we're going to take a look at a new development, which is our data visualization in PromQL. So think about uh any kinds of artifacts you're generating from the queries that you run against your data; you can actually create visualizations and things like that of it, uh, on the fly as well.

So, uh, without further ado, let me go ahead and jump into our first talk, and I'll bring Anush on the stage with me, who is our product lead for PromQL.

Hey man, how are you?

Hey Rob, I'm good. How are you doing?

Well, doing well. Uh, I have to start off by saying that I I saw in Slack that uh somebody, I think you were at, maybe, uh I won't say the company's name, but they go, "Oh, is that the is that the PromQL guy? That's the neon green guy?" Because how many how many of those shirts do you have?

I have four of these shirts so that I can keep recycling; I don't have to wear the same one every day.

But do you label them Monday, that I have to wear this every day?

Gotcha, gotcha. I'm glad that's not in my contract. All right, uh, before we jump in, there are just a couple things that I do want to say because I forgot to kind of do the housekeeping in the beginning before I turn it over to you. First off, folks, uh, if you'd like to, over in the chat, go ahead and tell us where you're joining from. Uh, we're out here in California, which the weather is looking pretty decent outside right now; hopefully, we'll get better throughout the day.

Additionally, if you have questions during any of the sessions for the presenters today, over on the right-hand side, there's a Q&A section that you can press. If you press and ask your question there, we can actually answer the question on stage with kind of like a little Chiron going across the bottom. So please do ask any questions that you would like responded to in the Q&A section, and then just, you know, the chats for general conversation, telling Anush you like a shirt, anything like that. But uh, let me go ahead and unshare my screen, and then I'm going to turn it over to you.

Today's topic is, um, are predefined agents really the answer, and how can we build accurate, task-specific agents on the fly? Um, so but before we go about answering that specific question, let's first think about uh AI itself, right? Um, I asked you, right, "Do you trust AI for business-critical tasks?" And uh, this trust is established when AI feels reliable. How do you trust humans, right? You trust humans when you feel that they are reliable; they're accurate in what they're saying; they're transparent in how they work; they are, they can do the same task repeatedly with the same level of accuracy; you can work with them; you can collaborate with them; you can steer them and guide them to the right answer; and then you trust them with safety and security, right? So these five pillars of trust not just apply to humans but also applies to AI, right? And, um, and this is what is hard to build. Any, any which claims it's 100% uh in the AI landscape should be taken with a grain of salt, and we are trying to claim that, uh, but at least we understand these are the five principles that we need to focus on to build highly reliable AI.

But why is it hard to deliver such AI, right, on enterprise data? It's hard because, um, you need to connect your AI to multiple different types of data sources, right? Uh, and the way you should be thinking about is listening to your colleagues, your leadership, your, uh, employees in the company and see if they're talking about that AI is not really ready to answer questions on their data; AI isn't able to connect the dots about your data landscape; AI is struggling to integrate with different uh data sources and stuff like that.

Then the second problem is that AI fails on answering complex questions or tasks, right? Like I'm not going to ask a simple question like, "How many hours are in a strawberry?" which also AI fails on, but uh, uh I need to get complex tasks done, and can AI really help me with that? Like which are like five-step, ten-step tasks which would typically take me a week to complete, but can I use AI to just get done within in a matter of minutes? These tasks can involve complicated compute, complicated math, data analysis, cross-domain joins, stuff like that.

Then give a freeform text box to a user, right? They will ask whatever they want, and unless you have systems under the hood, you will not be able to answer those questions, right? If you don't have that specific system under the hood, your AI will not be, either; it'll just completely hallucinate, uh, or it will um outrightly give, tell you that, "I can't answer that question." And finally, securing your AI is really hard; that your AI should not have any data leakage; AI should not be able to access any data that you as the user are not supposed to access to, right? So, uh, and you have to handle all of the company's data with high-precision security.

And these pains translate to each different level in the company, so from the business leader to the technology leader to the practitioner, the engineers and AI, AI engineers, data scientists, and then finally the AI consumer in the company, right? And these pains trickle down to each level. The business leader cares about that, "Hey, if I'm investing a significant chunk of money in AI, am I getting a clear ROI?" The technology leader cares about, "Is it actually solving the problem? Will it scale well? Will there be adoption inside the company or not? Will it be trustworthy or not?" The AI practitioner, the software engineer, the data engineer, they they think about, "Okay, how do I build such an AI system, right? What kind of tools do I use? Are they easy to use? Are they maintainable?" And then finally, the AI consumer, they want to know, "Whether I can consume an AI in my day-to-day work; does it help me with my day-to-day work? Do I really understand what it it is doing under the hood? Is it a black box or not? Can I trust it or not?"

And that is not the case today. Let's look at one example of an enterprise-grade AI, right, which is Einstein on Salesforce. And if you ask a question like, "Can you calculate our average sales cycle length?" I tried it once on our Salesforce instance, and it says, "Um, you have to do it yourself; here's are the steps of how to do it." So I refresh the U, uh, instance and asked it again, "Can you calculate the average length of our sales cycle?" This time it says 71.6 days. And when I asked it, "How did you calculate this?" It said that, "I took average of stage one to stage four age length field," but we have seven stages, so, "Can you try again?" But it did not continue responding. So refresh the page again, try it again, uh, this time it says 2.21 days, so different uh answer, but the same approach, U, completely inaccurate, completely non-repeatable, right?

But before the age of AI, right, if you think about how humans used to operate, if I were to ask this question, if let's say your business leader, let's say your CEO or head of sales wants to ask this question, they'll ask their a data analyst this question, right? And the data analyst will go to the database, or let's say it's coming from Salesforce, so they'll go to the Salesforce, write the right queries to pull the right data, then they might need to do some kind of compute or math or whatever they need to do to process that data, either it is in an Excel sheet or is it in code, but they will do some kind of data processing, and then finally shares a report back with the business leader, right, either a PDF report or a CSV file or something like that. So that's how humans operate.

But if I ask you as a human, "Here are a few numbers; tell me how many uh numbers are there? Can you sum up these numbers and give me an average?" And I did this; I I asked a friend, and they are no longer friends, but um, uh I I gave them 10,000 numbers, it just a short list of those numbers, and I asked them, as I tell you these numbers one by one, "Can you start counting them? Can you tell me uh the sum of these numbers, and can you tell me the average of these numbers?" They could not answer the question, right? But then I said, "Okay, but uh this looks like a Python array, right? So I'll I'll give you these numbers as a JSON, and then can you answer these question?" And then the person was like, "Yeah, I can; I can literally write one line of Python code and give you the answer," right? Perfect. So you understand that humans cannot really handle data in their own head; they need access to tools, right, to um uh to really uh get to the right answer, especially when it comes to large amount of enterprise data, different types of data, structured data, unstructured data.

So I thought, "Let me do the same experiment but with uh AI," right? So I took an LLM, GPT-4, and gave it these 10,000 numbers. It says that, "For demonstration purposes, let's take the first 10 numbers to process." Like, okay, uh, it did its math, write, it wrote its Python code, gave me an answer, but in its natural language response, finally, it said, "The count of numbers is 50, and the sum of all numbers is this, and the average is this." How? I don't understand; you literally just did the math; even for 10 numbers, you did the math, right? But your natural language response was wrong. Uh, so how does that even work?

So I thought, "Okay, low stakes, smaller model; let's use a fancier model." So I moved on to O3 Mini High, and O3 Mini High thought for 9 seconds. I was like, "Okay, the user has given a long list of numbers; this might be huge, perhaps even a thousand numbers, 10,000, but okay, since Python is available, I can use it to compute the values." Okay, it's smart, uh, so it tries to start printing these numbers into its Python um program. Now there are 10,000 um numbers; if it tries to repeat those 10,000 numbers on its Python program, there's no way it's going to get all of them right, um; I either it'll miss some of them, it'll get them incorrectly, or will completely hang, and that's what's that's what happened; it kept writing, kept writing for a few minutes, and then finally an analysis paused uh and did not even come back with the response.

So I thought, "Okay, a different AI could be better than that," came after that. So I looked at 3.7 Sonnet with extended thinking. Now 3.7 Sonnet with is really good at writing code; writes a lot of code; can handle massive amount of context. So I gave it 10,000 numbers, and it thought for five minutes, four seconds, for five minutes; it second-guessed its entire life, uh; it first started summing up these numbers one by one, and then realized it will take too much time, so, "Let me um start with a a different; let me use a different approach; let me uh take some do some kind of approximation," but then second-guesses itself again, saying, "Oh, no, the user has actually asked me to give give an accurate answer, so let me go one by one," then continues to a thousand numbers, then stops and say, "I know is going to take a lot of time," so then finally it says, "I've been counting systematically in blocks of 100, and I've determined that there are approximately 5,000 to 5,600 numbers in the array." Okay, systematically and approximately should not be in the same sentence, but okay, uh, and then finally says that, "Approximately 5,000 to 5,600 numbers, cannot accurately calculate the sum, cannot calculate the average," and then its natural language response finally it says that, "Are 4,800 numbers in the array?" How? Uh, you just your final response was 5,000 numbers, uh, and then somehow you have the sum, somehow you have the average, uh, but now see it can't handle the data.

So that's what I'm trying to say, which is uh we should think about LLMs just how we think about humans. Humans need to be equipped with tools to get uh work done; similarly, LLMs need to be equipped with tools to get work done. LLMs should not be directly given any data um in their own LLM context, because if they given the data, they will not answer; they will not be accurate. It's a different story for unstructured data because uh LLMs have to understand and look at unstructured data, but when it especially comes to structured data, U, SQL databases, NoSQL databases, API, SaaS applications, uh, we should not be handing the data directly to the LLMs, and that is a problem which is in-context data retrieval and processing, and that's what all the different approaches that we are trying to build today are, right?

And then let's take it to the next step. Uh, how do you communicate with your colleagues? Um, do you say everything out loud, or do you share like files and write code? Let's say the data analyst, the business leader asks, "Hey, can you give me a report of the 10,000 most active customers that we have?" Will the data analyst say the names of every 10,000 customer one by one in a meeting, which will last for seven hours? No, right? They will share a file; they'll share a CSV file or maybe a piece of code which they can execute or XY. So that's how you communicate, right, as humans. So why do we expect the LLMs to communicate in natural language, U? And we just proved that LLMs cannot handle data in their own head just like humans can't handle data in their own head. So now when it comes to communication of LLMs, why do we use natural language, right? Because uh when we have for these multi-agent kind of architectures, right, where you have an orchestrator agent and then you have a bunch of sub-agents which are designed for specific tasks, this orchestrator agent is talking to these sub-agents in natural language; they're responding back in natural language; and uh finally the orchestrated agent also responds in natural language, and all of the data is moving through these different uh U agents through their head, through their context, through natural language, and that is a very lossy process, right? And that is where they can't really handle data. What if I need to do a cross-domain join, and I need to use two different sub-agents to do this kind of a join, and how do I send uh thousands of rows of data to from one agent to the other in national language, right? I can't. So why do we build multi-agent architectures any differently than how we think about how humans communicate? Um, and that will be the recurring problem with all kinds of in-context methods out there, right? If you even if you do tool calling, if you if even if you do text-to-SQL, these are great; like these are great for making one specific task or one specific agent work well, but as soon as you start building an enterprise AI system where you stitch together these different agents, all of that starts happening in national language, so all the determinism of these underlying tools goes away, and the thing, and as as we realized, right, that the LLMs are not good at processing like this large amount of data in context.

So imagine an AI system trying to answer a question like this, right? Some of you who have attended the past community calls must have seen this that uh find tickets raised over the last seven days by enterprise customers. Think of like an internal customer support kind of use case. So customer support specialist wants to know that uh who are my enterprise customers who have raised support tickets in the last seven days, and that support tickets are about service timeouts, and uh where we are also missing out on their SLA. So can we understand whether these tickets were resolved or not, and if they were not satisfactorily resolved, can we issue $100 in credits to them? And um this is a difficult task, right, because multiple data sources, structured queries, structured queries, calculations, summarizations, and then actions, so a bunch of different things are happening. Now if I have task-specific agents under the under the hood, right, one agent which fetches tickets, one agent which fetches the customer details, one agent which does I don't know, SLA calculation, one agent that issues refund, how do these agents orchestrate with each other, U? They will, I might have hundreds of enterprise customers, right, uh who fit into this criteria; I might have thousands of support tickets; how do my agents transfer this information from between each other and ensure that there is no loss of information, there are no hallucinations, there's nothing? And that's what is hard. And even if a single agent just with tools which might have like a multi-step question asked to it, uh, even a single agent with uh tools will not does not work, and we tested it out; we tested it out with Claude 01 and O3 Mini with tool calling, and um we connected these different database tools and Python tools and stuff like that to uh LLM and asked these kinds of questions, and they completely fail; like the average accuracy was uh below 50% for Claude and less than 60% for OpenAI and U.

So we realized that just like humans, LLMs are good at planning. So instead of asking them to actually answer this kind of a question, if I would have asked the LLM, "Hey, can you just tell me how would you approach this problem? Do not actually answer my question; just tell me your steps," right? So LLMs will give you a very good answer; they are good at planning, and same like just like humans, right? Like humans cannot really read the entire database, keep it in their head, and then answer the question, right? No, but they can; they're really good at planning where they can come with, "Okay, to answer this question, first I need to fetch data from this database, and then I need to write some kind of these kind of data composition queries; I need to call these two other LLMs to do kind of summarization tasks and stuff like that," so I can create this query plan, right? So LLMs are really good at that as well. So that's what we should be doing, which is decoupling plan generation and plan execution, right? Let the LLMs come up with how to approach a problem, but do not let the LLMs execute that problem.

So coming back to this multi-agent kind of architecture, right? So what we are saying is that let them do the planning, uh, let them not uh do the execution. So I can put this philosophy in every single sub-agent, but now will I have very task-specific uh sub-agents, because as I said, right, one agent that fetches support tickets, one agent that fetches customer information, one agent that issues refund; I could have like a bunch of agents, but how many agents will I end up building, right? Uh, and there is no guarantee that I've covered all possible scenarios, and if I am building so many agents, is it really AI? Is there any intelligence there? Because now I'm just uh building these rigid pipelines under the hood, right? And there's no flexibility. Now business leader can ask whatever they want; will I have an agent to answer that question under the hood? Uh, so probably multi-agents is not the answer, right?

But what if you could have a 100% accurate, task-specific agent for any user query? And when I say any user query, I mean any user query; the user can ask anything they want; give them a freeform text box, and what if there was an agent specifically designed to answer that query? Um, and that allows us to reach a 100% an accuracy if you have very task-specific agents under the hood, and that is also the idea behind PromQL. So what PromQL basically does is that you get a data agent which is 100% accurate for that specific user query. So instead of having these multiple agents under the hood that you need to orchestrate using natural language, you just have one, um, one LLM which understands the underlying data landscape and then writes its own agent on the fly, which will be 100% accurate for that specific task. And instead of confusing the LLM with multiple different data pipelines, uh, we just give a single universal data access layer and just one way to access any kind of data, structured data, unstructured data, SaaS data, API data; just have one way of accessing any kind of data, and now I can have a highly accurate agent specifically on for that user query. So that allows us to uh surpass all of the problems that exist with the traditional in-context approaches, right? And get towards that 100% reliability that we've been talking about.

So uh let's quickly look at uh what I'm talking about, and then we can uh take some questions. So uh for example, I have, make this a little bigger, so I have this assistant which I built for like a typical financial firm, and they have an anti-money laundering use case where they want to answer a question like this, like, "We have received intelligence about potential trade-based money laundering; can you analyze the transaction patterns where there are frequent currency conversions between the sender and receiver accounts, especially where the payment currency differs from the receipt currency, and the transaction amounts don't align with the customer's expected behavior," right? So it's a complicated question, U, with a lot of data need to be fetched and processed. So what Prom does is looks at your question, looks at that underlying data landscape, and comes up with this query plan of how to approach this problem. Uh, it says that, "Okay, I will find the transactions which match our criteria, calculate the metrics like transaction frequency, transaction amounts compared to typical patterns, currency conversion patterns and stuff like that, and then identify high-risk transaction based on that." Makes sense; perfect. So now see what it's doing; it's implementing a task-specific agent under the hood. "Oh, I ran into an issue; okay, no worries; I'll adjust my approach to analyze the transaction patterns." Okay, it was using pandas, and I've not installed pandas in its runtime. And you see how it is implementing a very task-specific agent under the hood which will do exactly that; that's exactly how humans operate; that's how your data analyst operates, right? So um I let this continue and see uh what response we get. Okay, executed; it has identified 45 transactions with 13 different currency conversion patterns; nice, uh, okay, and uh significantly deviate from the sender's typical pattern. Okay, so now I have a good analysis of uh some, like this is the most suspicious transaction; I have different currency conversion patterns that I I'm seeing, and then uh there is no frequent uh sender-receiver pairs, uh, so there are no two specific accounts which are doing this a lot between each other. Okay, perfect.

Let's look at another example which is bigger, which is let's say, U, again, some of you must have seen this example, but I wanted to ask this question all in one shot, where uh I want to see if my, I'm an enterprise software company, let's say, and I'm I'm concerned that my highest-value customer is about to leave us, so I want to find out that how are they feeling about their product, and I want to take some action, right? So I asked that, "Hey, can you find the highest-billed organization that we serve? The organization ID data is messed up, so use users' email domains to find the unicorns. Then for this, or fetch all support tickets across all of their users," so you need to find these users, um, "Then for each ticket thread, which means ticket details and comments on those tickets, summarize it, and then use these summaries to extract this organization's sentiment towards a product," right, from based on all of the tickets across all of the users. So there may be hundreds of tickets; can you extract this organization's sentiment towards a product? And if the sentiment is positive, issue $1,000 to the highest individually billed users' most used project, so uh we can issue refund or credits to a specific project by a specific user of a specific organization, so um or if it's neutral, issue $2,500; if it's negative, issue $5,000. So now if I ask this question, let's see how Prom breaks this task down, right, and implements very task-specific agents under the hood. Like, "Okay, first task is to find the highest-billed organization based on email domains," so I get all the users and their invoice items, extract email domains from from the user emails, and then group invoices for each domain, and then find out the highest-billed; perfect; implements that agent; executes that agent. Now we know williams.com is our highest-billed domain; let's get all the support tickets. "Oh, ran into an error; no worries; I'll figure out what the issue is and uh reimplement my agent." Okay, so again, looking at what's happening under the hood, uh you can see it's U fetching all of the tickets for williams.com, then it's saying that for each ticket, get its comments, and then uh once it's done with that, it's going to keep all of this data in a structured memory so that it does not hallucinate the next time it needs to use that information. So as you can see, these are all the tickets by williams.com, and then uh let's analyze the support tickets to understand the sentiment and key themes. So now it realizes that's a semantic task; I can't really do it in Python, so what I need to do is delegate that task to an LLM. So what I'm going to do is instruct an LLM that, "Summarize the support ticket thread; focus on these criterias; be concise; and give me the information," and then finally what it'll also say is that, "Hey, another LLM, based on all of these summaries, can you uh uh can you analyze the overall sentiment, sentiment of williams.com towards a product and consider these factors and then give me a structured response in this JSON format," and that you see how it broke the task down; understood that it's a semantic task, so it needs to ask an LLM to do it, and now it's uh batching it in 10 summaries, like 10 10 LLM calls at a time, so there are 68 support tickets, so this should take seven turns uh to finish, and but now I know this is like all deterministic; there is uh all of this uh these SCK data is being handled properly; all of these uh uh summaries are being handled properly, and then I'm getting a very uh structured response. Okay, um let's see where it ran into an issue. Um, "I see the error; I'll get the owner ID for the project table from the project table." Okay, perfect. And now it realized the since the sentiment was positive, uh, "Let me issue $1,000 in credit," and it's the data layer says that, "Your AI is about to call this command; these are the parameters it's passing; are you okay with that?" And if I think that's okay, I just click on approve and let the data layer uh allow this uh command to work, and it'll call the Stripe API into the hood and issue the refund. And as you can see, I get this nice report of the overall sentiment is positive; this is what's going well; what's not going well; and this is my reason behind that; and yeah, the the refund was issued. See, this is how you can build highly accurate, task-specific agents uh on the fly without having to build multiple different uh domain-specific sub-agents under the hood which need to orchestrate non-deterministically.

Any questions? And drop back to you.

Awesome. Thank you. We do have a few.

Questions that have come through? Uh, this is kind of a pair for the first one, so we'll see if Harsha can throw them on the stage for us. Yeah. So essentially, what's being asked here is: how does PromQL make sure that 100% accuracy exists on the planning? And then, have we identified cases where we have 100% accuracy on the planning, but then the agent implementation may mess up with the result and the synthesis that comes out from that? Yeah.

No, that's a completely fair question. Um, so the way I like to answer this question is: how do you trust your data analyst is 100% accurate? Um, right. So you, you just have built some kind of trust with them, right? Um, and now you know that if you ask a question, your data analyst will most likely be correct. Uh, but how did we build that? Based on their experience, based on the interviews we did with them, based on their past performance and stuff like that, right? That's how you build trust. Um, if you're saying, "Can you make your AI 100% deterministic?" I'm like, "Can you make your human 100% deterministic?" You can't, right? So, um, and AIs are probabilistic models, just like humans are. So, uh, the way to think about accuracy is um, uh, two or three ways: one is, um, you test your AI out with all of our customers. What we do is uh, we ask them to give us a golden eval set of questions where on which they want 100% accuracy, and we allow, allow them to change these questions whenever they want. And, um, it's basically an SLA of sorts where we tell them that if we do not get 100% accuracy on that, we probably won't sell our product to you. So, um, and the customers do that; they, they give us a huge list of complicated questions, expected answers, and stuff like that, and test PromQL on top of that. Um, so, so that's how we evaluate accuracy. But the way we try to make this system very accurate is, is minimize the context it needs to handle, right? LLMs need to handle a lot of context when it really comes to enterprise data questions, right? So, right now, the here PromQL LLMs don't really have to handle a lot of data itself; they just need to handle metadata, right? And then they need to come up with this kind of an approach. Then the accuracy is in the system, not in one specific LLM response. We do not say that every single LLM response will be accurate, right? Uh, no, it's not; it errored out in front of you once, so, uh, of course, it won't be 100% accurate all the time, but it understands if it runs into an error, fixes itself; second, it tells you exactly what it's going to do, so you, as the user, have complete control. You can stop the execution, you can edit the query plan, uh, and you can ensure that you are reaching 100% accuracy. If you're technical, look at the Python code; if you're not technical, just look at the query plan, but you are equally responsible as this person, as the consumer, to, um, to ensure that the AI is on the right track.

So there's a couple of other questions that come up along those same lines that are kind of focused on the human in the loop. So, Har, I feel, throw up the, the first one that we had there from ABE, I think, uh, or the next one after that, Har, because that was kind of related to what we just had there. Uh, but a couple of questions have come through about the human in the loop, uh, and you just kind of explained, right, that you have the ability to edit the query plan itself, and that can change the execution that comes from there. I think some people have questions though around, as an example, like the mutations that exist or the commands that come through. So, like Stripe as an example, how do we, how are we triggering, or how do we know, uh, like that a human needs to be involved in the execution?

Oh, that's a great, uh, question. So, okay, any database, right? Operation, any operation that happens, a human will be involved. I mean, you can turn off that flag if you want, but a human will be involved because, and it's not determined by the LLM; it's determined by the underlying, the universal data access layer. So the universal data access layer, if it's connected to a database and there are right operations happening, it will, um, flag it as a right operation and ask the PromQL system that, "Hey, can you get user input on that?" Or a mutation is being called, or like a POST API is being called, right? Under the hood, but, um, uh, so, so that's one. The second place where the human needs to be involved is when the LLM realizes that it does not have enough context. Like if I would have asked this question, um, like my internal customer support agent, if I ask it that, "Hey, what are the highest-selling movies in the last 20 years based on the specific data set," like, "What are you talking about? I don't even have that data under the hood. So can you please ask me something?" So that's when the LLM determines when the, when the human needs to be in the loop.

Awesome. All right, I'm going to share my screen. I'm going to keep you on stage for just a second, uh, because we're going to be talking about something that you're going to be involved in, I know. Uh, so, uh, folks coming up next month, uh, on April 16th, we have an event in San Francisco, uh, we're calling it The Reliable AI Conference, AI Disrupt, for leaders that are daring to disrupt with AI. I'm going to leave this on the screen for just a second; you can scan that QR code and find out more information and register for the conference itself. Uh, it's going to be a day-long here in San Francisco, and the focus is really on how leaders are using AI to transform their business, uh, and how you can build AI that you can trust for your business. Uh, Ana, do you know what you're going to be talking about in particular, what uh sessions you're going to have?

Yeah, similar things we're talking about, um, a bunch of great, uh, new features that allow you, like incredible AI adoption in enterprises. So like we'll be talking about the kinds of problems enterprises should be thinking about, um, how business leaders should be thinking about AI, about AI transformation, how, um, many things that were seemed unsolvable are actually solvable if you have the right underlying AI architecture. So we, we'll be exploring all of these ideas and you seeing how it fits.

Awesome. Folks, we hope to see you there. Please scan that QR code if you like more information, and thank you so much, and I'll see you in the office. Awesome. Thanks, Rob. All right, folks, uh, additionally, we have some other events that are coming up that we're going to be at. We will be at Google Next in Las Vegas, April 9th through 11th. You can see the booth number right there, 1789. Uh, I've seen some of our folks on LinkedIn doing some clever posts with that number, 1789. Shout out Adam alone, but you can scan that QR code if you like to find out more information, uh, see the events that we're having there along with the actual booth itself, any speakers there, may be dinners, drinks, those kinds of things, so check that out, please.

Uh, next, next up, I am going to bring Prine onto the stage, who has been on the show many times. We're going to talk about chatting with our files. Prine, how are you, sir?

Thanks for having me. Yeah, I'm good. How are you?

I'm doing well. This, uh, may be the nerdiest title talk I've seen from you so far. We're gonna talk to our files. All right, man.

Yeah, awesome. Um, just quick, um, slide on the use cases here before we jump in. Um, so we've released this feature on the PromQL playground console, um, where you can now upload any file as an artifact, um, which now you can use to analyze existing data and database. The primary use case is that, hey, you already have some data in, in your database and probably across multiple sources. Now you have some new document or new text file or JSON, CSV, XML, or any kind of text format, uh, um, that you want to, uh, add as context, uh, real-time for that particular chat thread. Now you can do that with PromQL playground. Um, you can just add those as artifacts, um, and with the API also, you should be able to now upload these files as context, um, for that particular thread. And the second use case is that, hey, even if there is no, um, existing data that you want to like play around with, um, if you have like a large file, um, you now have access to like our Python runtime, um, which lets you like to any kind of computation, transformation of data pretty quickly. Um, although the UI currently has a 10MB limit, you can use API to upload files and then use the Python runtime and the execute program API, uh, to do any kind of, um, computation with, with files, um, primarily I think probably with CSVs, um, and, and any kind of, uh, Excel files and whatnot, but then you have access to this runtime which lets you to some of this, right?

Um, so to demonstrate this, um, I have like three use cases here and three different file formats, um, to showcase this. This is like a healthcare, um, supergraph that I have, um, and imagine you. So there's like patient data here, insurance, uh, details, um, and then case activity and whatnot, right? Um, the use case here is, hey, I have now gotten like a policy memo, like an insurance plan policy memo, which has a bunch of changes, uh, now I want to understand how does this affect the current cases that are there in the patient, um, records and whatnot, right? So I'm going to like quickly ask, "Can you summarize the policy changes?" And then as soon as you like attach the file, it will come up as a memory artifact, um, and PromQL will read this content from this artifact. So it will read this and it'll give a summary. Um, this could be, probably in the future we will have PDF as well, but then right now it should be like a text file format. So this is a quick summary of what has changed in the policy. Um, it looks like a couple of, uh, insurance plans has some changes with respect to pre-authorization, increased to like five days. I think previously the database had two days and whatnot. So there are a few changes here. Um, so I'm going to ask a quick follow-up question to say, "Hey, which of the current cases does this policy change affect?" And then we can now like map, uh, the existing data, uh, and then compare it with the policy memo to see what action needs to be taken for the existing customer list. So that would be like a use case for, or anybody on the admin side who might want to like make changes to the database or the workflow that they're looking at, um, in the API integration and whatnot. So now we've gotten a response, uh, it looks like all current cases will be affected, and then, and it also has a reasoning behind why. Um, so given the scale of impact, we should prioritize cases that have high urgency, uh, levels, and then I think you also have the urgency level, um, in the database, um, column. So you have the urgency being urgent, routine, critical. So you can now prioritize, um, updating the, the policy changes for these customers, right? And then, um, the pre-authorization changed from two days to five days, so those changes need to be reflected to the database. So now this could also be like a human under the loop step where, uh, a follow-up mutation happens which will change the data in the database to reflect that, uh, right. So this one use case, a simple TXT file you, with just like a policy memo, interacting with existing data.

Um, I have another use case where, um, this is pretty big supergraph and need to zoom in to, uh, show you what the different, uh, tables are, but this is for Telecom, uh, use case, uh, where you have customers, you have network performance issues being tracked, uh, and customer support, um, issues keep coming up, uh, for different devices that could be activation issues on the network and whatnot, right? So, um, again, pretty large schema, uh, you have network, um, you have support tickets here, and then you have customers. So again, there's a use case where, uh, you have, uh, an XML file, um, which would compile a few customer support tickets, um, being listed out. So I like to start with here, "Can you summarize, um, the support tickets here?" And then again, it would upload the XML file, um, then we'll use the Python runtime to pass this and then extract some key details here to get a summary of what the issues are. So let's understand what the XML file itself talks about before interacting with our data which is already in the database, just trying to extract, um, the case ID, customer information, issue details, status, and the stuff that's, that's in the XML schema. And of course, that could be some errors in, in, in the file, and then the retries, right? So we have a runtime, Python runtime, which will do retries. So just quickly looking at the code itself, you can see the XML, um, extractor. We would give any system, uh, tuning for this, which just uses a default Python XML library, uh, to import this data and then do some parsing, um, so it, yeah, looked like the structure is different; it's going to retry again. Let's wait for that to happen, and, and I think the common errors that typically happen with these is when the, the column names or the field names, um, in any kind of file format, could be JSON, could be CSV or XML, is if it is, uh, in camel casing, uh, for some reason, LLMs like snake casing, so, um, that's where the typical ambiguity comes in, uh, and LLMs just get confused with the type of data it's trying to parse, um, but yeah, eventually it is now parsed this file, and then it has some support tickets here, um, some categories of issues, um, which is broad here: network connectivity, billing dispute, device activation, plan upgrade, and whatnot, right? Along with the summary of the CBR. Now the, the job for us is to like see if there are similar issues in the database that are being reported, right? So, um, "Can you find if there are similar issues, uh, in our DB?" Pretty simple question, but then it now, uh, PromQL needs to understand the context of this, um, new file that we've uploaded, the new issues that are there in this, and then just, uh, check with the database again, um, to get back, um, a solid, uh, comparison. Let's see how this query takes, and Dropic has been behaving a little bit weird today, so we'll see. All right, so it is going to search the database for similar support issues to those found in the XML file, um, and typically with PromQL, it will use the, the already available artifact, um, so it will, if you look at the code, um, it will first get the artifacts, um, before doing any kind of processing, um, and again, as we already know, none of the artifacts that you see on the right, the memory artifacts, they are never in the LLM context. So the amount of data that you can process with this could just technically be infinite, right? So because you can just, you're just doing this outside of the context window of the LLM. Now let's see, um, yeah, there's some response. So it has found out that there are some initial issues, uh, similar to what we've uploaded on the XML file, uh, looks like there are like 20 customer plans, so they act to problematic status, um, and we look at the final artifact, these are the issue plans which have similar issues, um, so this is one of the use cases, uh, where now again, this is an XML file, and again, this will, this format could be whatever.

Finally, I wanted to showcase another use case. Again, I think this was, uh, a demo from Anush as well, where he looked at anti-laundering as a use case. Um, let's say we have another file, a JSON file, um, where we have a new set of customers, uh, being imported, uh, and before the import happens, we want to quickly verify if they have anything to do in our existing database, in our, in our sanction database, um, right? So the question could be like, "Hey, screen these customers against our sanctions database." So, "Screen these customers against, uh, sanction database," and the JSON gets uploaded as an artifact, um, and it'll try to query the sanction database here, um, and then compare the customer details if that is like an existing list, and then, uh, if there is like a pattern here where, uh, we're seeing a new customer getting, trying to get into the DB, and again PR retries with some quick, uh, SQL or program fixes. So yeah, it has found that there are no matches, um, so these are the existing customers, and then from the new list that we've uploaded, there doesn't seem to be any match, and it comes up with the match confidence score as well, uh, to be able to tell that. So yeah, that's, that's what I wanted to show, uh, broadly there are different use cases for different file formats, but then the, the idea is that you can now add that as context to your particular chat thread to compare it and analyze with your existing data that's, that's already has access to, um, quickly before dropping off, um, I wanted to also like point out that, hey, if you, if you have like three large files, you can now use the CSV connector that we have, uh, which just imports any kind of file, any number of files as well, um, and then you can just chat with that. Um, the other way is to like, if you have files in your cloud storage, which could be S3, Google Cloud, Azure, or from your file system, um, and you can use the storage connector that we have, um, and then just add your files, and then just authenticate with your buckets with the right configuration, and then we can start using PromQL to also like start asking questions. So yeah, that's, uh, all I have for today. Um, let me stop sharing.

Yeah, thank you, Prine. That was awesome. Uh, I think that before any of the questions start coming through on the chat over here, people are gonna probably ask, right? So the first demo that we looked at was around healthcare. What considerations are there for, you know, sensitive information or private information being put in the context of these LLMs?

Yeah, and as I mentioned before, um, the, the artifacts that, that you upload or the data that's there in the database is typically not available for the LLM, uh, in PromQL runtime use case, uh, and the reason being, uh, it's all handled outside of the context, and the only case where LLMs actually get access to the data is when they do summarization or classification or extraction where it actually needs that data access, but otherwise, PromQL is just writing query plans and writing some code to do the data retrieval and do computation, um, it doesn't need access to the data unless it's actually required for an AI primitive task. So yeah.

Awesome. That context is really helpful, play on words there, context or pun at least I guess. Uh, let's see. No other questions have come through, so I'll just ask you to hang around, and folks, if you have questions for Prine in particular, uh, you can still ask them in that Q&A section, and you can answer them off-screen for anybody that may have a question that comes through. Thanks, Prine. Thanks for joining us.

All right, folks, jumping to our last session of the day. We get to see this earlier in a company meeting that we had earlier in the week, uh, where we get to see some data visualization in PromQL. I'm going to go ahead and bring Siraj up on the stage, who is one of our very talented engineers who's been spending a lot of time on the front end, uh, so this is exciting, uh, we get to take not only these, you know, excellent pieces of synthesis and analysis that are coming out of these LLMs, but now we can actually visualize that data on the fly, we can make tweaks, uh, and I'll stop stealing your thunder and be quiet and let you take it away.

I am, too, actually. This is pretty amazing because, you know, even after building this feature, I myself was like amazed on the kind of things that it can actually pull off, uh, because, uh, you know, unlike, uh, you developing some visualization for a particular need, uh, this is actually, uh, like you can ask anything, right? Just like how PromQL is giving you answers to anything, um, you know, it's pretty amazing. I'll not kill the show. All right, here we go. Uh, so PromQL and visualization. So, so I just want to thought, start with a thought, um, I like a good visualization can give the greatest ideas in the shortest time. So, uh, yeah, just like the code says, like any, any answers to any particular PromQL questions, uh, you know, can be visualized, uh, and, and like grasped in the better way possible. So, uh, again, like what visualization means, um, the answers to any of the PromQL questions that you have can be visualized in, in, in the form of charts, interactive, uh, you know, visual elements, maps, and more. It's cool. Uh, so let's jump on to the demo. So for the demo, what I'm going to use is, uh, I'm going to use a sample PromQL project, uh, it's a public project. I can share this link, uh, later on, uh, visualization is still, uh, you know, not, uh, available on PromQL, but, uh, soon will be. And for the demo purpose, I'm, uh, having a bunch of Kaggle data sets connected to PromQL, so that can ask questions on top of, uh, what I'm going to do here is, uh, so without visualization, I'm going to ask an interesting question, um, maybe it's very common to everyone. So, "Give me a sense of all the data sets, how big." If I ask this question right now to PromQL, it does a very good job, uh, you know, trying to understand the data and then like counting the data, uh, and giving you a sense of what, uh, the data sets are. Um, so what happens is, um, after getting this data, like you, like, like, um, like any sort of data analysis, uh, tool that we, that, that you probably use, first thing is you, you get the data, right? And the next effort probably would be like, make sense with this data, right? Maybe like build dashboards, build visualizations, and stuff like that. What if I ask the same question, uh, when visualization is enabled? Okay, it started, uh, you know, processing the data. [Music] Um, it's, give me a second, it's, uh, a bit slow because it's, uh, running locally. Oh, uh, all right, let me ask the same question again. Okay, it started, uh, acting on your, on, on the data. So right now it has made a query plan, and then what, what it is trying to do is it's, it's trying to do a count on each of the data sets connected, and then, uh, you know, trying to make sense of the data. Let's, uh, wait and see how it, how it actually works. All right, so it actually, uh, you know, figured out how big the data, data is, and, uh, you know, now it has a, uh, details on how, uh, big each data sets are. Now my question is, "Help me visualize this in the, uh, best possible way." So if you have used PromQL and seen like PromQL in the previous, uh, demos as well, as you can see, like it basically started with a query plan, uh, and, and, um, you know, since the data artifact is already created, it's not querying again from the database; it's smart enough, and then it is, uh, you know, started with another action called visualization, and it also, uh, picked up, um, and, you know, it, uh, understood, uh, probably a bar chart can, uh, you know, work better, and also given the data set, uh, represent the, you know, scale and, you know, number of, U, rows underneath. It also realized, uh, maybe a logarithmic scale would be like, uh, you know, better enough to, better to represent a particular, particular data and it, uh, so imagine that you are actually passing on this, uh, you know, previously generated data, data set to your chart; you probably would, uh, you know, download the data and then like pass it on to, uh, a chart or maybe like a, uh, data, uh, visualization tool, and then you have to also think about how to visualize and what should be the x-axis, what should be the y-axis, and what should be the scale. All those things have been captured or automatically by PromQL, and like it actually made, uh, you know, the best possible visualization that it can. You can actually maximize this, uh, and then you can, uh, you know, zoom in, zoom out, pan, um, you know, pretty much do, uh, anything that you would do with, uh, a normal, uh, you know, uh, visualization tool. You can also, uh, do things like you can, you know, probably download and, uh, you know, take a screenshot and stuff like that. Now, um, it's amazing, right? Like basically you are getting empowered by, uh, you know, you know, BI visualization tool on the fly that also understands the context of your, context of your data, uh, and then creates visualizations in the best way possible. Imagine that you're prompting with, uh, you giving some prompt to PromQL, maybe you can add context like, "Hey, I am, uh, you know, this persona, and then I care about this particular, uh, you know, company objective the most, and hence I would like the visualization to represent data in that level," right? This is pretty amazing because you don't actually deal with the granular data; you, you want to see how the big needles are moving, and hence you can actually make sense of the big needles rather than, uh, like going deep into the granular thing, unless you really want to. Um, I would, I would want to, I'd like to show some interesting, uh, you know, visualizations, uh, that it can pull out. Uh, yeah, this is the same thing that we just discuss, discuss, just created another bar chart. And now let's say if I want to, uh, visualize with a, you know, different prompt. So in this case, uh, the data set is already connected with a wine quality data set, which, which contains the chemical properties of wine and then quality ratings. So here I have simply asked, "Hey, um, create a, uh, heat map showing correlations between wine properties." So PromQL part already processed the data and then discovered, uh, you know, this is the correlation, uh, matrix, and then, and then, uh, it has passed it on to the, uh, you know, visualizing part of it, visualizer part of it, and then, uh, it generated a dynamic visualization on its own. It can actually do more, as in like if you, if you ask, um, a further question, uh, can you, or maybe like I'll pick up the previous, uh, visualization and then say, "Can you, uh, make this a pie chart?" It's as simple as that, just like you're asking, uh, you know, your, um, um, probably, yeah, just like you're asking a real person to, you know, convert this into a different format and, you know, make things, make things on, let's see this in action. All right, it's, uh, yeah, so as you can see, like it's again not querying the data set, uh, it's using the existing data set, existing artifact to pull out the data and then creating a different artifact altogether, all right. All right, it generated it on the fly. So just like that, you can pull out different kind of, of, uh, you know, data visualizations. One thing would probably be like really amazing is, I was hoping you would show a map. Yeah, I reserved it, uh, you know, for this. So here I have asked a particular question, "Show the geographical hotspots of accidents from the most recent 100 accidents data." It dynamically generated, you know, everything on the fly, and then you can also like do, uh, you know, more analysis if you want to, like you can zoom in and, you know, pull out and see how things are. It's, it's pretty amazing, uh, and you can also actually interact with the visualization, not only with the data. Now, it's, it's pretty amazing; it's huge.

Thank you so much. Uh, unfortunately, we got to wrap things up just because of time. If you folks do have questions for Sage, you can ask them uh, in the chat as I'm wrapping things up here, and we'll see if we can get to them. Otherwise, uh, you can always reach out to us on Discord. Uh, you can find that information on the hass.io website, uh, and I'll link to get you there as well.

So, right, one thing I'll say before I shift to kind of our housekeeping for the end of the call: essentially what we're doing right with all these visualizations is there's like a miniature web application that's being spun up by the uh, by the L in itself.

Correct? Yes, that is correct. And that's a really big foundational piece for what could come, not just visualizations, right, but think about the underlying architecture there. Exactly. So you can not only create ch and visualization, you can also think about interactive charts, for example. If you can uh, you know, if you want to interact actually with the data and then render different charts on the fly, you can do that, just like you're getting a web app on the fly. It's exciting.

All right, let me share my screen. Looks like no questions have come up, S Ro, but again, if anybody has them, uh, again you can throw them in the Q&A or you can put them in Discord. Uh, one more reminder, and Hara, if you can move us out of the way of the QR code so people can see it there for themselves. Uh, we would like to be able to—I guess we didn't put it on that last slide—all I'll take it back for everybody that you have it there. We are QR code if you'd like to join us for AI disrupt on April 16th. Uh, and actually we do have one question that just came through, so we'll go ahead and talk about that. That'll be kind of our our last bit that we have here. But as a reminder, folks, right, AI disrupt April 16th.

So SJ, the question is: Are we going to have access to visualization data in the API in a fixed form format depending on the graph? So as an example, if the type of graph x is going to be equal to, and Y is going to be equal to, the use case would be to create our own UI where our customers could ask promql questions and we could provide dynamic dashboards. So maybe I guess a way of thinking about this question, Sur, is uh, with the API that we have with promql, can users access the visualization component? Uh, technically yes. We're still exploring on this path, as in like how do we uh, put out this uh, as an API and then what will be the schema like, or maybe the contract like, uh, but yeah, to answer your question, this uh, the data part is always like available, uh, just that it may not be like an XY X and Y scario because it's not only charts always; it can be like anything. So yeah, doable. That's my answer.

All right, awesome. All right. Uh, well, if you want to join me in saying goodbye to everybody and thanks for joining us on this community call, folks, and don't forget we've got Google Next as well, but we'll see you folks online. Thanks everybody.