Transcription
[applause] We're not talking about context engineering. We're talking about how to fix context. We will look at how AI agents you can provide context to them. We'll look at ways to do that. We'll see what the failure points are and what works. We'll look at how knowledge graph help filling the gaps where things are not working especially in context. And then I will give you tips on how to evaluate rag and knowledge graphs.
Contracts can be 500 pages if they are loan agreements and others. Then if you send a contract which is 500 pages the model can only take 200 pages because of the context length because in interest rate instead of telling us what the interest rate it gave us the formula to calculate interest rate. So this is where all the information if we send all this information then the model will be able to calculate how interest rate is connected to index and the market rate.
[music] Good morning. If you are in PST time zone, welcome, welcome. Thank you. Uh, today is another session where I'm going to push a little bit forward. Not what we talk about, not basic. This is a disclaimer, disclaimer. This is a 301 session. It's not 101, it's not 201. It's 301. So if you don't get it, it's okay. But my goal is to make sure that you get it. So please keep all your energies with me. I need all your attention. Today, uh, it will go by quickly and we can take a lot of things in question answer. But do understand that, uh, there's this idea that more and more people are asking product managers or they're just looking for generalists. And these generalists are basically have this builder persona. And the builders have this one thing in their persona that they know more than others. They know they go a little bit forward in solving problems. And when they go, they have some insights on tech pieces of what works, what doesn't work. And if you can convey those insights specifically in AI, you can bring that advantage to the company that you're going to work for. And that advantage pushes them forward maybe a year or two years. And they're ready to pay a premium price for it. And that is the kind of insights you will gain in this session today. So, but just understand if you're not there, you're starting, this is your first session, it's okay. We will take you there. But, uh, hopefully we can get started. And, uh, all of you know that it's like a little bit advanced topic today. Uh, we are discussing how to fix context. We're not talking about context engineering. We're talking about how to fix context. Okay, that being said, we'll continue to make this little easy, hopefully. I think it will flow by. I think you will understand. You all are awesome by this time. So, I'm, I'm very positive. But if you don't get it, we'll spend a lot of time doing Q&A. I have a few sections I have kept in between so that you can ask me questions and, uh, we can go for an hour if that works out better.
Philippa, thank you for sharing me to the community in UK. I forgot their name already. Uh, but there is a community in, uh, in UK which is Bloom, I believe. And, uh, they are helping people like in product managers to get to good connections. Uh, they have people who are like, little bit same builder persona. So this session came if you want to join that community, what it will take you. Uh, and, uh, I was talking to this, uh, guy who is recruiting for PM positions for startups. And I was just asking him like, what are they looking for? And I go for these interviews all along. I just want to make sure that whatever I'm giving you is helping you really today. And he was like, one thing that is constant at least in startups, and we have seen that also like here in US with these, uh, when you go for an interview, you get this case study. And in the case study, they will give you a lot of things that are builder specific. They will give you a task and they will give you, uh, problems. And they will say that, hey, this works, this doesn't work. And how will you design something or give me a system design of this? Because as with a new technology like AI, what happens is when it solutions revolution comes, there is a lot of things you can solve. And if you solve them, the customers are ready to pay for it already. Because some human is already paying for it. They're paying some human to do this job already. So they're happy to pay. But can you solve the problem becomes the challenge. So the PM they're going to hire, they want to understand at least they know what it takes to build AI specific solution and what are their limitations or what are the boundaries of each and every solution that is in the toolbox of builders. You know, not to build, but you need to understand the toolbox. And you, maybe if you can have like, which tool to use when, this just helps them hire you faster than the other people who they need to train on these toolboxes and then like limitations of each tool. So that was his insights. I just gave you what I learned in 90 minutes with him. So thank you for connecting me, Philippa.
Okay, then we, by the way, we are working with them. Uh, you will see us on Bloom on their web circle. Uh, we will try to create opportunities, especially people who have done the cohort for you working with them. If you're in London, and we're going to grow them in SF and New York. So that's the second piece of our puzzle. Okay, you did the best course in the world, then how can you get to jobs? And I'm looking at like solving that problem for people. So those are coming up. So I'm working on that. So just droo, welcome back. Uh, this is the session. I got all your attention today. I think I should start like this. This is a 301 session. I need everything. This is going to give you this or not. But if you're here for the first time, we're going to learn how to do context management for AI agents using graphs.
Okay, cameras on. I just love looking at you. And when you, when you look at me with these all weird faces, I know I have to double-click on everything I say. So please turn on the cameras, otherwise I will go like a radio broadcast, which I would hate myself for. So, thank you, uh, for your questions. Just hold on to your questions. I will give you windows, uh, to ask questions. And let's stick to one question and no follow-up. And, uh, then I can stay back for follow-ups and remove your confusion. If you have a question, once I will put you default back of the list. So just make sure if you are, raise your hand and ask a question. Think like it's your only chance to ask one question. One person would be a great idea today.
Okay, what are we going to cover? We will look at how AI agents you can provide context to them. We'll look at ways to do that. We'll see what the failure points are and what works. We'll look at how knowledge graph help filling the gaps where things are not working, especially in context. And I will do a quick case study of which I solved using graphs. And I will talk about couple of other companies which have used graphs. If you want to have more, then I will discuss limitations of knowledge graph. Because if they are so super cool, then why they are not as popular as vector databases or retrieval augmented generation called RAG? What, what is stopping you using knowledge graph everywhere? And then I will give you tips on how to evaluate RAG and knowledge graphs. Okay, because this last thing is important. Uh, maybe that's the only thing. If you can do correctly, everything falls behind. This is one thread you can pull, then all threads get pulled automatically. And that's pretty much it. If you want to sign up, you can sign up. I don't know why they put it here, but this is our form, uh, which allows you to get enrolled automatically to our newsletter and all the benefits we provide to our community. So if you're new here, just take a scan and enroll. Thank you.
This is a question that people throw or I got, uh, from somebody who was interviewing. I know I teach you context and then I teach you RAG. But when you are going in real interviews, especially like technical ones where they are looking for a technical PM, this is one of the questions they get, which is, hey, RAG is just the first step of any production journey. What kind of knowledge management you delivered in your last project beyond RAG? And you're like, we did the RAG. [laughter] That's all my taught me. Uh, I got to prompts, I did the RAG, and then, then everything just worked. Because all I did is put my things on that vector database, and then my query went to that database, and I got the right chunks. And then chunks plus the question, and I got my answers. What do you mean? You were looking up, down, and then you're making some crappy story which they are listening but not listening. And they're looking at their phone. And everybody's waiting for that 30 minutes to be over in that technical interview. And mostly these are engineering managers or somebody. So I'm talking about a technical interview here also. Like, uh, box CEO, I use this before in the context, I will use it again. And this is true all across. So if you are debugging AI agents, you will see that everything will fall back to that. If the AI agent failed in calling a good tool in answering the question, then after debugging, you will figure out that the context was wrong. We provided wrong context. 90% of your problems are context problems with AI agents. So if you can solve context, you can build differentiated products. That's the idea.
Okay, then if context is that important, then how can we provide context? How do we provide context, folks? This is where you talk. Okay. So this is my model. I can take any GPT5 or any other models. GPT5, Gemini. I can choose models. So that's the intelligence layer. I want to give it as much as context when I give solve my problems. What are the ways for me to give context?
>> Prompts.
>> Prompts. Good, good, good. Whoever said prompts, 10 points.
>> Prompts. I can add everything to my prompt. Let's say the use case I'm taking is a contract and I want to get risk in a key terms from a contract. Then I can send the whole contract in the prompt. Right?
>> This could prompt consistent prompt like
>> Is there a problem with this? Is there, is there a will this, will this fail ever? I start giving a prompt saying, you are a lawyer, you are the best lawyer. Give me key terms based on the contract type. Extract these 20 terms and their values from this contract. I will get a response. Life will be happy. Will it fail anytime this context problem? If I put it in production, what is the limitation on this?
>> The knowledge base
>> context.
>> Context size. Great. Great. Who will send context size?
>> Because, because this, this dude has a limitation. GPT5
>> Can take 400k
>> Right? That if you take to document, maybe it's like, let's say 200 pages. Contracts can be 500 pages if they are loan agreements and others. Then if you send a contract which is 500 pages, the model can only take 200 pages because of the context length. Then what is the problem? The problem is it will hallucinate. It will not read the last 300 pages and it will hallucinate the hack out of it. And your agent will be bad. And people won't trust you. And you will be out. Okay. How we solve this? Okay. So this is one way and we have a problem with it. How we solve this problem?
>> RAG.
>> With RAG.
>> Great. RAG. Great. RAG. So in RAG, what we do is we take this 500-page contract. We call something called an embedding model. We create a vector store.
>> Where we store these embeddings. And then instead of sending the whole contract, what we do is we send the user query to this vector database or some database. And then we only retrieve what is the relevant pages in the contract to answer that query. And then we'll say, this query plus whatever the prompt you have plus these relevant pages, let's say page number 37. And then you send it. And then you get responses. So you can fill only the context. What is your model takes? Good.
>> Everybody. Good. Is this how we are digging agents today? Like if you go, everybody says knowledge, upload document. And either they give you vector databases behind the scenes, or they hardcode your vector database. Like, or you can create this vector database in NAN or something. And then you can make this problem go away. The context problem goes away. Context length problem goes away. What is the problem? If I debug this, uh, so, so this is gone. And now we have second way to provide context, which is RAG. So RAG prompts and RAG. Then what is the problem with RAG?
>> The specific chunk stored there.
>> Relevant chunks.
>> Yeah. So, so there is some magic here which is query to these chunks or these 137. And what if if this goes wrong? These chunks are not coming right.
>> The user is giving a query. And sometimes when you debug it, what you see is you needed 1, 3, 7, 21. But you got 1, 5, 7, 35. Right? It can happen because some kind of some kind of logic behind the scenes of matching similarity search. And it can fail for you. So you missed, uh, three and you missed 21. And if you send this context, you will get responses as bad responses, correct. So then how you solve that problem?
>> Agentic.
>> Agentic.
>> This is a good, good audience. I, anyways, I, I wrongly said this, uh, that this is a 301 session. You guys are making it a 101 session. Okay. In agentic, what can you do?
>> Or orchestrate what you want.
>> Orchestration. So maybe you can say that before I send this, I won't just take these chunks. I will put some kind of intelligent model.
>> Agent.
>> GPT5 or agent. I will call it agent. And this agent, I will give a prompt and say, hey, check this query, do something, and make sure this context is aligned to this query. And if not, then let's try again with better prompting, better something. And somehow or believe that this is going to work, right? But this is also intelligence layer. There's no human involved. And this can fail too. Good.
And let's look at some examples. So I got some examples for you. So this is good. We have something. We have RAG. We have agentic RAG. These systems are working. And you can discuss all this when people talk to you about context. Now that, hey, I know about RAG. I know about agent RAG. Then I was looking and reading for this and I found more challenges. So this is Glean. Glean is this company which allows you to have knowledge management inside your company. Any everybody knows Glean? Glean is this company which I got the case study for. I, I have no affiliation. They don't pay me to promote them. I wish they do. Uh, they are just promoting some influencers. Uh, so their, their idea is this that what they Glean does is Glean is known for this knowledge management inside your companies. So Chat GPT is good at answering the questions for public search. But what if if you want to ask questions in your own company? So Glean says that, come give me all your documents. And now you can ask do Q&A inside your company. Good. Like we had internal search, remember Kendra from AWS? Uh, Azure has Azure search, which will allow you to search things inside your company documents. They are saying we can do Q&A. We are building chat GPT but for your company inside documents. Good. And they published this study which I was reading when I was preparing. And they said, we have applied RAG, we have applied agent RAG. But here are some problems. LLMs weigh the closeness of the terms. Because of that, what's happening is if somebody asks this question to the chatbot, what is Jennifer Armstrong's title? It says that she's head of customer supplier and employment experience. But that is because she actually conducted a case study or wrote a paper about these positions. And somehow our agentic RAG and RAG system found the similarity more to this. She's talking about this topic a lot in the internal document and she's conducting some workshops on this. But our title is actually Field Marketing Manager, right? And because we find more similarity to this, because she talks a lot about it in in internal documents, the answer came as this. Good. So that's one problem. Another problem with that is if you put similar names. Because if you learned embeddings that we did in the last session, you saw that embeddings are created with similarity, right? How the word appear to closer to other words. So then they ask this question, which is, how well did Claude 3 Sonnet perform on software engineering benchmark? And they think that Sonnet means the older Sonnet model, not 3.5, right? And they gave you some 70%, which is an older benchmark and not for 3.5. Actually, the answer was 49%. Why? Because they find more Sonnet, that word closer to the Sonnet, Sonnet. And then they found the last benchmark and they give you that result. And these are wrong results. And if you start solving these problems for enterprises, you will see these left and right in production. None of these questions will be correct. You will have your accuracy will be 30%. And I will be very proud of you if you can land a generic chatbot inside my company which just does RAG or agentic and can answer like 30% of questions correctly. Especially the edge questions, especially the detailed questions like these ones. It will answer questions like, hey, what is our policy for, uh, giving vacations? Those questions it will nail because those are just language understanding. But when you have to create these correlations, multi-hop reasoning, and, uh, the word weights and all, your RAG will fall for that.
Let me walk you through one more case study before I take you to knowledge graph. Let me talk to you about one more case study that actually I faced when I was building my first company outside Google. When I left Google, we started this legal graph. The idea is very simple that you give us contracts and we can give you key terms out of it. So we were working with a law firm. And they were doing loan agreements. Loan agreements are long agreements, 500 pages. So they called us and they said, you know what, somebody said you guys are great. Why don't you just query convert? Only works for 400 pages, 100 pages. We saw that Radhika said it, I believe. Then we did RAG. RAG solved that problem. Now we can handle 500 pages. We can break chunks to vector databases. Now we can get the question and relevant information. We got these answers also, name of the borrower, lender, correctly, everything is correct. We are so proud. But before the demo, the night before the demo, we looked at the 10 terms and we found interest rate. And this was the answer to the interest rate. We asked the question, what is the interest rate? Same overage. Because in interest rate, instead of telling us what the interest rate, it gave us the formula to calculate interest rate. Because that's what for model, it believed what the interest rate is. We cannot go to this company. We cannot do the demo now. We are stuck. And it's 10 p.m. The demo is 10 a.m. Okay. Then what can we do? We tried all the agent. We tried hard coding stuff. Didn't work for us. We can't hardcode because they can upload any kind of different documents. And then we don't know like which one to which one. And obviously, we can't fool them because the idea is that we want to give it to them and then leave. And they want to use it. And after that, they think that we have accuracy. And if we don't, that's the other problem. Good. You see the problem that we are not getting what is the interest rate? Then we said, okay, luckily we had human benchmarking done. So in in this document, the human has told that, hey, this is the answer. The spread is 4.75. We add to the initial benchmark rate. The index floor is 0.5. We have to calculate it like this. So page number 30 had spread information. Benchmark is London Stock Exchange. Page number 10, we have to know what is the stock market applicable for this. And then we have to calculate the floor which is on page number 70. So this is where all the information. If we send all this information, then the model will be able to calculate. But this is spread across 200 pages. This is how a human will go and answer this question. So for good. Okay. Then, then there is a formula also. The subject matter expert told us that interest rate is basically sum of value of spread plus interest benchmark. In this contract, it's LIBOR. That is the index floor. So find the value of spread and add it to the index floor. And that's your answer, which is 4.75. Okay. Then there is complex multihop reasoning required. First, we need to understand what is interest rate. Second, we need to know how it is calculated. And third, we need to go and find the values of how it is calculated. And four, we have to sum it and then answer. You see, this is how humans are doing it. First, they understand what is interest rate. Then they say, how interest rate is calculated. Then they say, for each item, they go on their journey to find information. And they found all the relevant information. And then they sum it up and answer the questions. So in RAG, whatever we do, we enhance the query, we miss these things. We will get these, but we will miss LIBOR and we will guess the index floor. So sometimes we answer 4.5% because these were there in the chunks. But this information was not there in the chunks. So what we did, we created graphs which allowed us to represent this information as entity relationship, entity, how these entities connect with each other. And then we can create a graph of it. And that's the first step we did. So what we did is we created this knowledge graph. How we created it? We took their documents, the data set. We created entity relationship entities. How interest rate is connected to index and the market rate. Or it can be, what is index rate made up of? Entity relationship entities. And interest rate is sum of spread and benchmark. Benchmark can be LIBOR or SOFR. Benchmarks have index floors or base rates. So now I can see a graph. And that's the graph you create first. And you require a subject matter expert to add it. We can use an LLM to generate these relationships. But and subject matter comes and adds it, maps or approves or rejects these relationships. They go into a graph. And then you can represent this graph into this form, which I gave you, which is an interest rate can be sum of these two values. This can be LIBOR or SOFR. And if it is LIBOR, then it has something called index floor. And then you can calculate the values of it from the contract. Okay, then we got that. So first, if the user asks a question, what is the index floor? We go, we query the graph. We extract what the entity relationships and how these things are connected. Then we have a query parser. What it does is it parses and dumps that information that to calculate interest rate, here is your plan. It creates a plan. But how it creates the plan based on what it got from the graph. And then it says, hey, to calculate the interest rate, you need to first make sure what is in this contract, how it is calculated. That information should be first find out what it is, which is it can be a fixed rate or it can be a sum of benchmark or spread. Then by the way, benchmark, if it is benchmark and spread, then you need to find value of spread and what kind of benchmark is it? Okay, for that benchmark, you need to look for an entity which is the index floor of it. Good, so far following me. I know a lot of, uh, few raise hands, but mostly getting it. And once you have those two, you can go and query the database. Because now you are giving a very enriched query to the database. And when we did that, we got all the relevant chunks out and we could answer the question. So this is how we solved the problem. This is all details. But now the answer is the spread is this, the floor is this. Another thing I can do is I can explain how I got this. The users want to trust these systems. If you can explain how you are doing the things, and they can edit it also. So it's a better taste and better user experience if you use graphs. Because the graph allows you to explain the answer. It's not like, what is the interest rate? Here is the number. Because that's what LLM will spread based on the chunks you give it. But if you have a knowledge graph, then you can share the plan on how you got here, which is first I calculated the interest rate. To calculate the interest rate, I looked at it is the sum of spread and base rate. Then I looked for the, uh, index floor. Index floor was LIBOR. Then I looked for the index floor base rate. And I combined it with the spread, which is 4.75 written on page number seven. And then you can give this explanation. If the user looks at it, they will trust you. And they will buy your products. And you have a differentiated product which nobody else has in market. And then hopefully this law firm will one day buy you. So that was the case study. Okay, got little bit idea.
Let me give you a little bit more on knowledge graphs before I take questions. So this is, uh, what I worked on also. This is Comprehend Events. So what it does is we allow you to have, we create knowledge graph based on news. The news where Walmart, where Amazon was acquiring Whole Foods. And if you put that news, these are the entities we can automatically capture. That there is a corporate acquisition. It's a merger with employments of these people. And this is the value of the stock. This is the amount. Who is the investor? Amazon is the investor. There's a transaction. This is the date. Today, which can be any date. And then this is our symbol. This is Whole Food symbol. And this is the merger that has happened. And then you can also go to employment on whose employees are who. And on top of it, now this is way before LLMs, by the way, before these all AI agents. And the idea of creating these or giving these APIs was that now you can build better search, right? So Thompson Reuters, LexisNexis provide APIs and gives you charge you for proprietary knowledge to do research. And if you can represent information in knowledge graphs, you can now query that, hey, is there a conflict of interest between Walmart and Amazon before the? Because if I sit on the board of Walmart and my wife is the CEO of Amazon, then there's a conflict of interest. And creating graphs like this can answer to those questions. Okay, I gave you two examples.