Transcription
Sometimes when you have a hammer, everything looks like a nail. And that is the world we live in with AI today. Everybody's like, "I'm gonna use AI for that and this and the other things." And what I'm here to tell you is don't do that, right? Stick to the things that work. If you have a process that's working for Q routing or deduplications or IOC lookups, you don't need to add AI to that if it's already doing what you want it to do. Maybe you could add AI downstream from it to make better decisions to do autonomous tasking, but you know, you don't need AI everywhere. But for this specific talk, what I wanted to focus in on is the forensics collection and analysis.
Thank you all for joining us. Uh, welcome back to office hours. Um, for those of you who have been here before, uh, our guest this week, Jimmy Assel, is probably a familiar face. Um, so he'll be uh, walking through uh, an abbreviated version of a talk that he did uh, at Sector um, up in Toronto. So, uh, super excited about that. Before we get into Jimmy uh, and and his presentation or kind of his walkthrough um, we're going to go through a couple things that are in the news. And I want to, like, you know, Dave, Jimmy, I'm not sure you know, how how in-depth you all have kind of like read about or familiarized yourselves with these things, but I'm super curious your take on both. So we're going to get right into it. On the audience side, if you all have questions for Jimmy um, or anything related to the news items, just drop them in the chat. We're going to try to pull stuff in as we go. Today, we want to make time for like Jimmy's topic and some discussion because I think it, it is absolutely awesome and I know one that like we've had a lot of engagement on um, at like, just AI and machine learning applications in general. So, we're moving right in.
So, first news item this week. Um, so for those who haven't heard yet or you know, have uh, have not been sticking your nose all the way into infosec news. So, an internal GitLab instance at Red Hat was breached. Um, GitLab put out what you see on the right, which is a very abbreviated, not super useful kind of like incident report. Um, but there have been a bunch of interesting developments. So, uh, one of the things that has happened, like, this is an internal instance that's used by GitH, sorry, Red Hat's uh, consulting group, meaning that like, when Red Hat goes and does work with another, you know, one of their customers, bespoke consulting, this GitLab instance is where they collect requirements, store a bunch of information um, and pro, you know, presumably store some custom deliverables. And so, uh, you know, the Red Hat disclosure says there's like, you know, it kind of lists what some, what minimal information they believe to be in there. I think the plot twist, and this is where I'm not probably 100% up to speed uh, that's occurred in the last couple days, is that uh, the threat actor who is responsible, Crimson Collective, who, you know, breached this instance and exfiltrated the data um, has partnered up with yet another threat actor that specializes in extortion, ransomware, things like that. And uh, they've now claimed publicly that, you know, despite the, you know, Red Hat disclosure indicating that it makes it sound like there's not super sensitive information in here um, these threat actors have said they've used what they've taken already to breach these downstream customers. And so um, not sure where the, you know, corroboration stands as far as all that goes. Uh, you know, you take all these things with a grain of salt, but that's where we are. Like, interesting >> um, and a little bit terrifying kind of case, like, just understanding >> how I'd say, in general, like, you know, customer-to-vendor interactions happen. Like, our customers are super candid with us. I know when the Okta, one of the Okta security incidents occurred, like, a really rich source of information for them was like, support cases and things that their customers had shared. So, despite what anyone thinks are in these support systems, I feel like the stuff that's actually in them is always a little bit more expansive and more concerning. So, uh, yeah, interesting incident. >> Yeah. So, Ken Kenneth Newton's asking about threats on the release of all Salesforce data from the Drift incident. I think these are related and it, and it speaks to like something I think we as practitioners should be thinking about, which is, when, when we receive sensitive information from customers, how are we handling it, right? Because support cases uh, or things like GitHub or GitLab repos are places where it's realistic to expect you're going to get sensitive information. Uh, my guess is that if those Red Hat actors are claiming that they've down, they've compromised downstream customers of from Red Hat Consulting, it's because they've pinched something like uh, access tokens or something from that environment. And like, I, I think like, for us as practitioners, the question we ought to ask is, how do our organizations handle sensitive information? If somebody gives us an access token in a support case, do we actually treat that like exposure of sensitive data? Right? Then back it out and say, "Rotate that cred. This is the approved way for you to hand this to us so that it ends up in a credential manager." And I'm not even sure if that's the right way to, to or the best way to handle that. But leaving things in clear text in a support repo opens you up not just to a bad actor breaching that, but an insider, you know, harvesting that as well. >> Yeah, I'll just dig into the vulnerability here. This was a 9.9 CVSS score. And what's interesting is they like announced the CVE the day this blog went out or the day before. So, was this a zero day? Um, which is interesting. I don't know, Keith, if you had any more details there. >> Uh, I saw a mention of a CVE in the Red Hat disclosure, but I think it was specifically because they, I'm not sure if you're there might be two different ones here. Like, I know that they announced one um >> the OpenShift, it's that one in the article down there, this OpenShift AI vulnerability that was announced yesterday, right? So it's like >> it sounds like it's unrelated. I mean, that's at least what they're saying in there, but okay, TBD. Um, so I don't think they've talked much about how they got into the GitLab instance, but um, >> yeah, >> I, I know and you know, the one of the scary things I think in >> uh, is that, you know, usually I think particularly like on the user side, like there's like things that, you know, are public which we rarely use to interact with customers, and there's like SAS applications and stuff that are in between, like provided by some third party, but, you know, you have to trust them to some degree to, you know, hold some semi-sensitive information. Um, and then there's stuff like this where like, it's a GitLab instance. It's presumably internal and so everyone unfortunately tends to be, I think, a little bit more free than you might want them to be, right, Dave? To your point, like assume it has sensitive information. And unfortunately, when these things are internal, might people might assume it has sensitive information, but because it's internal, they might also think like the risk is far lower. >> Yeah. >> It's uh, that's always >> So, one of the things that I've observed like in in a previous job, we had a lot of reservation of moving a system like this into the public cloud because it felt less secure. But what ended up happening is that that system ended up just getting cruftier and crustier uh, because the internal resources for keeping it patched and managing it and keeping eyes on it got pulled into other other areas. And so like this risk analysis of of whether it's safer to host it on-prem versus in the cloud is uh, like, in my mind, the evidence indicates it's safer to to have your stuff in the cloud because there's somebody that's actually been paid to keep track of it. >> Sure. >> Yep. Um, all right, moving along. So, we're just going to do a quick hit on this one because I want to make sure we leave plenty of time for Jimmy. Uh, so, uh, threat actors tampering with EDR, not a new thing. Um, but I will say like, there's been a lot more public reporting on this, it feels like in the last couple years, right? And I think part of that, and this is good news, um, I think you have a greater percentage of organizations that are deploying what you know, what we consider to be a top-tier like EDR or EP platform, right? So like, you're getting broader coverage, you're using it well. Um, and so rather than just doing an end-run around it entirely because the preventative controls aren't great, like people are running into these things, but what is, you know, there's a whole ecosystem now of tools to help kind of blind or tamper. Uh, so not new, but EDR Freeze is the one that's been getting some news lately and that's notable because a lot of these depend on, if you've heard of like, bring your own driver, right? So like, bringing your own vulnerable driver, introducing that onto an endpoint and then using that for privilege escalation or other forms of tampering. And uh, that can be effective, but comes with risks, right? You're introducing something new into the environment and that's an opportunity for defenders. And so this one uh takes advantage of the Windows Error Reporting framework, which kind of is like a semi-privileged process onto itself because it has to capture memory and do things like that. So, um, so this one was neat and I think the thing that was really notable for us is that like, not too long after this kind of hit the news, uh, someone for the community contributed an atomic red team test. So, um, that you can use to help kind of emulate uh, this technique and see if you can detect it. So, I thought that was, thought that was notable and um, like, Sam, fortunately, we don't have a lot of time to dig into this one, but we do have all the links over there, I think, in the resource area. So, um, >> So, I do have one quick question about this. Do you know if it was leveraging, uh, an error in the that binary or was it using it in its legitimately designed capacity to get this effect? H probably like some subtlety there in how you answer that, but I believe it's taking advantage of the fact that like the Windows Error Reporting framework, like when you invoke that, like when something crashes, that gets invoked and that thing that gets invoked has, you know, has a higher level of privilege. And so, you know, whether that's a feature or a vulnerability is I guess like subject to some interpretation. Yeah, this is a very gray area internal to Microsoft too because it's sort of working as designed, right? Um, and >> This is the living off the land access. Yeah. >> Right. This, that's the question. The question is, do we expect Microsoft to come out and actually patch that binary or if it's working as designed, is this a >> a feature of the system that we expect people to continue to be able to leverage that we're going to have to write detections around? Um, I would just say that usually these, Dave, um, we've worked on some in the past with the Microsoft team and, uh, we've all been told this is going into Windows next as the fix. So, which is Windows 12? >> There we go. All right, without further ado, uh, the U that come out before GTA 6. >> Someone just over there, too. So, uh, all right. So, um, Jimmy, talk to us, man. They're like, "Uh, I know we've done a bunch here recently. Like, I know we've had some really cool content up on the blog. I know we've tried to share a bunch of educational material like our use cases for AI agents." >> Um, >> I've seen firsthand just the tremendous source of leverage these things deliver like particularly to a team that's like already operating at a high level. Um, so gonna let you kind of dive in here. Uh, kind of set the stage. I'll drive for you. But yeah, let's kind of take folks through uh, I think this abbreviated version of your recent talk. >> Yeah, thank you, Keith. So, at a high level, right, um, all the talks you see, you see people talking about agents and all the great things they're doing in security operations, whether that's, you know, threat research, red teaming, defensive sides, etc. Um, but no one's talking about like the middle, like how do you make these things? How do you build these things? Everyone's talking about the outcomes. And so the spirit of this talk at Black Hat was like, let's just demystify all this. Like, let's talk through from the ground up. Like, how the heck do we build an agent? Let's pick a use case, go through it in detail, uh, and then open source all the code. And so we're going to give an abbreviated version of that today. >> Sweet. >> Yeah. Um, so this was the title of the talk. Uh, but what you, you need to set the stage before we talk about building agents and what they are. I want to level set with this group. There's lots of, you know, marketing terms out there on what agents can and can do. What is agentic versus an agent workflow? And I just want to just define some things. So, what is an AI agent in terms of what they're seeing in the stock? Here's my definition I put together, which is an AI system designed to think and act like a security analyst using reasoning and tools to autonomously achieve complex security goals. So, that's it. Using these reasoning models and giving them tools and letting the AI solve problems for us, right? That is the magic of LLMs. They turn to agents when they're solving tasks for you. And the key is layering these tasks and that's where you get a lot of that autonomy. And so what you see here is a really high level of of sort of how Red Canary looks at things, right? These are the jobs to be done in a sock. You have the intake, you have enrichment and context, investigate, um, response orchestration and learning and improving, right? And I wanted to call something out here. Sometimes when you have a hammer, everything looks like a nail. And that is the world we live in with AI today. Everybody's like, I'm going to use AI for that and this and the other things. And what I'm here to tell you is don't do that, right? Stick to the things that work. If you have a process that's working for Q routing or deduplications or IOC lookups, you don't need to add AI to that if it's already doing what you want it to do. Maybe you could add AI downstream from it to make better decisions to do autonomous tasking, but you know, you don't need AI everywhere. But for this specific talk, what I wanted to focus in on is the forensics collection and analysis. And so you'll, you'll see that light up here. Um, and then uh, and so let's dive into the next slide, Keith, to say like, okay, we know what this is. We're going to focus on forensics collection analysis. But before we do that, I want to just quickly talk about three different types of agents. This is how I see the world. Um, and so when I gave this talk, I said, a show of hands, how many folks have used a co-pilot? And the entire room raised their hand, right? I was actually pretty surprised by that. Um, and so a co-pilot, I'm sure many folks here have used them, have a high human control, meaning the human is the pilot. They're using a chat interface to interface with a bunch of data types that might that co-pilot might have access to and you can ask questions and it will help you facilitate a task. I think the biggest design flaw in co-pilots today is unless you're an expert using that co-pilot, your results may vary in the results you get from the co-pilot. If you're a novice and you're being asked to do a more advanced thing and you're using a co-pilot, you're going to struggle to to get sort of a good utility and a good return on that, right? And as the models get smarter, as you bolt on more things to sort of help reason and get more clarity on questions, I think co-pilots will will get better over time. Um, then in the middle, we have this thing called the the interceptor, as I like to call it. This is what Keith was talking about where he's seen firsthand how Red Canary is reaping the rewards from this. This is what we use. We use interceptors, meaning they're in line with human processes. So an alert comes in, an interceptor takes that alert and it goes out and queries all kinds of data and runs models and does an assessment and it builds this like nice summary for a human to look at, right? That is an interceptor. And then fin, and that has high human control, right? Because you're sort of telling the the interceptor follow this step-wise function, right? >> And then finally, you have the sort of terminator, right? The swarm coordinator where you give the AI a task with a lot of agency, it will go off and solve all these problems. I think these are actually super interesting to use. But in terms of a sock, when you need deterministic outcomes where you're trying to solve very specific tasks and make decisions on those, these swarm coordinators at least today um are not deterministic enough for us to use at the broad scale that we make decisions at Red Canary. So that's sort of the high level of uh of of where I see agents and where architecturally Red Canary sees agents and how we implement them. >> You know, when you uh, I love the interceptor kind of like uh, like labeling um, or or branding there. It's uh, and I, I think, you know, presumably like, I think that's the thing that when, when so, I always think back to like, when SOAR came out and this was like the promise of SOAR that everyone wanted, which is like, an alert will come in and, you know, this system will go pull all this context together, do all these things. It's a little bit more mechanical than I think we're getting with like the interceptor model, but, uh, yeah, like putting that in the hands of a team that has processes, knows what questions they want to task, but just doesn't have all the various systems and automation in place to do it, is like, that's so, so incredibly powerful. Love it. >> Yeah. >> Yep. So on this next slide, it sort of builds off of the interceptor versus the swarm coordinator being fully agentic versus these structured workflows. Bookmark this, screenshot this, whatever you want to do, but this is sort of your northstar when building AI agents or workflows in your organization, right? The key is what questions are you trying to solve? You know, how defined is this problem? How should I think about it? What's the human's role in this decision-making process? How predictable must it be? And then for there, you can dive into the sort of full-on terminator swarm coordinator mode or building out these structured workflows, right? Um, I don't think that these fully agentic things are bad. I actually think they're really good. They're really good for open-ended things like threat research, for threat intelligence gathering, for threat hunting. It's really great to use these and let the agents go off and reason and think. They're going to come back with stuff that you might not have thought about, right? A lot of our early agents that we've built at Red Canary are fully agentic, but what we do is we observe them and we go, "Oh, look, it keeps doing the same pattern over and over again. I can replace that with durable deterministic code and have some of the other components um use some of that reasoning and reaction." But for this specific talk, um, Keith, if you hit next, you'll see that we're going to talk about these things called structured workflows. Um, and there's a really great link that we can provide here that sort of like, I, I found this blog like a year ago and I was like, "Oh my god, whoever wrote this read my mind." This is like exactly been our approach at Red Canary. Uh, and so we're going to focus on this as part of this talk. So the next slide, what's our use case here? Um, one of my favorite tools that every defender should be familiar with is OSQuery. Uh, originally it came out from Meta. I believe now it's um, um, owned and maintained by Trail of Bits. But think of OSQuery as a layer of software on top of your operating systems that you can SQLize anything. You can write a SQL query to pull any data you want, right? And so the use case here is if we're going to build a concrete workflow and demystify building agents. Let's pick OSQuery data. You can see from this screenshot, this is a query that joins users and groups tables together and it gives you a really nice output of all the users and groups on a machine. But can you parse that as a human? No. I mean, yeah, this is pretty printed in a way where it's easy to see, right? But things that LLMs are really good at, right? They're really good at deriving data and patterns from structured things like this or even unstructured data. And so what we'll do in this example is, okay, if I want to do forensics on an endpoint, that's about 20 something different tables that you got to aggregate together. You got to organize, bring it all together and bring a concise report. So the idea would be, if I have an endpoint I want to run forensics on, I'm going to run 20 some odd OSQuery queries on that endpoint. I'm now going to get 20 some odd JSONs in a giant folder and I got to make sense of all of that. Aha, I got an idea. Let's build some agent workflows to do all of that. Right? So if you go to that next slide, Keith, um, when you're building an agent, you want to think about this sort of four-step recipe here. Um, you want to keep agents very narrow in scope. Sort of like you do with like someone's role in a job, right? You might be an expert in one thing but not good in another. You want to keep their jobs narrow. So, you want to keep a a defined goal. You say, "Hey, agent, you're the world's best driver investigator, and I'm going to give you the OSQuery drivers table results, and I need you to look at these things." And you're going to prompt it, and you're going to say, "Use Chain of Thought, which means step one, you're going to look at this. Step two, you're going to look at that." That's that clear instructions that you're going to give the agent. You need to prepare that data, right? A lot of the work that we do at Red Canary is how do we pull data in, it's called retrieval augmented generation. It's actually a really great uh alternative to fine-tuning models. Uh, because you can go and grab data at at the time you need it and use it at inference time and then it's gone. So there's a lot of good privacy uh things you benefits that you get from that. But we do a lot of data cleaning, a lot of compressing, a lot of enrichment. So that when we give that data to a model, it's not trying to make things up or infer things. We're asking the model to do something and we're giving it the sort of answer sheet. It might not be the answer we're looking for, but the answer is in that data somewhere. And it makes sure that the that the AI is making decisions only on that narrow set of data and nothing else. And then finally, you want to select the large language model that is driving your agents. Um, there's a whole bunch of decisions that you can do here. My advice would be start simple. The the workhorse of models is like GPT4o from OpenAI. It's super cheap. It works pretty well. I would say that it's stubborn in following directions. You need to be uh, you need to remind it a couple times of what to do. But we've been building agents over the last month or so with GPT5 and it is amazing. Uh, it follows directions really, really well. Its JSON output is is bar none the best. Uh, and you know, we've been seeing good results from a cybersecurity application, but don't take my word for it because NIST just last week came out with one of the most comprehensive studies of an AI model being used to solve cybersecurity tasks and they used um all the name big name models as well as DeepSeek. Uh, and GPT5 came out on like over 700 different categories. GPT5 was on top of all of them. And I think this is some foreshadowing of what's to come. Remember, the models we are using today are the worst they're ever going to be, right? And so as you see them evolve, it's kind of wild to see how good they're getting at specific tasks, especially in cybersecurity. All right, so I know we're going kind of quick here, but we talked about um, you know, using AI to go through a bunch of data. We talked about how you would build these narrow agents. And so in this specific architecture um, we're actually building eight specific agents. So we've built an agent to analyze system information and hardware and drivers and users and groups and network configs and service and processes and so on and so forth. Right? These are sort of the core foundational layer of doing a robust forensics analysis on an endpoint and making sure you have just enough data to get a good report from that endpoint. But you can't run these agents singularly, right? Like, or serially, like, let me go run the systems one first and then I'm going to run the hardware one first, right? The end of the day, these are just API calls to a web service. Well, this is where the power of these agent frameworks come in. I'm actually a big fanboy of not picking the LLM engines to drive your agents, but what are those back-end durable frameworks for helping you uh execute agents at scale, right? And so, in this specific case, this is actually a visual of a tool from the Langchain company. It's open source. It's free to use. It's called Langraph. And if you're a data engineering nerd, you've heard of this concept called DAGs where you can compile a graph with nodes and edges and then execute that. That's what Langraph has implemented. It's implemented the concept of a DAG and you can build these graphs with nodes and edges and parallelization and force a human in the loop and force retries and loops. And you get all this for free and so you're able to chain all of your agents together and you don't need to write code or anything to say, "All right, wait for this one to finish and then go do that one." No, you create all these in a graph. You compile them and in this case, we're doing the sort of fan out and do all all eight of these in parallel and then once they're all done, it'll fire off that aggregate node which will aggregate all that data together. And then from that giant aggregated report, what the three other agents are spawned off to create a quick look report. Right? This is your 60-second, "What do I need to know? What is severe? What is that?" It'll generate that executive report, right? And then it will generate an IOC in JSON format to feed to your TIP, right? So the idea is going through all the mountains of JSON that might take a human an hour uh, and we can do all of this in under a minute, right? That's, it's pretty cool. >> And so um, it's fun to build these things. It's really hard to talk about these abstractly and then say like, "I'm going to go build this agent. It's going to solve all of our problems." Well, great. Like, how do you quantify that? What is the business value to that? How do I, you know, get buy-in from my from my leadership to say, "Yeah, go and build that. I think that makes sense." Right? You've got to set some bars. And so, this is sort of my four-step recipe of for success and building trust in your agents. Speed plus accuracy plus consistency plus cost equals, you know, um, trust and resiliency and success in the agents that you adopt in your in your um, in your environment. Um, you know, in this specific case, we can take a 30-minute task and bring it down to two minutes. That was the rough goal. Well, we, we crushed that. Um, how is the agent correct? Like, is the findings in this actually correct or did it make things up? Uh, and there's some ways you can do that. Uh, we won't be able to cover them here, but you can use things like similarity where you have a control and you say, "Hey, when I'm building my agent, this is my control data set with all of the JSON that I want to test with and I know I should always expect it to mention this hostname and this username and these persistence mechanisms and these process names," and you can build uh checks and balances when you're updating your agents to make sure it's always mentioning those things to make sure it is correct. Um, the reliable output because we're using a durable agent framework like Langraph, it allows us to say, "When I run this agent, I know every single time it's going to produce those three reports. It's going to chug through all of that data and it's going to give me the same result every time, you know, free from being tired 24/7." Um, and then cost per investigation. Uh, with GPT4o, uh, this actually costs something like 30 cents to run and it does it in under a minute. Um, with something like Claude Sonnet, it still is uh under a minute. And it costs, I think it's like 50 cents to run, right? And so, you know, you bring all this together, it's like, great. So, I was able to just chug through 30 plus different JSON blobs. I'm able to assess what the heck's going on. I have these three great reports and I did it all in under a minute for a couple, you know, for under a dollar, right? That's sort of like the success that you that you want to see when you're adopting these things. >> Yeah. And if we have time, I would love to come back to that because there has been some kind of, you know, there's been some chat over here about like defining goals and kind of how to deal with quality control. Um, and in general, just like, you know, not not putting effectively creating Skynet, right? Like, how do you avoid the situation where the thing is making decisions and doing, you know, maybe to some catastrophic ends that you can't control. So, >> Agree, Keith, and that's like that's what the the industry calls like excessive agency, right? Um, you don't want to give your your agents too much because that's when Skynet happens, right? I think you could give a whole talk on just these these things, right? Outside of just building agents as a whole, there's a science and an art to this. >> Yeah. And I think you kind of touch on them here, I think, a little bit, too. >> Yeah. Yeah. Exactly. And so, you know, the key takeaways when you're building these things is it's not just about that single call or copy and pasting a chat GPT. The real value is when you chain the results of agents together. And the key is, if you keep your agents narrow and do a very specific job or a very specific task, I know it might seem tedious over time, but when you build up that stable of narrow agents, when you chain them together, that's when the magic happens, right? And so a lot of the work that we do, uh, the foundation we've built over the last 18 months is building these narrow agents. And now we're really starting to reap the rewards of overlaying sort of these larger workflows and smarter workflows with those agents. Um, and we're, we're starting to see some some really cool results from that. Um, but you got to build trust. You know, one thing that I would say is you can't build these agents and just go deploy them and forget about them. They will break. They require care and feeding. And there's this concept of human-machine teaming, right? That is what we're going for here, where both the humans and the machines work in unison to get better, to be more performant, right? To help scale a sock. But when that turns into human versus machine teaming, that is really, really bad and toxic and you don't want to do that. And the way you prevent that is you continue to keep that high level of trust in your sock analysts. And so one of my favorite things is when one of our analysts says, "Hey Jimmy, I got this result. It doesn't seem right." Or, "Hey, I see this thing right here. It's not quite right." I'm like, "Thank you." Like, if you're taking the time to to give uh, to give us these assessments, like I love it. Like, let's take that in because remember, humans labeling data is really expensive and getting that feedback is exactly what you want when you're building these agents because you're trying to encode those standard operating procedures into those agent workflows so that you can continue to scale your sock and offload that busy work to the robots and let the humans go do that higher level um, you know, tasking and thinking, right? So that's, that's sort of the the the grand goal here. >> And I, I want to leave this up and just kind of let folks make sure everyone has an opportunity to check this out because like I think, you know, we uh, we may have buried the lead in that like Jimmy has put all this out there so you can go do it, which is the most important thing. Uh, there have been a million cool talks, you mentioned at the start, like talking about AI and agents doing this work for us. Really cool. Uh, the concept of like an exoskeleton for an existing sock analyst. Awesome. But like, this is, I, I love that you were able to put this together so people can go do it. And I want to dig in like, while people uh, you know, we'll kind of wrap up here if everybody wants to hang on for a minute or two of overtime. But, you know, like some of our um, internal like talks we've been doing for like our Red Canary Live series, we've been going around just trying to educate people on how we're using these things and the results that we're seeing. A lot of the discussion in the chat was like the difficulty in defining goals and just generally like ensuring quality and ensuring safety. And I think that has been, you know, when you think about, since we love a good tortured analogy here, but like, we have talked a lot about when we, you know, explain this to people, treating them like employees, right? Like, you wouldn't hand like a new entry-level cybersecurity analyst a bunch of alerts and say, "You figure out which ones are good or and tell me if you think you're on to something," and then just walk away. Like, you would never do that with a person. And like, you should really be taking that exact same kind of model of like coaching, trust but verify, like when you said building trust, like I think that's like the most important thing is like, supervise these like you would any employee who's doing a new job for the first time and keep a close eye on it until you know you have some confidence in like where it's weak and where it's strong, right? >> Yeah. Yeah. You know, despite what the benchmarks might say, a lot of these models are no smarter than like a high school level education, right? Uh, I would, I would put GPT5 at that. And so remember, you wouldn't put, you wouldn't hire a high school intern and have them sit in your sock and say, "Have at it. Good luck." Right? Like, that is not how this should operate. And yeah, you're spot on, Keith. Like, uh, I think the term is like, you want to anthropomorphize the AI. You want to treat it like a human and give it human-like traits. You want to measure it like a human. You want to do performance reviews like you would a human. Um, and you'll see a lot of success there. Um, and in driving adoption of of AI in your socks. Yeah, it's uh, you know, a lot of times teams will come to me and say, "Jimmy, I want an agent to do this." And I'm like, "Cool. Show me the process map. Like, what, what are we supposed to be automating?" Well, I don't, I don't have that. Well, great. Well, how, how can I automate that with an agent if we don't have it written down somewhere, right? Um, and yeah, it all, it all starts there, Keith. >> Yeah, that's very cool. And like I, uh, I do love that concept. Than it is like, and honestly, like the fact that people, encourage everyone here, go try this. Um, we would love feedback. I know Jimmy in particular, like would love, we'd love to see like, hear how you're using it, if you've thrown like, you know, new data um, or come up with new use cases for it. I think that really is like where the magic happens. It was as you were walking through the OSQuery thing where it's like, you have all these different tables and you join them. I'm kind of thinking back to when we started looking at like, you know, turning this on identity data and whenever I think about like identity investigations, you have like your Okta proper data as an example, right? Like authentication attributes and all the IDP specific data you would expect to be operating off of, but you have all this other rich data that's effectively like the Nginx instance that fronts it is like user agents and IP, like just very basic information about like what device you're interacting with, where it came from, and things like that. And like just merging the two of those can be super telling and interesting. And like that's a great example, analogous to that OSQuery case where you've got different tables. They're all interrelated. They all need to be looked at together. And, uh, I don't know, there seems like there's so many amazing applications for this that just kind of open up things that like were previously impossible without disgusting SIMs and other solutions. Yes. >> You know, and Keith, part of the motivation with open source and this code is like, maybe someday we can have like an atomic red team equivalent of agents, right? Where folks are able to contribute different agent jobs into a repo, right? And and sort of build off of that. And my idea was like, let's just give this out, see what the community does and and sort of see what, you know, comes from there, right? >> Yeah. Absolutely love it. That's awesome. Well, thank you so much, Jimmy. Always love having you on here, man. Like the work you're doing is uh, I don't know, super exciting. It has been again, you know, I kind of mentioned earlier, is like this, it's been an incredible source of leverage for our team and like the success we're seeing operationally and just the outcomes we get from this uh, we're just able to like look at more, look at it faster and more effectively and like, at the end of the day, the goal is to, you know, find more bad things quickly, right? And so, uh, >> this stuff um, you know, >> impose cost. >> Yeah, that's right. I love it. Well, thanks for being here. Thanks everyone for joining. Thank you for those of you who could stick around for a little while longer. I really appreciate it. Um, thanks everybody. Have a great day and we'll see you next week. [Music]