Transcription
Chaos Engineering. We've all heard about it, but how's it being used in 2025? Find out in a moment.
Welcome to Slip Reliability, the show where we learn about reliability, observability, and resilience one day at a time. I'm your host, Steven Townsen. Welcome back to Slight Reliability. I'm Steven Townsen, and this is the show about reliability, observability, and resilience. And today, chaos engineering.
Today I'm joined by Colton Andress. He's the CEO and he's the founder of Gremlin, the world's first reliability and chaos engineering platform, helping companies avoid outages and builds more resilient systems. Prior to that, he focused on building and operating reliable systems at Netflix and Amazon. At all companies, he's managed systems at scale, companywide incidents, and built reliability programs and platforms. Colton, it's awesome to have you here. How are you doing today?
>> I'm doing great. Thanks for having me, Stephen. very excited to chat as was part of that opening right on brand with what you talk about on this podcast. Can't believe we haven't had this chance yet, but excited to have it today.
>> I know my exposure to resilience engineering and the resilience community has been pretty recent, but uh this seems this is great timing. I think a great opportunity to talk about this. So I guess the first thing is before we go any further and I think most people have a sense of what chaos engineering is about but could you give just for anyone who maybe hasn't heard about it or maybe doesn't isn't very clear about what it is a quick overview of what it's all about.
>> Yeah. So I think most people are familiar with chaos engineering in terms of Netflix's chaos monkey and the purpose of Netflix's chaos monkey was to randomly reboot hosts because that's what happens in the cloud and engineers need to be prepared for it. But that really birthed the movement that's around going out and testing failures in real systems to understand the real failure, how it's going to behave, what are the side effects, what are the knock-ons. One of the my favorite misconceptions about chaos engineering is that everyone thinks it must be chaotic that because of chaos monkey, it has to be done randomly and it's got to be done in this surprise you kind of approach. And I think that's the good news. The evolution of chaos engineering is we treat it like engineering. We want to do thoughtful experiments. We plan them out. We're careful about how we run them so we don't cause any unintended side effects. But we do think it's very important to go cause a real failure in a real system to see not just how that software behaves but how the platform, the network, the infrastructure, the whole system behaves in the face of that failure.
In my career, I've worked for a lot of organizations where the thought of experimenting, especially in production, is terrifying. And the mindset is, shouldn't we just stop all bad things from happening? Why would we voluntarily let something bad happen? Now, I understand why it makes so much sense to experiment with things going wrong because it's the only way we can practice and build this kind of muscle memory and to see the unpredictable ways in which things could go wrong. But have you got any advice on how you might change the mindset of someone in an organization like that or someone who's terrified that says, "No, no, why why would we do that? That's that's scary."
>> It's a bit like saying, "Hey, we shouldn't crash test cars because we really don't want cars to crash."
>> And the truth is cars are going to crash whether we like it or not. So why don't we invest the effort to make it as safe as possible when it does occur? I think that's one of the things I learned early on in my engineering career. the chance of an individual component failing is relatively low. But when you look at all the components of a system, the chance of something failing becomes quite likely. And so something that could only happen once in 10 years happens every day in a large enough data set. And so failure is a constant. It's going to occur. The question is, do we want to understand it and mitigate it where possible or do we want to be surprised by it? And so I think a little bit of that attitude is like, hey, it'd be better if we just stuck our hands heads in the sand and ignored this risk. And if we aren't aware of this risk, then maybe we don't have to mitigate it.
I think when it comes to production, the way I like to explain is like there's this pie chart of everything that can go wrong.
>> And only like half or twothirds of it lives in staging. You're never going to find a set of failures until you test in production. Why? Production has a plethora of customers and customer behavior. And customers always do something unexpected. So you may require a customer to trigger that unexpected behavior. We have security groups, we have load balancers, we have DNS, we have multiple layers of racks, we're doing deployments. Every service behind the scenes is changing regularly. Some of those things you will just never catch in staging.
>> And so should we test in staging first? Yes. Go find everything in staging or in dev that you can fix all of those. That's just smart engineering. But you got to go to production in the end. The other piece I hear a lot is like we're not mature enough to test in production.
>> I was just about to ask about that.
>> Yeah, we're afraid. And to me, the the kind of pathy response I have here is, you know, if you're afraid to get started, you're never going to get better. This is a bit like I have a New Year's resolution and what I want to do is get in better shape, but what I've decided is I need to lose 20 pounds before I start going to the gym. That's a bit like what we're saying here. And the answer is it's a chaos engineering practice and you need to go out and you need to start doing it because that's how you're going to gain maturity. That's how you're going to gain comfort. The thing that really scares you is you don't understand your system as well as you think you do.
>> And you're going to be surprised. And guess what? That's okay. Everybody's surprised. Nobody understands everything perfectly. But until you go out and start poking at it and asking questions and and analyzing the result, you're never going to have a better understanding. and therefore you're never going to be more comfortable and you're never going to get anywhere and we unfortunately we see that happen sometimes.
>> I can absolutely see staging although it's not perfect a staging environment as a good place to get buy in and say you know it's safe to do it here we can practice it here and see what the benefit is do you compl in a staging environment do you partner with so performance engineers to sort of run load through systems so that to get that more realistic experience is that something you've seen before?
>> yeah I think that's an important component to especially any large scale testing, but typically failures are more interesting when you see the system under load or under common behavior. And so yeah, I'd say some failures you can test just functionally without load, but really we see that quite often c we have integrations with load testing tools. Customers will kick off a load test. They'll go run a set of failure tests and they'll basically be doing both of those activities at the same time.
I was thinking about at the start you said that chaos engineering people get caught up in the chaos bit but but it doesn't have to be chaos. I was just thinking right if you're let's say I don't think this would ever exist in 2025. You would hope but there's a business critical system that literally sits on one physical server or one virtual server. You're not going to go into production and shut that server down right because it's just what's going to happen. I think it feels like chaos engineering comes into play where you've built some kind of robustness into your systems and you expect them to be able to cope with certain conditions and it's about validating that not just trying to destroy your business.
>> The purpose is not just to see if we can break things to break things and that could probably be done fairly easily. The purpose is to be thoughtful about the failure and learn something or to validate a resilience mechanism. And so yeah, when you're thinking about this single host, what you should do is some failure mode analysis. Hey, what do I expect to have happen? And so this is another part of what I think is really the discipline that gets missed when we talk about just chaos engineering. And that's hey, we want to sit down. We want to plan out these experiments. We want to have a hypothesis. Here's what we think will happen. We want to have ways to measure it. We might want to have a fallback plan if things go wrong of how we're going to clean it up and restore back to steady state. We probably want to notify some people if we're running in a shared environment so it's not a total surprise. And then yeah, and if we do the analysis and we say, "Hey, we're testing the redundancy of our server and this application only runs on one server." Boom. You don't need to run the test. You know, you're at risk. Go fix that. go get it running on two or three hosts and then bring one of them down and see if then you're validating this resilience mechanism you put in place.
>> But there's no need to put the system at risk if you already know the answer.
>> Yeah. Do you have any frameworks or categories or advice on on how people can design chaos experiments? Are there like different categories or types or some sort of not playbook, whatever the word is?
>> Yeah. Well, if you're if you're setting me up to pitch people to go to my website, you know, and here we put together a lot of how do I test this? What's where do I start? What do I begin with? Yeah, I think look, there's the same 10 things that go wrong in computers. Do I run out of CPU? Do I run out of memory? Do I run out of disk? How much IOPS capacity do I have? Time matters. Certificates care about time. So, what happens if time changes or we lose track of the NTP server? And then the network and the network the classic adage the network is not reliable. The truth is we've all built these distributed systems. That's what's in vogue. And so everything has to function over the network. And everything relies on a dependency on another server that might be critical, might not, might bring us down, might not. And so that's the next thing we test is not just how do I handle losing the network, but how do I handle losing things over the network?
So, you want a trick, a subtlety here? A lot of people, okay, I want to test what happens when AWS S3 has an outage. Do I call up Amazon and ask them if they could take S3 down for an hour? Nope. They're not going to do it. That doesn't work. And that's not necessarily how the failure would look like. Like, Amazon gracefully shutting something down or dropping access isn't how a real failure would occur.
>> And so, what I want to do is I want to understand how my application, the code in my control, handles that failure. A side point is a lot of people want to test does Kubernetes do the right thing or does AWS do the right thing and the real question you got to ask yourself if it doesn't what are you going to do about it because you don't own that software that's their responsibility so a lot of people want to test what is AWS's responsibility or Kubernetes responsibility and they can't really fix it so what they should focus on is what's in their control so the way we would do that is we would drop all network traffic to S3 selectively And that would tell us to from our applications perspective, S3 just disappeared. That's what it's going to happen in the real world. What happens? How do we respond? How do we handle it? And once we run that experiment, we can go learn, hey, are we comfortable with that? Did we find anything to fix? And if things are going wrong, what we do is we just stop impacting the network. And now we've got this great roll back switch. Instead of like waiting for some other server to come up or come down or some other change to make, we just restore the network traffic. we're back up and running and we've been able to mitigate the risk of the experiment quite quickly.
A lot of people aren't sure what to do or if the right thing happened and that's why we built reliability intelligence. That's our new product where we analyze your test, we tell you why it failed, and we tell you how to go fix it. One of our mantras at Gremlin is make it easy to do the right thing. That ability to really guide people to like, hey, here's what happened and here's what you should do about it. We think that really hastens the process and makes your average engineer much better at reliability than they might be without the help.
I >> I work in this very complex space where there are maybe 30 different systems. They range from mainframes to microservices and everything in between. And there's so many different teams involved. And I cannot imagine an AI agent being able to work within that wider context, but within an individual technology or system, I can see it being valuable. Am I missing something? Is AI capable of more than I think it can?
>> So what role does AI play in like SRE in building reliable systems? And I think do I trust an AI agent to go make production changes today? No. And I think part of what's missing is there's a judgment call that's made. The best on call engineers, the best SRE, when things are going wrong, they have limited information, limited time, and they need to make a judgment call to recover things. And I think that's a stretch for AI today. I've spent 20 years being on incident calls, dealing with outages, navigating a complex system to figure out what changed and what went wrong and fix it. Could AI do that? Maybe someday, but I don't think it's there today. And part of the problem is the non-determinism, especially the LLM approach. You may get a different answer every time. And we're talking about correct distributed system correctness. One of the hardest computer science problems. It needs to work. It needs to be correct. You know, if you let a consistency bug fly by, you could undermine the credibility or the value of your whole system.
>> Uh the other piece that comes to mind is accountability. The truth is the people that we trusted at Amazon to make these decision calls, we also held accountable and they came to meetings and we they talked about the judgment call they made and it was high stress. A lot of people did not want that position because they didn't want to be the one that made the wrong call when things went wrong. Now I think you know if an AI bot makes the wrong call the business is going to be mad but who are they going to hold accountable and the truth is you need some level of accountability to ensure that the right behavior is happening and the right behavior happens over time and I think people's tolerance for like oh yeah the AI bot shut down this critical service because it didn't understand the outage and the outage was twice as bad as we thought it would be like that only has to happen once before AI isn't allowed to touch production for years.
Isn't that nice though? You can say it's not my fault. It's a AI bot. You shift the accountability onto the the technology. Isn't that the
>> you blame the tech, but if you're the engineer that made the choice, then it's your choice. That's like blaming the library for a bug that you didn't know about. It's chose to use that library. That's that's on your shoulders. That might not be something we get mad about. That might not be something we fire someone about, but we need to be able to have that discussion.
I I've seen a trend in certain probably organizations who are not traditional tech orgs but uh obviously rely on tech to run day-to-day of replacing these systems that were traditionally on prem or self-hosted systems of record really big complex monolithic systems and the vendors that supply those systems are now moving to a kind of SAS cloud offering and I've seen this quite a bit which is an interesting context to work in because as a person who works In an organization with these SAS services, you have limited control and visibility and access and things and they were somewhat of a black box. I was wondering if you know of any of those type of vendors who are building in chaos engineering capability into their products or is that not nothing you've heard of before?
>> I think I definitely have heard of a variety of the observability vendors having chaos engineering initiatives and reliability initiatives. I think monitoring in particular tends to be a tier zero service. If monitoring goes down, you lack visibility in the system. You can't really fix it. So, I think they have a natural incentive to stay up. Paged your duty or the incident management. That's another one that comes to mind. If your system goes down but no one gets paged that they've really dropped the ball and so they have an incentive to prevent outages there.
>> Yeah.
>> I think your comment about SAS is like that's a bit of just the world we live in. back to distributed systems. Whether it's a dependency of yours internally owned by another team or a dependency externally owned by another company, it's the same to you. You have limited influence and control, maybe a little bit more internally. But really, you just have to you have to ask yourself what happens if this system stops working and can I live with that? And that's part of how you identify what's really critical. I think that the cloud providers and the monitoring providers they tend to be critical infrastructure and if they fail they cause wide impacts to their customers and so yeah I think it's important for them to test it but I think it's also our due diligence as engineers to make sure we can handle it. The one that always irks me is a a SAS company will come up and their postmortem will basically say it was Google's fault or it was Amazon's fault they went down don't blame us. The answer is the customers are paying that company, not Google, not Amazon. They didn't make the technology choices. They didn't have a say in it. So to them, you failed. And the truth is, you need to go try to turn a cloud failure into a nonfailure for you and your customers rather than passing the bug.
>> Absolutely. That reminds me of um Steve McGee I interviewed a couple years ago talking about how Google's whole philosophy from the start was it's just really cheap hardware at the bottom but you build so many layers of reliability on top of that and uh contingency that the whole things the system collectively is very very robust and I think as a if you're consuming SAS services you need to make your consumption of that reliable that's your accountability right yeah I agree with that huh so obviously companies make money by running services delivering features is a lot of the time that's the priority that's where the money and the funding is. If I was trying to bring chaos engineering into an organization, the first thing that comes to mind you've already mentioned is what is the cost of things going wrong? What is an outage? Easier said than done though to calculate in a complex business. Do you see people having success in calculating the impact of outages and then using that to convince leadership teams?
>> I have. I've seen customers do this well and companies do this well. I've seen companies struggle with this and I think this is a bit of how mature you are in this space. Do you have a a firm definition of what an outage is? When I was at Amazon retail, it was seeing 25% less orders than we expected. That was back in the 20ou late 2000s early 2010s. We tightened that up to 10% threshold over time. But basically, if we saw variance under that, things were okay, acceptable. If things went beneath that, it was an outage. Now, at a in an e-commerce site, you got a pretty good sense to the amount of orders you lost. And we typically see like this resurgence when it came back up where we'd recapture some set of those outages, but some customers would not come back and we would lose that revenue. So, we could track that. There's usually a direct revenue cost that you can correlate in some regard. I think there's also a brand impact. You know, if a customer comes to your website on Black Friday and it goes down, my wife has said this to me multiple times. She's, "Yeah, I don't shop at that website anymore because didn't work when I needed it to. I lost trust." And the last element I think about is engineering time, which is sometimes just as expensive as the outage itself because something goes wrong. We page a dozen engineers. They're on a call for an hour. We triage. We fix it. And then some three or four engineers go spend the next three days looking at logs, looking at metrics, following up, writing the retrospective, getting the things ready. Then we got to go fix whatever was wrong. That takes a little bit wire longer, too. And so that's where we end up investing a lot of time after the fact.
>> And this is one of my big pitches. So, yeah, we need to go to the business. And if we want the business to fund our efforts and take them seriously, we need to show the return on investment and the value we're adding. But the flip side is we go to the business and we say, look, we can either a we can lose three days randomly throughout the year when an outage happens or we can plan an hour every two weeks with our teams and not have that outage. And the truth is by advertising that cost up front, we save the developer time. We break even at worst on the developer time, but then we've saved all that revenue loss and that brand impact. And so that's typically the advice I give is let's help the business understand that we're going to spend this time and money either way. Would you rather budget it and plan for it and not have the negative press or would you rather be surprised, scramble and then have to go piece everything together afterwards?
>> Now the harder bit. See, I see the value of an exercise like that going well beyond the financial, beyond brand reputation. I think that doing work like this has a a massive positive cultural impact I think on the engineers involved because they're being proactive. They're experimenting. They're learning about complex systems. They're learning how to work together. They're building respect and confidence that the together people can work through really challenging situations in a safe environment. Now, I find that very hard to articulate to a leadership team that this is going to benefit not just the loss from outages and those kinds of things. You're also going to have ultimately a better engineering team, probably going to deliver better quality products, probably going to have better staff retention, more engaged staff. That's hard for me to articulate, but it
>> Yeah,
>> it feels like that this is the perfect thing to do.
>> You know what's funny? I'll tell you a little story from inside Gremlin. 5 years ago, we had all the tooling expertise and knew how to do it and we were bad at it. It was spotty. We weren't doing it regularly. We struggled a bit. This is one of the problems we saw across our customer base and we knew this was some of the one of the things we needed to go improve in the product. And so we built this thing called reliability management. It had reliability scores. It had test suites. We tracked history. We were able to see progress. We built executive reports. And that was all great. But the thing that changed the behavior within Gremlin is I started showing it in our engineering meetings. I pulled it up and I said, "Here's our score. Why did it go down this week? And then we made it part of our on call rotation. So every engineer at Gremlin takes a turn being on call. Every on call runs all of our reliability tests. Actually, most of them are scheduled. So often they just go through and they make sure they ran correctly or they address any errors that came up and it became part of our on call handoff ritual that we looked at our scores and if the scores went down week over week, we all talked about it as a team. Well, the result of that is we have very few outages at Gremlin. We've been running in 59's territory overall for the last decade and we've been running solidly in that territory for the last 3 years. But it's also just baked into our culture. We believe in reliability. We care about our systems. We care about our customers and we don't want to be surprised. Nobody wants to get paged on the weekend. No one wants to get woken up in the middle of the night. And that very rarely happens because we make it a dedicated time and effort as part of our team culture.
>> That just sounds awesome. I have to say that's just it's actually very heartening to hear about that kind of thing happening. Okay, going back to the the AI discussion. Yeah, which I know people who have been listening to me for a while know I'm generally a big skeptic but let aside from S, what kind of value are you seeing in terms of AI augmenting engineers in the in the work that they do?
>> Yeah. So, one of the things that I've always prized at Gremlin is we want to make it easy to do the right thing. We want to build a tool that's easy to use, easy to understand, easy to get value out of. And we've had to learn a lot along the way. We started with just this platform and kind of all the building blocks and said, "You put it together, you run the test, you go figure out what happened." And with the reliability management platform, we did a lot of integration into monitoring so that we could tell you what happened. Here are the tests you should run and we'll tell you if you passed or failed, but it's still up to the engineer to go fix it. the classic academic, the solutions left as an exercise to the reader. And that's really the last milestone we wanted to tackle is how do we start telling people how to fix the problems they find. Let's not just find them, let's fix them. And so we leveraged 10 years of millions of chaos experiments, all of the expertise our team has had in being on call and serving and solving problems and we built it into the product. Now is that LLM based? No, actually most of it's machine learning and like data analysis. There's a little bit of LLM here or there, but not we're not just like, "Hey, chatbot, tell me how to fix this problem." We're leveraging the data we see. Hey, we saw you run a network attack. We saw your dependency failed. We saw your latency went up and we saw your error rate spike. So, from those signals, we're able to tell you that it failed, but we're also able to give you a good idea. Here's what failed and why, and then give you concrete suggestions. Here's what to do to go fix it. And to us that's what is really important is being credible and actionable. I don't need to just say hey this dependency failed. Don't depend on this dependency. Not a real useful observation. Hey this dependency failed. You don't have it marked as critical. Is it a critical dependency? Do you need a circuit breaker around it? Do you need an exponential time off so you don't overwhelm the service when it goes down? So that ability to really give people good actionable advice on how to fix the system. That's what we're launching just this just recently with reliability intelligence. So that's our take on how do we make every engineer a reliability expert without having them sit on 30 outage calls and live the lives that we've had to live over the last decade.
>> I mean that that's good. That sounds like a great use case. There's something I've been talking a lot about recently about cognitive load and as our systems get more complex the engineers who he look after them it's just overwhelming at times how much you need to know and be an expert in it's hard to be a a deep expert in every technology and also a wide expert in understanding the overall system and how things interact and all the different teams involved and the reasons behind things. So anything they can we can do to take away that burden I'm all for.
>> So that kind of agent you're talking about it's so it's drawing on learning from previous data. Is it also are you feeding in specific information about the technology that you are like experimenting on as well? How does it work?
>> Yeah. So, we're pretty thoughtful about like how we use our customers data because we got a lot of large enterprise customers that care deeply about protecting their customers data and how we use it. So, we tend to operate on a bit of a meta level. How did the system behave as a black box? Not necessarily understanding the intricacies of every bit of the technology underneath it. And that allows us to have a pretty good sense of most systems look similar if you squint and therefore the advice around most systems is similar. But I think that's a place that we expect to deepen it over time. We actually are pretty big Kubernetes experts within Gremlin. We have a lot of customers that run on Kubernetes that test Kubernetes and there's a lot of subtleties between pods and containers shared network resource allocation. Does this happen at the host? Does it happen at the container level? who controls what? Why didn't I see this number I expected to see? And so I think that's a good example where actually a lot of our Kubernetes advice is informed by that expertise by us debugging a 100 different failures to understand why it didn't behave the way the customer expected or to help them explain what's going on so they don't think Gremlin broke or they don't think oh everything's fine. It's like oh no here's what's going on and here's why. And so I think that's a good example of like let us teach you something along the way and guide you, but it's also a place that we can just get better over time.
>> I'll put a link to the Gremlin website and the reliability intelligence in the description of this episode if anyone wants to find out more about that. And putting you on the spot here. All right, let's say you're a person like me who's in an organization. There is no chaos engineering at all happening at the moment, but I think it's something it really could the organization could benefit from. How do I start? What's the first step that I could take?
>> Yeah. So, that's a great one because I think we've tried a lot of approaches over the last decade of what works best and it's going to be based on your organization. I think early days we tried a lot of grassroots approach. Let's just teach a lot of engineers. Let's empower them and they'll they'll go demand better reliability and the business will listen and things will get better. I think this is a little bit what I call field of dreams DevOps. if you build it, they will come and engineers just want to do the right thing and they'll have the time and the right thing will just happen. And my experience is that's not reality. Yeah. Uh, everyone's busy. Everyone's time is spoken for and it's really got to be time has got to be created and it's got to be seen as valuable for the business. So, if I were going to do this today, I would go to my CTO or my VP of engineering and I would say, "Hey, I think we're leaving money on the table and I think we're running things in a way that isn't as efficient as we could for the business or the engineering team."
And I think we need to rethink how we're approaching reliability. And so, I think the TLDDR there is if you don't have an executive champion that says, "Yes, this is important to the business." and that your champion should be the one that is looking at the metrics, looking at the progress, looking at the scores. That's also the person you need to hold the business accountable. Everyone's busy. So, if nobody holds anyone accountable, it's just not going to get done. And so, you need a like the job I did in the on call handoff, I ask, why is this change? You need somebody to hold the team accountable. But if you do that, I think the good news is you can have your cake and eat it too, as we described. I think there's ways where you can do this in a pretty lightweight fashion that fits into your software development life cycle. It's part of your build and deploy process. It's one of your gates before you deploy to production. It's part of your just regular test that you run every week or every build to make sure the system is behaving correctly. And when you treat it like that, it's not that foreign from engineers. It doesn't have to take a lot of time, but you can save five, six hours of downtime. What's that worth to you? probably millions of dollars at scale.
>> That's awesome. All right. Thank you so much for taking the time to come on the show today, Colton, talk about chaos engineering. Can't believe I haven't done a conversation about it before. It seems so on on brand for the podcast, but uh yeah, thank you so much. I hope it was I hope you had fun.
>> Absolutely. No, this was great. I I love chatting with you, Stephen. And look, you want to do a follow-up? We want to go into detail. I've got opinions. I've seen a lot along the path. And one of the things that's important to me is sharing sharing what I've learned. really trying to convince and persuade people that this is the best thing for them, the best thing for everybody. Let's make our lives easier. Let's make the internet reliable. We want our airlines to run on time. We want our banks to be up. We want to be able to use our phones and and have the internet available when we need to. And that's a solvable problem.
>> Absolutely. Awesome. Thank you once again, Colton. And to everyone else out there, hope you have a wonderful week, and I will see you next time.