Episode 985 ·

Reliability Without the Guesswork: Foresight AI with Kolton Andrus, CEO of Gremlin

The best reliability work is the kind nobody notices, and that's exactly why it deserves more credit. Today, we're talking to Kolton Andrus, founder and CEO of Gremlin. We discuss why the engineers who prevent fires deserve as much recognition as the ones who fight them, how testing failure on purpose turned a dreaded holiday on-call into a quiet one, and why teams that pair AI speed with real-world testing can build software that holds up. 

All of this right here, right now, on the Modern CTO Podcast!

Want to know more about Gremlin? Check their Website here

About Kolton Andrus

Kolton Andrus is the CEO and founder of Gremlin, the world's first Reliability and Chaos Engineering platform, helping companies avoid outages and build more resilient systems. Prior, he focused on building and operating reliable systems at Netflix and Amazon. At all companies, he has managed systems at scale, company-wide incidents, and built reliability programs and platforms.

Transcript

(Colton Andrus at 00:00:00) The other part of the equation is people is the hard part of technology. And getting people to do the right things, more than just the SRE team or the platform team, but the engineers across the company, really are required having the right incentives and the right measurement in place. I'm told not to say this, but we want boring systems. We want systems that just do the right thing without us having to worry about them. Plans have dates.

(Colton Andrus at 00:00:24) Otherwise, it's just wishful thinking. And one of my favorite ones is one time, people were talking about an issue they encountered, and Jeff said to them, "This sounds like it's happening to you."

(Joel Beasley at 00:00:43) Can you explain a little bit to start, like, what Gremlin does, and then we can jump into that shift?

(Colton Andrus at 00:00:50) Yeah. So Gremlin, the analogy I typically use is that of, like, crash testing a car or a vaccine. You know, we want to go out, we want to inject a little bit of harm, we want to reproduce real world scenarios to understand what happens when they behave, and we want to go fix them so they don't impact customers, so they don't actually end up resulting in harm in production.

(Joel Beasley at 00:01:16) Okay. And so you guys were on the show years ago. I talked to Matt. Tell me a little bit about, like, where you... You've gotten...

(Joel Beasley at 00:01:25) That's where you got your start, but where are you now?

(Colton Andrus at 00:01:29) Yeah. So I think one of the things we found is what we built was an expert level tool, and you needed to do a bunch of homework, essentially, in order to get value out of it. You need to really understand the system. You need to design these experiments. Our platform makes it safe and secure, but you had to really understand how everything fit together.

(Colton Andrus at 00:01:49) So that was the technical piece. There were some aspects that we have adjusted over the years to make the technical bar lower. How do we give people more out of the box? How do we integrate with their alerting, their monitoring, their logging so we can tell them if their system behaved correctly? That was part of the equation.

(Colton Andrus at 00:02:06) The other part of the equation is people is the hard part of technology. And getting people to do the right things, more than just the SRE team or the platform team, but the engineers across the company, really are required having the right incentives and the right measurement in place. One of the ways I like to phrase this is, "Hey, if I prevent a big outage, am I going to get promoted?" Because you can get fired if you create a big enough outage.

(Colton Andrus at 00:02:32) But will you get promoted if you prevent one? And oftentimes, if things just run smoothly, no one notices. So how do you convey this to your leadership or to your boss that you've done a bunch of work? The reason the system went smoothly was not because you got lucky. It was because you went in and put forth a lot of effort to help ensure that things went that way.

(Colton Andrus at 00:02:55) And that's really where you have to start measuring the types of issues you found, the issues you corrected, the improvements that were made over time. And that's where we came up with a reliability score that really gives leadership a clear idea why did things get better and help them to track it. The other side of that equation is accountability. If leadership doesn't have any visibility into what's happening and what's improving or what isn't improving, then they can't go and ask teams to go prioritize work or lean on them a little bit to make sure that work gets done. So that was the social side of things.

(Colton Andrus at 00:03:31) I don't know if you want me to get into the AI side of things as well. Yeah. So then right now, the thing we've been working on the last six to nine months is around how do we just do all of that for people? That was really the vision. One of my favorite... I'm a Simpsons nerd.

(Joel Beasley at 00:03:47) Mhmm.

(Colton Andrus at 00:03:47) And one of my favorite Simpsons episodes is when Homer becomes the garbage commissioner. He's got this great song, just like, "Can somebody else do it?" And sometimes that's how we feel about some of our systems. Wow, can somebody else just do that for me? So that's really been our long-term vision at Gremlin. How much can we do for people? Can we analyze their systems? Can we tell them what would go wrong?

(Colton Andrus at 00:04:11) Can we tell them what risks they have? Can we design the experiments that they should run? Can we run them for them? When we run them, can we analyze the results and tell them what didn't go well or what went well? And can we recommend the fixes and the remediations out of those?

(Colton Andrus at 00:04:28) And so I think that's what's exciting about today's era is now, much more than five years ago, we can do many, many more of those things. We can do that whole lifecycle that I described, start to finish with, you know, a little bit of human oversight, a little bit of human nudging and guiding here or there, but in a pretty much autonomous fashion.

(Joel Beasley at 00:04:48) And is that the new Foresight AI product?

(Colton Andrus at 00:04:51) Yeah. That's our new Foresight AI product.

(Joel Beasley at 00:04:54) Is it beta? Is it actually out? Where is it at?

(Colton Andrus at 00:04:57) It's currently in beta. We've had a bunch of our customers that have been hammering on it, making sure it's good enough. A couple have said they can't live without it, which, you know, as somebody who writes software, you always love to hear. And we're going to be launching it. I think, you know, by the time this airs, we're going to... it's going to be live.

(Joel Beasley at 00:05:16) How is this going to impact, like, the AI SRE world?

(Colton Andrus at 00:05:21) Yeah. So I think the AI SRE world is interesting. But, basically, you're still waiting until things break in order to go back and fix them. So first of all, something's already gone wrong. And we'd love to live in a world where we find and fix those things before they go wrong, before they impact customers, before you have an issue.

(Colton Andrus at 00:05:41) The other aspect that AI SRE takes is, you know, they look at the metrics, the logs, the data, what happened, and then they guess what was the root cause, why did this occur, and how do we go remediate it. When you take that proactive approach, you've already analyzed the system, you've designed the experiment, you've caused the failure. So you know what the root cause was. So you don't have to guess about what the root cause is.

(Colton Andrus at 00:06:08) You're able to look at what the side effects and knock-on effects are. But I think the most important part is it gives us an opportunity to close the loop and verify that the system is truly fixed once we've remediated it. So we can go back and we can run that test again, and we can ensure that the system is truly hardened against that failure mode. And so that way, we're not kind of duct taping the system together and moving on and hoping that we fixed it. We're actually able to prove that we fixed it.

(Joel Beasley at 00:06:38) You almost are extracting, like, a core function a lot like the SaaS companies did. It's almost like reliability as a service. Right? You're just pulling that out now. Is that too much? Did I go too far?

(Colton Andrus at 00:06:50) We toyed, we toyed with reliability as a service. Did you? Reliability is... you know, means different things to different people. You get into this, like, reliability versus resilience. What exactly does it mean? But that's what we want to provide. You know? We want to really understand and fix all the issues in the system so that things behave as expected. We want...

(Colton Andrus at 00:07:12) Really, it's... I shouldn't... I'm told not to say this, but we want boring systems. We want systems that just do the right thing without us having to worry about them.

(Joel Beasley at 00:07:21) You do want boring. As an entrepreneur, I can tell you, you want a boring business with boring systems that generate boring profit. Yeah. That's what you want. It's unattractive. It's not sexy. It's not Instagram highlight reel. It's not... But it's steady, it's stable, and it's a beautiful thing. It allows you to explore other parts of life too. So...

(Colton Andrus at 00:07:42) Yeah. That's part of the... comes back to the incentives. You know, if somebody goes out and there's a fire and they... and the fire gets to a bad spot and they put it out, they're a bit of the hero.

(Colton Andrus at 00:07:53) And everybody saw the fire and everybody knows the risk of the fire. And so everybody, you know, gives accolades to that person that put it out. They're the hero. They're the person that saved us all, and they're super excited. But what about the person in the background that prevented the fire from ever occurring in the first place?

(Colton Andrus at 00:08:11) Wouldn't it be better to just not have the fire? And I think that's one of the challenges you run into in this space is you have to get people to think a little differently about the problem.

(Joel Beasley at 00:08:23) They hate the codes guy. The guy walking around with the building codes, we hate him with the checkbox, and he won't give us the permit and the CO and... But he's doing that. He's stopping the fire before it happens through the right systems in place. Yeah. Exactly. That's interesting. Wow. Well, you know what? I bet fires starting are a large reason why people spin up reliability teams.

(Joel Beasley at 00:08:49) Right? You get enough fire starting, you're going to create a little fire department. And then if you can get them to put codes in place to prevent the fire, it's kind of hard. Have you noticed any resistance? You're basically saying, "Hey, industry, you guys have this continuous problem. You have entire teams on it. We want to kind of, like, automate most of what you're doing."

(Colton Andrus at 00:09:07) Yeah. I think the pushback we see is people are busy, and they don't want to invest the time. And they're used to the status quo. They're used to having on-call. They're used to getting paged.

(Colton Andrus at 00:09:19) You know, maybe somebody else has to deal with that fire. And so they're focused on "How do we get the new product feature out the door?" You know, "How do we go get the new hotness out in the market?" which is great, but they're... I think they discount that pain that comes when you launch the new feature and it falls on its face. Or you're scaling up with a bunch of customers, and all of a sudden your service is unreliable, and people start feeling that pain.

(Joel Beasley at 00:09:45) So we've got this market. It was mostly, like, AI SRE, and I think you guys fit into that type of tooling as well with your earlier product. Am I wrong?

(Colton Andrus at 00:09:56) Well, pre-AI, we've definitely fit into the SRE DevOps space. Yeah.

(Joel Beasley at 00:10:01) Yeah. So we've got that. That is all fundamentally reactive. Right? But then this new product that you have is taking all the tools that you've built to help the teams react to all of these things and essentially prevent it as well.

(Joel Beasley at 00:10:16) Because I want to touch a little bit on your background because it's fascinating. You were at some large companies. I think I heard a story about you running through the server room, like, yanking power. Was that you? Were you doing that?

(Colton Andrus at 00:10:27) I've never gone through and yanked power cables in a server room, but I've been part of many, many incidents and outages as a call leader at Amazon and as an incident commander at Netflix. So I have my fair share of war stories.

(Joel Beasley at 00:10:43) But isn't that how one of the core reasons why the company got started is they were, like, intentionally creating these, like, little gremlins in the system so they could basically... And it's like a controlled fire. They're, like, controlled. Like, "We're going to create this little fire and then see how to react to it." Am I wrong? Because it has been, like, six or seven years since I talked with you guys.

(Colton Andrus at 00:11:03) Yeah. Back in, you know, back in the early 2000s, 2005, 2007, that's how people were doing these tests. Somebody would go into a data center. They'd unplug a rack. They'd unplug some cords.

(Colton Andrus at 00:11:16) And there's a story before my time in Amazon where they went in and did that, and it turned out that all of the Seattle office's internet ran through that data center. And not only did they take down the website, they took down the entire dev office's internet. And they had to, like, rush to go fix it, and people couldn't do work, and it was a bit of a debacle. And so that kind of approach, you know, when we first built this tooling, I built a version for Amazon, I built a version for Netflix.

(Colton Andrus at 00:11:46) It was really about "How do we make it safe and revertible?" You know? If something goes wrong, we want a big red undo button that we click that puts us back to steady state. We don't want to have to go send someone down to the data center to, you know, swap out a rack if things aren't going the way we expect.

(Joel Beasley at 00:12:04) Yeah. We want to imagine Homer and Marge did this deploy. Homer and Marge. Sorry. That's not a... Dude. All right. So... Look. I've got a question about comfortability with this. Right? So the Tesla, when it first started with the self-driving, I was really like, um, and then I did a mile with it and then I did 10 miles. Now, like, 70% of my driving is on autopilot and I'm very comfortable with it. But it took me some time to warm up to this autonomous thing with my life in its hands. And I'm curious, how do you help a team build confidence with this tool from the beginning?

(Colton Andrus at 00:12:45) Yeah. It's interesting. I also have a Tesla, and I tried the self-driving. I like... I love the auto steer, but I tried the self-driving a couple years ago, and it wasn't to the point where I was willing to trust it. And I took a trip this summer, and I said, "Well, let's try it again. We've got it." And then it drove 90% of that trip, and it did an excellent job. And since then, I've left it on, and I use it all the time. So I think, you know, that's really...

(Colton Andrus at 00:13:09) The proof is in the pudding. You've got to go actually do it. You've got to see it work, and you've got to build that trust and confidence. One of the things we care about is actionable, credible advice from our systems. So we don't want to guess at what we think the problem is. We don't want to guess at what we think the remediation is. We want to have a high degree of confidence. And it's funny. I say, you know, I think a lot of systems out there kind of working in the 80/20 land. My CTO will say he thinks a lot of people are working in the 20/80 land, like, 20% accurate, 80% guessing.

(Colton Andrus at 00:13:44) And we really want to measure accuracy in nines, like, nines of availability, nines of uptime. I want a system that's 99% accurate, 99.9% accurate. So we've done a lot of work in the way we built our harness, the way we built our tools, the way we built our skills, and the way in which we've used all the data within Gremlin. Over the last 10 years, we've run millions of experiments across tens of thousands of systems. They're all labeled.

(Colton Andrus at 00:14:11) We all know how they performed. We knew what the outcomes were. So that allowed us to build this failure atlas that was a data source that we built into our product, and that allows us to have a high degree of confidence in the results we're giving. The other way we approach it is, you know, the LLMs are great at the human interaction, at summarizing, at conversing. They're not always the best at some of the engineering or mathematical operations you need to take.

(Colton Andrus at 00:14:39) And so that's where we delegate to our control plane, to our tried and true software for the really scary actions, for the really important ones that must be correct, that must revert correctly, that must be safe to operate. And that way, we're not relying on hope as a strategy when it comes to executing those experiments or interpreting some of the results from them.

(Joel Beasley at 00:15:04) And you're right. I've... As an engineer in this modern day, what I have come to find true for myself and my projects that I'm working on is that it's almost like good music. It's almost an art to understand how to mix deterministic and non-deterministic systems together. At what point do you bring in which for what job? And that is an art. It's not necessarily an exact science right now.

(Colton Andrus at 00:15:31) Yeah. Yeah. Yeah. Some tuning, and that's where you put it in the hands of customers, and that's where we use it ourselves. Mhmm.

(Colton Andrus at 00:15:38) It's funny. An Amazonism is not we eat our own dog food, but we drink our own champagne. And Jeff Bezos got up in an all hands, and he said that once, and he gave us all a bottle of champagne and said, you know, we don't eat our own dog food here. We drink our own champagne. But we use Gremlin on Gremlin.

(Colton Andrus at 00:15:55) We run Gremlin in production. All of my team members are on call. They all take a turn analyzing the results, understanding what happens, and we use that to tune the system and to make sure it's trustworthy and it's working correctly. And we run it about five nines in production. We've got a very stable and robust system as a result.

(Colton Andrus at 00:16:16) And so as we've gone out and we built Foresight, we're playing with these new AI technologies, we really want it to be genuinely useful. We want it to solve our problems. We want to have a high degree of confidence in the outcomes. And so we've gone out. We spend a bunch of time, you know, on our systems in real environments testing it before we gave it to customers.

(Colton Andrus at 00:16:36) And then we went to beta, and we had a bunch of customers do the same thing. And we said, look, we want this to solve your problems, and we want you to feel really good about the outcomes. And if at any point you don't, we want to hear about it because we're going to narrow in on those, and we're going to go fix those issues, and we're going to go understand why you didn't get the answer you wanted or why you don't have confidence in the result. And we're going to tune it in and make sure it's as accurate as it can be.

(Joel Beasley at 00:17:01) So that champagne thing came from Bezos?

(Colton Andrus at 00:17:04) Yeah.

(Joel Beasley at 00:17:05) Oh, you know, the first time I heard that, I was interviewing Archana, the CIO of Atlassian, and she said it. And I was like, I thought it was her.

(Colton Andrus at 00:17:17) I was...

(Joel Beasley at 00:17:17) Like, it's brilliant. But some of the best... There was, like, the small Amazon book that I picked up a long time ago, a bunch of things about Amazon. So many good nuggets of leadership and insight came out of that company.

(Colton Andrus at 00:17:29) Yeah.

(Joel Beasley at 00:17:30) I love it.

(Colton Andrus at 00:17:30) I'll give you two more of my favorites, and I don't know how public they are. Some of them came from, you know, people that were in the room that heard them. But plans have dates. Otherwise, you're really... It's just wishful thinking.

(Colton Andrus at 00:17:45) And one of my favorite ones is one time, people were talking about an issue they encountered, and Jeff said to them, this sounds like it's happening to you. You need to be happening to it.

(Joel Beasley at 00:18:00) Okay. I like that. That's very Jocko Willink extreme ownership.

(Colton Andrus at 00:18:07) Yeah. Yeah. Yeah. Well, I think sometimes, you know, sometimes we want to say, oh, it's hard. Oh, this thing happened.

(Colton Andrus at 00:18:14) Oh, this was out of my control. You know, what do I do? I did the best I could. And, you know, if you want to be top tier, you've got to find ways to get in front of that, to mitigate it, to go address that problem, to go fix it, to understand the underlying issues. So I've taken inspiration from that quote over the years.

(Joel Beasley at 00:18:33) Have you come across that book from Jocko?

(Colton Andrus at 00:18:36) Yeah. Yeah. Yeah. Actually...

(Joel Beasley at 00:18:38) I have...

(Colton Andrus at 00:18:39) I have it right here on my shelf.

(Joel Beasley at 00:18:41) Oh, me too. Yeah. My business partner in my... Maybe ten, fifteen years ago, my business partner at the time, he gave me this book, and he says, I think this is going to help you, like, long term. And I said, okay.

(Joel Beasley at 00:18:53) And it was just, you know, I wasn't taking ownership the way that I should have been. And I read that book. And when he said this sentence, and I'm paraphrasing, he said, the moment that you take control of it, like, the moment that you say this is my fault, you immediately are granted the power to create the change to fix it and to get it to stop happening. And I was like, that is fascinating and it has worked so... Doesn't work well in the marriage.

(Joel Beasley at 00:19:20) It works... You've got to be careful in the marriage with it.

(Colton Andrus at 00:19:23) You can take ownership of all the failures. And let me tell you what. There's some wisdom in that as well. When I first got married, one of the bits of advice we were given is you can always turn to the other person and say, you're probably right.

(Colton Andrus at 00:19:38) And you can say that sarcastically, genuinely, you know, as a bit of a joke. You're probably right, dear. And truthfully, you know, there's a lot of gray area in there. And if you just own that, you know, maybe the other person's right, maybe you got it wrong, maybe you made a mistake, that's also, I think, something we all have to learn. Something I had to learn.

(Colton Andrus at 00:19:58) There's a humility of, yeah. You know what? I'll own part of the problem. Maybe I'll own more than my fair share part of the problem so we can just go fix it and move forward because we're a team, and in the end, we've got a long journey to go.

(Colton Andrus at 00:20:12) So, yeah, I'm not going to get all bent out of shape about being right this one time if it's going to cost me, you know, our long term harmony.

(Joel Beasley at 00:20:21) That's right. Yeah. You can... What is it? I think my pastor at the time, he said to me, he goes, Joel, you can be right and lose the relationship.

(Joel Beasley at 00:20:28) And I was like, oh, it's not about being right. And so that was something I learned in my twenties. And then after I started more of the ownership stuff as you said about looking to the other person saying you're probably right, I learned there's an art to that too. You can't do it too quick. If I don't want to argue, I can't just...

(Joel Beasley at 00:20:46) Yeah. Yeah. You're right. Because it has... The art is finding how much resistance to put up.

(Joel Beasley at 00:20:52) You do it too quick, they're upset at you because they want to, like, hash it out.

(Colton Andrus at 00:20:55) Yeah. Yeah. Well, they know when you're just, like, laying down and don't want to fight over it. It's like, my wife, she's a redhead.

(Colton Andrus at 00:21:02) She likes a good fight, so it's a tough balance.

(Joel Beasley at 00:21:07) All right. Chaos engineering. You guys kind of coined the term or you came in right about the time this phrase was coming about because that's how I learned about the concept was through my first Gremlin interview. For people who don't know what chaos engineering is, can you just give the brief understanding of it and then why it's transitioned to proactive reliability?

(Colton Andrus at 00:21:26) Yeah. Chaos engineering is thoughtful planned experiments on systems to understand how failures impact them and impact the users. And one of the keys, I think, is really needs to be live running systems. One of the ways I like to think about it, there's unit tests, there's integration tests. We need distributed system tests.

(Colton Andrus at 00:21:48) What happens when this dependency of mine over here fails? What happens when I lose a host or a zone? What happens when my cloud provider has an outage? What happens when my database disappears or gets really slow? Those are the type of questions you want to answer.

(Colton Andrus at 00:22:03) I think the problem with chaos engineering is it's a bit of a misnomer. When people hear it, they think, oh, I have to do this chaotically. And there's a reason it was called chaos engineering. In the beginning, Netflix, as really an organizational approach, they were moving to the cloud.

(Colton Andrus at 00:22:21) Hosts were being rebooted from underneath them, replaced by their cloud provider, and their engineers had to learn to adapt to that environment. And so... And what I think was a brilliant idea, they said, great. We're going to make this happen to you in your staging, in your dev, and in your production environment. We're just going to remove hosts from underneath you.

(Colton Andrus at 00:22:40) And if you thought that host was going to last forever, you could put all your files, you could name it, you could treat it, you know, like a pet. It would be there forever. That's not going to happen. It's going to disappear from underneath you, and you're going to have to find and fix those issues. So that's where the chaotic piece came in.

(Colton Andrus at 00:22:56) But, truthfully, a lot of folks are scared about doing things in a chaotic fashion, in a random fashion. And what we really care about is engineering for the chaos. There's enough chaos in our systems. We want to make sure that we can handle that chaos. It's going to be there innately.

(Colton Andrus at 00:23:12) And so we want to treat it as an engineering discipline. We want to plan about it. We want to think about it. We want to do it in a way that mitigates the most risk. We don't ever want to go out and cause an outage.

(Colton Andrus at 00:23:24) We want to prevent outages. We don't want to cause customer pain. We want to prevent customer pain. And so that's really where the name becomes a bit of a distraction. People think that it's scarier than it is.

(Colton Andrus at 00:23:37) They think the approach must follow the way it's named. And that's really why I'm a much bigger fan of just talking about reliability or reliability engineering. That's what we... That's the outcome.

(Colton Andrus at 00:23:48) That's the goal is reliability. And the way we accomplish it is by engineering this set of tests and approaches that teach us about our system, help us understand the sharp edges and the side effects, and allow us to go mitigate them in a deterministic way or a nondeterministic way. That part, you know, there's multiple approaches, and it's not really material to whether or not you should do it or how to get the most value out of it.

(Joel Beasley at 00:24:16) You've been selling this for a while and working with a lot of customers for a decade on this. If I'm an engineering leader and things are running, you know, relatively smoothly, you know, management doesn't bother us too much, things are mostly up. And, like, how would they sell that to their team? Like, if they want to be better, how do they take it when it's not on fire and turn around and get their peers or their C level to buy into this?

(Colton Andrus at 00:24:43) Yeah. So, well, I think of my own experience as an engineer. When I was... First of all, I was on call, and I wanted to not get woken up at two in the morning. I had a bunch of young kids.

(Colton Andrus at 00:24:55) I liked my sleep. Life was hectic enough. How could I avoid that? And people, you know, even if you don't get paged, people are worried about going on call. They know it could happen.

(Colton Andrus at 00:25:06) So there's a bit of apprehension that comes from that. So I think it's about mitigating that risk in a way that gives you confidence. My first winter at Netflix, and they have a big peak around Christmas. A lot of people are off. They're able to stream.

(Colton Andrus at 00:25:21) They're sitting at home. They're catching up. They're binging their shows. And my VP comes to me, and he says, Colton, what's your confidence we're going to make it through the holiday season without having an outage? And I'm like, 25%.

(Colton Andrus at 00:25:34) Like, we're probably going to have an outage. And sure enough, we had a couple, and I spent part of Christmas day on a call dealing with some issues that had occurred. Well, we really brought this approach in that following year within Netflix. We did a lot of work, and we mitigated a lot of issues. We put a lot of solid testing in place.

(Colton Andrus at 00:25:53) Ultimately, that year, we went from three nines of uptime to four nines of uptime. So we reduced from eight and a half hours of outage over the course of the year to less than an hour, close to forty minutes of outage over the course of the whole year. So that next Christmas, my VP comes up to me and he says, Colton, what do you think the likelihood that we have an outage over the holiday is? And I'm like, 10%.

(Colton Andrus at 00:26:15) You know? I should say the percentage is the same way. 90% we're not going to have an outage. And we didn't have an outage that Christmas.

(Colton Andrus at 00:26:23) It was smooth sailing, and I was on call for part of the time, but I didn't have to get on a call, didn't have to, you know, deal with anything. Things went as expected. So I think there's a bit of that thought of what do you think might happen? What's the perception? And then what's the likelihood?

(Colton Andrus at 00:26:39) And how do you build that confidence so that you can turn it into things that aren't going to occur?

(Joel Beasley at 00:26:45) Netflix is up because Colton loves his family.

(Colton Andrus at 00:26:49) I love it. Yeah.

(Joel Beasley at 00:26:51) That's my takeaway.

(Colton Andrus at 00:26:52) That's better than me saying I'm a lazy engineer, and I just don't want to get woken up at night.

(Joel Beasley at 00:26:57) You know what, though? That's how we work. Like, you know, there's a phrase out there about how engineers... Like, the best engineers are the laziest. It's because I want to just automate things.

(Joel Beasley at 00:27:06) I just want it to be easy. I mean, I want my experiences to be easy with my interfaces, everything.

(Colton Andrus at 00:27:14) I think that's why we saw a lot of success early on with the DevOps approach. And, again, at Amazon, it was you build it, you own it, you operate it. And if you break it, you're going to fix it. And if you did a shoddy job, you're going to get woken up in the night. And so you had this incentive to go the extra mile and build that quality in because it was your time.

(Colton Andrus at 00:27:35) It was you on the line if things went wrong. And I guess a little soapbox here. I think we've kind of drifted back to the ops team deals with things. We're a little bit back to where we started from in many companies, where a bunch of engineers will work on the code, but the SRE team or the platform team or someone else will be the ones who get paged. And I think that's a bad alignment of incentives.

(Colton Andrus at 00:27:59) I think, you know, it's great if you don't... You know, there might be pros to that. There might be some efficiencies for consolidation. But when you know that the code you write, you might be called upon to fix in the middle of the night, or you might be on a call with a bunch of engineers whose other teams you impacted, their services went down because yours went down. You treat it a little differently, and you give it a little bit more thought, and you put in a little bit more effort to mitigate those issues.

(Joel Beasley at 00:28:28) You absolutely do. I mean, me personally, building software, I, you know, I was like, okay when, you know, when I was doing it. I wasn't doing best practices necessarily. I could make it run, but then I started to get revenue and I had to do that better. So I learned about testing and I learned about these other ways to make it run better.

(Joel Beasley at 00:28:48) And then I heard... I think it was Martin Fowler or Sandi Metz or someone, and they talked about, like, most of the coding you'll do is, like, rewriting your code. And when I heard them say that, it gave me a new perspective on how I crafted the code that I wrote and so I went in from I'm going to make this work mentality to I'm going to make this easy to make changes later when I know I'm going to have to make them. And when I did that, my... I went from, okay, I can make things work to, like, actually a pretty good software engineer.

(Joel Beasley at 00:29:19) So... Yeah.

(Colton Andrus at 00:29:21) And I think there's a lot of levels as well. We stand on the shoulders of everyone who's come before us. We didn't write the operating system, the assembly code. There's hundreds or thousands of libraries we depend upon. We make network calls out to all these other services.

(Colton Andrus at 00:29:36) There's all these moving pieces, and I think that's why the unit test, integration test, distributed system test is so important because we think a lot about our code and the unit tests we write. We think about maybe our direct dependencies and the integration tests we write, but then we deploy it into a system. We've got to worry about load balancers, security groups, traffic routing, Internet outages, cloud providers, host... There's just this plethora of things that happen all up and down and around us. And if we haven't accounted for those, we've probably made the wrong trade offs or the wrong assumptions in the software we built.

(Colton Andrus at 00:30:11) And I think that's part of what really facilitated the birth of this movement is, you know, I remember giving a talk and we talked about, well, why don't you just unit test everything that could go wrong? And the answer is the combinatorial explosion of testing everything, the sun would burn out before that test suite finished. So you simply cannot test all combinations of things and all variations. So what's the most effective approach? We'll take a real system and test the things you're...

(Colton Andrus at 00:30:40) That actually happen. Test the things you're most worried about and see what the side effects are. See how it behaves. And that way you're not guessing, you're not trying to foreordain all the things that might occur. You're really going out and you're seeing for yourself how it behaves, and you can fix those.

(Colton Andrus at 00:30:58) And it becomes a much more efficient approach to getting to a reliable production system than trying to guess everything that could go wrong.

(Joel Beasley at 00:31:07) You're absolutely right. And then people in their testing journey, they'll learn about unit test. And then, you know, it was funny because when I discovered testing, I thought I discovered it, right? But when I learned how to do it, I was like, oh, unit tests. And I was like, oh, unit tests suck.

(Joel Beasley at 00:31:22) I'm gonna just do integration tests so I can get to the point. And then I realized you kind of have to find this balance between working and creating unit tests while you're writing the functions and then doing the integration test. But then you guys got involved with that third one and you're calling it distributed systems test and that's the reliability engineering. Right?

(Colton Andrus at 00:31:41) Yeah.

(Joel Beasley at 00:31:41) Yeah. That's brilliant. And people know they have to do it, but it's like exercise. It's like writing tests. Sometimes they just don't do it.

(Joel Beasley at 00:31:49) But is this AI that you guys have now, this foresight, is the problem solved? Can I just press the button, dump all my production keys to it, and it just make sure that it's reliable? How does this work?

(Colton Andrus at 00:32:01) I mean, that's the vision. As much as you trust us, as much as we trust, you know, giving all your... You know, we don't... We actually don't want all your keys. That's not a good idea. But, yeah, that's the approach. You know, bring in Gremlin.

(Colton Andrus at 00:32:18) Now, you know, used to be an engineer would sit down. They'd whiteboard. They'd think about these failure scenarios. They'd do a failure mode analysis. They'd think about what could go wrong.

(Colton Andrus at 00:32:27) Let us just do that for you. We can look at your system. We know what good systems look like. We can look at a lot of the configuration. We can look at a lot of the pieces.

(Colton Andrus at 00:32:35) We can just discern from that what are the risks, what are the things that we think are likely to go wrong, come up with a set of tests we want to go run to exercise. And look, some things we can find without running a test. We see something that's blatantly misconfigured, we're done. Here's an answer.

(Colton Andrus at 00:32:51) Here's how to go fix it. Go mitigate that problem. No need to go put the system at risk or do extra work if you already know the answer. Then we have that set of experiments. Okay.

(Colton Andrus at 00:33:01) Here are things that we are concerned about. And I'd say a little sidebar here. You know, it's the same 10 things. The same 10 things that go wrong in computing.

(Colton Andrus at 00:33:11) You got CPU. You got memory. You got disk. You got networks. You got hosts.

(Colton Andrus at 00:33:16) You know, what happens if one of those goes away? What happens if one of those is maxed out? What happens if you're at capacity? That tends to be, from a black box perspective, how the system works. So go test those fundamentals.

(Colton Andrus at 00:33:30) Dependencies are my favorite to pick upon. What happens if your dependency fails or gets slow? And I think that's the low hanging fruit. That's the 80/20 for most engineers out there because you've probably done a pretty good job testing your system and testing its performance and its capabilities. But then somebody else's system surprises you in the middle of the night and behaves differently, and all of a sudden you gotta work out what happens if it goes away?

(Colton Andrus at 00:33:56) What happens if it's a second slower than it used to be?

(Joel Beasley at 00:33:59) Exactly. What happens when that API is not available right now because they're down for maintenance? You know? Yeah. You gotta have...

(Joel Beasley at 00:34:06) I wanna talk a little bit about the titles that you've held at the company throughout its history. So you were CEO, I think, in 2021, 2022, you stepped down, then you were CTO, and then you're like, no. No. I'm going back to CEO. Can you explain to me this beautiful journey you've been on within the company?

(Colton Andrus at 00:34:25) Yeah. Well, you know, so as we've discussed, engineer by trade, you know, technical person. But when it was time to found the company, I had the most expertise in the space. I cared deeply about it. I was the one out pitching the venture capitalists, pitching the early customers.

(Colton Andrus at 00:34:41) So it made sense for me to be CEO. And that was a journey, and there was a lot I had to learn along the way. Remember screwing up my first payroll and having to go fix it or understanding, you know, all the things you gotta learn around HR and recruiting, go to market, sales, marketing, all the different parts of the job. Did that, had a lot of success in the early days of the company, and then the pandemic hit. And the world got turned on its ear, and people went a little crazy.

(Colton Andrus at 00:35:09) And we had a lot of things to deal with. And, frankly, that was, you know, five, six years into the journey, and I was a bit burned out, and I was a bit frustrated. And I had some doubt. I thought, you know, maybe things would have been going better if I was a better CEO. Maybe somebody else could come in and do it better than I had.

(Colton Andrus at 00:35:29) Maybe somebody who had really... This had been their job longer or had learned more would be a better fit for the company. And I went out and found somebody that I had met, that I trusted, that I got to know, that I thought could come in and really help me out and brought him in as CEO, and I moved over to CTO. There were some niceties to that. Really helped me focus on the product, build the second evolution of the product, streamline the engineering team, make sure we were running a really efficient factory, not in today's AI sense, but with people, making sure we were just building software in an efficient way. And I think that was very useful, and it allowed me to get much closer to our customers and the problem and understand the direction the product needed to go.

(Colton Andrus at 00:36:14) And the CEO came in and did a great job running a variety of things. But along the way, I think what I learned is there's a certain passion and a certain commitment that a founder brings that you don't always get from an outside CEO. And, ultimately, we came to the point where we decided it was best to part ways. He wanted to go off and pursue other things. It was time for me to get back into the CEO role. And it was great because I got to watch someone else do it for a few years.

(Colton Andrus at 00:36:43) I got vindicated on some of my early choices. I learned a few things. I picked up a few skills along the way, and I had a little bit of time to recuperate and, you know, in the background, find a better work-life balance, take better care of myself, have a better balance of family time and work time so that I wasn't as burnt out. And that allowed me, when I came back as CEO, to really do a much better job and be able to keep that closeness to the product and the technology, but to be able to have the energy and the commitment to come in and drive the business forward and help us build the third iteration of the product and continue on the journey.

(Joel Beasley at 00:37:23) That's brilliant. So you found someone that you trusted that had the experience. You're like, here's the reins. You got some rest, got to focus on the new version of the product, the new iteration of it. And then at the same time, watch this expert person that you hired operate it, and then you could relate that like, oh, okay.

(Joel Beasley at 00:37:41) I could do this differently. I could do that different. I like what he does here. I don't like what he does there. And you can learn from that.

(Joel Beasley at 00:37:48) And then you went back into it refreshed, renewed, and with an understanding of, yeah, if I brought an expert in, place them in there, I can now measure myself against that and know where I sit in the stack because before then, you're just a first-time CEO. You don't really have a measurement like that. Yeah. Yeah. Good stuff, man.

(Joel Beasley at 00:38:09) I like you. You are my type of people. Burnout. Let's talk about a little burnout stuff because you obviously experienced that to some degree. A lot of founders don't like to talk about it openly.

(Joel Beasley at 00:38:23) Some do. But for you, what did burnout look like?

(Colton Andrus at 00:38:28) Yeah. When your phone buzzes, you panic. You know, I used to joke as CEO, there's a fire every day. There's always just some emergency, some problem you gotta deal with. Being grumpy at home, just being... just not being, you know, not being fun to be around, I think, was part of it. Being stressed a lot, you know, having some unhealthy habits, skipping the workouts, you know, finding other ways to decompress that weren't really the right long-term goals. And all of that, you know, really came to a head during the pandemic when you're stuck at home and you can't go anywhere. And you're just... you gotta sit there with your mind all day and worry about...

(Colton Andrus at 00:39:14) You know, and then you get stuck in your thoughts. You get stuck worried about all the things that are not going well or could be going better. You start reflecting on all the decisions you made that you think weren't the right decision or that led you down the wrong path or that you wish you'd handled better. And, yeah, ultimately, I had to learn... I had to find a better, healthy schedule.

(Colton Andrus at 00:39:35) And I get up every morning, I lift, I bike, I walk the dogs, I meditate, I have a little bit of spiritual time. You know, it's that opportunity to kinda set the tone for the day before I jump into things, helps me keep that larger picture. You know, when you're early in on it, you know, this is the most important thing in your life or it feels like the most important thing in your life. And I think level setting that, it's not the most important thing in my life. Most important thing in my life is my wife and my children and the long-term relationships with them.

(Colton Andrus at 00:40:09) This is an important, you know, thing that I built. I wanted to be successful. I wanted to be able to provide for my family. I wanted to be something that my children can learn from, but it's not ultimately the most important thing in my life. And so reorganizing some of those priorities and being able to put that into perspective.

(Colton Andrus at 00:40:29) And I still have to remind myself some days and put it into perspective when I'm focused on the mission. I'm a get it done, work hard, you know, plow through kinda guy and stubborn in many ways. And so, you know, some of that step back and remember the broader vision, you know, have a life worth living and a journey that is worthwhile along the way regardless of the outcome or regardless of how things are going that day are some of the things I've had to learn.

(Joel Beasley at 00:40:58) Yeah. I've worked quite a bit on the morning routine. I get up, and the first couple hours are just with my fam... We decided to homeschool the kids because...

(Colton Andrus at 00:41:08) We just wanted that when my kids were younger.

(Joel Beasley at 00:41:11) Did you?

(Colton Andrus at 00:41:12) Yeah.

(Joel Beasley at 00:41:12) Yeah. Yeah. We... They went to school, so they're nine, eight, and four. So we homeschooled them all the way up until, you know, a year ago.

(Joel Beasley at 00:41:22) Last year, they made friends in the neighborhood that went to school, so they wanted to. We didn't wanna fight them on it because, you know, if you fight them, they're just gonna resent you. So we said, okay. They made it through, like, two to three months into the school year, and they were like, take us out. We made them stay the whole school year.

(Joel Beasley at 00:41:39) Yeah. But they're not... They're homeschooled again, back in homeschool again this year, and they love it because they get to wake up with dad and mom and we have breakfast and we go on a walk. And those first two hours are just like us bonding as a tribe. And then I open up the door to the chaos.

(Joel Beasley at 00:41:55) I open up the email inbox and I get to work and start doing things. But I look forward to those moments the most, Colton. Like, this could all crumble, and I'm still gonna have waffles with my kids. I'm still gonna have that morning walk. Like, those are the most valuable things to me.

(Colton Andrus at 00:42:10) Yeah. Yeah. No. That's great to hear. I'm a big fan of that.

(Joel Beasley at 00:42:16) Let's see what else we got here. AI hype versus results in 2026. You had an article that you wrote, and you said it's gonna be a reality check year for AI. What do you mean by that?

(Colton Andrus at 00:42:28) Yeah. Well, you know, being a pessimist, looking at all the things that go wrong in systems, doing a lot of failure testing, you might not be surprised that I'm a little skeptical. And I think we've seen a lot of growth and progression in the AI space to the point that I become a believer to some degree. I don't know that I'm 100% in as some folks are, but I see a lot of value in the tooling. I see a lot of value in the approach.

(Colton Andrus at 00:42:54) But I think we are running headlong, you know, into the abyss here, and we're gonna have to learn some lessons along the way. And I think when we... It's the checks and balances that need to be developed in order to build, you know, stable systems that make it the distance. And on one hand, it's great. People are kinda treating AI... It's like the Trojan horse. They're just opening the door, letting it in. Let's see what happens. Maybe it's all good. And that's great if you wanna get people to change their behavior.

(Colton Andrus at 00:43:24) And that's one of the hard parts about reliability. It's back to the health analogies. It's like trying to get people to go to the gym. And if somebody doesn't wanna go to the gym, it's really hard to convince them that, hey, this is in your best interest in the long run. You don't wanna wait until you have a heart attack to decide you wanna get in better health. You'd like to be in good health and prevent that heart attack from ever occurring. But sometimes it takes those wake-up moments for people to really change their approach, change their behavior, understand the severity of it. And I think we're starting to see some of those with AI. I think there's some that are almost kind of funny.

(Colton Andrus at 00:44:01) Hey, AI, what happens if my database goes down? Well, I deleted your database. Let's find out. Well, hold on.

(Colton Andrus at 00:44:07) That's not what I wanted you to do. But then you have some of these more serious events where AI is taking the initiative and finding ways to do things that we haven't asked it to do or bypass our safety or security guardrails to be able to, you know, accomplish its task. So I think there's that angle. There's kinda malicious AI, but I think there's also just slop code. And, yeah, I think there's, you know, some good code being written.

(Colton Andrus at 00:44:35) As an engineer, I've got a little bit of hubris. I think, you know, a great engineer probably still writes better code, but doesn't write it as fast or in the quantity that we're seeing with AI. But we've got so much going out the door that we can't check it all. We can't watch it all. We don't know all the side effects of what's gonna occur.

(Colton Andrus at 00:44:55) And I think in... It's the unit test, integration test, distributed system test analogy. Again, in the unit test world, we can have a pretty good understanding of how it behaves. In the integration test, we've got an okay view. Now we throw it out the door into production, and it's the wild west.

(Colton Andrus at 00:45:11) And there's a thousand variables we haven't accounted for. What happens for that system? What is gonna happen when something it never, you know, envisioned occurs and it's not trained in the model, or it hasn't been accounted for in the way that it wrote that code. And so I think... And we're seeing this to some degree.

(Colton Andrus at 00:45:30) We're seeing outages go up, and we're seeing, you know, issues arise as opposed to go in the other direction. And I think it's gonna get a little worse before it gets better.

(Joel Beasley at 00:45:41) Yeah. It's... Well, I've been building like crazy the past year, and I've ran a couple different projects. I wanna get your take on this.

(Joel Beasley at 00:45:50) I've ran a couple different projects, and I had rules for each code base. Like, this code base, all the way down to the method, I'm just, like, having it help me write the specific functions. This code base, I'll give it basically, like, an architectural layout of what I want. I'll review the plan and hit build.

(Joel Beasley at 00:46:10) And then this third code base, I'm just going to talk to it, let it build in the cloud, never look at the code and then see what happens. And I've been running all three of these products in production. And I will tell you what, man, it is surprising to me where I've landed just naturally over the past nine months.

(Joel Beasley at 00:46:32) I landed in the whole... Just the project that I think is the best is the one where I just review the plan and I'll catch little things about like how it's associating models or something that it mentions. I'm like, oh, if you do that, I know that's a mistake. But all the projects are running smoothly. And it's kind of like, I'm not suggesting this to people.

(Joel Beasley at 00:46:54) I want to be really clear. I am playing on very... There is like... There's like no repercussions if these projects explode.

(Joel Beasley at 00:47:01) This is very shooting from the hip, having fun, but I am pleasantly surprised at how smoothly things are going and how much I'm starting to trust it now.

(Colton Andrus at 00:47:12) Yeah. What a fascinating experiment. I'd love to hear how that goes. And I think... But the question is then, you know, is it...

(Colton Andrus at 00:47:21) Has it been put through the ringer? Is it facing a lot of duress? Is it dealing with, you know, spikes and load? Is it dealing with... Does it have a lot of dependencies?

(Colton Andrus at 00:47:31) What happens if those dependencies start failing or going down? Have you tested one... You know, have you... Has one of those survived a cloud outage? Or, you know, has...

(Colton Andrus at 00:47:40) Have they... Have you lost an availability zone or a region yet? Those would be the things I'm curious about because I think it's pretty good at the happy path. I think it does a pretty good job for when things go as expected, better than I would have thought. And we trust it quite a bit more than I would have thought I would a year ago.

(Colton Andrus at 00:47:59) I know. So I'm with you on that front. I just, yeah, I don't know. The engineer in me has taught me to keep a little bit of skepticism.

(Joel Beasley at 00:48:06) Oh, I'm skeptical. I'm... I think we're in a very similar path. Like, it's gained more trust than I'm comfortable.

(Joel Beasley at 00:48:14) Like, it shouldn't have been able to gain that much trust from me. But it's also... I would say my takeaway from the entire experience, and I want to get your thoughts on this too because you're running a large engineering team right now. I am not. I'm starting to see this shift when I'm watching myself.

(Joel Beasley at 00:48:30) My wife's the main user. So I build like applications to manage my touring schedule. I build like all these things. So she's the user primarily. So I interact with her on this.

(Joel Beasley at 00:48:40) But what I have noticed is running these four unique applications and them being very feature rich. What I think is now going to happen in the future is that you'll have this role that's like feature owner and or maybe you own two features, but you know everything about them, how they relate to the other features, all the code historically, how it's written. You're probably using multiple cursor agents to manage it, but it's the act of like loading like an operating system loads into memory. It's like loading that feature and how it interacts with the world and the other features into your mind and then interacting with someone else to process a request like another engineering team or someone from the business world. And so when I have started to imagine like how this is going to affect engineering at scale, I...

(Joel Beasley at 00:49:29) That's where I'm leaning to. That's what I'm seeing with myself is my biggest problem, Colton, is forgetting how features interact and work with each other in order to even tell it what I want it to do.

(Colton Andrus at 00:49:41) Yeah. Well, I think one of the hardest parts about software engineering has always been really good program project management. You have to really understand all these details. And the best program managers are technical because they can't just understand what you hope happens. You have to understand a lot of the underlying details in order to build it in a way that works and is...

(Colton Andrus at 00:50:04) Works well with other systems and in these edge cases. I think it's a bit funny because what we're coming back to is the people that are well written and the people that are good at articulating what they want, I think, are going to excel in this environment. And this idea that you can give AI, you know, a one sentence description of something and get a good output, I think, is where the confusion lies. Some people feel, oh, you know, and, you know, if you need something that's fairly well understood, maybe you can just give it a sentence or two. But if you're building enterprise software and you need it to work in, you know, very specific ways in certain environments, have you documented those?

(Colton Andrus at 00:50:46) Do you understand those? Have you told the AI the trade offs? What the boundaries are? What you're willing to accept? Because there's not a lot of black and white answers out there.

(Colton Andrus at 00:50:55) There's always a lot of gray. You've got to muddle through and figure out, well, what should it do in this situation? So I think, yeah, that ability to articulate it well and that ability to write it down and to document it. Or if you can, keep it all in your head and just never forget any of the subtleties or the little pieces. If you can do that, props to you.

(Colton Andrus at 00:51:17) You've got a better memory than me. I've got a pretty good memory.

(Joel Beasley at 00:51:20) My memory is not right.

(Colton Andrus at 00:51:21) Little details.

(Joel Beasley at 00:51:22) Yeah. Well, that's where I ran into, like, in this tour management software, there's different types of tours, like, there's one where I'm taking the car, there's one where I'm flying and renting a car, there's one where I'm in the tour bus, and so there... And then there's different rule sets under each and just remembering just the logic of how we forget software, just you and me, Colton, just remembering the logic on how we make these decisions and then to be able to ask it to change or improve or how it would handle. So that is going to stay human, I think, for a long time because as long as the purpose of the software is to serve the humans, it's serving some function for the humans. There has to be some connection between the human that has the need and the software and then the memory of the history of how it came to be.

(Colton Andrus at 00:52:08) Yeah. Well, I'm not a doomer as some folks are. I think that it's a great tool, and it will allow us to do a lot more as previous tools and technological innovations have allowed. But there will always be things that are valuable for people to do, for people to understand.

(Colton Andrus at 00:52:27) And if anything, it feels like we've all gotten busier in the last couple of years, not less busy. So it doesn't feel to me like, you know, this world where the bots do everything and we sit on a beach is anywhere close to reality.

(Joel Beasley at 00:52:40) No. I think it's the opposite. I'm doing more work now, like, in... As far as, like, output of efficiency. I'm creating more efficiency now than ever with my life because I'm able to build these tools that drastically improve it.

(Joel Beasley at 00:52:56) And there's all sorts of reports coming out about, like, decline of companies purchasing different SaaS softwares because they're building it themselves. And that's one of the edges I think Gremlin has as far as being a strong product long term is one of the biggest values you guys have is this decade of data of these experiments being run that you can see all of this information so you can be the best version of reliability. People are going to plug into you and not think about it. They're going to ask their LLM, hey. Who do I...

(Joel Beasley at 00:53:26) This has gone down three times. The LLM is going to say you're going to need a service like Gremlin, and then they're going to MCP into it, and then it's just going to run it for them. Like, that's the future I see. And you can tell me I'm wrong because I'm not in it every day with you. But...

(Colton Andrus at 00:53:37) Well, I think it's interesting. I was with a group of CIOs this summer, and they did a poll, you know, who's building versus who's buying. And the answer was not as lopsided as you would think it is. And I think in part, generating the code is half the battle, but you still have to own it, operate it, improve it, maintain it, make sure that it works the way you expect.

(Colton Andrus at 00:53:58) And that's still time and effort that has to go into that. And so if it were just the code, I think open source would have killed all these companies a decade ago, and it hasn't because you still have to have all those other concerns. We also see all of these things. There's this term I learned called citizen engineer, and it's when one of your sales guys starts writing code or your marketing people start writing software systems. And they're great at the POC phase, they build it, it kind of works, and then they're like, hey, engineering, you own this now, figure it out.

(Colton Andrus at 00:54:29) And that's where, again, rubber meets road. You've got to go figure out how to make it effective. You've got to go figure out how to make it efficient and scalable. You've got to maintain it. You've got to update it.

(Colton Andrus at 00:54:40) You've got to go do all the features that come. So I think there's a lot of value we get in getting to that point, but I think there's still a lot of value in understanding the software well. And then there's the people side. As we were discussing earlier, it's not just the software. You also have to go get the people to change their behavior and do the right things.

(Colton Andrus at 00:55:01) And one answer is have the AI do it for you, and great, that'll solve some percentage of the cases. But how you actually get the organizations to adopt and change their behavior, that requires some real expertise and guidance as well. So, yeah, I think... Look. Like every software business owner, I'm concerned.

(Colton Andrus at 00:55:20) You know? I think you've got to adapt. You've got to understand what happens. But having a lot of folks on my team that have owned and operated these systems, that have real experience dealing with outages and failures, I think is a real advantage. Having a decade of data about how the systems fail that the LLMs haven't gotten to yet, I think that's a valuable part of the equation.

(Colton Andrus at 00:55:42) And, ultimately, if we are continuing to find how to use technology to adapt and stay on the forefront, then I think that we will be providing better results, and there will always be a home. Or at least for the foreseeable five or ten years, there'll be a home for companies that are providing better than average, better than what you get out of the model if you just throw some junk at it and see what comes out the other end.

(Joel Beasley at 00:56:06) A hundred percent. Because we're already at... A lot of us are already at capacity. We're just like, who's the person who's been doing this for a decade who I can offload this to, who I can trust to make sure this goes right? That's what we're looking for.

(Joel Beasley at 00:56:20) Gremlin, what's the website?

(Colton Andrus at 00:56:22) Gremlin.com. The new foresight stuff lives under gremlin.ai. But, you know, we got we got the six letter domains over here.

(Joel Beasley at 00:56:31) I love it. As we start to wrap up, I like to ask people, what's the best piece of leadership advice that you ever... Like, you heard it, you saw it, and then you tried it, and then it stuck with you for a long period of time.

(Colton Andrus at 00:56:45) We've kind of touched on some of this because I think it's around ownership. You know? Own the problems. Understand them. Go fix them.

(Colton Andrus at 00:56:54) Go address them. I think that's part of Amazon's core culture, you know, ownership and customer focus. Are you solving real problems? Do you understand the pain? And can you go out and really improve people's lives and make things better for them?

(Joel Beasley at 00:57:08) Beautiful.

(Joel Beasley at 00:57:11) Thank you so much for listening. And if you found this episode useful, please share it with a friend or colleague who you think would get value from it. And if you have topics that you'd like to hear discussed on the podcast, either add me on LinkedIn or send me an email [email protected]. Every time I get an email or LinkedIn message, it absolutely makes my day and inspires me to keep going.