Episode 616 ·
Navigating the PR-Engineering Divide & Building Trust for Success with Nora Jones, Founder, and CEO at Jeli.io
Today we’re talking to Nora Jones, Founder, and CEO at Jeli.io. We discuss the awkward nature of company incidents; the gap between what PR says and engineers have to fix in the fallout of a company incident; and why trust is the bedrock for growth.
All of this right here, right now, on the Modern CTO Podcast!
For more about Jeli, check out their website: https://www.jeli.io/
Produced by ProSeries Media.

About Nora Jones:
Dedicated and driven technology leader and software engineer with a passion for people and reliable software, as well as the intersection between those two worlds.
I truly believe that safety is pivotal with software development nowadays. If we don’t think about the safety and reliability of the product we are building for the humans that use it everyday (from the inception of our idea), then we are presenting a giant problem for humanity. From this passion, I decided to pursue a part-time, advanced degree in Human Factors and Systems Safety; which provides a unique complement to my understanding of how complex software systems fail.
Learning more about how other industries approach incident analysis, I thought a lot about how the software industry could learn from them -- and ultimately founded Jeli.io, the first incident analysis platform.
I've had the fantastic opportunity to co-write the book on Chaos Engineering, and how a product’s availability can be improved through intentional failure experimentation (available for free through O'Reilly: http://www.oreilly.com/webops-perf/free/chaos-engineering.csp).
I've also had the opportunity to share my experiences helping organizations large and small, reach crucial availability. In November of 2017 I keynoted at AWS re:Invent to share these experiences with an audience of ~40,000 people. Since then I have keynoted at several other conferences throughout the world highlighting my work on topics such as: Chaos Engineering, Human Factors, Site Reliability, and more.
Reach out to me if you’d like to talk about software, availability, how cultures impact products developed, or English Bulldogs (please leave me an intro when connecting).
About Jeli:
Jeli.io is the first dedicated incident analysis platform that combines more comprehensive data to deliver more proactive solutions and identify problems.
Jeli is the only platform that goes beyond technical errors and collects data on contributors that are so-often overlooked, like team coordination and human factors. Let’s face it—we’re not all people-people. We’re software-people.
And Jeli is a people-software.
We work with the tools you already use. If you're ready to see more insights from your incidents, reach out. We want to help.
Transcript
(Intro Narrator at 00:00:00) Today, we're talking to Nora from Jeli about the nature of company incidents and her journey as a company founder. You're listening to Joel Beasley, Modern CTO.
(Joel Beasley at 00:00:15) I read all about you, and I've got so many questions.
(Nora at 00:00:18) Cool.
(Joel Beasley at 00:00:19) The first one being, tell me about this railgun you worked on in the Navy.
(Nora at 00:00:24) So that was right as I was graduating college, and it was an undergraduate research project. I had a period of time where I kind of thought I was going to go into the Navy. But basically, what they did was they took someone from every type of engineering. So they had someone from computer engineering, which was me, someone from electrical. They had an ocean engineer. So one of every kind, and we were all designing this reduced-scale electromagnetic railgun. And essentially, we were trying to fire it in a test environment to prepare them for the real environment. And so that kind of got kick-started in my career around, you know, test versus production, how incidents occur, what it's like to work with other disciplines that don't understand things the same way you do. And so it kind of all relates to what I'm doing now with Jeli, but it was a very cool project. We actually got to fire it on the drill field at my school, so that was a lot of fun.
(Joel Beasley at 00:01:27) So you successfully got it to work?
(Nora at 00:01:29) Yeah. Yeah. We successfully got the reduced-scale version of it to work and then reported the findings back to naval officials and the university officials too.
(Joel Beasley at 00:01:39) How big was this thing physically? Was it three feet, five feet?
(Nora at 00:01:43) I am honestly trying to remember now. It's hard to recount all of it, and we also, you know, we're not allowed to talk about all the aspects of the thing we were working on. But yeah, the actual ones are much bigger. But the one we were working on was a lot smaller of a version.
(Joel Beasley at 00:02:03) That is so cool. Yeah. It's always fun doing it.
(Nora at 00:02:06) Really cool.
(Joel Beasley at 00:02:08) Second thing I was wondering about you from a personal level is, how do you handle the fact that you will never win the SEO game against Norah Jones?
(Nora at 00:02:19) Honestly, it's kind of my life mission to win the SEO game against Norah Jones. I think if you type Nora Jones plus chaos engineering or incidents, there was one day where I was kind of ahead of her. That was after my re:Invent keynote, but I'm still constantly chasing her. She has some sort of YouTube video where she's talking about incidents. I'm like, you can't give me just this one thing? You have all this other stuff. Like, let me be the Nora Jones that's associated with incidents.
(Joel Beasley at 00:02:51) Well, Josh will make sure in your episode page to make sure the anchor is your name only. Right?
(Nora at 00:02:57) Yeah.
(Joel Beasley at 00:02:58) So you started this company. It's called Jeli. Where did that name come from?
(Nora at 00:03:02) Yeah. So we backronymed it to "Jointly Everyone Learns from Incidents." We're a big learning-from-incidents platform and incident management platform in general. But, you know, if you want the real story, it was four letters. We own the domain. It was pretty cute, and it's easy to remember. So that was kind of where it came from. And it has this nice acronym. We internally call ourselves jelly beans, and, you know, we do a lot of jelly bean puns with customers. We actually made dog toys as swag recently that are these little peanut butter and jelly dog toys, which is really great. But yeah.
(Joel Beasley at 00:03:39) I love that. Everybody loves the dog stuff.
(Nora at 00:03:41) Yeah. Totally.
(Joel Beasley at 00:03:43) So you started this company, but you worked at other big companies. So you did the railgun, you did some stuff at Slack, I think it was, some other big companies. Was the incident response there that you were working on as well?
(Nora at 00:03:56) Yeah. So after the railgun, and the railgun, you know, was us developing a test environment and seeing how well tests can match production, if we could find issues there before it turned into production. I was really focused on hardware at that point. I actually thought I was going to go into hardware. And I went to a company called Alarm.com, which is a home automation and security company. So I was working on the hardware side there on the quality side, and I was working on, you know, making sure all these different components working together didn't cause an incident, didn't cause someone to break into your house or anything like that. Yeah. I've always just kind of been interested in incidents. But after that job, I got recruited by a company called Jet.com, which I'm sure—I think everyone in the world received a little purple flyer in their mailbox saying, you know, we're cheaper than Amazon. Right? And so I was really excited about that mission. Mark Lore was kind of bringing Costco to the internet, and they had done more marketing spend on Google Ads than anyone in the world at that particular point, but they hadn't hired any SREs or anyone to manage incidents yet. And so they were looking for someone that could help with that. And so I actually ended up being the first hire in the developer productivity area that allowed us to kind of move fast but make sure we were not delivering a bad experience every time someone showed up on the site. And what I realized—and it was my first software job—but what I realized was how much the philosophy is carried over between hardware and software and other industries. You know, there's a lot of organizational psychology involved in making your sites reliable and easy to use for both your customers and employees, so that you're spending your money efficiently and effectively. You're not having to overhire. And so I started getting really into not only the software of it, but the organizational psychology of it, which is kind of where that interest really sparked. Jet ended up getting acquired by Walmart. I went to Netflix to do chaos engineering and incident response and analysis work. And then after that, I did a short stint at Slack around the time of their IPO. And so I was, you know, I had this pattern of getting hired by companies right when they're having these massive scale challenges. Right. So they're having a lot of customers that are complaining that their website is going down and a lot of employees that it's also impacting. And so I would kind of get hired as this go-between, between helping everyone figure out organizationally what was going on from an engineering perspective and then also from a PR and comms perspective.
(Joel Beasley at 00:06:39) You can tell too, because when I went on your home page and was looking at the screenshots and the features and everything, you can tell someone who has a ton of experience with how humans communicate made this software. Because I've seen a ton of incident response softwares and ticket systems and, you know, all that type of stuff. And we've talked about chaos engineering on the show, and I've seen products that do all of that, but I've never seen one that got the communication between humans as beautiful as what you guys have done. And so I said, okay, I usually don't talk a lot about the products on the show because people can't necessarily see it, but it was enough to stand out.
(Nora at 00:07:17) Thank you. I really appreciate that. It's certainly an aim of ours. I, you know, I've been a developer my whole career, but I've always—there's so much of a human component to what we're shipping and an organizational psychology component. And I've seen all these tools on the market that try to automate everything and, you know, reduce your incidents and reduce everything, but they're not really focusing on what makes things hard for the humans, which is where a lot of your answers are lying. And so I wanted—I mean, Jeli is essentially a tool I wished I had had when I was in these roles to help all the departments kind of understand how they played a role in the incident and how it affected each other so that they can be better in the future. I mean, incidents, especially highly emotional ones, can be so awkward internally afterwards. Right? You know, the company will put out a public statement, but everything is kind of emotional and on fire internally. And someone's feeling really bad, and someone, you know, is trying to make them feel better, but they're also upset at what they did. And it ends up being a whole thing, and you end up losing a lot of the data because it's so emotional and awkward. And so our tool is really trying to help focus on the facts and just like, what happened between this communication? What was hard for folks? Why I had to bring Joel in when Joel wasn't on call and he hasn't worked on the system in five years? Like, why was he the expert here? You know, there's a lot of data in that that I don't see a lot of companies unpack. And I think it's mostly because they don't know how. And so we're trying to give that—we're trying to make that accessible to more folks so that they can make those conversations happen and make them happen faster.
(Joel Beasley at 00:09:07) And because you have this experience working at these companies scaling, and I'd say chaos engineering is a newer concept. At least it wasn't around as obviously or as—wasn't called chaos engineering, you know, fifteen, twenty years ago when I started. Right? But it's definitely been emerging. And so you've taken your learnings from these different companies. And at what point did you think, okay, I've seen enough that I can go create my own thing?
(Nora at 00:09:39) It's a really good question, Joel. I think the seeds of this idea really started forming when I was leaving Jet. But, you know, I went to Netflix after that, which was a very different type of company, and I got hired on the chaos engineering team. And we had built this really cool platform organizationally that could inject failure into production systems without causing customer pain. And it was really cool from a technical perspective, but the problem was no one was really using it internally besides me and my teammates that had made the tool. And so I was like, what value are we getting out of this if the people that are using these platforms aren't exactly learning from these failures? And so I started really looking at past incidents when I was at Netflix, but I was doing that to try to drive more traffic towards my chaos platform. Like, I was trying to find patterns. So say I was trying to get the search team to use the chaos platform. I would try to find patterns in search incidents because they could relate to that. Like, oh, you've had this incident here with search recently. Let's run an experiment on it to make sure we're still good there. But I realized there was so much more data in these incident patterns besides just trying to drive traffic to my chaos platform. Like, I realized it could be used to inform headcount. I realized it could be used to inform build-versus-buy decisions. I realized it could be used to inform roadmaps and financial things. And incidents can actually be a mirror into what is hard in our organizations and what is helping us ship fast versus helping us ship slow. And so I was like, wow, there is a huge potential for a tool like this, but I also wanted to make sure it wasn't just a Netflix thing I was seeing. And so I actually sent out a tweet because I also wanted to build more community around incident patterns and learning from incidents and looking at the coordination and cognition problems behind them. And I tweeted back in, like, 2017, 2018. Like, hey, is anyone else thinking about this in their organization? And I see other industries doing stuff like this. I feel like tech is kind of behind in how we think about incidents and how we use them to improve. And I got like a hundred random DMs just in that night. And I formed a little Slack community called the Learning from Incidents community, where we all started talking about this. And it was companies from all over the place. Like, I was seeing startups. I was seeing Salesforce. I was seeing IBM, Stripe, all kinds of different organizations where someone was like, there has to be a better way, and we can use this to influence so much of our organization. And then, you know, a lot of how we communicate during incidents happens over Zoom, and it also happens over Slack. And when I was at Netflix, it was really happening over Slack. So what I would do is I would manually take the incident channel in Slack, and I would look at who was talking in there, how long they were talking, when we needed to bring other people in to fix things. And then I would correlate that with what was happening in PagerDuty, and I would correlate that with past incidents as well. And I was doing that manually. And so when Slack offered me a position, I was like, wow, that would be great. You know, I could build something maybe on top of Slack, but I realized it was so much bigger than just Slack. And so that was kind of when I left to go form Jeli. But our initial product was built on top of how people communicated in Slack.
(Joel Beasley at 00:13:02) That is so cool. And then because it's emerging, do you find yourself—you've got all this knowledge, right, on incident response and management.
(Nora at 00:13:08) Mm-hmm.
(Joel Beasley at 00:13:09) The community, is it still alive? I know you have the conference now, but—
(Nora at 00:13:13) Yeah. The community is very much alive. And the conference was built to open-source what was happening in that community because a hundred strangers basically got together. Like, I knew a few of them, but we've been—now we're friends, and it's like a real community. It's like over four hundred people at this point. And I try to keep it intentionally small because people are talking about their incidents that are happening in their companies within this community. And so it's sensitive subjects. Like, I don't, you know, I don't want to invite anyone in that's just going to lurk and not contribute to the conversation because that makes people less likely to feel safe to share what's happening in their organizations. But so much of what we experienced through the tech industry is the same incident, just in a different flavor. Like, it's not the technologies themselves that are triggering these incidents. It's kind of how we're organizing around them, how we're talking about them, our attitudes towards these technologies. And so there's a lot we can learn from each other. But that community is still very much vibrant. The reason we ran the conference is to open it up beyond the community, to invite other folks in, to hear what we've been learning and implementing through the past three or four years with this community and open it up a little bit broader so that folks that want to get interested in it and maybe join it themselves after trying a few things can do that too. But, yeah, I hope we get to keep running it for many years. It was an absolute blast.
(Joel Beasley at 00:14:42) What was the most powerful takeaway from the conference?
(Nora at 00:14:46) I think there's a lot of interest in thematic analysis right now. So what patterns are happening across our incidents over time in our organizations. You know, I think there is a lot of thematic analysis you could just, you know, do and blanket across every company. But every company is a little bit different. Like, hey. How have our incidents been with this particular system when we had one person in charge of it? And just seeing how that changes over time when now we have two people in charge of it, and we have those two people in different time zones. And, like, is it still critical to our system? Like, if this system goes down, does our business go down? And just seeing how that maps out over time and how seeing how various changes we make to our organization impact the incidents internally over time. So I feel like that was a big theme. And then one of my favorite things that we did—and all these talks are recorded except what we call the hallway track. In the hallway track, we had people sharing actual incidents from their organizations that weren't allowed to be recorded. And so we had, like, probably fourteen talks or so in there that were really wonderful and sparked a lot of great conversation because it was so emergent.
(Joel Beasley at 00:16:06) Did you have Southwest there?
(Nora at 00:16:08) Maybe next year. Yeah.
(Joel Beasley at 00:16:12) They're not doing a lot of media right now. We reached out to them, and we're talking to them this week. And they said, well, reach back out in two months and we'll talk. And it's always super PR driven when they do it because whenever an emergency happens, everything then gets scrutinized—every media appearance, everything that you do.
(Joel Beasley at 00:16:31) Because, you know, they're a publicly traded company. But did you learn anything from watching that happen publicly?
(Nora at 00:16:36) Yeah. I mean, I was watching it really closely, and it was fascinating. I actually know some folks that were front desk agents on that day, so I got a chance to kind of talk with them as well. But it was an amalgamation of things that happened, and it was things that were blatant in their system for a long time. There was a lot of manual updating in their system for their point-to-point flights. Right? Like, our flight attendants and our pilots are supposed to be at this place at this time, and they're not. You know? And they're not manually updating where they are. And so the system just kind of slowly collapsed.
(Nora at 00:17:12) And I think, you know, there's a lot of—I imagine things are really hard there organizationally right now. And some of what we do is we help understand why things were hard for them and how things were hard for them in that moment. And I don't fully see a lot of their media capturing how things were hard, and that's where they can really get the most learning. But, you know, like you said, you know, if someone's outward facing, it's more of kind of a safe face campaign. Right?
(Nora at 00:17:45) Like, we're going to be doing a—
(Joel Beasley at 00:17:46) Lot of emotional campaign.
(Nora at 00:17:47) It's really emotional. Yeah. And, you know, after big incidents, I really encourage companies to spend time analyzing small incidents that don't hit the news. Because when these big incidents hit the news, now not only is engineering involved, but PR is involved. And PR has to make a statement faster than engineering can diagnose what happened, and it drives this huge wedge.
(Nora at 00:18:11) And then PR ends up committing things, or customer service ends up committing things, that engineering now has to do but maybe are not as feasible anymore. And this happens all across the industry with a lot of different companies. But if we're practicing these small events sooner—like maybe incidents that don't quite hit customers but they interrupt someone's day internally—we can actually prepare ourselves for these big ones a little bit more.
(Joel Beasley at 00:18:38) Yeah. Because it's emerging as a field of its own, do you find yourself, you know, being a tool builder here, having to educate people during the purchase process?
(Nora at 00:18:50) It depends. You know? We have a lot of customers that have committed SLAs to their customers, and so they get it. Like, they know that they have incidents, and they're calling them, and they're reviewing them all the time.
(Nora at 00:19:06) And then we have customers that are having a lot of incidents, but they're not taking the time to review all of them. You know? They're constantly in firefighting mode, and so we're trying to help shift them from this reactive to proactive mode where they're doing even little reviews for their incidents afterwards that allow people to share and understand their perspectives. So there is some education we're doing. You know, myself and about five or six people on the team have been at large organizations where we've done this change management process in terms of getting people to think about incidents differently internally.
(Nora at 00:19:43) And so we're baking a lot of that thinking into the tool too. So we do a lot of education through the product itself.
(Joel Beasley at 00:19:50) Yeah. Do you love what you do?
(Nora at 00:19:52) So much. Yeah. You know, I think I've always wanted to have my own company and build an organization—like, build the organization I've always wanted to work in—but this is my passion. Like, I am really passionate about learning from things that seemingly went wrong to learn how we can be better, to learn how we can work better together.
(Nora at 00:20:15) I'm really passionate about language after incidents. We bake language into our product as well. Like, you know, instead of the word incident, you'll see the word opportunity everywhere in our product. And little nuances like that help people think of this not as a checkbox exercise after something happens, but more of a journey that they have to go on and figure something out. And so they have a little bit more fun with it.
(Nora at 00:20:40) They're a little bit more engaged. And I just love seeing that light bulb moment go off for people when they're like, "Oh, I'm not writing my incident report that it happened. I am writing it to learn from it." And that's just—it's why I do what I do. It helps people like their jobs better.
(Nora at 00:20:59) It helps them stay at organizations longer because they feel like they're being challenged and they're learning more rather than just kind of checkbox filing every time. So, yes—long answer. I love what I do.
(Joel Beasley at 00:21:13) Yeah. I liked when you were talking about thematic analysis because—
(Nora at 00:21:17) Mhmm.
(Joel Beasley at 00:21:17) You know, as an entrepreneur, you're always looking at what you need to invest in or you need to spend time and money, what systems you need to improve, whether it's a system in sales or in marketing or in hiring or, you know, in the product itself. And that would give you—if you have trends across your incidents for the course of a year or two years or something—you can have who knows what will be in there. I mean, have you seen something come up that's actionable that really helps people make investments? Can you give me an example? Like, anonymize it a little bit.
(Nora at 00:21:49) Yeah. So we're working with this one company that had a particular cluster that had gone down, and it really impacted this organization. And, you know, when that kind of stuff happens, when it really impacts an organization, it's an incident that is going on for like three days or so. Obviously, it gets the attention of upper leadership. Right?
(Nora at 00:22:11) And upper leadership can only see a high-level version of what's happening. Right? You know, the individual contributor employees have their opinions on what's happening, and their opinions on what's happening—it's not this cluster that's having issues. It's our attitudes towards this cluster. It's the way that we're formatting the schema around this cluster.
(Nora at 00:22:31) But this VP was kind of seeing, "Wow, this cluster is in every incident. It's always a huge pain. We need to rip and replace it with something else." And, you know, what I see a lot of the time is individual contributors will know that that's not the right approach, but they only have so many hills they can die on. They only have so many battles they can fight.
(Nora at 00:22:50) And so, you know, they might spend some time trying to convince this VP, "Hey, this is not something that, you know, should happen. We should do this instead." And what we do is we help tell that story a little bit better, and we help them mark and tag, "Hey, here's where our attitudes towards this cluster were wrong."
(Nora at 00:23:09) "Here's what was hard about understanding how this cluster worked. We were treating it like a SQL database. It was really not a SQL database." You know? And so, like, helping them show that story to that VP really short-circuits those conversations.
(Nora at 00:23:25) So we're saving folks, like, you know, eighteen months of ripping and replacing and having to hire a whole new team to understand a new technology, when really it's educating the current team on how to use this piece of software a little bit differently. And so that was a really, really powerful one. I saw recently that they were able to use past incidents to kind of group together and package up that narrative for someone that wasn't experiencing it as on the ground as they were.
(Joel Beasley at 00:23:57) And then what's your sweet spot as far as customer? Fortune 500s, startups?
(Nora at 00:24:03) We have both, honestly. And, you know, I'll tell you we have two different products. Right? We have an incident analysis product, which helps them create a timeline efficiently. Once they do that with an individual incident, it gets grouped into the thematic analysis.
(Nora at 00:24:18) And, you know, I want to touch back on the thematic analysis because I see a lot of companies trying to do it. Right? People are going to do it anyway. They're going to find themes behind their incidents. But the problem is when you're putting garbage in on an individual level, you're going to get garbage out.
(Nora at 00:24:32) And so if you're making decisions based on themes from your incidents, but you're not actually spending time to review them on an individual level, you're going to be making the wrong decisions for your organization. So we help them make those individual incident reviews efficient and effective and focused on the coordination and collaboration like we were talking about earlier. All companies like that, but, you know, we see a lot of success with the really large Fortune 500 organizations that have so many people and so many different viewpoints, and it's really important to help keep everyone aligned and on the same page after incidents, especially when they're globally distributed, when they're really distributed by tenure. That's something we bubble up in the tool too—like, how long has this person been at the organization? Have they been on call for the system before? And how those different dynamics play out?
(Nora at 00:25:25) And so our larger organizations really like that. But our smaller organizations—and these are smaller startups and such—really like our free incident response bot. And so this incident response bot is built right on top of Slack, and it essentially allows the responder to do what they do best, which is respond. There's so many different aspects during an incident where I need to keep in touch with my customers. I need to keep in touch with my CEO.
(Nora at 00:25:51) I need to keep in touch with sales to let everybody know what's going on, but I'm also trying to fix the thing at the same time and keep track of how much time has gone by. And it's a lot of cognitive load for that person. And so what our bot does is it actually helps them keep track of those communication mechanisms and also how much time has passed and easily broadcasting things out for people. So I see a lot of success from our smaller organizations on that. And then once they do a few incidents with our incident response bot, we now have patterns for them to look at.
(Nora at 00:26:24) Right? We now have themes for them to look at themselves. And so that's when I see them start getting interested in some of the incident analysis platform.
(Joel Beasley at 00:26:32) So if you could have—I want to try to get it to where the audience thinks, "Like, oh, that's me." Your ideal customer. Can you tell me just, like, a couple bullet points that would match, like, ZoomInfo filters? Like, who are they? The people that are your ideal customer right now?
(Nora at 00:26:50) Yeah. So if it's for the incident analysis platform, it's a bit of a larger organization, globally distributed. The more complex their system is, the more appropriate Jeli is for them. But it's also really useful for companies to use this when they're starting out because when they get it starting out, they don't have to sprinkle it on later.
(Nora at 00:27:15) It's not something to unwind and use later. And so getting it—it's like hiring a designer. You know? You don't want to hire a designer for the first time when you're like 300 people and just start sprinkling design in. Like, it doesn't work like that.
(Nora at 00:27:28) And so getting us in as early as possible is really helpful for folks, but we find a lot of success in organizations that are larger, that have product-market fit, and care about their incidents. So, like, our sweet spot is when companies are just approaching their product-market fit themselves and are starting to get a handle on things. They're really ramping up quickly, and they're having incidents. Right? They're having things that interrupt their day that they need to do something about.
(Joel Beasley at 00:27:59) What's the pain point of the decision maker—the VP of Engineering, the CTO—
(Nora at 00:28:04) Mhmm.
(Joel Beasley at 00:28:04) —that they're going to say, "I'm experiencing this, so now this tool solves that problem?" What is that?
(Nora at 00:28:10) Yeah. I mean, they're having incidents, and they're having to make decisions based off of them. The pain point that we solve is we help them do it efficiently and effectively. And so we're faster, but we're also enhancing the quality of the output. So we're saving them time and money, and we're also short-circuiting that decision cycle like I was talking about earlier.
(Nora at 00:28:32) Like, the VP I shared in the story earlier was really happy with this because they saved about eighteen months. And we're also cheaper than the cost of an SRE. And so companies really like that as well because we help give their SRE shoulders to stand on. We're kind of like a buddy, basically. We're like a teammate that allows them to move faster, allows them to manage how their incidents are doing—those themes over time, like I was saying before—and helps them be more efficient with putting together their timelines after the incident.
(Joel Beasley at 00:29:07) That's really smart, the SRE association, because we found out—we started a new product about fifteen months ago where we started making podcasts for other companies.
(Nora at 00:29:16) And we—
(Joel Beasley at 00:29:17) —were selling them based off of, like, number of episodes that would be produced. And then I had gotten some feedback because we were working on our business model, and they would constantly want to piecemeal the different parts. But they were all people that had roles, and so they were dependent on each other, and it was kind of awkward. So we ended up switching the entire sales process to say, "You're getting five people on your team. Each person costs this, and this is their responsibilities."
(Joel Beasley at 00:29:43) And if you take them away, you know, then those things don't happen. And when we made the shift from "You're buying this product" to "You're having people on your team" type deal, that was a huge advantage. So I love the idea of comparing, like, "Okay, you could hire another SRE, or you could have our system." Is that how you are pitching it?
(Nora at 00:30:03) Yeah. Sort of. It's like we're escalating the SREs they already have. Like, some of our most successful customers—we got brought in right after they hired their first SRE, and we've been described as, like, a team for that SRE.
(Nora at 00:30:18) You know? And, like, as a product, you know, we really help them ramp up because if you're coming in as a first SRE, you have a thousand things to do. You know? Incidents is just one of them. And so we really help them educate the rest of their organization on how to do incidents more effectively and in a way that is better for the customer and better for the company at the end of the day.
(Nora at 00:30:39) So that's one way we're looking at it. And then another way that we're sort of pitching it to folks is after an incident takes place, you know, if you want to do an effective incident review without Jeli, it takes like three to four hours to put something together. You end up putting it in Google Docs where it goes to die. Like, how do you do the analysis on incidents throughout Google Docs? It's not pretty.
(Nora at 00:30:59) And I've tried it before. You don't. Yeah, it turns out really badly. But the thing is there's always going to be someone asking you to do it. Right? There's going to be—every job I've been in, some CTO or VP has come to my desk and been like, "Hey, how many incidents is this system in? Like, I need to go talk to the board about it in an hour." And I'm pouring through Google Docs to try to find something that's not going to lead to the board making a faulty decision based on this bad data.
(Nora at 00:31:27) And so that's one other thing. And, you know, people can do an effective incident review in Jeli and pull effective themes within twenty minutes. And so we're a really big time saving as well.
(Joel Beasley at 00:31:40) How long do you think until there's a ChatGPT-style version of Jeli where I can just push my entire infrastructure data unstructured—
(Nora at 00:31:45) Into it, and then the little Jeli—
(Joel Beasley at 00:31:45) Is this a fox? Like, a Jeli Fox? I don't know what that is. But on your icon, is that a fox or no?
(Nora at 00:31:56) You know, it's actually intentionally abstract because incidents are abstract as well. So they look like different things to different people. To me, it looks like a flower, right? But to you, it looks like a fox.
(Nora at 00:32:07) And neither of us are wrong.
(Joel Beasley at 00:32:09) It does.
(Nora at 00:32:09) Right? But both of our perspectives are—
(Joel Beasley at 00:32:12) Oh, no. You're wrong, too. Yeah. You're right.
(Nora at 00:32:15) Right. And that's honestly how people feel after incidents. They're like, "No, it's a fox." After every incident, someone has an ax to grind. They're like, "Let me tell you about this fox that I've been complaining about for so long." And I'm like, "But what about the flower portion of it?" And they're like, "What are you talking about? That's not actually important." But we're coming at it from two different angles.
(Nora at 00:32:35) And that's what Jeli really helps do too, is helps everyone understand their angles a little bit more. Yeah. So back to your question on, could I just dump all my incident data into ChatGPT and organize it that way? And I think you'd still run into some issues. And I think there's a lot of power in having AI play a role in our incidents and having us understand our incidents, but the human still needs to be involved.
(Nora at 00:33:00) Right? Otherwise, we're just gonna be making these faulty decisions again. So, you know, very carefully is how I would suggest doing it. Yeah.
(Joel Beasley at 00:33:09) I would think that the way I would use it right now, having little experience writing software for those language models, is I think it would be a cool experiment to push a lot of unstructured data at it and then ask it questions, you know. And not necessarily saying that, like, "Hey, this system's gonna replace it or whatever." I just think that if you can prompt it correctly, you can almost have analytics and insights without actually writing the algorithms yourself.
(Nora at 00:33:38) Yeah. Totally. And that's actually—we do, you know, we're not using that specifically, but in our tool, we have something called Learning Center. And you don't really have to do anything with Learning Center, and we show you which incidents had the most people that weren't on call. We show you which incidents happened to people that are experts outside of their business hours.
(Nora at 00:34:01) And so having data on things like this helps you organize your organization a little bit more effectively so you're not overhiring or burning people out, and it's helping you really retain your best people. So that's what you can do right now without actually doing anything, but there's more we want to do with it as well.
(Joel Beasley at 00:34:20) As a founder, super hard to start a company. How are you dealing with that in your family?
(Nora at 00:34:26) Yeah. I mean, it's all about balance at the end of the day. I mean, I live in Denver, and that's actually been really good for my mental health. I moved here a few years ago, and I am outside hiking and playing with the dog pretty regularly. And, yeah, I think it's just all about actually taking time away from the computer and also making sure you're really focused when you're at the computer as well is how I do it.
(Nora at 00:34:56) And it helps me show up better for my family. It helps me show up better for my company when I'm doing stuff like that.
(Joel Beasley at 00:35:02) 100%. What is the best leadership advice that you've ever gotten, you've implemented, and it was actually useful?
(Nora at 00:35:11) Oh, man. I've had so much good leadership advice. And, you know, forming Jeli has also been so meta for me, right? I've always wanted to work in a learning organization, and, you know, we're making a tool that helps companies get there, but we're also having to walk the talk ourselves.
(Nora at 00:35:27) Right? Like, you know, I'm now this VP during an incident or a CEO during an incident that's being like, "Hey, what's going on? When are we gonna get back up again?" And I think the best leadership advice is advice I used to give leaders too, which is trusting your people to do what they do best, to do what they do better than you.
(Nora at 00:35:47) You know? You hired experts for a reason. They have expertise in the area they're in. You know? Give them context, give them boundaries, and let them run is really, you know, I think some of the best advice I've gotten.
(Nora at 00:36:01) But it's about actively staying involved with that too and constantly redrawing those boundaries and constantly giving new context as the business is changing and moving faster too.
(Joel Beasley at 00:36:13) What's the most misunderstood aspect of incident response?
(Nora at 00:36:19) Oh, I think that it's actually not about the incident. I think people think the incident itself is about the incident, but it's really a mirror into how we work together. Because during an emergency, all rules and procedures are going out the window, and you're just trying to stop the bleeding. You're just trying to fix everything as fast as possible. And there's so much data in what happens when you're trying to stop the bleeding that can be used for everything else.
(Nora at 00:36:42) So that's my whole thing is, you know, you're leaving a lot of ROI on the table if you're not looking at what's happening when people are trying to stop the bleeding.
(Joel Beasley at 00:36:53) I love it. Let's direct people to the website. What's the website?
(Nora at 00:36:57) Jeli.io.
(Joel Beasley at 00:36:59) Jeli.io. You can install the Slack chatbot, right? And—Mhmm. Why do people install that chatbot?
(Nora at 00:37:06) One, it's free. And we help you coordinate during the incident, so we help you really focus on responding. That chatbot will help you spin up a new channel. It will help you broadcast updates to stakeholders. It will help you keep track of time during an incident.
(Nora at 00:37:20) And then after that incident is complete, it will ingest it into our analysis platform where we can show you some fun things about the incident as well.
(Joel Beasley at 00:37:29) But can it change its accent or talk like Elon Musk or—No. Like the ChatGPT stuff. Can't—that's next version.
(Nora at 00:37:38) Yeah. Totally.
(Joel Beasley at 00:37:40) Oh, that sounds fun. What else did we want to get out there to the world?
(Nora at 00:37:45) Um. Yeah. What else? I mean, check out the videos from the Learning from Incidents conference. You know, we open source these for a reason. There's some really great talks out there.
(Nora at 00:37:54) I highly recommend checking out the IBM one where they implemented a learning from incidents program in their organization and really got their organization to start thinking about incidents differently. I highly recommend checking out the joint talk between Indeed and Salesforce. That was also a really great one as well. Mhmm. And, yes, watch this space.
(Nora at 00:38:14) You know, we're just getting started. We have a whole guide on how to do incident analysis called the Howie Guide. That's also free on our website. And, you know, I love talking to people about incidents. So if you just want to talk about incidents, you know, reach out.
(Nora at 00:38:28) Like I said, we're huge nerds about this space. Like, I've wanted to be an entrepreneur, but it's like, this is the thing I really care about. Like, this is the thing I really want to do, and I want to exist in the world. So—
(Joel Beasley at 00:38:42) I love it. We did it, Nora. We made a podcast one step closer to winning your SEO battle with Nora Jones.
(Nora at 00:38:48) Thank you so much for helping the cause. It's very important.
(Joel Beasley at 00:38:53) Thank you so much for listening. And if you found this episode useful, please share it with a friend or colleague who you think would get value from it. And if you have topics that you would like to hear discussed on the podcast, either add me on LinkedIn or send me an email, [email protected]. Every time I get an email or LinkedIn message, it absolutely makes my day and inspires me to keep going.