Episode 750 ·

Exponentially Expanding Software Testing Coverage with Steve Semelsberger, CEO at Testlio

Today we’re talking to Steve Semelsberger, CEO at Testlio. We discuss how modern approaches to app testing are helping engineering leaders increase quality coverage, why the stakes for accurate app testing are higher than ever before, and why the mindset behind testing is crucial for uncovering tangible business benefits.

All of this right here, right now, on the Modern CTO Podcast! 

To learn more about Testlio, visit their website here.

Have feedback about the show? Let us know here.

Produced by ProSeries Media.

For booking inquiries, email [email protected]

About Steve Semelsberger

Steve Semelsberger is the CEO of Testlio, the originator of fused software testing. His 25 year career spans leadership, board, consulting, and investing work with high-growth technology and services companies. He’s the Founder of Alder Growth Partners, was a President at SYPartners, helped grow two startups to acquisitions (Pluck and iChat), and served as an executive on two IPOs (DMD & MOTV). He holds an MBA from Duke University along with a BS from Binghamton University.

About Testlio

Testlio is the originator of fused software testing. Our innovative approach leverages a proprietary software testing platform, partnerships and integrations with DevOps leaders, and a global services delivery team in 150+ countries. Together, Testlio powers quality assurance in any location. On any device. In any language. Via any payment instrument. Testlio clients include Amazon, Microsoft, NBA, Netflix, SAP, Paramount, PayPal, and Uber. Collectively, they serve nearly 2 billion users— and have awarded us an industry-leading 4.7 G2 rating.

Transcript

(Intro Narrator at 00:00:00) Today, we're talking to Steve, the CEO of Testlio, about their modern approaches to app testing, payments testing, and more. You're listening to Joel Beasley, Modern CTO.

(Steve at 00:00:16) It's good to see you, Joel. How have you been, man?

(Joel Beasley at 00:00:18) Dude, it's great to see you. I'm super excited. I was working out this morning, and something popped in my upper back, and now my voice sounds a little off. But I feel fine. My back feels fine too. So I don't know what to do about that.

(Steve at 00:00:34) So you popped your back and got a frog in your throat? I didn't know that was a thing.

(Joel Beasley at 00:00:40) I'm not a doctor. I don't know, but I asked my wife. I was like, what should I do? And she's like, maybe you should go to a chiropractor. So my back feels fine. And she's like, wait. Just wait 24 hours. That happened this morning. I was like, alright. I'll wait. I'll see what happens.

(Steve at 00:00:54) It's turned into a little bit of Seth Rogen is where I'm going with it, which I like Seth Rogen a lot. So I think you should just roll with it, man.

(Joel Beasley at 00:01:05) So tell me about Testlio. Am I saying that correct?

(Steve at 00:01:09) Yes. Yes. Well done. It's a made-up word. Test, L-I-O, Testlio. So the company's about 10 years old. It's a testing services company. It's technology powered, and we work with some really big and innovative companies to help them solve really tricky problems in the area of quality engineering and quality assurance.

(Joel Beasley at 00:01:33) Where do you draw the lines? You seem like you test so many things. Everything from, when I was prepping, I saw Super Bowl, World Cup, giant digital events, the OTT devices, apps before they go on the app store. My mind was very narrow to testing because I was an enterprise software engineer for lack of a better term. And so I just considered testing as far as the tests I would write with my Ruby code and test-driven development. But this is a whole other level. This is testing multiple things, multiple industries. So where do you draw the line?

(Steve at 00:02:07) So as you probably experienced, as you were developing and automating unit tests yourself and responsible for the code you were checking in, and then as you've worked with lots of CTOs and VPs of engineering, shift left is a real thing. It works. A pyramid-centric approach towards software testing is logical. The majority of testing should happen at unit and API levels and end-to-end systems as they come together. At the same time, though, the stakes for end-user experiences are higher than ever before.

So the industry, I think, has advanced a whole lot in shift left. What we see is that there's also a similar shift right that's occurring where companies are investing more in exploratory testing, visual testing, usability testing, accessibility testing, location-centric testing, payment verification testing. And so when companies have experiences that hit the glass, especially if they're hard to roll back. So pure web, I'd offer, is different. I know LaunchDarkly was a guest of yours. We love LaunchDarkly and systems like LaunchDarkly that allow teams to intelligently try things out, feature flag, roll back, which works really well for web. But when you're releasing a native app or you're integrating a payment experience or you're hosting the Super Bowl or you're about to deploy a beta release to drivers, you're a global mobility provider, the stakes are really high and things need to work. And I think that's where Testlio is really, really well suited. We think of these as moments that matter of software releases, where you want to put energy and emphasis to truly ensure a great end-user experience. Does that help, Joel? Did I get to what you were looking for there?

(Joel Beasley at 00:03:56) Yeah. Well, we're headed in the right direction. So I'm going to push it from a lot of different aspects because this is how I understand something. I attack it from a bunch of angles, and then I leave with my understanding. So if I was going to do integration tests or some type of feature tests to replicate an end-user experience and then check that certain things are working, well, first of all, typically I'm checking just for the functionality of the post actually happening, the response actually happening. I'm not actually looking—I don't have any systems like actually look at the screen and look at aesthetics or end-user experience. So that's one huge benefit. The second thing is when I am doing those types of tests, I'm typically bumping up a Selenium browser or some sort of headless type browser or something. And there's so many different types. When I send an email campaign, I have a cool little tool that'll show me what it's going to look like in 100 different platforms. Right? And so to get that, I couldn't imagine myself writing and implementing tests. And there's so many test cases to test. So this part of the question is, how do you help figure out how many test cases to test when there's an infinite number when you're working with a client?

(Steve at 00:05:06) And then as you think about how many test cases, which of those should you automate? And which of those should you intentionally use manual testing for? And so, like many things, it really varies broadly depending upon the team, the maturity of the team, the pace of deployment, what sort of software deployment approaches the different teams are using, how big the problem is, what the risk profile is. So I'll give you an example that might be on the opposite end of the spectrum of what you were describing. So we do work with large software companies like Microsoft. And there's a team at Microsoft that has a product that many of us know and love, and the stakes are super high for that product. And the team has measurements as well as compensation systems that are tied to a completely bug-free experience. And the approach that that team takes is a multi-release gate approach. They have, I believe, four now different release gates that they go through, and Testlio is actually the final release gate in working with this team. So what they've said is because the stakes are so high, we want to do all of what you described at scale in different ways and more. And then where Testlio comes in is this machine and human approach to ensuring that the end-user experience is great. And a lot of times, that tends to push out towards multiple device OS combinations. So on average, and this varies quite a lot too, but when we work with companies, each test run covers 24 device OS combinations. And we have 1,200 different hardware devices in our global network. So we can test against a lot of different types of hardware. And for some companies, that really matters a lot to them beyond just a standard, say, Chrome browser instance. So let me pause there and see if that went in a direction that helped.

(Joel Beasley at 00:07:09) Yeah. Things now are becoming very hardware specific. I was having a conversation the other day with a guy that's building an offline ChatGPT, which is using specific hardware that's only available in certain phones. Right. And so being able to pick the specific hardware and understand how it works across, I didn't realize how many different devices there were and how many different types of hardware and also how in different platforms, they only let you target phones in specific ways. Right? You can't, from what I understand, you can't target a specific version of iOS and make it available or not. But you can target a hardware segment, which was interesting. So now that we're getting such a huge array of devices, hardwares, and systems to test that at scale becomes even more complicated, right?

(Steve at 00:07:58) Very much so. And devices vary so much by different countries. So for example, we've helped Sky launch into Africa. And we work with a large African messaging, or African continent-centric messaging company called Ayoba. And the device profile, oftentimes, these are $10 smartphones that are being sourced from small providers in China, available from communications providers and retail outlets. And so those devices are really different from what might get tested on the National Basketball Association, for example, is a client of ours too. And so we'll test in places like Rio, because basketball is super popular. And the experience in Rio, there you have a translated, localized experience of the NBA app that has to be amazingly authentic for what's happening in now time in Rio. And so the device is different and the experience matters. And, you know, maybe at some point, too, we could have a conversation about localization beyond just a linguistic translation, because that's also a big driver for companies that are really operating at massive scale and thinking about how authentic that they need to be in different markets.

(Joel Beasley at 00:09:17) What does that mean, authentic?

(Steve at 00:09:19) That means that—so, for example, we work with a top three social media provider. And for that social media provider, they want to ensure that the experience of their product is true to the culture, which is a combination of history, sociology, trend, and language. And so that when they localize, they will go through translations, and they will go through different screen modifications and even feature modifications. And then they'll use Testlio with people on the ground in those countries who are a combination of linguistic and testing experts who actually know what the experience of that product should be because they're authentic users of the product. So maybe to even take a step back, in order to tackle something like that, we had to go through about 2,000 potential testers to build an initial team of 200 people to serve this customer. And what the customer demanded was that not only, you know, do people really know testing, but that they have a very high expectation for the experience of the product. And this particular company may only release in some of these countries—we work with them in 62 countries right now—they may only release two or three or four times per year. So it's unexpected when these releases are coming. They're not scheduled. A lot of times things come together and there's a decision that's made. And so that's where that elasticity or that burstability of testing really comes into play too. Did that help give you a sense of—

(Joel Beasley at 00:11:01) Yeah. You guys do a lot. It's huge. What is happening within a company for people that are listening that might be considering using a service like yours? What type of conversations are typically happening in their company before they come and find you?

(Steve at 00:11:17) So we're thrilled to be here, Joel. We love the podcast. Thank you. Thank you for hosting us. Our clients are typically chief technology officers or their equivalent. And so many companies have quality teams internally, yet the quality team works in close partnership with the CTO, SVP of engineering, who really sets a release and a bug tolerance and a testing strategy for the organization. So we often come when companies are ready to experiment with a new testing strategy. And in some instances, it's been, hey, we tried 100% automation for a couple of years, and we realized that that has some pros and cons associated with it. So there's a large fintech company that their CTO came to us about a month ago and said, you know, we sense that it's time now to integrate a new approach to manual testing into what we've been doing in a world of 100% automation. With other companies, it might be that, well, we've gotten pretty good at manual testing. Maybe they're using a crowdsourced technique as well. And they realize we're not automating enough, and we're not automating quickly enough internally. So can you help us accelerate the pace of automation? And then for some companies, there's this sense, we see this quite a lot, that we've been doing pretty well in automation. We started with our unit tests, and we're all the way up to now visual and functional automation. And you know what? We've been doing some manual testing too, and that kind of works well. But the two things are completely disconnected. You know, manual testing happens over here, and automated testing happens over there. And there's got to be a better way. Right? And so we have a methodology that we call fused software testing that brings automated and manual testing together. And it really handles the fluidity and pace that's needed in a modern DevOps-driven release environment. So it's usually this sense of we want to go faster. We want to try something different. And in some cases, Joel, it's cost too. So what we see with world-class companies is that they'll spend between 10 to 20% on quality as a whole. And maybe we could have a conversation about how do you think about that? How do you measure those economics? But sometimes, quality costs can creep up, and companies feel like they just might be spending too much. And so they want to see, can I ensure great releases, world-class end-user experiences, and actually save some money in the process, especially in the world that we're a part of right now where companies are so cost conscious?

(Joel Beasley at 00:14:00) Ten to 20% of what?

(Steve at 00:14:02) Of overall software development spend. So you can look at that as, you know, how much do we spend on our core development team? Do we have any dedicated quality resources? And some companies don't. So they may think of a percentage of time that's spent on quality as an element of their overall software development spend. Some organizations have an outsourcing partner. Some have an in-house team and an outsourced partner. And so thinking about this sense of how much do I spend to help my technology be great at the 10 to 20% mark? And it's different. Some companies are significantly less and some companies need to be significantly more depending upon the world that they're in. If you're in a regulated industry, you know, if you provide medical technologies, your ratio might be much higher than 20%. But that's a range that oftentimes helps companies enter into the conversation and the consideration of what's my spend tolerance, what's my bug tolerance, and, you know, what do I want to do strategically? And how open to experimentation might I be in the world of quality too?

(Joel Beasley at 00:15:05) How do you measure the cost of not doing this?

(Steve at 00:15:09) So there's a concept of escaped issues. And an escaped issue is arguably something that hits production that should have been caught somewhere along the way. And so when we see companies testing, whether we're testing for them or they're testing in other ways, we hold that for every thousand hours of testing, you should see one or less escaped issues. It's a 0.1% escaped issue to hour of testing ratio. And so there's this sense of, if you're doing things and you're seeing more escapes, then maybe you need to change the approach that you're undertaking. Or, again, if you feel like, you know what, it's taking too long, measured by clock time or calendar time. You have a sense of speed and pace is another indicator. And if, you know, your releases are delayed based on the approach you're taking to testing, that may also be a hidden cost of not considering other testing strategies.

(Joel Beasley at 00:16:15) How are escapes different from bugs?

(Steve at 00:16:18) Yeah. An escaped issue would be a bug that should have been caught through your testing standard. So it's a way of categorizing your bugs. So let's just pretend you build code and you just ship it. Well, then, you don't really have escaped issues because you're arguably not testing at all.

(Joel Beasley at 00:16:37) But if you—you have all sorts of issues if you're doing that.

(Steve at 00:16:40) All sorts of issues. But if you're a big company and you have four gates of testing and you still have issues or bugs escaping there, then, you know, something is probably broken along the way. And sometimes it's a process break. Sometimes it's a technology break. You know, but depending upon your level of investment and the sophistication of your combination of people and machines, you know, you can have very different thresholds. That's why we just give a basic escaped issue per hour of testing rate.

(Steve at 00:17:13) We also see that when humans are doing testing pre-production, code's pretty good, processes are pretty mature, humans will pick up about 0.3 issues per hour of testing. And that number varies dramatically as well. But what we'll often see as a signal of a good partnership with our clients is that number might start higher when we begin working with a company. It might be between one and two issues per every hour of testing. And if it stays high like that, well, then perhaps there's something wrong with the feedback loop. Processes aren't changing. As bugs are being caught, they're not being addressed or remediated, fixed, merged, et cetera.

(Steve at 00:17:54) Because typically a healthy partnership for us is that issue per hour of testing rate is going down over time. And then, depending upon release velocity, how much we're regressing, how much we're doing new feature testing, it tends to stabilize in what seems to be a healthy zone. And again, that differs based on companies. But it's interesting how often we see this 0.2 to 0.5 issue per hour of testing rate across the business.

(Joel Beasley at 00:18:24) It seems like payments would be an area where people really want to test, right? Because that's when the financial transaction's happening. How much of the testing you do is related to payments?

(Steve at 00:18:34) It's one of the fastest growing areas of the business. So over the last five years, we've been doing a lot of OTT testing with media companies and sports leagues and teams. The challenge there has been 12 different streaming devices—Rokus and Samsung TVs and PlayStations. It's a hardware-centric problem with a bit of location. Over the last two to three years, the fast growing piece of the business has been in payments testing.

(Steve at 00:19:01) You may have seen, Joel, that amplified by Black Friday and Cyber Monday. Something like one out of 10 digital payments are failing currently. Now, it's gotten better. I think two to five years ago, it was like two to three out of 10. But still, one out of 10 digital payments are failing. And a lot of times, it's the intersection where things fail.

(Steve at 00:19:27) So maybe you've had the experience where you're traveling and you go to buy something and your location is picked up and your credit card company puts an interstitial in between the transaction. You have to verify sometimes via a text message. And it's those kind of payment handoffs between different systems that oftentimes cause the challenge. And so companies then have to determine, well, do we go into the field and send people into real world situations to try and see, using all sorts of different devices, if we can catch some of these potential payment failures? We do a lot of work with a large payment network provider, and they are about to enter into a new partnership with a mobility provider.

(Steve at 00:20:11) And they decided that it was so important that the whole experience from reservation through car pickup, through lot departure, through car usage, to return, to ultimate finalization of invoice—that that whole process had to be absolutely perfect. And so they had to send people all over the world to airports. They've been building this partnership. It's very mature technology. And as we send people into airports, we found sometimes that the popup that's supposed to happen as you're leaving the airport lot was coming in a day later. Okay, well, why was that? Where was the problem in the handoff? What was delaying it? And so those sorts of things are really hard to pick up if you're only doing lab-based testing or you're doing testing with just an in-house team. You can do testing like that at big scale for these—again, I use the term—moments that matter, these critical moments that matter.

(Joel Beasley at 00:21:08) How much work does it put on us when it comes to testing? Do we just give you our product and you tell us all these things that should be tested, and then we add some extra special things on top of that? Or do we have to do 100% of the work? Where does that lie—coming up with test cases, I guess?

(Steve at 00:21:25) There's a few different approaches which seem to work well for a lot of our clients. One is exactly what you described. So we have a team of engagement managers and testing managers and testing coordinators, and they take builds, and they ensure that machines and humans are equipped to test in smart ways. And developing the instructions to test is a perfect LLM problem. So earlier this year, we started applying generative AI to this whole test case development process.

(Steve at 00:22:01) And it turns out that using—for us, in this case, it's the Azure instance of OpenAI because we have a partnership with Microsoft, and Microsoft ensures that data is extractable and that it's protected and that it's confidential. So we didn't want to just throw things into OpenAI ChatGPT since test cases are our clients' data. So we put this data into this segmented version of OpenAI, and we can create test cases 30 to 40% faster. And we can also refactor test cases because there's not only the initial set of instructions that you develop, but then anytime the product changes, your test cases, if you're doing regression testing, for example, should also change along with that. So refactoring test cases is also much faster using large language models.

(Steve at 00:22:50) And just to maybe go a little bit more on AI, there's also the opposite end of instructions, which—an issue for us is something that needs to be captured in a way that's demonstrated, reproducible, all sorts of metadata, tags, prioritization, severity against it, then sent into our customer system. So about 70% of our clients use Jira. And our platform, where we do all the work, hooks into Jira. There's a bidirectional integration where we push things into Jira, then we get updates as things are happening in Jira. And it turns out that if you give a tester and a test lead the opportunity to run the issue through a large language model that's tuned for what we're trying to do, that the issue is—we estimate—about twice as trustable, which is grammatical clarity, clear structure, no typos.

(Steve at 00:23:45) Because it turns out that when humans see an issue and it's messy, they don't trust it as much, which means that they then may deprioritize it, which could actually be a mistake. So this speed of the test case and trustability of the issue, those are really good AI challenges. Now to your original question of how hard is this for companies—the one model is just give us the build. The second model is some companies have test cases that sit in systems like X-ray or TestRail, for example, and we have integrations with those systems. So we can take a build in.

(Steve at 00:24:17) We talk to CI/CD pipelines. We use TestFlight, lots of different ways to take builds. We also take test cases into our platform from third-party systems. If companies are already maintaining those, we don't need to duplicate the efforts. And then I'd say the third model is that companies will sometimes come into our tools with us. We have this partnership, sometimes with product leaders, sometimes with quality leaders. And we're really co-creating how to test things and when to test things. And it's not a fully managed service partnership in that case. It's kind of a co-managed testing experience. I know, Joel, I threw a lot at you there. Hope that wasn't too much.

(Joel Beasley at 00:24:57) No, no, no, not at all. What was interesting to me is I didn't know quality managers existed. Tell me about those types of people.

(Steve at 00:25:06) So many companies have directors or vice presidents of quality engineering. So these are companies that are operating at scales of hundreds or thousands of software developers. And the vice president of quality engineering oftentimes reports directly to an SVP of engineering or the CTO. And they're thinking a lot about process, tool, and partnership in order to ensure the velocity and integrity of the release. And so there's the quality of the engineering process, and then there's the quality of what the engineering team produces.

(Steve at 00:25:44) And those teams will sometimes have quality managers and quality analysts—people who are doing manual testing full-time, dedicated. They'll sometimes have software developers who are building automated test scripts. Sometimes they're called quality engineers internally. And in some cases, these teams will really choose to crowdsource or outsource the majority of that work.

(Steve at 00:26:06) And so we love when teams have quality organizations internally because it's a demonstration of a commitment to quality. We have some clients where we are their quality solution, but usually this partnership model where there are in-house members embedded within the engineering team that are driving quality standards and testing quality hypotheses—those are awesome folks for us to partner with. So while the CTO might set the strategy and decide to bring Testlio in, it's that internal quality team, if it exists, who can often ignite great experiences.

(Joel Beasley at 00:26:39) That's what all the CTOs are doing right now. They're going to testlio.com and they're signing up. I'm just putting that out there. Okay, so how do you test Super Bowl and World Cup type events? What do you even begin to do to test something like that?

(Steve at 00:26:54) Yeah, those are fun. So it's been an honor to work with—this is public—Fox and Paramount. And Paramount is now the parent company for CBS. And so for those of us in the U.S., we know CBS and Fox as primary broadcasters. So they've had Super Bowl rights for the last five years. And now we've also publicly disclosed that Peacock, which is the NBC streaming service, is a client of ours.

(Steve at 00:27:18) So if we continue to do work well, we should test every Super Bowl going forward with the three primary broadcasters who rotate the rights each year. And each company does it a little bit differently, but what you can imagine is that there's a distributed war room that starts to build up to game day and that uses techniques of development, monitoring, and testing on increasingly large-scale, important events. So you might start with a Sunday game and say, well, what if we treat it as if it was the Super Bowl? And then you may have an NFC playoff game, and then you may have the NFC championship. And then you may, depending on the year, have a World Cup game, or you might actually have a pay-per-view fight that has a massive audience. And a lot of times you're preparing for a big front door challenge.

(Steve at 00:28:07) People are coming in all at the same time. You're preparing for, is your advertising firing if you're selling it at all by location? Many advertisers still sell local TV markets. And so how do you apply that to digital so that you don't get these weird lags or just, "Wait, we'll be right back with you." That's wasted ad revenue.

(Steve at 00:28:29) And then also, how are you tracking everything? So are your analytics systems operating well also? And how is that operating for hundreds of millions of people across—most broadcasters we see are focused on 12 OTT devices. That gives you a lot of coverage, at least here in the U.S. But there's all sorts of different firmware and operating system issues even associated with those device families.

(Steve at 00:28:53) And first-generation Apple TV is really different from a recent Apple TV. We live in Texas, so we actually have—it's a little embarrassing to say, but we have five Samsung TVs. I have three kids and we're lucky to have some space here in Texas. And each of the experiences of running digital apps natively on those Samsung TVs is really different in just my household.

(Joel Beasley at 00:29:17) Oh, me too.

(Steve at 00:29:18) Yeah, OTT device. But yeah, you get it, right? It's really different.

(Joel Beasley at 00:29:22) We have one—I have an 85-inch screen right now in front of me. Back a couple of feet.

(Steve at 00:29:29) I'm sorry that you have to see me in 85 inches. I should back up a little bit there, Joel.

(Joel Beasley at 00:29:36) So what is signal-driven testing?

(Steve at 00:29:41) Oh, signal-driven testing. So when companies start to integrate the different tools that they're using to build, release, and test products, and even to support their customers and to pay attention to what people are posting on things like review sites, they start to have more and more signals of the end user experience, the quality of their technology. And so signal-driven testing is a mindset and a methodology which says you should be tuning your testing constantly on where you're getting signals that demonstrate that maybe something is escaping—going back to our escaped issues conversation from before. And so a signal could come from the App Store. And it's not just a one-star review. It's often the detailed content that sits within the app review.

(Steve at 00:30:31) Because you may get a five-star review that actually tells you that the person had a problem. Like, you might love Roku, and you give Roku a five-star app review, and then you would talk about this horror movie problem and the inability to filter out content. Well, that would be a signal. And so if we're mining the App Store intelligently and we're interpreting and using sentiment analysis and other forms of intelligence to pull these signals out, well, those signals then should be hooking back into your test cases to determine, okay, do we miss something in testing? Or in certain cases, signals can very rapidly be turned into issues. You don't have to do more testing. You can see that the signal is truly a problem. So companies use systems like Zendesk, and too often customer support is fairly disconnected from quality. So just like DevOps and quality, I think, are coming together, I think support and quality are coming together as well.

(Steve at 00:31:24) And so how do you take advantage of what Zendesk might tell you? If you use LaunchDarkly and you find that you're feature flagging things on and then frequently turning them off, well, that's a signal, perhaps, that something that is creating a candidate for a feature flag-on isn't going through the appropriate testing. So how do you then interpret that and change some of your processes? So we think right now there are nine different categories of signals that can be used to drive quality processes.

(Steve at 00:31:57) And really, it's probably pretty old school. It's all about continuous improvement. When quality—we see this a lot with our clients—when it seems to be kind of dialed in and it's working, that's usually an indicator that we're not thinking of what more we should be doing or what we should be changing because products are changing, technologies are changing, and end user expectations are changing. And the signals are out there. It's just how do you interpret them, prioritize them, and then help teams also who are always just way too busy. Most of the engineering organizations we work with—maybe all of them—have backlogs that are impossible to get to. So then, as more and more things are coming through, well, how do you really take an intelligent approach towards remediating the things that have, ideally, a high return? Help the company grow revenue, help retain users, lower costs for the company itself, and then have a reasonable investment to fix. And that ROI of issue remediation is an important thing as well. So that's how we think about signal-driven testing.

(Joel Beasley at 00:33:01) That's interesting. So are you actually doing the sentiment analysis off of App Store type reviews? Are you partnering with somebody that does that?

(Steve at 00:33:11) So we partner with a company that has rights to take our customer-owned App Store review data and to have that sit within the Testlio system with permission. We have this notion that our clients' content—we don't own it. They own it. We keep it highly secure and confidential, et cetera.

(Steve at 00:33:30) So then we've been using some third-party technologies, some appropriately used open-source technologies, and some proprietary technologies to turn the review into a signal. And we're still learning along the way. I think we're quite good at taking certain signals, and other signals we're trying to better understand what they mean. And a lot of it we do in partnership with our clients too. So we may introduce the idea and then run an experiment to see, like, if we pay more attention to this particular type of signal, what might it do for us?

(Steve at 00:34:02) And another example are flaky automated tests. So you create automated tests, and then they're giving you positives or negatives, but are they false or true? And so a flaky automated test could be a type of signal that you really want to pay attention to. And so the journey might be not about the software itself, but about your automation journey and how do you make test automation of a higher quality standard, if that makes sense.

(Joel Beasley at 00:34:29) Yeah. The more bugs we can find, the better. Right?

(Steve at 00:34:33) That's it. That is it.

(Joel Beasley at 00:34:34) That is it. It's all fine.

(Steve at 00:34:36) Well, we like to say the first goal is to not ship bugs. We can have a great partnership with some of our clients. It's really an assurance. If you think of quality assurance, it's really about helping engineering leaders sleep well at night by not shipping any bugs. So we have very successful partnerships and we report like zero issues.

(Steve at 00:34:56) Whereas in other cases, especially for nimble companies, emerging companies, high velocity companies, companies that have a higher risk tolerance, well, then it's about, you know, how quickly do you catch? What's the cost per bug effectively? How much does it cost to remediate? Are we sending too much noise? Because we'll see sometimes that based on prioritization and severity marking, the way that things get tagged, things may turn out to be heavily skewed towards what our customers perceive as low priority issues.

(Steve at 00:35:25) And clients then might say, well, should Testlio send us these low priority issues? And most of them will say, yes. We still want to know. As long as it's a valid reproducible issue, we want to know that it's an issue, even if it's something that we're not going to tackle it now from a prioritization standpoint because we know we have this next major overhaul and probably this thing is going to get merged into a whole other thing. But, you know, certainly anything that might be a P1 or a SEV-one, like those are the kinds of things that most companies are looking to address really quickly.

(Steve at 00:36:00) And we're trying to help companies better measure the addressability of issues and the timeliness to address as well. So one metric that we're experimenting with some clients is what percentage of issues that Testlio provides are addressed within 90 days. And we hold it something like 90% within 90, partially because that's just an easy heuristic to remember. You know, are 90% of the issues getting addressed within 90 days? But then are your P1s getting addressed within hours? Not within months, but literally super rapidly. And they may get deployed a little bit later, but how quickly is that feedback loop happening? And some clients are picking up issues, and I mean, they're tackling things in minutes. It's really cool to see the pace of addressability happening now, Joel.

(Joel Beasley at 00:36:47) That's cool. Well, yeah. My butt's always on fire when there's something wrong. You gotta get it taken care of.

(Steve at 00:36:52) Yeah. That's it. That's it.

(Joel Beasley at 00:36:53) Dude, Steve, man, we made a podcast. How do you feel?

(Steve at 00:36:57) I feel great.

(Joel Beasley at 00:36:58) Thank you so much for listening. And if you found this episode useful, please share it with a friend or colleague who you think would get value from it. And if you have topics that you'd like to hear discussed on the podcast, either add me on LinkedIn or send me an email [email protected]. Every time I get an email or LinkedIn message, it absolutely makes my day and inspires me to keep going.