Episode 955 ·

Why Kubernetes Still Feels So Hard & What to Do About It with Steve Francis, CEO of Sidero Labs

What if you didn’t have to solve problems? You can just remove them entirely.

Today, we're talking to Steve Francis, CEO at Sidero Labs, about why the most dangerous thing running in your infrastructure might be the operating system itself. We discuss how eliminating features rather than adding them is the real path to security, why the promise of multi-cloud portability turned out to be a lesson in what customers actually care about, and why "extreme ownership" remains the most empowering philosophy a leader can adopt.

All of this right here, right now, on the Modern CTO Podcast! 

To learn more about Sidero Labs, check out their website here.

About Steve Francis

Steve Francis is the CEO of Sidero Labs, the company behind Talos Linux and Omni. Before Sidero, he founded LogicMonitor, a SaaS-based data center monitoring company, and spent years running data centers as an SRE before the role had a name. He holds degrees in both computer science and law, and has been building and leading infrastructure-focused software companies for over two decades.

Transcript

(Intro Narrator at 00:00:00) Today, we're talking to Steve Francis, CEO at Sidero Labs, about what makes Kubernetes still so complicated and what he's doing about it. You're listening to Joel Beasley, Modern CTO.

(Joel Beasley at 00:00:18) I was excited to talk about the company and the Kubernetes and all that, but the first thing I want to talk about is how do I pronounce the name of the company?

(Steve Francis at 00:00:27) So we call it Sidero Labs. It's based on ancient Greek, so we may well be pronouncing it completely wrong.

(Joel Beasley at 00:00:36) What does it mean in ancient Greek?

(Steve Francis at 00:00:38) So Talos Linux was our first product, and the company used to be called Talos Systems. So Talos, I'm maybe jumping ahead here, but Talos, in Greek mythology, was an automaton, kind of like the first robot, who used to patrol the island of Crete and keep it safe and secure. So Andrew, the founder, decided that was a great kind of totem for software that keeps you secure and is kind of always on guard. And Sidero in Greek means both to the stars, which is a good, you know, aspirational goal, and it also means iron. And so you build your robots out of iron. Most of our customers run bare metal, so it has that kind of connotation.

(Joel Beasley at 00:01:26) Dude, that is actually one of the better names I've heard. That's brilliant.

(Steve Francis at 00:01:30) Yeah, but on the other hand, we have no idea if I'm pronouncing the ancient Greek correctly at all. Sidero, Sidero. Yeah.

(Steve Francis at 00:01:37) Nice.

(Joel Beasley at 00:01:38) And then so what's the main problem that you solve over there at Sidero Labs?

(Steve Francis at 00:01:42) The main thing is that Kubernetes, especially if you're running it yourself—like, you know, you can use a cloud provider, they all have their services—but if you want to run it yourself on your own infrastructure on bare metal, it's complicated, it's hard to get right, and it's easy to make insecure. So mainly, what we do is we have a product, Talos Linux, named after the automaton, that is a version of Linux that can literally do nothing except run Kubernetes. So it's basically the Linux kernel. It's not a Linux that we took bits away from. It's a Linux kernel, and then we wrote just enough to start Kubernetes, and that's literally all it can do. So there's no systemd, there's no Bash, there's no shell. None of those things exist. There's programs that we wrote that start the kubelet, that start the container subsystems, and then get Kubernetes up and running. And so it's all written in Go. It's all memory safe. That has some applications because it's so small. It's like, depending on which version of Talos and what file systems are supported, it's less than between 15 and 50 binaries, the whole operating system. Whereas your standard Ubuntu, a small install, has like 3,000 executables. So that's 3,000 executables that you could try and attack. You can't attack programs that aren't running if they're not even installed on the system. So Talos Linux has the point of being very secure, very small attack surface, and it uses very small resources like memory and CPU because it's not running an HTTP daemon in the background that happened to get installed by default that is just consuming cycles for no reason. So it runs on small form factor devices, great for edge use cases, or if you're running, you know, big GPU farms where every bit of CPU makes your training costs less, it's great for that too. And the other big thing is it's only managed via an API. So you can't SSH into it at all. There's no SSH daemon running, and you manage it just like you manage Kubernetes. You create a configuration file and you apply it through an API and it reconciles its state, and there you go.

(Joel Beasley at 00:03:51) That's unbelievable. And how did you guys figure out this was a thing?

(Steve Francis at 00:03:55) So the founder of the company, who is not me, Andrew Rynhard—so back up a bit. I used to be an SRE before they were called SREs. I ran data centers. I did that for some SaaS companies, and then I always had the problem of monitoring SaaS companies. So I started a company, LogicMonitor, which was a SaaS-based data center monitoring company. Andrew worked for me at LogicMonitor and did our Kubernetes migration probably ten years ago, when Kubernetes was pretty early. And he loved the technology, and he decided that was kind of the field he wanted to work in. But he got sick of the fact that to run Kubernetes, where you have this great orchestrator, you know, scheduling jobs and dealing with access and who can do things and deploying things, you'd have to do a lot of the same work on the operating system. So he's like, why the hell am I configuring Ubuntu to give user management and, you know, basically package management and patch management, which is kind of like doing Helm deploys? Why am I doing the same thing on two different systems when really the only one I care about is the Kubernetes layer? So he, it was his idea to just, like, write an operating system from scratch to just run Kubernetes, basically, to save himself work. And then he posted about it on Hacker News and Reddit, and it was popular. So he came to me and said, "Hey, I think there's a company here." So I was one of his seed investors and came on as CEO six years ago, and here we are.

(Joel Beasley at 00:05:31) Is it rocking and rolling?

(Steve Francis at 00:05:32) It is. Yeah, yeah. Open source software is a different beast than commercial SaaS-based software, which was my last company. The first couple of years of the company were a bit rough. We came out with Talos Linux and got it to 1.0, and we got an enterprise customer, which was Nokia, which was amazing. And I thought, "Oh, this is going to be easy." And then a year later, we still had one customer, which was Nokia. I was like, "All right, this isn't as easy as we thought." Lots and lots of companies were using Talos Linux because it had value and benefit and made their life easier. But, you know, I did the same thing when I was running data centers. I used a lot of open source software, sometimes bought support contracts, sometimes not, depending on, you know, how reliable it was. We often joke internally that our problem as a company is we've made Talos Linux so stable and reliable and secure that people can just run it in production without needing support contracts or help, and that's a good thing. Makes commercial life harder for us, though. So then our next thing was, "All right, let's add more value." Talos Linux makes it super easy to run a Kubernetes node securely and easily and turn that into a cluster. But if you're doing that at scale where you have more than, say, three clusters—some of our customers have thousands—if you're doing it on large scale clusters, that's where you still need some tooling around that. It's like, how do I orchestrate my upgrades around the nodes in the cluster and which ones to upgrade when? What versions? How do I integrate it with my enterprise identity provider, stuff like that? So that's where we have Omni, which we operate as a SaaS, which provides all that tooling. It literally makes installing a cluster—it's kind of ridiculously easy. You kind of have to see it to believe it. Like, you sign up for Omni, you go, "I'm going to download my installation media," and you get an ISO. You boot however many machines you want to boot off that. They all connect back to Omni over a WireGuard encryption tunnel and say, "Hey, I'm a machine that's available." You can go into your Omni UI and say, "I want these three to be control planes, these five to be workers. These ones that have GPUs, you know, install the GPU," with one click and create a cluster. So you can create a production-level secure cluster running Kubernetes on your bare metal in like three minutes. It's ridiculously amazing. So that part of the business is also going well. I do feel, though, we kind of do the world a bit of a disservice because it makes it so easy to get a Kubernetes cluster up and running that often companies see this and go, "Oh yeah, this is great. Of course, this is what we want to do," and it is what they want to do. But Kubernetes itself still has so many layers and so much complexity. It's not an easy topic. So we take one part of it and make it super easy: make your operating system stable and secure, get Kubernetes deployed, get the cluster secure. But then you've still got, you know, which CNI do you use? Which storage systems do you use? How are you going to deal with ingress? Kubernetes is still a very complicated ecosystem in itself, which we're working on. We'll get to the other layers.

(Joel Beasley at 00:08:45) Yeah. Are you guys solving any of those secondary problems, or are you just focused on the first problems?

(Steve Francis at 00:08:50) Yeah, so we've been focused on the first problem. We're not going to be a complete—you know, OpenShift does everything. They kind of have their mandated way of doing the complete stack. We're not going to take that approach. We are going to try and have opinionated defaults that will solve most problems. Like, I've been thinking about why Kubernetes is so complicated compared to, like, what I did in a data center. You would, you know, spin up servers that would be web servers and other servers that would be databases, and your routers would be doing your routing and your DHCP server's doing your IP address management and so forth. And there's complexity in that. But Kubernetes has a great dream, which is kind of like a declarative data center effectively. Like, everything goes on this fleet of computers. Your IP address management and your storage systems and your load balancers and your firewalling—it's all integrated into one system, which is a great ideal, but it's very complex. So even when you get something working, it's like, but now when there's a problem, "I know that this IP address on this container is supposed to come and follow it over here," but, you know, Cilium may be doing a conflicting IP address management than what you're doing from the Kubernetes layer, and they may be overlapping with your CIDR ranges. And then figuring out which controller and which stack is responsible for the blocking, it gets very complicated. So it's a great aspirational ideal. It's not easy to operate in production at scale yet. You still need a very skilled team.

(Joel Beasley at 00:10:38) What is this concept that you guys talk about, control by design? What does that mean?

(Steve Francis at 00:10:43) That's kind of the genesis of what we do. You know, Andrew's idea was get the operating system out of the way as much as possible. So whereas most operating systems are, "I'm a general-purpose operating system. I can do whatever the hell you want. I can be a web server. I can be a database. I can, you know, get your coffee"—well, they probably do run Linux on coffee machines, but Talos Linux does one thing, and it's not starting from a general default and then trying to configure it. You explicitly have to enumerate: this is your interface, this is your certificates that you have for the cluster that you're going to join, this is the roles, this is the privileges that these roles have, this is whether you have bonding. So everything is declarative rather than trying to take a system that's generally applied, shape it to what you want, and then apply controls to it, which is, you know, that's what most systems do. Certainly the way operating systems worked back in my day—this is why I was so excited about Talos Linux when Andrew first told me about it—the standard approach is you have an operating system. You try and configure it the way that you think you want it to be, and then you try and enforce that through a tool like, you know, Chef or Puppet or CFEngine or whatever your flavor is. But that enforcement is never going to cover the complete range of every package that's installed, every configuration file that's installed on the system. It's just not practical. If you start from the other end and say, "Literally, everything that I'm going to do has to be specified in this configuration file and nothing else. If it's not in that configuration file, it doesn't get configured. If it is in that configuration file, there's no way for it to drift or change or deviate from what is designed in it to be." So it's just a way of approaching it from, I guess, bottom up rather than top down. Like, this is actually one of the reasons that Nokia, who was our first enterprise customer, came to us. When I asked them why they wanted Talos Linux—because, you know, we were a two-year-old company with probably, I think we had five people in the company when they signed on, which is amazing, so they were clearly taking a risk with us—and I asked them why, and they said, "We do use configuration management tools, but there's always something that is not in the configuration management tool that some engineer will go in and fix something. The configuration management tool won't notice it or revert it or control it. And then when there's an upgrade, which may be, you know, three or four months later, something will break in an unpredictable way because of this one parameter that we didn't catch. We'll then put it in the configuration management tool, so we'll catch that one. But this is a constant job that they were always chasing." They could never get a complete control by design system. And they're really smart people with great engineers, so they were battling this for a while. But they saw the concept that we could do control by design from the bottom up, completely eliminate configuration drift, and just not patch or handle that class of problems, but make that class of problems not exist. So different approach.

(Joel Beasley at 00:14:06) Yeah. Just eliminate the vector entirely and gain the security improvements.

(Steve Francis at 00:14:10) Yeah. It's like, you know, you're not putting more locks on your door. You're taking away the door entirely.

(Joel Beasley at 00:14:16) Yeah. Just the cement wall. Yeah. So what—there's a couple different angles here. There's security, there's uptime or reliability, as you mentioned, with the configuration drift. But what are some problems that CTOs listening might be experiencing where they would say, "Oh, okay, maybe I should take a look at this"?

(Steve Francis at 00:14:38) I think the most common one is things taking longer than you think they should. Like, that's a very common problem that we see when people are getting into the Kubernetes space. Because, you know, there are so many ways to do so many things, and figuring out the right way to do it, especially when you're a team that is new to Kubernetes—the right way to do it, and why there are trade-offs in many of the configuration options, trading off security for, you know, ease of use, like whether you're going to use the pod security policies or not. You know, it's much easier to turn them off. It's like, you know, when you use SELinux, everyone plays with it and then says, "Oh, let me turn it on," and then nothing works, and then everyone just turns it off or puts it in non-enforcing mode. But if you have your operating system that's designed to keep all the things secure by default and install Kubernetes for you so you know it works—you know, every version of Kubernetes is tested with a series of Talos Linux versions, so they're going to work correctly. There isn't going to be any configuration management needed to, like, "Oh, I'm upgrading Kubernetes. What do I need to change to make it work with the way my operating system's configured or vice versa?" It's just, you know, you just take away whole classes of things, which just makes—just lets your engineers work on the business-level problems, which is really what you want them to be solving.

(Joel Beasley at 00:16:13) That's right. Yeah. Absolutely. Now that this new technology exists, this Talos exists, now you might be able to look at it from the perspective, if I were a CTO, that running a general purpose operating system of 3,000, 10,000 binaries, that might become a liability.

(Steve Francis at 00:16:34) Yeah. I think it will be. Like, one of the newsletters that I like to subscribe—that I do subscribe to—is Matt Levine, The Money Stuff. He's a really clever writer about finance things, and I'm a bit of a finance geek. But one of the things he always likes to say is everything is securities fraud, by which he means, if there is any event in a public company and, of course, the stock price goes down, someone in the public is gonna say, you failed to disclose that this was a risk. The risk happened, and your stock went down.

(Steve Francis at 00:17:11) Therefore, I can sue you. I am very convinced that someone is gonna get hacked because they're running a general purpose Linux operating system and there's an exploitable CVE that's maybe even zero day. But some company is gonna suffer a financial loss because of that or at least even just a publicity hit because of that. Their stock will go down temporarily, and they will get sued because they failed to disclose that they were running a general purpose operating system that is vulnerable to many, many more CVEs than a special purpose operating system, and there will be a lawsuit as a result of it. So I'm convinced that'll happen probably in the next year or so.

(Joel Beasley at 00:17:50) Refunding the lawsuit?

(Steve Francis at 00:17:53) Actually, that would be a good idea.

(Joel Beasley at 00:17:55) There you go. Oh, goodness. Yeah. Because, yeah, a lot of the conversations they're having is about AI, and now with the AI advancing, the number of attack vectors becomes unbelievable opportunities.

(Steve Francis at 00:18:09) The number of exploits that, you know, Mythos, which has just been released, is gonna find—it's—and I'm sure there are vulnerabilities that, like, we run AI tools against Talos Linux, but for us to put Mythos against, you know, 15 binaries and figure out what exploits are there, that's a lot easier than how many can you chain together in a system running 3,000 binaries. That's a very different scale of problem. And so defending—it's much, much harder, figuring out what to address and remediate, and just the methods of attack on a specialized operating system are—most of them aren't there. Like, you know, most attacks don't actually—most attacks in the wild use people's credentials.

(Steve Francis at 00:19:02) We don't have users. There are no users on Talos, so you can't have a user's credential log in to the system. Most attacks use things like, I'm gonna load a kernel loadable module. You can't load a kernel loadable module on Talos Linux. It's just not permitted unless it was signed by the exact same signing key that was used to build the operating system and laid down on the disk at install time.

(Steve Francis at 00:19:27) And the key that was used to build the operating system is ephemeral. It's used to build it and then destroyed and thrown away. So you can't just say—have an attacker say, here's a kernel loadable module that's gonna obscure my tracks and hide my processes from ps aux because no one can load kernel loadable modules on Talos Linux. It just can't be done. So those attacks just don't exist.

(Steve Francis at 00:19:48) You can't chain things together with bash and cron and wget because they're not installed. So even if there's a vulnerability, the chance to exploit it is almost zero. Not definitively zero, but almost zero. So I do think that not running a specialized secure operating system will be a securities fraud action. Yeah.

(Joel Beasley at 00:20:14) Is there any of the big ones that you can point back to, like Southwest or CrowdStrike, any of the big meltdowns that could have been prevented by having this type of specialized operating system?

(Steve Francis at 00:20:25) Well, like, the copy fail Linux vulnerability that was around, whatever, two months ago, that didn't work on Talos Linux because you had to chain together binaries that weren't there. You know, CrowdStrike one was all on Windows, so you could probably say that wouldn't have happened if they'd run on any Linux system. But I actually think most security tools, you know, they fill a gap at the moment because general purpose operating systems are hard to secure and have so many moving parts and so many complexities and so many ways of configuring all the different modules that may or may not be installed. But they also act as a proof of concept for exploit code. So if you are running the CrowdStrike agent, it needs to do things like load a dynamic kernel loadable module so that it can elevate its own privileges and do the things that normally are exactly what an exploit code would do. So often, one of the resistances we do run into in some enterprises is they say, oh, we've standardized on this, you know, endpoint management system.

(Steve Francis at 00:21:36) And said endpoint management system may not run on Talos Linux because it tries to do things like, oh, I'm gonna load a dynamic kernel loadable module to elevate my privileges so I can inspect all traffic and see what processes are running and reasonable-sounding things to do if you're securing an insecure place. But we just don't let any of that stuff happen. And the fact that your security tool can't do it also proves that attackers can't do it. So in effect, you're secure. If you are insecure enough to need a security tool, the fact that you are running a security tool proves that you're vulnerable to attack. So the better approach is validate that none of the attack vectors are even there, and then you don't need that whole class of security. There are certainly other security tools you need, audit logging and stuff like that, which, you know, we have all of them. But you don't need the ability to load root-privileged host-level operating system things that anyone can kind of allow.

(Steve Francis at 00:22:33) So that's just not good practice.

(Joel Beasley at 00:22:35) Are most of your customers—they already are using the open source version of it, and then they want more? Or, like, are you having in your sales calls to just continuously explain to people that this concept of just having less reduces everything?

(Steve Francis at 00:22:53) No. Most of our customers are using the Talos open source Linux, and then they are talking to us either because they're like, alright, we've deployed this at a large enough scale that we do need support, or they are using Omni, our fleet management tool because it's like, yeah, we're running thousands of clusters, tens of thousands of clusters in some cases. And, you know, we have customers that have built their own tooling around that to manage for Talos Linux because Talos Linux is all API managed.

(Steve Francis at 00:23:20) That's doable, but we have—I think we have the best way of managing Talos Linux because we know its intricacies in and out. So that's what—

(Joel Beasley at 00:23:31) Yeah. If I'm a CTO and I've got mission critical workloads running on this, I'm probably gonna get—provided it's reasonable, I'm gonna get a contract with the people who know it inside and out. Because how could you not? Like, if you're running a large organization like a bank or Nokia or something, like, you have to, as a professional, make sure you've got the best people there to handle things that are mission critical.

(Steve Francis at 00:23:54) You would think. Surprising amount do not. They—there have been—there have been banks that the only reason that we found out they are running Talos Linux internally in production is because they gave a talk about it at KubeCon.

(Joel Beasley at 00:24:12) And I guess they could talk about it because it's so secure. Yeah. You know?

(Steve Francis at 00:24:16) Yeah. That one is—they're actually in procurement now two years after they gave the talk, but so they now are in procurement for support. But they were definitely running it in production for the last two years.

(Joel Beasley at 00:24:28) That is awesome, though.

(Steve Francis at 00:24:30) Yeah. Yeah. That goes back to what we were saying. We just made a product that's too easy and too secure.

(Joel Beasley at 00:24:37) Shame on you. Yes. A lot of the conversations come up this past week or two about panic, about AI taking jobs. I wanna hear what your thoughts are on that.

(Steve Francis at 00:24:52) Well, I think for smaller companies, you know, we're still a small startup. I think every piece of efficiency we get from AI is gonna lead us to employ more people. I think this is probably true in most companies. Most small companies at least is like, you're not constrained by, oh, I've done all the work I need to do. So if I get AI to make the work more efficient, I'll let people go.

(Steve Francis at 00:25:18) Your constraint is there's so much more I wanna do, but I can only afford to hire, you know, thirty, fifty people at the moment. If those 30 or 50 people are more efficient, that will mean we get more marketing material out. We make more features in the product. We have, you know, better, more efficient customer support. We can sell more more easily.

(Steve Francis at 00:25:39) If we can sell more more easily, that means we get more revenue. I would 100% use that to hire more people. That is entirely our constraint is that there's a lot more we wanna do. We're constrained by revenue and cash burn rather than anything else. It's not, oh, AI's let us do everything we need to do.

(Steve Francis at 00:25:56) I can get rid of all these people because I've achieved all I want now. It's like, no. AI will just make everything more achievable, which means I wanna hire more people. And, you know, I use AI all the time. I happen to have a law degree as well as a computer science degree.

(Steve Francis at 00:26:16) But if I didn't have a law degree, I could not be using AI to, like, review contracts and write master service agreements. You really have to know your domain to use AI effectively. So you still need people that know computer science to look at the code the AI has written or people that have law degrees to look at the legal contracts that the AI is suggesting. It's like, wait. Doesn't that contradict that?

(Steve Francis at 00:26:36) And then it's like, oh, yeah. My bad. I'm stupid. I'll go fix it. But you gotta point it out to it where the—you really have to know your data. It's like I was looking—doing financial analysis. It, you know, there's places where it's not making up stuff as much anymore, but it will just, like, miss things. And unless you really know your data and your domain, you're gonna have bad outputs from the AI. So I don't see any foreseeable future where AI is gonna make us hire less. It's gonna make us hire more, and depend on those people more.

(Joel Beasley at 00:27:10) Oh, yeah. Yeah. And I talk about that all the time with my friends and family and stuff and different software engineers as well because, you know, there will be times where I'm building an application, and the AI generates the code. And if you didn't—if, you know, I'd been a software engineer for twenty years. But if you didn't have that experience, you'd be like, yeah.

(Joel Beasley at 00:27:30) That looks okay. Right. But I could tell you the relationship between those two models is gonna change drastically the moment you add this next feature. So if you do it like that, it's gonna screw everything up in twenty minutes down the road when you're trying to do the thing that you know where it's going.

(Joel Beasley at 00:27:46) And so that's just experience because I just manually did that ten years ago, and I was, you know?

(Steve Francis at 00:27:54) And even knowing how to phrase the—what to ask AI to do is, like, talking about models and controls and stuff. Like, you need to know that to even have the AI try and do something useful.

(Joel Beasley at 00:28:05) Oh, yeah. It reminds me the first time that I, like, looked over someone's shoulder and watched them use Google. I was like, oh, you don't know how to use this thing. That's—

(Steve Francis at 00:28:17) Yeah. Yeah. So—which oddly, probably what they were doing is exactly the way you use Google now. You're like, you now do give it complete sentences and—

(Joel Beasley at 00:28:24) I know. I know. I know. It's fun, though. I like your view, though.

(Joel Beasley at 00:28:30) You know? I haven't heard that because a lot of times, the mental map that people have is that everybody loses their jobs. Yeah. Yeah. That's a lot of the mental maps out there.

(Joel Beasley at 00:28:44) But this idea of there's obviously a lot of fat in the huge mega corporations.

(Steve Francis at 00:28:50) But corporations are a different thing. Like, I've always been amazed with—like, you know, I've worked at SaaS companies my entire life. Some—maybe what was the biggest? Maybe it's 2,000 people. I've always been amazed.

(Steve Francis at 00:29:04) Like, what the hell does 200,000 people do to run a website, like, at Twitter or whatever it was? It's like, it's one website. How the hell do you need that many people? Apparently, you didn't need them all. I know.

(Joel Beasley at 00:29:17) I know. I was, like, I was all excited. I was like, keep going. Let's see how efficient we can make this thing.

(Steve Francis at 00:29:24) Yeah.

(Joel Beasley at 00:29:24) You know? Oh, good.

(Steve Francis at 00:29:26) Yeah. The mega corporations, I, you know, I've never worked in one, so maybe everyone has valuable productive jobs. But, like, just from my perspective of knowing what it's like to build a SaaS company and run a SaaS company at a reasonable revenue scale—not certainly nothing in the billions—but it's like, yeah. Even if you're scaling out your teams, I—you know, I've heard stories of teams of 10 software engineers and product and a product manager and technical product manager whose job was to work on one form on one page of one application.

(Joel Beasley at 00:30:00) It's—

(Steve Francis at 00:30:00) Like, that just sounds like a very inefficient and horrible job.

(Joel Beasley at 00:30:05) It does. And it's gotta be boring. Like, as a software engineer, it's like, what are you doing?

(Steve Francis at 00:30:10) Very boring. And, also, literally, how can it take 12 people to do one form on one page?

(Joel Beasley at 00:30:16) I don't know. Never underestimate the power of incompetence. What happens is you're so focused and driven in founder mindset that, like, it's actually hard for you to imagine that reality.

(Steve Francis at 00:30:28) It is. For sure.

(Joel Beasley at 00:30:31) Okay. There's a couple questions here about Steve's hot takes. Is there a popular belief in the infrastructure world that you just completely disagree with?

(Steve Francis at 00:30:41) Well, I will say, I don't know that it's that popular, but I think it's pretty clear that multi-cloud is kind of—it is not done for fault tolerance and resiliency. It is done for it's too hard to move off. We got into two clouds. We acquired another business that's in a third cloud. It's too complicated.

(Steve Francis at 00:31:04) We're too enmeshed. There's no way—in theory, running everything in, you know, one application here and the same application in a different cloud is a great concept. It gives you high availability and redundancy. In fact, you know, in theory, theory and practice are the same. In practice, they are not.

(Steve Francis at 00:31:22) But running multi-cloud, you end up being enmeshed into the identity systems and the API calls of Amazon and different ones for Azure. And when you need to move from one to the other, it's just like, oh, this stuff just doesn't work over there because all this system that we inadvertently used requires these API calls that don't exist on that one. They've got different way of doing it, different identity managements, different platforms. That is something we actually thought was gonna be the value proposition of Talos Linux. Talos Linux, you can deploy the exact same operating system on all the cloud providers on VMware, on bare metal.

(Steve Francis at 00:32:02) And because it's API managed, you get the exact same API to manage it on all the different cloud providers and the same Kubernetes versions and with the same Kubernetes APIs. Turns out no one cares about that. That was what we thought was gonna be the value—that you could run a Kubernetes cluster on one place and would be identical to the way the Kubernetes cluster runs on, you know, Azure or Google or Microsoft—Amazon—or your bare metal or VMware. But people are just like, that's not why they do it. They run in this place because that's where they run it, and they don't need to migrate from one to the other.

(Steve Francis at 00:32:38) They should, perhaps, if they want to do true disaster recovery and test proper failover, but that's not what people do. They just like, we're in there. We're accepting that cost. It's too much work, so we'll just leave it there. So that was actually a mistake we made.

(Steve Francis at 00:32:52) I really thought the portability was going to be a value proposition, but no one just no one cares.

(Joel Beasley at 00:32:59) But it is still a pro because if you're an engineer and you really like Talos Linux and you move to a different company or a different team that's using a different cloud, you don't have to relearn this. Your tool will work regardless of what cloud you're in.

(Steve Francis at 00:33:15) Yeah. Yeah.

(Joel Beasley at 00:33:16) So there's a benefit for that because, like me, most like 30 plus percent of my business is repeat customers, and that's awesome. They're just like, hey, we want to do this or that. And so the fact that you can stay as the solution with different engineers as they move around, which they do, is a huge pro.

(Steve Francis at 00:33:34) Yeah.

(Joel Beasley at 00:33:35) Now we just have to convince them all to get support contracts.

(Steve Francis at 00:33:38) Right. Exactly. Yeah. I also think that the cloud providers and the hyperscalers are—I don't know if this is true, but it just seems like it is—that they are not incented to make Kubernetes or other systems simple. Like, complexity works to their advantage.

(Steve Francis at 00:33:59) For them, it's like, oh, no. Kubernetes is scary, and you don't want to run it. Don't worry your pretty little head about it. Just run in our, our hyperscaler version of EKS or whatever.

(Steve Francis at 00:34:11) We'll take care of it for you and take care of the magic behind the scenes. So they are not motivated to make Kubernetes simple and reliable and secure. It's in their interest to keep it on bare metal, to keep it fragile and insecure and hard to manage. So I don't know that they're actively doing that, but it's certainly, you know, as they say, show me the incentives, I'll show you the outcome. Their incentive is not to make it easy.

(Joel Beasley at 00:34:39) Yeah. And you always follow the incentives. Yeah. That's the most important thing. What is the most overrated thing in platform engineering right now?

(Steve Francis at 00:34:50) It's probably the idea that you can build a big developer platform to shield your developers from all the complexity, because you're fundamentally taking a complex operating system, putting complex Kubernetes on it, and then a complex developer platform on top of that to try and shield your engineers so that they can just write and deploy code. And that can work, but you've just got so many layers of complexity now, one on top of the other, that it ends up being quite fragile. From what I've seen, the places that have tried it haven't had huge success with it. I haven't seen everywhere, and I've certainly heard stories of people that are doing it well. But the building those kind of platform for developer environments always take longer, like years longer from what we've seen in cases, and don't always deliver the feature set that they were envisioned with.

(Steve Francis at 00:35:52) So I think that's just a—I think it'll get there, but I think it'll get there a lot quicker and a lot better once people start realizing simplicity is—you know, the solution to more complexity is not always to add more complexity on top. It's to rethink what is complex, how can I make that simple, what do I really need out of this, have it just deliver that? And then it makes—oh, now I've got a much more simple set of API calls that I need to make, that much less ways to go wrong, much less inputs needed. Things work better that way.

(Joel Beasley at 00:36:26) One of my favorite things on your website was this, like, anti-hero infrastructure concept. How did you guys come up with that?

(Steve Francis at 00:36:34) So I think Tyler, our head of marketing, came up with it. But Jeff, who's our chief product officer, he's—heroics is kind of an anathema to him. It's like, if there's something that can only be solved by someone doing extraordinary work because they know all the secrets and they've done the things, your system is bad. It's set up wrongly. You want no heroics.

(Steve Francis at 00:36:58) You want—if there's a failure, you want to be able to be, you know, you want to be able to be unreachable out surfing and just know that everything's going to just work because no heroics are required, because things should automatically fail over. If they don't, it's simple, and the on-call engineer knows exactly what to do because, you know, there's only like three things you can do. You don't need to know—it's like, oh, this one, when it's using, when it fails over to this VPN route, I've got ICMP filtering on it, and it's blocked the do-not-fragment bit. And so my MTU path discovery is not working, but only over this route.

(Steve Francis at 00:37:33) So that's why it's looking weird over here. It's like, you don't need to know that stuff. You want things to be simple, have no complexity, which means you don't need heroes. You want people that can productively get stuff done as part of a team and contribute to the business moving forward. Heroes are like the opposite of what you want.

(Joel Beasley at 00:37:53) You want to not need heroes.

(Steve Francis at 00:37:55) Yeah. Totally. Heroes often in the software world, they often have ego problems, which means they're not always great team players.

(Joel Beasley at 00:38:07) Oh, I mean, we all know some of those. Yeah. Yeah.

(Steve Francis at 00:38:11) Often, they're very smart and great people and valuable in a way, but it's like they're not who you want to work with on an ongoing basis.

(Joel Beasley at 00:38:18) I remember the first time I got to work with one of them. I was like, I've heard about you. It was like seeing a unicorn. I was like, oh, it did not disappoint. I was like, this is amazing.

(Joel Beasley at 00:38:29) Yeah. I had the time of my life. And then, that's the beauty of my career path being consultive or consultative and then my own business owner is I wasn't stuck with that person, but I got to see them and meet them. And I was like, oh, no.

(Joel Beasley at 00:38:45) What part of Kubernetes stack do you think is going to feel outdated in five years?

(Steve Francis at 00:38:54) I would hope storage becomes much, much simpler. Like, there's a variety of ways of doing storage. They seem to come and go. You know, at the end of the day, it's like, all right. You want some block storage.

(Steve Francis at 00:39:10) You want it to move with your virtual machine, container, and you may or may not want it, you know, replicate it in high availability. These aren't hard problems. Like, you know, whether you've got RAID 0 or RAID 1 or whatever with high performance disks has been around for a long time. The ability to map it to the different places in your Kubernetes cluster, I just think there's a lot of opportunity there to make that whole system much less complicated and much less to worry about. Like, setting up a Ceph cluster right now, you can do everything with it. It's amazing.

(Steve Francis at 00:39:50) You can get great scale-out performance. Not easy. You know, iSCSI and NVMe over TCP, it's like there's lots of different ways of doing it, but I think it'll kind of consolidate down to just be like, we've got, you know, file storage and object storage. I'm optimistic that that whole system will just get much simpler.

(Steve Francis at 00:40:15) Like, there's always been lots and lots of development in storage, and it's been a great wave where companies have added value. You know, NetApp used to be my favorite storage system from back when I ran data centers. I loved their technology. I thought it was super cool, the way they did their own parity calculations and the WAFL file system. But I don't know.

(Steve Francis at 00:40:39) I'm just optimistic that things will get much simpler in the storage world.

(Joel Beasley at 00:40:43) What's your advice for CTOs that are listening in and they're actually trying to simplify their infrastructure?

(Steve Francis at 00:40:49) Well, start with the operating system and look at what you really need. Like I said, OpenShift does a great job for when you need everything because they have, you know, opinionated ways of doing things. There are certainly downsides to that. To upgrade an OpenShift node can take, you know, multiple hours. To upgrade a Talos Linux node—it does an image.

(Steve Francis at 00:41:11) It's an image-based system. There's no patching of components, so you know exactly what state you're in. We do a kexec and a reboot in five minutes. If you're running in a data center on this particular cluster, you don't need all the capabilities, do as minimal as possible. The minimal you have on whatever you're doing is going to be simpler.

(Steve Francis at 00:41:30) It's going to be more secure. It's going to be more reliable. So that's kind of the high level philosophy that I would say for everything. Just figure out what you really need. Do that. Don't do anything more.

(Steve Francis at 00:41:41) Don't put in capabilities because you might want to do some system that you don't need at the moment. You don't need—if you're running at the edge and it's you're running single node clusters, you don't need, you know, fancy load balancers because you have load balancing in one amongst one node. So put in as little systems as you need as you can do for your current capabilities. Doing that will just mean you'll have less to manage, less configuration, less complexity, less security issues, and everything will be faster, and that's what everything needs.

(Joel Beasley at 00:42:15) Yeah. It's like people shopping when they're hungry. It's like grabbing everything. Exactly. It's like just what meal are we making tonight? Let's just make that meal.

(Steve Francis at 00:42:22) Yeah. That's a good analogy.

(Joel Beasley at 00:42:24) So you're a leader. You're running the company as CEO. What are you learning right now as a leader?

(Steve Francis at 00:42:31) Well, as going into this, because it was my second startup, I knew it would be harder and take longer than I thought. I'm still learning that. Even though I knew it, it's still true. It's still surprising me. But the complexity of—like, we're getting to the point now where we are getting the relationships with NetApp and Cisco, and they're paying attention to us and Dell and these companies.

(Steve Francis at 00:42:59) So we're getting partnerships with them. The time for that took a lot longer than I thought. You know, and what you were saying, getting back to focusing on what actually moves the business, we had a period of time where it's like, oh, we'll try these smaller partnerships, which, you know, they would come to us. We were, you know, roughly the same size in the business, and it's like, all right. Well, Dell isn't paying attention to us, so we'll try these smaller VAR and reseller channels.

(Steve Francis at 00:43:30) But that just kind of sucked up time and didn't lead to anything on either side. So getting back to testing what works, sticking to figuring out what problem you're trying to solve, which KPI will measure it, backtracking from that as to what experiments you're going to run to influence that KPI, then seeing if it works. All right. What we should have done is instead of investing—you know, we didn't invest that much, but we invested more than we should in smaller channels and didn't get any benefit out of them. So we should have done an experiment, one or two of them.

(Steve Francis at 00:44:00) It's like, nope. That's not working. Even if we can't get the attention of the big ones, it's not worth working with the smaller channel at the moment. So we'll just hold off on that and focus our attention on what is more productive. So nothing new and earth shattering.

(Steve Francis at 00:44:11) It's just, you know, the fundamentals of focusing on what works, coming up with a theory to test what you're doing, do the test, and learn from the results. Yeah.

(Joel Beasley at 00:44:22) Very Eric Ries of you.

(Steve Francis at 00:44:24) Yeah. Exactly.

(Joel Beasley at 00:44:26) He just had a new book out. They just sent it to me called Incorruptible. I haven't read it yet, though.

(Steve Francis at 00:44:30) Oh, you haven't read it yet?

(Joel Beasley at 00:44:30) No. No. I liked his startup book before.

(Steve Francis at 00:44:34) Yeah.

(Joel Beasley at 00:44:36) And then one of the things I always like to ask leaders that I interview on this show is for a piece of leadership advice, but I have some constraints. The constraint would be that you heard it a while ago, you implemented it, and you're like, oh, this is good, and you've kept it with you for a long period of time.

(Steve Francis at 00:44:57) This is kind of general leadership advice, but also general everything advice. Like, the Jocko Willink book on extreme ownership. Like, that I think is my kind of go-to book on the philosophy of business and personality and that. It's like, whatever is—whatever you're not happy with—I mean, this is especially true as a CEO, but it's like I think it's true of everyone in every level of every organization. Whatever you're unhappy with, it's your responsibility to fix it, change it, identify it, come up with a solution, or change the situation somehow. And maybe, you know, if you try and if you're a product manager or developer at a company and there's a problem and the company is not moving fast enough or you identify issues, you try and raise it, you have the communication and it doesn't work.

(Steve Francis at 00:45:51) Maybe the only thing you can do is leave to a different company, but, you know, there's always the things that are within your control. But to me, that's like everything that I see in the company is like, oh, this person is not interacting well or getting enough progress done on this particular issue. It's not that person's fault. It's my fault because they need help or I haven't coached them or I haven't pointed out the problem to them clearly enough or I haven't explained the consequences of the problem. Like, everything is my fault, and that sounds negative, but it's actually very empowering.

(Steve Francis at 00:46:28) Because it's like, oh, that means there's something I can do, you know, about everything, about if the company is not growing fast enough, that's my fault. Great. I can do something about that. If the engineering is not developing fast enough, I'm not an engineer, but that's my fault. Great.

(Steve Francis at 00:46:41) I can do something about that. I can have a discussion. I can set KPIs. I can change priorities. It's like that I think is just, you know, sounds negative, but it's actually really empowering.

(Steve Francis at 00:46:51) And that took a while to sink in, and I still hold that true.

(Joel Beasley at 00:46:55) I love that because that specific book was a pivotal moment in my career. When my business partner handed it to me, and he's—I was, we're building some financial software, and he's like, I want you to read this book. I think it would help. And I didn't read a lot at the time, especially those types of books. And I read it, and I adopted that thought process, and, whoa.

(Joel Beasley at 00:47:21) Yeah. Some people find it very uncomfortable. Some people, it's like, don't like it. They're allergic to it. Yeah.

(Joel Beasley at 00:47:28) So I run it on my—I run it in myself. That is a mental process that's still running to this day. And it does give you an interesting perspective too, because now when I'll look at—you know, we had—I don't have a lot of employees now. But when I did, I'd say, oh, you know, it's my fault because I hired them, and I didn't do a good enough job in my hiring process to weed out this, you know, behavior or whatever that we don't want. Yep.

(Joel Beasley at 00:47:56) And, yeah, it just—you can't explain it to your wife, though. It's very hard to explain to your wife. So that's one of the things I've learned. And I still find myself, like, how do I teach this to my eight-year-old? You know?

(Joel Beasley at 00:48:11) It's like you just got to kind of wait for them to mature in life. But—

(Steve Francis at 00:48:15) Well, I had to wait for me to mature into that. I first read that book, and it's like at first, it was like, yeah, it's interesting, but it didn't quite sink in. And then I reread it. I don't know. I've read it many times now.

(Steve Francis at 00:48:25) But yeah. The second time I read it, it was like, oh, shit. I get this now.

(Joel Beasley at 00:48:29) Oh yeah, and like you said, when you come around to it and actually start trying it, it's incredibly empowering. You're like, oh, I have control.

(Steve Francis at 00:48:38) Over everything.

(Joel Beasley at 00:48:40) Over everything. And I can now choose how to shape my life, and now I'm responsible for me. You know, I wanna be a—

(Steve Francis at 00:48:47) Scary and empowering. But it's like, yeah, you go from being a character in your story to being the author.

(Joel Beasley at 00:48:55) Yes, absolutely. So as we start to wrap up, let's see here. Where can people learn more about Talos Linux and Omni?

(Steve Francis at 00:49:06) siderolabs.com is the best place.

(Joel Beasley at 00:49:09) Okay, we'll put a link in the show notes. And then if people wanna reach out to you, what's the best way to do that?

(Steve Francis at 00:49:14) Easy. Email is [email protected] or LinkedIn, Steve Francis, of some variation.

(Joel Beasley at 00:49:23) Yeah, just tell him you have some offshore developers and you wanna help him. That's like the most common LinkedIn message I get every time. Like, nope, I get that 10 times a day, Steve.

(Steve Francis at 00:49:32) Yeah, so do I. Yeah. I mean, I don't check LinkedIn all the time because I get so many of them.

(Joel Beasley at 00:49:38) Oh, I know. I'm checking it every couple weeks, but it's just everyone. And then what's the thing that you are most excited about at Sidero Labs?

(Steve Francis at 00:49:53) Oh, we've been doing this for six years. We've finally, I think, we've got great recognition. The open—like, if you go on to Reddit and you ask, you know, what operating system should I use for my home lab or my bare metal cluster, you will get 20 people suggesting Talos Linux, and they're not us anymore. Like, three, four years ago when people would ask that question, I would go in there and suggest it. But now it's like, we don't need to. We've got great community recognition. So that's really exciting. And now we are getting that into the enterprise. One thing I'm actually oddly excited about is the regulation that's coming into the EU with the Cyber Resiliency Act and the NIS 2. They're gonna require companies that are using open source to show that they have secure supply chain management, which means they're gonna have to have a commercial relationship with us. So that's great.

(Joel Beasley at 00:50:51) Oh, awesome. So if you guys need that relationship, give them a call. Steve, we made a podcast. How do you feel?

(Steve Francis at 00:50:58) Oh, I had a great time. Very exciting.

(Joel Beasley at 00:51:01) Thank you so much for listening. And if you found this episode useful, please share it with a friend or a colleague who you think would get value from it. And if you have topics that you'd like to hear discussed on the podcast, either add me on LinkedIn or send me an email, [email protected]. Every time I get an email or LinkedIn message, it absolutely makes my day and inspires me to keep going.