Episode 404 ·

Edward Kmett, Head of Software Engineering at Groq - Incredible Processing for AI Models

Today we’re talking to Edward Kmett, the Head of Software Engineering at Groq. And we discuss how Edward came to be one of the most prominent members of the Haskell community. How Groq’s chips are able to provide incredible processing capabilities for AI models, and how Category Theory gives us a different way to think about mathematics. 

All of this, right here, right now, on the Modern CTO Podcast!

To learn more about Groq, check them out at https://groq.com

In case you missed it: check out our episode with Groq's Founder and CEO Jonathan Ross

About Edward Kmett:

I work at Groq with an eye towards building a better future. I also strongly believe in effective altruism and serve on the boards of the Topos Institute and the Haskell Foundation.

I spent most of my adult life trying to build reusable code in imperative languages before realizing I was building castles in sand. I converted to Haskell in 2006 while searching for better building materials. I served as the founding chair the Haskell core libraries committee and collaborate with hundreds of other developers on over three hundred projects on github.

I am obsessed with building better tools so that in seven years I won’t be stuck solving the same problems with fundamentally the same tools I used seven years ago.

About Groq:

Groq designs radically simple elegant processor architecture technology to accelerate complex workloads in artificial intelligence, machine learning, and high-performance computing.

Transcript

(Intro Narrator at 00:00:03) Hello, my friends. Today, Joel is talking to Edward, the head of software engineering at Groq. And they discuss how Edward came to be one of the most prominent members of the Haskell community, how Groq's chips are able to provide incredible processing capabilities for AI models, and how category theory gives us a different way to think about mathematics. All of this right here, right now, on the Modern CTO podcast.

(Joel Beasley at 00:00:32) Here we go. This is the Modern CTO podcast.

(Edward at 00:00:44) I started very young. I had an uncle who used to work for RCA back in the day, and so he was a telecom guy, or a television guy rather than a telecom guy. But he had worked in the Air Force, like, flown around from base to base playing chess with generals. That was how he basically dodged the draft or, like, sidestepped going to Vietnam or what have you.

(Edward at 00:01:10) And then when I was like six or seven, he helped buy me my first computer. So I first started off on his PET, and then he got sick of me borrowing his PET. And so later on, he helped me get a VIC-20 and a Commodore 64. And I started in on those as my sort of introductory platform. And along came the Commodore 128, and I thought that the CP/M mode that you could boot into on the Commodore 128 was the serious business way that people were supposed to program, but it was boring. It didn't have any programs for it. So I went and wrote a whole lot of software just by porting things over from the Commodore 64 mode into the CP/M mode, and then sold a ton of shareware early on. And when all my friends had paper routes and whatnot, I was able to just write some code this weekend, and there'd be, you know, $2,000 or $3,000 or whatever. It was like this money tree, and my mother never really understood it.

(Edward at 00:02:15) So that's how I got started, was through that, then into the bulletin board scene. Then right as Netscape Navigator shipped, I built an anime import-export business named OtakuLand.com. It's long since dead and has been purchased by other people or whatever. I was running that for a couple of years and then helped build a phone company. So the ISP that I had been working with was an ISP named mish.com. They had been hosting my anime import business, and I really got to know the owner of the business pretty well. And so this was around 1996 at this point. This was right at the time that there was the last wave of breakup of Ma Bell where they were trying to—like, back in the day, Congress believed that monopolies were a bad idea, and they kept trying to break up anything that looked too big. They got away from that mindset recently, it seems. But back then, it was still not considered cool that Ma Bell was this gigantic monopoly.

(Edward at 00:03:22) So they broke it up into lots of little monopolies by geographical region, which didn't do anything to actually serve the purposes. And then they kept trying to fix that by making it so that if I'm a phone company and you're a phone company, and I have calls terminated in your—like, basically, they kept trying to foster a competitive local exchange carrier market, and they never really quite got it right. And so I was part of the crowd that realized that they could take advantage of that last wave of laws to foster the creation of ISP businesses. Because if I'm getting paid by the hour for the calls that terminated my hardware and I'm an ISP and I only ever receive calls, then it's just free money. There was a whole wave of them at the time.

(Joel Beasley at 00:04:06) So did that company make you a billionaire?

(Edward at 00:04:08) No. I made and lost a lot of money. In the end, I lost a lot more money than I made. So I wound up crushingly in debt. And so I decided to do what everybody does when they're that deep in debt is add more in the form of student loans. Because I'd been a CTO or a CIO, technically, at that point for like seven years, and no one would believe that I would stay for anything entry level. And no one wanted a CTO without a formal education right in the dot-com crash. And I was completely stuck. I spent like a year on my now wife's couch just doing nothing. And then I said, okay. Fine. I'll go back and fix this education slash work-life experience imbalance for me. And yeah. So I went and started and finished my undergrad in 2004. So I did a double bachelor's in math and computer science.

(Edward at 00:05:03) The next semester, I did a master's in math. The next semester, I did a graduate certificate in AI. Met all the requirements for a master's in bioinformatics. Didn't take it. Did a grad certificate, met all the requirements for my master's in computer science, audited a master's in computational linguistics. And by the time I was done with that, it was three years later, and I had fixed my education imbalance. But I was really bored with the person that I had wanted to be before that, if that makes sense. I had changed a fair bit in that time. And at the very end of that, I discovered this programming language called Haskell. It wasn't actually part of my curriculum in any way, shape, or form. It was just a community of folks that were doing very interesting things with the way that they compiled their programming languages. They had better theory than I did. I had been building—cut up with this. When I started doing the phone company thing, I had been very active in a community called the demoscene.

(Edward at 00:06:02) A bunch of graphics programmers who like to do flashy graphics, kind of grew out of the sort of cracking scene for games. They would try to put little files in there that would be like, here's a little .com file that you'll actually execute that happens to scroll past some fancy graphics to their friends and what have you. So I got very active in that community for a while. You know, I was wearing the suit in the day and then coding frantically at night, and I kinda kept that cadence to my entire career. And it's how—how to put this? So when I found the Haskell community, I had been—I'd gone through this weird time of being completely self-taught, then having gone through a formal education very late. So I had kind of two different perspectives, and I came into Haskell having been trying to build little toy programming languages that all looked like the bastard child of C++ and Perl and Python and whatnot, because they were what I knew, with a little bit of immutability thrown in. Like, if I can't mutate a thing, then I can share it across threads. I don't have to do any locking around it. There's a lot of extra benefits.

(Edward at 00:07:09) And I had taken it so far. Like, I think I can move more of my calculations to compile time and what have you. But the Haskell community had really embraced this idea. Like, what is immutability? What does it mean for everything that you learn in the undergraduate computer science degree about amortization of algorithms, all this concept, and really push it in a new direction. And I kinda had to decide, do I make their tricks my own, or do I stick my head in the sand and pretend this never happened? You know, I never saw this. And I eventually just decided to sort of reinvent myself as a Haskell person, and it seems to have worked. So I had been—because of the whole demoscene, like, doing all my stuff, kind of coding in the evenings—I had always logged into IRC or whatever it is under pseudonyms, and like, even when I first started doing that in the Haskell community. And so when I took all my fumbling missteps, it was all under random nicknames and whatnot.

(Edward at 00:08:13) And then I looked around the channel, and I started getting more and more self-conscious about the fact that everyone there was under their real names. So I eventually just put down my real name, and no one saw me take those fumbling missteps. I just sprang forth fully formed from the brow of Zeus knowing Haskell. Right? And I don't know, about a year or two after that, they kept starting putting me in charge of more and more things. So that's how I got to the point where I'm at. I had been trying to write reusable code. It was a sort of mission that I had had from when I was a little kid, right, like, transcoding all that code from the 6502 to the Z80 for the CP/M mode for the Commodore. Like, it was really so crushingly, depressingly boring to rewrite the same kind of code over and over again. Whereas all I really wanted to do was be able to build a pile of code that I can keep standing on top of to build the next layer.

(Edward at 00:09:08) And I kept feeling like whenever I would do that with OOP in other languages, that I was like building castles in sand, and it was all kinda crumbling around me that I was trying to climb up, but could never really get to that next level. I was just kinda failing for structural material reasons. And Haskell seems to be, I mean, for reasons I did not anticipate when I found it, very good at letting me build reusable pieces of code that all interlock together. So now I maintain some ridiculous amount of code.

(Joel Beasley at 00:09:37) Is that what was on SchoolOfHaskell.com? I saw you had a framework there. You said it was changed to like read-only now.

(Edward at 00:09:44) Oh, so School of Haskell is a site that was started by FP Complete, which is a company that kinda came into the community and tried to be the commercial face of Haskell. And to some positive extent and some—it didn't quite achieve all of their initial aims. So School of Haskell was a website they put up, which originally had a very interactive blogging framework where you could write Haskell code and it would run on the website. And so I did a bunch of tutorials on there. Most of my original content is on an older website called comonad.com. Yeah. They were just one of the places that I was putting stuff on top. It's really the short version of that. But I had been trying to write reusable libraries, and I got to Haskell, and I thought everyone in Haskell knew this branch of math called category theory. And I just assumed that it was endemic to the culture, because the one person that I was really seriously talking to about Haskell knew it. And then we were sitting in IRC, and there were 300 people in the channel, and they were all quiet.

(Edward at 00:10:44) And I just assumed that meant that they knew everything that this guy was talking about. I didn't realize he was, you know, 40 years old living in his mom's basement just reading papers all day. And so I spent the next six months reading everything I could on category theory and learned it and started blogging to an audience that I imagined existed. And it took a couple of years for it to show up, but then it did. And so now Haskell people are kinda known as the category theory people. Sort of the wish fulfillment seems to have worked out.

(Joel Beasley at 00:11:13) I wanna stop you right there. Okay. Because you did something that's interesting, and you kinda glossed over it. You started doing this thing. It was very niche. And it didn't work out at first. You did it for several years and then it paid off. What was your mindset like going through those several years where there weren't a lot of people?

(Edward at 00:11:33) Yeah. So for me, I joined the Haskell community. It was like it was the smartest group of people I've ever met on the Internet. And I felt like a complete impostor the entire time that I was hanging out, talking to these people, and trying to write down a bunch of category theory stuff, and doing it poorly. And at the time, so this was after I had finished that journey through education, and I was doing defense contracting for a while in Boston. And then after I was doing defense contracting, I landed at Standard & Poor's. And I was doing, like, I helped them build a programming language for them that was used for financial reporting that's used inside of the way they do their portfolio evaluation and factor back-testing engine.

(Joel Beasley at 00:12:19) When you said portfolio evaluation, I think I did that for like three years. I had to build a back-testing engine. You know, we back-tested market crashes and then showed how the portfolio would perform in these market crashes. And then we would do things like how you would withdraw your money in the most tax-advantageous way over the course of your retirement, because you're taking a little bit from different accounts. And we had to model all of these different products, right? Because you get annuities, you get all of these different types of income streams and products. And then it's like how you take what you need from the different streams in the best way. I don't know. It was like the nerdiest thing I've probably ever done. And it's weird because I go into those things and I would learn them and learn all the intricacies of them. And then I go to the next project. Like, after that, it was like a fitness project, and I just completely dropped like 90% of that knowledge. My brain just doesn't retain it.

(Edward at 00:13:09) I will admit it's mostly gone from mine as well. I had—so I went and gave a talk at MIT on like reducing big datasets in parallel using very weak algebraic structures, groups, and monoids, and what have you. There were a couple of people in the audience from Standard & Poor's, Paul Chiusano and Rúnar Bjarnason, who decided they were gonna take my approach that I was describing in Haskell and reimplement it in Scala so they could use it on top of the Java backend that they were using for all of this financial data processing at Standard & Poor's. And they, because they had to replace their existing portfolio evaluation and factor engine and all that. And when they did so, it took them from taking all of the memory on the biggest machines they could throw at the problems to running in a constant amount of memory, you know, with a tenth of the amount of code, it was more easily maintained. Basically, no matter what metric you pick, it was on the Pareto frontier. And they brought me in once they got in kind of over their head as consultant. And so I stuck around there for several years, actually, until I had a cancer scare at the tail end of it. And they were very, very nice about, you know, like, oh, like, go do research and all this kind of stuff. But I really wanted to sink my teeth into a problem.

(Edward at 00:14:21) And so I eventually left there to go work for a company called Digital Asset for a while, built another—or worked on another programming language for them, the Digital Asset Modeling Language, which is used to run smart contracts on blockchain-like things. The major clients there are stock exchanges. So the Hong Kong Exchange and the Australian Securities Exchange are doing trials, as I understand it, at the moment of kinda switching out settlement. So instead of taking T—you know, the time of transaction plus a couple of days of people shuffling things around in the back office—if you could turn around and settle it in five seconds, that's a huge amount of float removed from the market.

(Joel Beasley at 00:15:02) Oh, yeah. We've talked extensively about that on the show. We've had a number of different financial people come on, from the creators of different cryptocurrencies that are trying to solve that problem, to different financial software companies. It was so fascinating to find out, if you do an international transfer, this concept of there's like hops between these multiple ancient technologies and they're taking little pieces of it each hop. It seems absurd if you tell it to a modern software engineer.

(Edward at 00:15:32) It really is. It's horrifying. If you were just to do settlement for stock exchanges, I believe it's something like $600 billion worth of float that could be removed to sort of narrow bid-ask spreads and all that kind of stuff without the downsides of high-frequency trading and the like. So Digital Asset got their foot in the door at the Australian Securities Exchange, which was, I think, the largest exchange that was sort of correctly positioned in a way that would actually want to adopt new technologies in the space, that didn't have incentives misaligned with that. So that's how that happened.

(Edward at 00:16:14) And so after Digital Asset, I landed at an organization called MIRI, the Machine Intelligence Research Institute, that does AI safety research. If you've run into LessWrong or rationalists on the Internet, there's a whole halo around that that is fairly polarizing one way or the other. Their executive director had reached out to me when I was still working for Digital Asset and asked me what it would take for me to stop working for finance companies and just work on the things that I believe need to exist in this world, to continue to do mostly my open source work. And I gave him a number, and he went away for six months and then eventually came back and said, "Okay." And I was not expecting that to happen.

(Edward at 00:17:00) It was more or less a "go away, don't bother me" kind of situation. And I had to kind of decide, well, do I carry on in the direction I'm going, or do I go off to California and do all of this? And I have to say, that was probably the most rewarding experience that I have had in my career, this sort of blank check to do research and to sit there and interact with a scarily smart group of people. So I was there for, I think, four years or so. And then this summer, we kind of—I mean, what happened was while I was still there, I was approached by Jonathan Ross.

(Edward at 00:17:38) I believe Jonathan did an interview with you guys at one point.

(Joel Beasley at 00:17:41) Oh, yeah. He's like alien-level smart. Yeah.

(Edward at 00:17:44) Yes. And so he, not quite cold-called me, but pretty darn near, through a mutual friend of ours. And I spent a scary amount on compute last year just trying to do my own research. And the idea of anything that will help me mitigate that made it through the shields that I had up. So eventually, he talked me into kind of coming on a day a week to help them out, understanding the finance space, the defense space.

(Edward at 00:18:12) You know, they're reasonably heavily invested in Haskell as a technology stack, and I guess I'm the Haskell guy. And so there was enough overlap there that that got me a foot in the door. Let's just say that, you know, due to the way that crypto went this year, I kind of transitioned from a place where maybe I shouldn't be taking MIRI's money. Maybe I should be donating to MIRI-related causes and things in the effective altruism space. And right around the time I was kind of deciding to make that jump in my life, I talked to Jonathan and kind of decided to come on full-time.

(Edward at 00:18:48) And so he offered me a role at Groq as a fellow in the sort of Google Fellow sense. And then I looked around at Groq and said, "Okay. Well, if I'm going to be here, voting with my feet, mainly because one of the things I've been working on at MIRI had been, how do I make functional programming, logic programming, formal methods scale? And the Groq hardware was the closest thing I had found to being able to execute the things that I wanted to do at scale." Then maybe what I should do is vote with my feet, join Groq, make those couple of changes—couple of small changes to the silicon or whatever—that I feel need to exist for that to work, because I'd gone as far as I could outside the tent.

(Edward at 00:19:30) And so I joined Groq, looked around, realized that there were some places in the software stack that could use some attention, and started trying to manage that a little bit from the side and then realized that it wasn't going to work unless I just accepted the fact that I was going to have to run the software engineering side of Groq and told Jonathan, "I guess I'm going to do this." And now that's where I'm at. So that's a very long-winded path to get here, but that's the story of how I landed at Groq.

(Joel Beasley at 00:20:05) Are you loving it?

(Edward at 00:20:06) I really am. It's a very good culture. And I know that people say the word "culture" as it's some touchpoint. "Oh, yes, and we have a good culture," or something. In Groq's case, I think some of it is—almost everywhere I've been, there's a system of climbers or people who are just trying to clamber over each other and get to where they want to be with as much headcount under them as they can get.

(Edward at 00:20:32) And I think it's almost accidental. It's sort of part of the teething pains that Groq had, that it had to pass through this very narrow valley where they crushed down to 35 employees and almost ran out of money. All the people who were there with those kind of ambitions seem to have been kind of boiled off. And then since then, there's been a strong focus on sort of maintaining a healthy culture, which, given the core nucleus they were able to build, has actually seems to have worked. So I've been very impressed about that.

(Edward at 00:21:07) And I think Jonathan really leads that charge more than anyone at Groq.

(Joel Beasley at 00:21:12) I'm excited about Groq. What you guys are doing with chips is—well, you want to explain it? We went into this assuming, because I know so much about it, I've gotten to talk with Jonathan, but people who are just hearing this for the first time, can you explain what Groq is, what they do?

(Edward at 00:21:25) Yeah. I can. It's a little tough given that I'm a very visual person, and this is going to be a podcast. So if I look at the Groq chip, it is actually very visibly different than, I think, pretty much anybody else who you're going to label as a comparable just because they happen to be a tech company working on mostly machine learning models that have many hundreds of millions of dollars invested in them at this point. Most everybody else seems to be taking a huge pile of cores and stacking them on top of each other—a core is something like downtown Manhattan, and let's tile Manhattans next to each other.

(Edward at 00:22:02) And you can imagine the sort of traffic congestion problems that you get if you wanted to sort of look at the emergent space of other offerings in the space. Whereas I would say that Groq is the closest thing to a pure play that I know of. One of the things I was very interested in was, how do I run functional code at scale? And the approach that I had started to evaluate was to use techniques, something called SPMD on SIMD. The Intel SPMD Program Compiler is an example of that by Matt Pharr, which is more like compiling like shaders do, but for a CPU. And then you can keep extending that.

(Edward at 00:22:45) If you look at AMD's Graphics Core Next architecture, that has much of the same flavor to it. NVIDIA has a little bit less, but their warps are pretty close to that. And I kept looking in that direction. And if I look at the Groq chip, it's one gigantic SIMD unit. You have the matrix multiply units on either side of the chip.

(Edward at 00:23:06) You have memory in the middle, and at the very center, you have this sort of vector processing unit. There's a couple of shuffle units and stuff like that in there. But the thing is that every operation is deterministic. It takes exactly the same amount of time every time you run it. There's no caches getting in the way.

(Edward at 00:23:22) All the memory in there is sort of cache-speed-style memory. It's all SRAM. So since everything runs on a very completely deterministic clock, we can do things like predict exactly how much power your model will take when it's run. During compilation, I can tell you how much it's going to take to run your model. So if we take one of our chips and put it in an edge deployment scenario in some kind of self-driving car and you want to put a capacitor nearby that's just large enough to draw the charge for your big model that you expect to have to have on very low latency infrequently, you can do that, which is a fairly unique point in the design space compared to all of our competitors who are all, "Here's a pile of regular chips."

(Edward at 00:24:08) They have their own L1 caches or whatever, and they're talking to everybody else around them. And so when you have lots of chips and they're all taking a variable amount of time to do a thing, well, if to get the answer for the next layer, you need to basically have all the answers from the previous one, you wind up paying the worst-case cost. So any sort of variability there becomes that you pay the upper end of the range. And when we start looking at these bigger and bigger models—let's say I want to do GPT-3 or GPT-Neo or whatever we just saw with Megatron, or these other models that have ridiculous numbers of parameters—and you want to run those things at scale. If you want to be able to run those with any decent amount of latency, being able to control that variance, because we can lock it down to zero, is pretty key, at least in my eyes.

(Edward at 00:24:59) So I would say, go ahead.

(Joel Beasley at 00:25:02) It sounds like you're giving them some certainty as far as the power consumption argument or the time to take it. I don't think about this a lot, by the way. I'm usually higher up in the stack, but as you're describing this, it seems very chaotic to have a chain of chips all with variable timing and then trying to do something that's time-sensitive.

(Edward at 00:25:27) And then you have the fact that, really, when you're using GPUs for this, there's just a lot of things that GPUs are really bad at. And we've kind of duct-taped GPUs together to do machine learning because they can. They're the closest thing to write, right? You know, TPUs came along. Jonathan designed the initial version of those in his 20% time at Google back in the day, in a dialect of Haskell, no less, in Lava.

(Joel Beasley at 00:25:51) Can I use a Groq chip to mine some Bitcoin? Would that be advantageous or no?

(Edward at 00:25:56) In theory, that seems to work reasonably well. I don't have full numbers yet, but crypto seems to—how to put this? We are better than an FPGA in a lot of spaces here. I don't know that we're—we're not going to beat dedicated ASICs or something.

(Joel Beasley at 00:26:12) So they're out there. They're in use. They're in the wild. They're being sold. So if people message me and they're like, "Joel, I heard Groq, you guys talking to them. Can I buy some of their chips?" I can just connect them with Jonathan or the team over there.

(Edward at 00:26:25) Yeah. I mean, we have chips. We are currently working at different strategies to get customers access to chips easier than having to buy them from us in bulk, right?

(Joel Beasley at 00:26:38) Let's talk about people.

(Edward at 00:26:40) Let's talk about people.

(Joel Beasley at 00:26:42) Jonathan's a great human being, right? I really sincerely enjoyed him. I don't think you were there last time I talked to Jonathan, but now you've come over. There's a bunch of other interesting people coming over. Tell me about those other cool people.

(Edward at 00:26:55) Oh, gosh. Let's see. So I recently joined—I joined full-time this summer. I'm trying to think who is new, because to me, they've almost always all been there, right? We're just now—Groq has this sort of notion of mega-hires, people that we should sort of chase after because they will, on their own, sort of propel us to be an industry leader in some space, right?

(Edward at 00:27:08) So I would say the person that I am personally most excited to work with now is a guy named Satnam Singh. So Satnam is actually the person who introduced me to Jonathan. He did a lot of work at Google on the TPU and around proving correctness of various RISC-V systems, work that they've been doing.

(Edward at 00:27:47) So he's an old hat in the Haskell community. He is the most Scottish Indian or most Indian Scotsman you'll ever meet, simultaneously. And he is an absolute joy to work with. So if I have to say, who am I most excited about that is brand new to Groq, I would say Satnam.

(Edward at 00:28:10) I also brought in Tim Sears with me from the Haskell community. Tim joins us. He used to run the machine learning team over at Target, actually. So he built a whole data science team in Haskell at Target, which kind of helped propel them to the point where, you know, Target knows that you're pregnant before you do. That kind of level of—

(Joel Beasley at 00:28:32) I was going to bring that up. I was like, is that him? Is that his doing?

(Edward at 00:28:37) So yeah, Tim's joined us, and he's—I kind of view Tim as the other half of my brain. He is my grown-up friend. He is the one who, when I'm trying to figure out how to frame something or to sort of speak to business, he is the body that supplies that input to me. And it's one of those things where I've been looking for an excuse to work with Tim for many years, actually. He helped me get the Topos Institute started, which is an institute dedicated to category theory. It's a nonprofit that I'm still on the board of. And he also helped get us the Haskell Foundation off the ground, another nonprofit that I'm active in. So we've been sort of dancing around working together for many years now.

(Joel Beasley at 00:29:25) You've mentioned this category theory a couple times now. It seems to be—what is it? Why should we care about category theory? What does it do for us?

(Edward at 00:29:35) Okay. So, look, you might have heard of set theory if you've heard a little bit about math. If you start asking about what is math made out of, people often go down and, well, at the very bottom, you find sets, and then you can build everything in math out of sets and set theory. Set theory is sort of the study of the nouns of things. Even the way that you talk about functions in set theory, you just talk about the set of all the pairs of the inputs and the outputs of the function, right? Everything is in nouns. The way I like to think about category theory is, category theory is sort of the study of relationships between things. It's the arrows between things.

(Edward at 00:30:10) That's how things are connected rather than the nouns. It's more like the verbs. And you can go through and you can go back and sort of rebuild math on top of this verb-like structure rather than the noun-like structure that we had on the set theory side. And I think it's a more informative way to talk about math, right? Set theory was the first thing that people found that worked. It wasn't necessarily the best thing that they could have found. And I realize I might be doing this a little bit of a disservice by kind of trying to use more populist math terms than diving into objects and arrows and morphisms and functors and adjunctions and Kan extensions and whatnot. But there's a whole language of this.

(Edward at 00:30:29) Because once you start talking about, okay, well, I have objects and the arrows between them, and then I have the arrows between—then I can have arrows between my arrows, and you kind of grow higher and higher dimensional structures like this. How to put this? In lots of areas of math, there's something like, "What is the natural map between this and that?" This was an old piece of just vocabulary that was floating around very heavily in the thirties, okay? And category theory was initially an attempt to sort of make that rock-solid, to actually make that a rigorous statement.

(Edward at 00:31:29) We can talk about the natural map between two groups or some other kind of mathematical constructs. And it worked as a Rosetta Stone for a very large number of areas of math to let them understand each other and how to talk to each other. Right? If I go through and I learn, like, oh, that's a category. Okay.

(Edward at 00:31:46) I can ask, what are the initial—what's the initial object? What's the terminal object? You know, what are the tensors? You know, is it a monoidal category or whatever? If I learn about, I don't know, quantum computing one day, and all I know about it is it's a symmetric monoidal category with some extra structure, then everything I know about symmetric monoidal categories transfers to working with quantum computers as an example of something.

(Edward at 00:32:10) And symmetric monoidal categories have as their internal language the logic, the world of linear logic. And there's a whole area of philosophy and logic that I can bring to bear on quantum computing using category theory as a sort of bridge, as a translator between these two different approaches or these two different sets of vocabulary. And I find that endlessly fascinating. I find that the fact that I can sort of come into a new area of math and stomp around simply because I've gained a lot of faculty for working with these very broadly applicable category theory tools is really shockingly useful to me. And then if I look at computer science, there's really just nothing in the world of computer science that has that sort of pedigree.

(Edward at 00:32:54) Think back to the nineteen thirties. How much of what was invented in the nineteen thirties do you use day to day as a programmer? Well, in math, category theory has been kicked around since 1930, 1934, whatever. And it's continued to grow and has been hammered on by basically every area of math. So it's not terribly surprising that computation is another area in which you can ask the same kind of questions you can ask about any other area of math and get useful answers on it.

(Edward at 00:33:20) And then as a mathematician, I might learn, oh, this is isomorphic to that. But as a computer scientist, well, I can actually learn that while these two things do the same thing, they have different asymptotic costs. One of them moves some of the cost to run time, some of it moves it to construction time. And maybe I can use those as rules of the game for how to optimize my code. And so I write a bunch of those things down, and then it turned out that a whole bunch of people were able to tell me a lot of things about the meaning of my own code that I didn't know.

(Edward at 00:33:46) Right? I believe very strongly in kind of using the right names for things. Because then if I'm kind of standing on the shoulders of giants or whatever, I can give you the path I followed to get here. Right? Because you can Google for the words that I used.

(Edward at 00:33:59) Right? You can find papers. You can tell me stuff about my code that I don't know. I make up the names for everything. If everything's appendable or something weird that I made up, and I pull the laws out of my butt, then nobody has any frame of reference unless they also happen to have seen that other thing and fully know that literature well enough to transfer the results themselves.

(Edward at 00:34:18) Whereas if I use the right terminology, they can go find it. I don't want to deny people, you know, up to ninety years worth of literature. So that's sort of what category theory is to me, if that makes sense.

(Joel Beasley at 00:34:28) What are you personally excited about right now?

(Edward at 00:34:31) What am I personally excited about? So at the moment, a lot of what I've been trying to do with Groq is get us to a point where—I think that the way I'm defining success for Groq personally is something like this. If I can get our tools to the point where I really want to use them—it's not like I can use them to solve my problems, but when I go, oh, man, I have an idea, and the first thing I want to reach for is the Groq toolchain, I will feel like I will succeed.

(Edward at 00:34:56) Right? And we're so close. We're very much on the path to that being the case with our code base today. And that's got me relatively excited. The other thing I'm really trying to get, organizationally—how am I sort of defining success myself is I really would like to get to the point where if I took a day to do research, the company wouldn't fall apart around me, which is a lot about how do I make myself replaceable in more situations?

(Edward at 00:35:25) How do I make sure that the organization can survive without me? The exact opposite of how every company winds up with the one sysadmin in the corner who is the only person who can do his job, and so that person can never go on vacation and is miserable for the rest of their lives, but has job security. Right? I would really like to get Groq to a point where I don't need to be the person in my seat. Right?

(Edward at 00:35:45) And I hate to say that the thing that has me excited is getting to a point where I don't need to do my job. But it is very much a thing that gets me out of bed in the morning. Right?

(Joel Beasley at 00:35:56) You're speaking my language. I'm an entrepreneur. That's the goal. That's the point. The point is to build autonomous auto scaling systems that yield profit.

(Joel Beasley at 00:36:05) Because that means it's useful to other people, because the only way you can get money is by being useful to other people. The only legal way you can get money is that way. And that's what you want. You want to build something and watch it grow. I've got two small kids, and I get the same satisfaction as building an autonomous business and watching it grow as I'm getting to start watching my kids grow.

(Joel Beasley at 00:36:23) It's like I can't control them, but I can provide them an environment that will be conducive to them having success, just like a business. You can't control every little thing, but you can put the right people in the right places.

(Edward at 00:36:34) Yeah. I'm very pleased with—I came into Groq. There were a couple of things where when I joined the software team, folks were kind of throwing over the fence what they thought each other needed, but not actually checking that that was actually the case. So, you know, a lot of what I came in to do was more or less just make sure that there was an architecture in place. Right?

(Edward at 00:36:55) And then to deal with the messy parts of a startup that is now sort of past the initial startup stage. We're well into 180 people. It's getting to a point where it's a reasonably healthy sized company. But it's not to the point where it's a gigantic matrix organization of, oh, you've got all the electrical engineers and the computer engineers and the software engineers on a bench, and you can pull them into a project and then let them go off and then grab the next bench, right, like you would have in the defense space. We're somewhere in the middle.

(Edward at 00:37:24) And from a management standpoint, that's a very interesting and challenging place to be. Right? Because you know, we have the bones of a much larger company. We have a lot of people who have come in from the Intels of the world or from other hardware vendors or from other big software companies. And they're very used to being able to do a completely product first approach to management.

(Edward at 00:37:49) You sort of sit down and you do the MRD and you step into the PRD. You specify all the OKRs, and you've got a project and you're able to run it. Right? You now have the things you're going to hold everyone accountable to. But almost all of those sort of ignore the headcount, the depth of the bench that you have to have in order to be able to run a business entirely in that top down driven manner.

(Edward at 00:38:10) Right? The trade-offs of, oh, I have to trade this to do that are still a thing you feel at about 180 people. You have to be about maybe six, you know, five times larger before that sort of fully product driven MRD, PRD kind of approach really works. And so a lot of what I've been trying to offer is sort of that mid-scale corporate perspective of being somebody who's at that very uncomfortable position of being at the intersection of, okay, here's what's technically possible and here's what all the sort of corporate objectives and marketing and sales and product people want to build. So a lot of what I'm trying to offer Groq is that sort of pragmatic mindset or someone in that seat who can try to see a way to steer through those rocks.

(Joel Beasley at 00:39:01) Yeah. I like it. I like you. I like this conversation. I want to wrap it up with something kind of fun.

(Edward at 00:39:07) Alright.

(Joel Beasley at 00:39:07) Any recommendations on good sci-fi content? You seem like a good person to ask for this.

(Edward at 00:39:13) Oh, gosh.

(Joel Beasley at 00:39:15) If you're not, if I'm in left field, just tell me.

(Edward at 00:39:18) Good sci-fi content. Current sci-fi content?

(Joel Beasley at 00:39:22) Any content.

(Edward at 00:39:23) I mean, I'm a big fan of a lot of Vernor Vinge's work. It's now getting fairly long in the tooth, but it's, you know, still a classic. I've been reading a lot of—I admit I've kind of drifted away from sci-fi as my genre and have been reading lots of Chinese Xianxia and Wuxia novels lately, which are their own brand of crazy. So it's like Chinese fantasy with heavy Daoist elements and stuff. That's where most of my reading has gone, and it's mostly awful.

(Edward at 00:39:52) So I don't know that I can actually make positive recommendations in that space. I really just need to stop.

(Joel Beasley at 00:39:57) Do you read Chinese, or is it translated?

(Edward at 00:40:00) So I actually did spend a fair bit of time in the last couple of years learning Mandarin. I'm terrible at it. It started out as—I had been reading all these novels and then I ran out of things that were translated. And then I started trying to read the machine translations. And at that point, now I was getting to a point where it was frustrating enough that I started trying to read it myself.

(Edward at 00:40:21) And then because MIRI and the Center for Applied Rationality and all these different organizations are associated with all these effective altruism organizations, and those organizations were having a hard time getting inroads into China, I was like, well, I have the excuse, the big leap of, well, maybe I should learn Mandarin for these kinds of purposes. And, you know, because I was getting invited to give talks in China and stuff like that before the sort of statepocalypse kind of kept us all home. And so I went pretty hard down that road.

(Joel Beasley at 00:40:49) My sister lived over there before the whole virus thing for several years. And she was trying to teach me not the language itself or specific words, but the concepts of the language and how they do things differently compared to ours. And it seemed vastly different, but it also seemed efficient in some ways and fairly interesting. But I was curious, you're pretty smart with machine learning and things of that nature. There's the famous GPT-3 examples of you can put some stories in and then it'll spit some similar sounding stories out.

(Joel Beasley at 00:41:21) My question to you is, would that be just as effective in one language as it is another language? Like, is it agnostic to language?

(Edward at 00:41:28) The GPT-3 approach is pretty language agnostic. It doesn't have a sense of sentence structure or anything like that. It's literally just a predict the next token in sequence kind of engine. So yeah, that would work perfectly well in Mandarin or whatever, if you have a large enough corpus to feed it.

(Joel Beasley at 00:41:46) So the fluency and, I guess, for lack of a better term, the smoothness in which it reads. You know, we've watched the different models improve, like how they read smoother now, right? Like more natural language. And so is that smoothness similar in both languages? Because that's just fascinating to me.

(Edward at 00:42:02) I mean, as far as I can tell, I have not seen a GPT-3 scaled model for Mandarin. Like, I just haven't looked. But there should be no particular issues with it in terms of—I mean, other than having an Internet scale database to feed it and, you know, the whatever, $4,500,000 or whatever to train it.

(Joel Beasley at 00:42:25) After Groq, it's going to be like $2.

(Edward at 00:42:27) I mean, we do need to get to a point where Groq is better suited to training. We've mostly targeted ourselves at the inference market, which is great because so many of our competitors went all in on training. And then once you've got the thing trained, now you have to play it somewhere. So in some ways, we sort of locked into a nice market segment for ourselves.

(Joel Beasley at 00:42:46) Do you guys go into the area, like, consider yourself somewhat of an energy company or use sort of the incentives that the government offers for energy efficiency at all? Have you done any of that, or have you thought about that?

(Edward at 00:43:00) I would say Groq is rather well positioned in a couple of ways, but I don't think it's really from—I mean, we're very solid in power. As an example, someone who I'm very happy to work with—we've got this guy, Omar, who did a lot of the work on getting Apple's M1, their original the power design for the Apple M1, is someone I talk to fairly regularly. And Groq's positioning to me speaks well to—if you look at, say, what's the size of the competitor? Right? The legacy competitor in this space is NVIDIA.

(Edward at 00:43:36) Right? They're the one that folks, you know, measure themselves against just because they're the 600-pound gorilla in the room. Right? You know, and they pull, what, $300,000,000, $500,000,000 worth of revenue a year or something like that. You know, we're a billion-ish dollar company.

(Edward at 00:43:54) They're significantly larger than us. And we don't need to win every battle. We need to find the niches that they don't have well defended, the places where—because they have to be good at everything, and we have to be good at one thing. We can pick our battlefield. Right?

(Edward at 00:44:11) We can pick what portion of our market we're going to turn around and nibble off and go after for ourselves. There are spaces where we're just 300x in terms of TCO or in terms of power. We should find and finish identifying those and then knock those out of the park. Right? Just take ownership of that corner of the battlefield.

(Edward at 00:44:31) At that point in time, then the business changes in a couple of years to, well, now we have to remain sufficiently efficient that no one else can come around and take that segment from us. But it's a very—I don't know. It's a very target-rich environment. It's a great place to be at this stage of the company, right, where we can sit here and call our shots one after another. And that's sort of where I've been getting excited about Groq.

(Joel Beasley at 00:44:57) Has the chip shortage—everyone's talking about the chip shortage in the media. Right? You guys make your own chips. Is it the manufacturing of the chips, or is it farther downstream, the raw materials for the chips? Is that affecting you guys at all?

(Edward at 00:45:10) We have been fairly fortunate, I would say, with regards to the way the chip shortage worked. We have a very bright guy in charge of our sort of supply pipeline, and he went through—he bought a few parts in significant bulk to make sure that we wouldn't run afoul of sort of limits in this space. But I think we are good for demand for the midterm here, which means we can go after business that was otherwise being bogged down by—well, their vendors just can't ship. And it's been a very surprisingly effective way to run a business, is to be able to ship the chips when your competitors can't.

(Joel Beasley at 00:45:53) I love it. Well said. I saw recently that you had a press release about growth and being geo agnostic, right?

(Edward at 00:46:00) Ryan Vaughn (3nine thirty seven):

(Toby Crabtree or another Grok representative (likely from HR/People Operations) at 00:46:01) Yeah, and there was a lot of strategic discussion around it. But a lot of times people think like, oh, the whole geo-agnostic thing, it's like, oh, you're comfortable with remote work. And everybody's been comfortable with remote work even before the pandemic. It's been more about, like, how do we support truly the initiative around work-life balance and, you know, a scenario classic kind of old scenario, right?

(Toby Crabtree or another Grok representative (likely from HR/People Operations) at 00:46:25) It's like, imagine you're an amazing engineer and you want to go work. You probably have to go to a major metro hub like Silicon Valley, but your partner is a surgeon or a doctor that works at a hospital. So which one of you gives up your career? Right? And that's a classic story in the old brick-and-mortar model.

(Toby Crabtree or another Grok representative (likely from HR/People Operations) at 00:46:44) So what we're recognizing now is if we want to meet talent where it stands, we actually help complement that work-life balance. Like, hey, be where you need to be for your life, and we're going to work with you on that. So we've got these amazing programs. And Edward talked about the real nature of our culture.

(Toby Crabtree or another Grok representative (likely from HR/People Operations) at 00:47:03) We've got this amazing director, Toby Crabtree. Her and her team are spectacular about what they've done for keeping us all connected as remote workers. But the policies that she's written has really been kind of groundbreaking, and that's why we put that press release out.

(Joel Beasley at 00:47:17) I love it. I love it. And we'll just go ahead and credit the growth that you guys are appearing on the podcast.

(Edward at 00:47:21) Okay. Thank you so much for—

(Joel Beasley at 00:47:25) listening. And if you found this episode useful, please share it with a friend or colleague who you think would get value from it. And if you have topics that you'd like to hear discussed on the podcast, either add me on LinkedIn or send me an email: [email protected]. Every time I get an email or LinkedIn message, it absolutely makes my day and inspires me to keep going.