Episode 849 ·
Revolutionizing AI as a Physicist with Guy Gur-Ari, Co-Founder at Augment
Today, we're talking to Guy Gur-Ari, Co-Founder at Augment. We discuss Guy’s journey from physicist to founder, why AI is becoming both creepier and more useful than ever, and why the intelligence AI exhibits is unlike anything humanity has ever encountered.
All of this right here, right now, on the Modern CTO Podcast!
To learn more about Augment, check out their website here.
Produced by ProSeries Media: https://proseriesmedia.com/
For booking inquiries, email [email protected]
About Guy Gur-Ari
Guy Gur-Ari is the co-founder of Augment, a company specializing in AI-powered coding assistance for large-scale software development. With a background in theoretical physics, including research on black holes, Guy transitioned to AI and machine learning, bringing a unique perspective to the tech industry. He previously led a team at Google focused on understanding and improving machine learning models, contributing to significant projects like PaLM. Guy's work bridges the gap between complex scientific concepts and practical AI applications in software engineering.
About Augment
Augment puts your team’s collective knowledge—codebase, documentation, and dependencies—at your fingertips via chat, code completions, and suggested edits.
Get up to speed, stay in the flow and get more done. Lightning fast and highly secure, Augment works in your favorite IDEs and Slack.
We proudly augment developers at Webflow, Kong, Pigment, and more.
We are alumni of great AI and cloud companies, including Google, Meta, NVIDIA, Snowflake, and Databricks.
If, like us, you believe in augmenting and not replacing software developers, join us on our mission to improve software development at scale using AI.
Transcript
(Intro Narrator at 00:00:00) Today, we're talking to Guy Gur-Ari, co-founder at Augment, about how he's used his experience as a physicist to apply AI to large-scale code bases. Thanks to Augment for sponsoring the show. To learn more about how their AI can deeply understand your code base, go to augmentcode.com. You're listening to Joel Beasley, Modern CTO.
(Joel Beasley at 00:00:27) The thing I was most curious about is you're not just a tech guy. You're a tech guy, but you're also a physics guy, researched on black holes. I'm so curious, how did all of this come together for you?
(Guy Gur-Ari at 00:00:42) So all of this is the physics or the AI? Or which part of it?
(Joel Beasley at 00:00:46) Yeah. How did it— Did you want to become a physicist and then AI caught your— How did you go through this experience?
(Guy Gur-Ari at 00:00:53) Yeah. When I was, let's say, high school, I was fascinated by physics. I was reading about relativity and quantum physics and was super curious about the philosophical implications of it and so on. And I knew I wanted to study physics to understand what it was all about. I did my undergrad in physics. I really enjoyed it, then did a PhD in string theory, then went to Stanford for a postdoc. It took a while to kind of get the hang of what physics research was about, and I really enjoyed that journey.
(Guy Gur-Ari at 00:01:18) I think toward the end of it, when I was doing the postdoc, over time I kind of became less interested in continuing to do research in string theory. The reason is that string theory, the theory is probably, I don't know, I'm guessing hundreds of years currently ahead of experiment. So we come up with all these beautiful theories that are all self-consistent, but it's really hard to test them experimentally because the energies that you need to run these experiments are just out of reach for us and will be for a long time.
(Guy Gur-Ari at 00:02:09) But at the same time, I was at Stanford, heard a lot of chatter around machine learning. Back then, people were calling it more machine learning than AI. At some point it kind of shifted. And me and a group of friends started looking into it. I have a background in programming, and so that made it somewhat easy to get into at least the practical side of things.
(Guy Gur-Ari at 00:02:34) And so I kind of went through this transition of becoming less enthusiastic about continuing a career in physics and more enthusiastic about what I thought was a scientific revolution, really. The kind of thing that happens once in a lifetime, what we're seeing now with AI. And so I started learning about the subject, started coming up with very bad research ideas, and then over time, getting better research ideas. And then through all of that, ended up joining Google to do research in machine learning full-time.
(Joel Beasley at 00:03:09) What was the thing, the moment where you're like, "This is going to be a scientific revolution. I'm going to hang my hat on this"?
(Guy Gur-Ari at 00:03:16) Yeah. So for me, back then, it was a lot of talk about vision tasks, image recognition, semantic understanding of images and segmentation, things like that, because things were fairly focused on convolutional neural networks. That was, at least from my perspective, the hot thing. And it took me back again to high school where I actually learned about neural networks. They were fairly basic at the time, small, not a lot of data.
(Guy Gur-Ari at 00:03:54) And I just remember thinking, "Well, is this really the path to artificial intelligence? I mean, am I ever going to be able to show it an image of a cat and it will tell me that it's a cat and not a dog?" I just, in my mind, could not connect the dots between these neurons and how they interact with each other, and how would this ever perform a task that to us humans is kind of trivial. But then, you know, fast forward quite a few number of years, and yes, this is what these convolutional networks were doing. And actually, image recognition was kind of the simplest thing they were doing.
(Guy Gur-Ari at 00:04:36) So to me, it went from, "Oh, there's no way this is ever going to work," to "Oh, this is a reality now." I found that just astounding. And so I thought, "Well, if I can make such completely wrong predictions and in my lifetime get to see that the technology actually catches up and surpasses all of these wrong predictions that I made, then probably we're kind of at the tip of this exponential growth of capabilities from these models."
(Joel Beasley at 00:05:06) Yeah. And personal growth too, right? Because realizing that you're wrong and still being cool about accepting it and wanting to push it forward is great.
(Guy Gur-Ari at 00:05:15) Well, to me, it's exciting because I mean, I feel like these days, it kind of feels like we're living in sci-fi world. It just keeps getting better. So to me, that's all super exciting, and I just feel very fortunate to be able to be a part of this.
(Joel Beasley at 00:05:32) I'm curious about Google. So you start doing some work at Google. What was your role there? Were you managing people? Were you just individual contributor?
(Guy Gur-Ari at 00:05:40) Yeah. So me and a few colleagues from physics, we pitched this idea that, "Why don't we start a team at Google devoted to opening the black box of machine learning models, understanding how they work?" We'll bring together folks from physics, from engineering, from computer science to work and do basic research to better understand these models. And they liked it. And so we joined. I led the team, so we grew it. There were three of us in the beginning. We grew to about 10 people.
(Guy Gur-Ari at 00:06:09) And in the beginning, our focus was really on, yeah, let's try to understand deeply how these models work. What are the training dynamics of these models? Let's try to borrow ideas from physics, mostly theoretical physics, to try and make some of the calculations easier. And so, yes, in the first two years we focused a lot on vision tasks and on model optimization, trying to understand when we train these models, what's actually happening there, and how can we make that process more efficient. So that was the premise.
(Guy Gur-Ari at 00:06:26) And then about two years in, GPT-3 came out, and this was to us a revelation because suddenly there was this model where you could actually talk to it and do experiments through this new language model interface and start trying to understand how the model actually goes about solving tasks in a way that was much more approachable and easy and fast than before. You didn't need to train full models from scratch to run experiments. You could really just talk to the model and iterate that way. Yeah. So to us, that was a pivotal moment. But until that point, we were really focused on more, I would say, more theoretical work trying to combine physics with machine learning.
(Joel Beasley at 00:07:45) And so from a role perspective, is it just you on a small squad of five people working on a project, or was it teams of teams?
(Guy Gur-Ari at 00:07:54) It was a team. We had— With research, typical collaborations at that time were fairly small. There were like three, four people, maybe with interns. And so we would have usually on the team two, three projects running at a time, roughly.
(Joel Beasley at 00:08:13) Okay. Cool. And what was— You're obviously good at growing yourself and progressing your career. What was something that you learned in that project about managing people or at least working together with other people?
(Guy Gur-Ari at 00:08:28) So I think on that project, one thing that was, I think, fairly clear is that, you know, in these roles, especially in AI, you have folks who are more on the research side, you have folks who are more on the engineering side. There is a spectrum, and there are some people who can kind of do both well, but often you have people who specialize a bit more. And what I found there is that it's quite important to not silo people into these specific roles, to let people try out things that— If you read their CV, maybe it's not obvious that they would be good at a different kind of thing, but to also let them experiment and grow. And I was often surprised at how great people ended up being at roles where, again, based on their previous experience, there'd be no reason to think they'd be good at that. So that, I would say, is one thing.
(Guy Gur-Ari at 00:09:30) The other thing that was pretty clear is letting engineers and researchers work closely together. Again, not introducing barriers, organizational barriers within a team, or if we work across different teams, just have as few barriers as possible. If you have excellent smart people, they're going to figure it out and just let people from different roles work closely together rather than say, for example, "Oh, we have a research team, we have an engineering team," and kind of things have to get thrown over the wall all the time. Let people work as closely together. And this is also how we've structured Augment.
(Guy Gur-Ari at 00:10:06) We have people who— Research engineers, back end, front end design, all working together to develop certain features. And that was definitely a lesson that I learned at Google.
(Joel Beasley at 00:10:18) The most frustrating thing for me, Josh, is when I call support and I can't understand them.
(Intro Narrator at 00:10:23) Yeah, man. I hate that.
(Joel Beasley at 00:10:24) That's why I like U.S. Cloud. Not only is it better, faster support, but all the engineers are U.S.-based engineers and it's also a lot cheaper. 94% of U.S. Cloud's clients report saving a third or more when switching from Microsoft Unified Support to U.S. Cloud. So now you'll just have to figure out what to do with all that extra money. If it were me, I'm responsible, so I'd reallocate the money to improve my team. Josh, what would you do?
(Intro Narrator at 00:10:49) I'd probably just buy more guitars.
(Joel Beasley at 00:10:51) More guitars. Visit uscloud.com to book a call and figure out how much your team can save. And so what are you building at Augment? I have a good idea of what it is, but for people who don't know, what are you building over there?
(Guy Gur-Ari at 00:11:03) Yeah. So we build AI coding assistance aimed specifically at large organizations. Think hundreds of developers, let's say 100 developers and up. So aimed at developers whose day-to-day job is working within large, complex code bases. This is an area where we see existing AI assistants fall short. They've been falling short since the beginning of Augment. Frankly, they still mostly are. And so we're building our AI assistance for professional software developers who work in these large code bases.
(Joel Beasley at 00:11:38) What's the adoption? And sorry to just put you on the spot with stats. You might not know this, but what's the adoption rate as far as in these companies? Are 50% of developers using assistance? Are 10%? Do you know?
(Guy Gur-Ari at 00:11:51) Yeah. So this is actually an existing pain point for users of AI products, or let's say specifically for developers. The typical adoption rate for these products in companies tends to be fairly low. So for some competitors, I would say it can be as low as 10 to 20% adoption internally. And when we talk to customers, this is one of their top pain points with AI.
(Guy Gur-Ari at 00:12:24) You know, there is all this talk about the big promise of AI and how much it's going to improve productivity, and then organizations adopt these solutions, and they don't see it. They don't see the adoption. With Augment, we've structured the way we sell the product to be really aligned with customer— So basically align our incentives with the customer's incentives. If users at a company don't use our product, the customer does not pay us for that. So we don't have a seat-based pricing. We have consumption-based pricing. And what that does internally is really push us to build a product that developers actually want to use day-to-day. And so we see much higher adoption rates as a result of that.
(Joel Beasley at 00:13:17) Why do you think the adoption— So in my circle, the adoption's 110%. I mean, we're like— It's— I was having this conversation with my buddy Derek yesterday, and I said it's the equivalent of we've all been sawing wood with a hand saw, and this dude's walking around selling chainsaws, and some people are not interested. I'm like, "What are you talking about not interested? We can cut the tree down so much faster. It would be crazy to do it by hand once you've seen that the chainsaw exists."
(Joel Beasley at 00:13:46) That's how I feel about it. But then when we talk about adoption rates or I talk to large companies, it's, you know, some 10, 20, 30, 40. It's low. Do you think it's the companies telling them no? Or what do you think it is? Are they scared?
(Guy Gur-Ari at 00:14:00) Yeah. So here's the thing. What we're learning about AI, I think, is it is a very powerful tool. I'll talk about AI specifically for software development because that's the world I live in. It is a very powerful tool, but it is a tool that one has to learn how to get value out of.
(Guy Gur-Ari at 00:14:28) We saw that with completions. We saw this even more so with chat for code. And now with agents, it's even more true. And so with agents specifically, it takes a while to get the hang of how to get the agent to do what you want in a way that actually saves you time. There's some skill that goes into understanding the strengths and weaknesses of this tool, understanding how to prompt it, understanding where and when it can mess up, and building trust over time with the tool so it doesn't kind of, you know, destroy your laptop or destroy your work environment.
(Guy Gur-Ari at 00:15:09) There's, I would say at this point, a fairly steep learning curve. And the more powerful the feature is, the steeper the learning curve is. And I think if you're a developer at an enterprise, you know, you have your queue of tasks you want to work through. Do you really have time to go and experiment with a whole new way, really, of building software and learn how to do it well? I think today, you'd need to be fairly motivated to get into it and spend the time and really get the value out of these tools.
(Guy Gur-Ari at 00:15:49) So I'd say on the one side, it is still at a point where the learning curve is steep and users to get value out of these tools do need to really be motivated to go through the trouble of understanding how to use these tools. And on the other side, you can say, well, in terms of as a product, we are working hard on making these tools easier to use, easier to learn, more accessible, but they're not there right now. And one thing that makes it tricky is the fact that the capabilities of these tools are growing exponentially. And so in some sense, we are trying to make these tools as approachable as possible and as easy to use as possible. But then pretty quickly, we will build a new feature that's much more capable.
(Guy Gur-Ari at 00:16:40) Again, just going from completions to chat to agents, each one of those things is a huge step in terms of capabilities. And so as we're working on getting chat, let's say, as easy to use and as friendly and approachable as possible, in come agents, and we need to go and do all of that work again until the next thing that comes out. So currently, we are really still on this exponentially growing curve of capabilities, and that poses a challenge for us and others in the space building products, and that also poses a challenge for users. One last thing I'd say is we clearly see this dichotomy between the users like yourself who are very early adopters and jump on this technology and pretty quickly get a ton of value out of it. And the more you give users like yourself, the more value they get versus users who are not necessarily as early adopters.
(Guy Gur-Ari at 00:17:42) And we try hard to also cater to those users to kind of bring them along on this journey and show them how you can get value out of these tools and also make the tools more useful over time for them.
(Joel Beasley at 00:17:56) Yeah. So I agree with you and disagree with you on the steep learning curve. So I agree that, yes, if you start at nothing and you go to brute force it, meaning someone says, here's an LLM, you get a prompt console, and you just have to go figure it out. That's a huge learning curve. You can dump hours and hours and hours into figuring out, like, what's the right way to prompt them to get this result.
(Joel Beasley at 00:18:16) And that's what I was doing at first. And to be honest with you, the first month or two, I kind of dismissed it. My buddy was talking about it. And I was like, alright. I tried it out and tried to get it to do some stuff. I was like, you know, whatever. So I kind of dismissed it. And then one day, I was scrolling through X. And there was this YouTube video somebody had posted of them, like, doing something specific with it. And I got to monkey see monkey do. I got to watch this person achieve it. I'm like, oh, you have to talk to it like that. Oh, you have to use it in like, this is little tiny bites and use it in this specific way and then apply my, you know, almost twenty years of software engineering on top of that. And like, that's how you and then all of a sudden now it just clicked. It's like, okay. So I do think, you know, yes, a steep learning curve if you're brute forcing it. But if you get, like, two or three good ten minute videos of people, like, at your experience level in software engineering using it, you can just start running. Like, you just have to let those things click.
(Guy Gharari at 00:19:15) Yes. I definitely agree with that. It's not rocket science to get value out of these tools. I will say, though, that the larger and more complex the code base is, the harder it is. So these models, they do really shine. You can really see them shine easily when you're doing, let's say, a zero to one project. Right? You want to build a website from scratch or an app from scratch. And that's where, yes, if you spend more time prompting, you give a lot more detail on what you're trying to do, you certainly chunk up the tasks into bite sized chunks, and there's some calibration that goes into understanding for a given model how big the chunks need to be, you can get results pretty quickly and very good results. In large code bases, you can still get good results, but it takes more effort, simply because when you work in a large code base, there are more constraints that you or the agent need to solve for. You need to respect the coding guidelines. You need to respect design. So the code needs to fit in with an existing design. All these things do make it more difficult to get value within large code bases, but that's the problem we're excited to tackle.
(Joel Beasley at 00:20:44) It's a good problem. I mean, a lot of people are in that position where they're working with these large code bases. I personally that's, like, my limit of experience. All the projects I've done are less than 20 engineers. So, like, I don't I can talk about it with people like I do on the podcast, but you don't really understand something until you actually do it. Yeah. No. That's good. That's good. So should we trust these AI assistants? How do you can you trust them? Should we trust them?
(Guy Gharari at 00:21:14) Yeah. I think building trust is going to be a theme for a while. So, and we see this also internally as we rolled out the agent. Some people are quick to adopt the agents, and not just from a from the perspective of figuring out how to get value out of them, but also trusting them. Because what happens is that these things are extremely good at generating a lot of code quickly. And then the question becomes, okay, what do you do with all this code? You're presented with often, like, hundreds of lines of code. You ask it to do something, you get hundreds of lines. How do you, over time, build trust that it's doing the right thing? And on a task to task basis, how do you understand, like, did it actually do what I wanted? Did it do the right thing? I think we're still learning how to do that. Certainly reading the code is useful, but it is a lot of code. And so one thing that I've been doing is figuring out how to quickly read the code to get a sense of whether the model is going in the right direction or not. Like, did it put the code in the right place? Are there any glaring issues, to first get a sense of are we on the right track? And then rely very heavily on testing. What I found is that the more you can test, the easier it is to establish trust in what the agents do. And so there are the place where this becomes clearest is when there are areas that are not easily testable. If you have the agent write some parallelized code, right, some multiprocess or multithreaded code, what I've seen is that because those things are pretty hard to test, there will be necessarily bugs in there, and you have to read the code very carefully. But if you're building something which is not that complex and you can automatically test it, then it's often a very good idea to write a lot of tests, to use the thing yourself, and to make sure it works. That makes it a lot easier to build trust in what in what the agent outputs.
(Joel Beasley at 00:23:42) Yeah. So tell me about the agents you guys are building at Augment. Is this new?
(Guy Gharari at 00:23:47) Yeah. This is a feature we are releasing. It's kind of the next evolution after the previous features we've had, completions, next edits, chat with your code base, and then the next step is agent. So this is meant to be a feature where you can, if you want to, take a step back from being super hands on with the code, ask the agent to do certain things, and then the agent will go and edit code for you across the whole code base. It will run commands, and so it can run tests and iterate on them. It can help you with version control. It's connected to to GitHub and other things like Linear, Notion, Jira, Confluence. So that opens the door to whole end to end workflows that you can do with the agent, like go from issue to PR or go from code to spec in Notion, for example. And so it's really a next level in terms of capabilities that we provide developers. And we do have Augment's context engine built in to the agent. And so this is the thing that provides all of our features with full code based understanding. It's built into the agent as well, and that's what really makes the agent powerful when you use it to perform tasks in large code bases.
(Joel Beasley at 00:25:09) Is it one agent or multiple agents?
(Guy Gharari at 00:25:12) So right now, it's one agent per ID window. The thing and so you can open multiple windows if you want multiple agents, but it's one agent in the window. We have found that once you get the hang of using the agent, you quickly notice that it's slow in the sense that when you ask it to write a piece of code, it will go and edit a file. Maybe it will run the test. It will find a problem, and then it will iterate on that for a while until it gets to working code that passes the tests. And so we do believe that we are headed toward a future where we don't just have a single agent, but we actually have multiple agents performing all kinds of tasks. But for now, what we have in the product is a single agent, and I think that's for where a lot of users are, that's already plenty to get value.
(Joel Beasley at 00:26:13) Yeah. Have you been using Grok at all?
(Guy Gharari at 00:26:16) We've tried various, you mean Grok with K. Right? You mean Grok three?
(Joel Beasley at 00:26:21) I know. I know. Our buddies that make the chips are like they've gotten into this now too, and now it's like, alright. Which Grok is it? Elon's Grok. So the thing I like about it, the user experience that I enjoy the most being a product person my whole life is how it shows you how it's thinking. When I saw that, I was like, oh my gosh. That is the best experience ever. Are you guys doing things similar to that?
(Guy Gharari at 00:26:48) Yeah. That's interesting. So we, the main model we use for agents right now is Claude. And Claude, they just released their thinking mode, so we don't have that yet in the product. What you say really resonates with me because one thing we found with the agents is that users really love observing what the agents do at, so the the more detail we provide, the better. Now, of course, like, by default, it looks like a pretty big wall of text. So you need some UI to to cover for it because there's maybe there's the chain of thought that we can introduce, but there's all the tool calls that it makes. Right? When it edits files, run commands, those things generate a lot of output. But generally, users want to see as much as possible about what the agent is doing and to have a lot of control over what it's doing. So I think the chain of thought makes makes total sense to add. Yeah.
(Joel Beasley at 00:27:48) Yeah. I just love when it's when you can watch it think to itself and be like, oh, no. They probably meant this. Or, oh, no. I corrected this. And I'm like, yes. This is brilliant because I always felt like GPT and, look, I love all the services, so I'm just going to speak openly. But I always felt like GPT was lazy at first, like, when, like, 3.5 or whatever. Like, it would it would try it would get out of, like, trying to do any work. It's like either I know this or I don't, but I'm not going to go, like, try harder or, like, try to understand better. But now the tools, even the GPT ones, they're so much better. In the past two years, we've made leaps and bounds of improvements.
(Guy Gharari at 00:28:23) Yeah. For sure. And I think I think they're the those folks who are trying to train the models, they're trying to thread the needle, I would say, between not having the model be too lazy versus probably not having it be too aggressive. Because one thing we've seen, for example, going from Sonnet 3.5 Claude Sonnet 3.5 to 3.7 is that they clearly train the model to be more kind of forward leaning and and proactive about going and taking actions. And for some users, that's great. But for some users, so we talked about, like, adoption of these things. For some users, that can be pretty off putting. Like, oh, I just asked you could ask it, like, do I still need this code? And then it'd be, like, well Delete. Yep. Yeah. No. You don't. I deleted it, and I opened a PR for you and assigned a reviewer to to check it, so don't worry about it. You know, this is like it's a bit of an exaggeration, but that's sometimes how it feels like. And so we are we're still trying, I'd say, to all of us to kind of try and then we try to correct for that with prompting. We're trying to find find the balance between these two extremes.
(Joel Beasley at 00:29:32) Oh, and and there's not going to be a direct answer for everybody. Eventually, it's going to come into some sort of option or feature where you, as the individual engineer, get to tell it, like, more about how you want it to work with you. Exactly. So these are
(Guy Gharari at 00:29:45) These are the kinds of questions that we're we're asking today is how to build all those preferences into the product in a way that feels really delightful.
(Joel Beasley at 00:29:55) Are we learning anything about it's like we have intelligence. Right? We call it intelligence, humans interacting with each other and whatnot. Are we learning anything about intelligence as a whole through the advancements in this artificial intelligence?
(Guy Gharari at 00:30:11) I think to me, what I think I've learned is that there are multiple forms of intelligence. I think comparing how we're looking at these things now versus how we were looking at them before large language models, a lot of the discussions that I remember before large language models, they really focused on, you know, we have from a research perspective, what does it look like if you're trying to get to AGI? You have this proof of concept that it's possible. Humans exist. Humans have intelligence. We know it's doable. Right? And then the question is kind of, okay. How do we do that but artificially? But because we only have this one example of how it works, a lot of the discussions were around how to kind of replicate human intelligence in these machines. And a lot of discussions on, well, you know, there's been there's been a lot of discussions on how these models are not, they're just parroting what they what they were trained on and so on. So it's like not real intelligence because it's true. This is not how humans learn. And it's like the way we train these models looks nothing like the way humans learn. The way these models work at inference time looks nothing like how humans solve problems. So maybe just to give a little, like, a small example that that we've seen in research, we've at Google, we've trained a model that was that was very good at solving math and science problems because we were very interested in reasoning and how to get these models to reason better. And we ended up with a model where it could pass university entrance exams in math, and it could also do very basic things in math. So for example, this model was capable of adding two numbers, and I think you could you could let it you could let it add two numbers, multiply, subtract, like basic arithmetic, up to eight or nine digits, and it would still do incredibly well. And it didn't it didn't use long addition or anything like that. You would just give it a few examples in the prompt, and then you would tell it, okay. Now add these numbers, and it would spit out the answer with something like 80 something percent accuracy on an eight digit number. Something like that. Better certainly better than I could do in my head. But then as you cranked up the number of digits and these numbers higher, the performance would drop pretty sharply. And if you gave that problem to to a child who was in elementary school, if they learned long addition, they just performed long addition, then they could essentially solve that problem for any number of digits. So their the human approach generalizes to any number of digits, while the model's approach to learning how to add numbers, at least the way we train that model, and that's kind of the common way, did not generalize. And for many people, the discussion was, well, then it's not true intelligence because it didn't really learn how to add these numbers. But on the other hand well, but up to eight and nine digits, it's actually better and faster than a human. So to me, the way these models learn, it's just different. It's not I don't think it's better or worse. I consider it a form of intelligence, and what we have now are, to me, are just two different examples of what intelligence looks like.
(Joel Beasley at 00:33:45) That's exactly it. I think the intelligence is is partly an expression of, like, the substrate it's on. Like, so we've got this biological expression of intelligence and that they have a silicon based expression of intelligence. And they're not perfect, but they can come up with, like, very similar results. So then the question is, like, well, then what is intelligence? Is it, like, is it an exterior force, like, applied onto a substrate? Like, what is that?
(Guy Gharari at 00:34:13) Yeah. I don't know that I have an answer to that, but but I think I think one thing that's very well, I would say highly speculative, but one thing that I find very interesting is that sort of you can look at the there is the substrate, which is, of course, like, important. There's also the you can think of it as the the environment these things grow in or are developed in, and the environment, you know, human intelligence and everything else about us was was shaped by evolution. And evolution is a pretty tough environment to be in, very resource constrained. And so humans, animals in general, we all have to be very resource efficient. And so we have to learn quickly.
(Guy Gur-Ari at 00:34:58) We don't have a lot of opportunities to learn, right? Because it's about survival, really. And so we have to learn quickly, and so we've adapted through evolution to learn very quickly from examples. Versus these models, there is, as we train them, there is nothing that pushes them to be efficient.
(Guy Gur-Ari at 00:35:19) We kind of shower them with tons of compute, tons of data, and there is nothing in the algorithm that pushes them to be efficient in any way. And so to me, that's kind of why they're, on the one hand, able to acquire a lot of knowledge, right? Probably these models, any one of these modern models know way more than any single human on Earth because we do shower them with data and compute. But on the other hand, there's nothing pushing them to really distill their knowledge into, let's say, an algorithm like long addition.
(Guy Gur-Ari at 00:36:01) There's nothing pushing them to compress their knowledge into something that's a lot more efficient because we're not even trying to get them to do that. And so there's no reason to expect them to generalize to harder problems like humans do. So again, it's just a different set of constraints we're putting them in, but both are leading to some form of intelligence.
(Joel Beasley at 00:36:22) Yeah. And then it's like I'm not necessarily worried about the AI algorithm being rogue or whatnot. I'm more interested in what's going to happen when somebody builds the AI to be that thing. Like, you prompt the AI to be aggressive and to go accomplish this goal. That to me is scarier than the AI waking up one day and being like, "I'm going to just do this crazy thing," you know?
(Guy Gur-Ari at 00:36:45) Yeah. For sure. These things, I agree, these things are just very powerful, and a bad actor can definitely steer them in those kinds of directions. And I agree. That's definitely a risk.
(Joel Beasley at 00:36:58) Yeah. Sorry, I'm a little bit off topic, but I just find it so interesting because, you know, a couple years ago, I started looking into what we know about consciousness, looking at people that do sedation and all that type of stuff for surgery and like, what do we actually know? And I was a little bit disappointed to find out we know very little about consciousness.
(Guy Gur-Ari at 00:37:20) And so—
(Joel Beasley at 00:37:20) I was hoping, I was like, well, maybe these advancements in artificial intelligence will help us better understand intelligence and help us better understand consciousness as a whole.
(Guy Gur-Ari at 00:37:28) I personally hope so. I think that's the part where if you talk to people, should we go in this direction? A lot of people will find that, let's say, creepy to some extent. But I think from a research perspective, yes. I think this is—I agree, it seems we know very little about consciousness, and this seems like an interesting place to explore that concept, for sure.
(Joel Beasley at 00:38:01) How does it get creepy? At what point does it get creepy?
(Guy Gur-Ari at 00:38:04) So I think even with coding, some folks already find it creepy because, you know, when we talk to customers, often when we show them the things we have cooking, right, the response is often both, "Oh, this looks amazing. I really want to use this," and also, "Oh, this is happening way faster than I expected." Let's say it like that. There is, I think, a very natural response when you see these tools become more and more powerful.
(Guy Gur-Ari at 00:38:44) You start thinking about, oh, what is my day-to-day going to look like? Am I, basically, am I only going to be driving these models with natural language? And am I not going to be coding anymore? What does my role evolve to as these things become more powerful? I think that's definitely in a lot of people's minds as they see the rapid progression and capabilities, especially in code.
(Joel Beasley at 00:39:12) Yeah. But we're smart. We'll figure it out.
(Guy Gur-Ari at 00:39:13) Yeah.
(Joel Beasley at 00:39:14) Yeah. And plus, like, hey, one of my best arguments for this is there's still people that own horse farms that make millions of dollars. Like, horses are still a thing. If you're in the horse industry, you understand this. It's not what it was before, but it still exists today, and people are still running successful businesses in the space.
(Joel Beasley at 00:39:33) So it's not all doom. There is always a silver lining there. If you're willing to work hard and put effort in, you can stay.
(Guy Gur-Ari at 00:39:40) Yeah, that's true. I'm actually even more optimistic than that because I think, you know, just looking back, I think there's kind of this maybe implicit assumption in all of this that we need kind of a finite amount of software. And if we automate all of that, then there's no room for developers to do anything anymore. But I think, at least historically, every time we've been able to improve productivity on software development, all that happened is that we just developed more and better software.
(Guy Gur-Ari at 00:40:14) I expect the same thing to happen here. I don't see that there's a finite need for software. I don't think we're even scratching the surface of the systems we could be developing if we had much higher productivity. So that's my take on it, on why I think software is a bit different than other historical examples, like horses and cars.
(Joel Beasley at 00:40:38) Yep. And I agree with you. Historically, that's what happens. Abstractions all the way down from the electrons up through the ORMs, right? Like, when ORMs started to become a thing, people were like, "Ah, this is like cheating. This is crazy." And then you're not writing SQL as much with some of the tools up there, and then now you're not writing as much model code because you're chatting with an agent, and then, you know, it just keeps—we keep stacking and stacking and stacking and just making the job easier and making it more accessible to more people. Like, I'm on the side of things, Guy, where I love when I see a non-super-technical person code something up with an LLM and put it out there. Yeah, people—a lot of the senior engineers, "There's bugs in there. There's going to be—" Yeah, yeah, yeah, yeah, yeah. But he got it done, and he didn't have much experience with it at all. I think that's an amazing thing because that's going to make—that's selfish for me because it's going to make life better. Make more cool stuff, people. Like, we need more cool stuff.
(Guy Gur-Ari at 00:41:38) Yeah. Let's go. Absolutely, absolutely. Yeah.
(Joel Beasley at 00:41:41) All right. So I had in my research prep doc that there was something about GPT-3 being a pivotal moment in your career. Can you touch on that for a second?
(Guy Gur-Ari at 00:41:50) Yeah, for sure. So when GPT-3 came out, we were at Google. We were working on training models, or trying to understand how models train and trying to improve that. We were also thinking about, okay, what's going to be the next step here? We knew that getting models to reason and understanding how models reason, how can we improve their reasoning capabilities would be a very interesting direction.
(Guy Gur-Ari at 00:42:27) We had some initial attempts of doing that, playing fairly simple games with image models and trying to get them to see if we can get them to reason. But with image models, all these experiments are pretty limited. Like, there's only so much you can do with pictures and trying to run, like, basically classify this picture, is it this or that? You can imagine that the kinds of reasoning experiments you can do on top of models like that are pretty limited. And then when GPT-3 came out, it showed a few very interesting things.
(Guy Gur-Ari at 00:43:08) One is that language can be an interface for—well, basically, an interface for interacting with a model. Because when you have image models, your interface is essentially you train the model to perform a certain task, and then you go and test it on that task. And the output that you get is output in either what did it classify an image as, or how did it segment an image, things like that. With language models, the input is language, the output is language, and so the range of experiments that you can do is much greater, which was true, of course, before GPT-3, but with GPT-3, the capabilities became such that the experiments you could do were really interesting. And the other advancement there was showing that few-shot prompting was possible.
(Guy Gur-Ari at 00:44:04) Basically that—or, a better way to say it is in-context learning. So that's where they showed that you could take these models, show them a few examples of the task that you wanted the model to do in the prompt. So a few examples of inputs and outputs, and then show it another input. So this could be question, answer, question, answer, and then another question. And just showing it these examples in the prompt meant that the performance on the question you finally asked it was much higher.
(Guy Gur-Ari at 00:44:36) And so it was a form of learning because without doing that, the model could not really answer a question, but it did enough pattern matching off of the examples you gave it that it just performed the task better. And what that means for research is that it dramatically lowers the cost of doing research because if you want to test the model on a task, previously, you had to, let's say, fine-tune it on a task. You had to go collect a large dataset, train, make sure your training works well, and then evaluate. Within-context learning—and that could take, let's say, anywhere from hours, best case, to weeks or sometimes months, depending on how complicated the data collection is. Within-context learning, that brings it down to minutes.
(Guy Gur-Ari at 00:45:21) And so you can do experiments and iterate a lot faster. That really changes how you do research with these models. And then finally, the fact that, yeah, when language is an interface, you can actually start doing much more interesting reasoning experiments. And so the first project we did there was called Big Bench. We saw how model capabilities became so good that existing benchmarks for evaluating these models kind of got saturated pretty quickly, got solved basically. And so we tried to come up with a new crowdsourced evaluation benchmark to really push the limits of these models and see where they are from a reasoning perspective, but also from other perspectives.
(Guy Gur-Ari at 00:46:08) And then after that, we started working hard on actually improving reasoning capabilities.
(Joel Beasley at 00:46:15) And then as we start to wrap up, I want to make sure that we get a couple points out here so people can really wrap their mind around Augment. Is Augment model agnostic? How does—where does it sit in the stack? Like, if I hear this episode and I'm like, "You know what? I want my team—we've got a large code base. I want them to experiment with Augment." Like, where does that sit in the stack? How does that change my workflow? And what is it exactly?
(Guy Gur-Ari at 00:46:37) Right. So Augment is an AI assistant that lives in your IDE. So we support VS Code, we support JetBrains, we support Vim. It is, in terms of the product, it's run by a variety of models. So some of the models we train in-house, some of the models are external. We try not to bother the user with those choices. We try to find the best mix of models to give the best possible user experience without the user having to fiddle and understand the differences between different models. And so you download Augment, you install it, and you're basically off to the races.
(Guy Gur-Ari at 00:47:22) You get the best possible coding experience out of it, whether you're using completions or NextEdit or chat or the new agent feature. And the thing Augment is especially good at is understanding your code base. So if you're doing zero-to-one development, you will likely not see a lot of differentiation between Augment and other products that are out there. But if you're working on a code base that's, let's say, a million lines of code or more, or even 100,000 lines, I think you would already see the difference. It should feel like Augment, from the beginning, really understands your code base and gives you better predictions as a result.
(Joel Beasley at 00:48:06) And that's on top of your framework too, right? Because, like, if I—let's say I'm a Rails developer most recently. So you've got the code I've written on top of the Rails framework, but you have the actual Rails framework in its entirety behind it. So are you taking into account the entire code base, everything in the git?
(Guy Gur-Ari at 00:48:27) Yes, for sure. So the models are all familiar with Rails, React, all the popular frameworks. And then on top of that, you have your proprietary code that you're developing in, which is the whole repository or multiple repositories.
(Joel Beasley at 00:48:45) Okay. Cool. Dude, that's awesome. That's so cool what you guys are doing. How can people try this? Is this something they can just go to augmentcode.com and try?
(Guy Gur-Ari at 00:48:53) Yeah. Just go to augmentcode.com, download, install, and start using it. The onboarding is super simple. And there's a free trial, so you can try it out for free.
(Joel Beasley at 00:49:07) Oh, that's amazing. Augmentcode.com, we'll put a link in the show notes. I've got one last question that I was really excited to ask you about. You did this research paper on LLMs. 2022 got a handful of citations. But in '24, it got like 3,000. Today, it's almost 6,000. What is this paper you wrote, or coauthored? Why is it getting so popular?
(Guy Gur-Ari at 00:49:28) So is this the PaLM paper?
(Joel Beasley at 00:49:31) Yep.
(Guy Gur-Ari at 00:49:32) Yeah. So that was—yeah. That was work—so that was a very large team effort. That was—so PaLM is Google's predecessor to Gemini. So Gemini is the name of the current language model.
(Guy Gur-Ari at 00:49:45) There have been quite a few models, large language models trained at Google before that. And this was one of them. So I was involved in some of the work on PaLM and then PaLM 2, its successor, and I think that was the last one before it switched to Gemini. On that specific paper, I was involved in evaluating the model on benchmarks to see how good it is. That's a part of research in building these models. You train something and you need some way to know if what you got is good or not.
(Guy Gur-Ari at 00:50:25) And so you need some way to evaluate these models and understand—and then during the course of research, you can also do hill climbing and try to, as you make changes to the model to make it better, you rely on evaluations to understand how good it is. And so that was my part of that, or I was working on part of the evaluation. But these kinds of projects, they are—well, as you can see from the papers and the amount of authors and people involved, these are very large efforts. So, you know, when I started working on this stuff, we would do, as I mentioned, three-to-four-person collaborations. And now these days, we're talking about every one of these projects has hundreds of people, probably tons and tons of compute. So these things have really scaled up.
(Joel Beasley at 00:51:10) Dude, well, that is so cool. I think I just found my resident expert on language models. Thank you, Guy, for doing this. Any other things that we need to get out there to the world?
(Guy Gur-Ari at 00:51:19) No. Please try Augment out. I'm sure you'll love it, and let's get coding.
(Joel Beasley at 00:51:26) Thank you so much for listening. And if you found this episode useful, please share it with a friend or colleague who you think would get value from it. And if you have topics that you'd like to hear discussed on the podcast, either add me on LinkedIn or send me an email, [email protected]. Every time I get an email or LinkedIn message, it absolutely makes my day and inspires me to keep going.