Episode 743 ·
Fusing Generative AI & Video Production with Anastasis Germanidis, Co-Founder & CTO at Runway
Today we’re talking to Anastasis Germanidis, Co-Founder & CTO at Runway. We discuss the groundbreaking AI technology that Anastasis is working on, how generative AI is impacting the industry for media professionals, and how Runway cut The Late Show’s edits down to 5 minutes.
All of this right here, right now, on the Modern CTO Podcast!
For more about Runway, check out their website here.
Have feedback about the show? Let us know here.
Produced by ProSeries Media.
For booking inquiries, email [email protected]

About Anastasis Germanidis
Anastasis Germanidis is the Co-Founder & CTO at Runway, an applied AI research company shaping the next era of art, entertainment, and human creativity. Runway was one of Time Magazine’s “100 most influential companies” in 2023. Runway has been a persistent viral sensation in recent years, and is behind many of the most famous AI demos online.
About Runway
Runway is a research company pioneering new tools for human imagination. Runway has been at the forefront of multi-modal AI systems ensuring that the future of content creation is accessible, controllable and empowering for creatives. Runway’s mission is to ensure that anyone anywhere can tell their stories. We believe that deep learning techniques applied to audiovisual content will forever change art, creativity, and design tools.
Transcript
**Cleaned Transcript:**
(Intro Narrator at 00:00:00) Today, we're talking to Anastasis, co-founder and CTO at Runway, about the intersection of generative AI and video production. You're listening to Joel Beasley, Modern CTO.
(Joel Beasley at 00:00:16) And I did see that you were the co-founder, and so I was hoping you could set me up with just a brief explanation of what Runway is.
(Anastasis at 00:00:25) Runway is an applied research company. It's how we like to call ourselves. And what that means is that we do fundamental research in new AI techniques and AI models, and then we deploy them to a series of tools for creative teams and creative individuals to be able to work faster, to make things that weren't even possible before. So the most recent effort at Runway has been around video generation. So we released a series of models that allow you to generate a short video from just a text description or initial image or a driving video.
(Anastasis at 00:01:07) And we're just very interested in how generative models, like AI techniques, can really help speed up the processes of creators and allow them to try more ideas faster or to accomplish things that traditionally required a really large amount of resources and budget.
(Joel Beasley at 00:01:26) So what are people building with this that they can justify? I know you have a free version, but people spend money with you. What are people generating with this?
(Anastasis at 00:01:36) Yeah. So Runway is being used in a lot of different parts of the production process. If you take the pre-production, production, post-production aspect of building, let's say, a video for a short film or an ad or any kind of video, Runway is used on the production side directly in some cases where you can generate some portions of your final video with a model like Gen-2, which is our text-to-video model. But also, it's being used a lot in the post-production process where we have a series of VFX tools for rotoscoping or for inpainting, which are very part of the traditional video editing process, and they can be really time-consuming tasks. So we've had those being used in films like Everything Everywhere All at Once.
(Anastasis at 00:02:29) They used Runway for some of the VFX shots in the film, and it allowed them to perform those tasks way faster than more traditional techniques. And Runway is also used in the pre-production stage. And in pre-production, you're basically ideating on ideas you're trying out. You're basically trying to figure out, previsualize how the final output would look. And Runway makes it really fast to basically generate any 80% version of the final result and show how the shots, each of the shots of the final video, would look.
(Anastasis at 00:03:04) And you can use those internally to pitch a specific idea, or you can use them to define how the shots will look before you actually go into production. So we see Runway being used all across this spectrum of the creative workflow.
(Joel Beasley at 00:03:21) How are these—so I have friends with Animal Logic. They make The Magician's Elephant and Peter Rabbit, DC League of Super-Pets. I got to visit their studios. I've had them on a couple times over the years, and they're really great, brilliant people. How would the interaction with your product work?
(Joel Beasley at 00:03:40) Would some of their video engineers just go to the website and just start using it? Do they create this big enterprise contract between you guys? How are people actually implementing the use of your technology in their companies?
(Anastasis at 00:03:54) So both options are there. So we have, as an individual, you can go to Runway, you can sign up with a free account or subscribe to one of our plans, and use the product directly without needing to sign up for a large team account. But we have large companies using Runway, and in those cases, there's more needs around collaboration, around being able to share assets more efficiently, around training models that are specific to your use case, which is something we provide, where if you have a data set of images and you want to train a model just on those images, you can create a model using Runway. So as individuals, you can go and use Runway today. But if you are a company that has needs around collaboration and more fine-tuned use cases, then there is an enterprise plan where you can have those additional options.
(Joel Beasley at 00:04:44) So your tool is crazy cool. You can literally type a text prompt that says, "claymation monkey swinging through the forest," and it will create this incredibly high-resolution, detailed, beautiful video clip of this claymation monkey swinging through the forest. It's unbelievable. How did you train that data set? Yeah.
(Joel Beasley at 00:05:09) I guess that's the question. How did you gather all the video and train this, or did you build on top of some existing model? How did that go about?
(Anastasis at 00:05:18) Yeah. So Gen-2 is essentially the culmination of research that we've been doing almost since the beginning of the company. The first problem that we needed to solve quite well was the image generation part. So just getting a shot to be as high fidelity and high quality as possible. And we did a lot of research on the image generation side.
(Anastasis at 00:05:41) And at some point last year, we had this breakthrough moment where image generation was finally possible to use the outputs of those models into production, into build tools around those models. But that took years of research to basically get to this point. And since then, we leveraged some of the findings and some of the learnings, some of the architectures for image models into the video domain. And what's different is that now it's not about generating a high-quality initial shot, but you essentially need to train a model that can learn how to generate temporal sequences, frame sequences. So in a sense, you're starting from an image model, and then you're retraining the image model on video sequences, on millions of videos.
(Anastasis at 00:06:34) And from those, the model essentially learns temporal dynamics, learns how things move in space. So there definitely, some of the learnings from the image domain have been transferred to the video domain, and there's a lot of similarities there. But we needed a new approach in architecture to tackle video generation.
(Joel Beasley at 00:06:55) I'm gonna only get a little nerdy for a second, then we can come back up to a higher level. So if I have a model and I ask it to generate something, I prompt it to generate an image, that's one part. And then it taking that image and allowing it to move, as I think you said, in temporal space, is there additional data attached with that image because it was generated from a model that you own? Or can I just take any image and send it to the video model and it would just understand that image?
(Anastasis at 00:07:24) You can take any image, and the model should be able to complete it as a video.
(Joel Beasley at 00:07:30) Whoa. That's crazy cool.
(Anastasis at 00:07:33) So we've seen it being used with real images. We've seen it being used with illustrations. We've seen it being used with generated images. And the model is quite versatile. So it is able to handle and generate motion that corresponds to what you would expect from that initial frame.
(Joel Beasley at 00:07:52) I've got 700 hours of video of me interviewing people, high-quality video and audio. Can I pump some of this into a model and say, you know, act like Joel hosting an interview, and it just creates new images of me talking to somebody?
(Anastasis at 00:08:07) Yeah. So today, you can go into Runway. We currently only have, in the app, the option to train an image model. So you can train it on images of you and be able to generate—and the model can generalize quite well. If you want to change the setting or change aspects of the final output, you can do that with the trained model that you've created based on frames from your videos.
(Joel Beasley at 00:08:38) So I could take a bunch of video of me and then I could say, you know, put Joel in a tank top, and it would put Joel in a tank top.
(Anastasis at 00:08:45) Yeah. So we have an inpainting model, which is what we call the process of you're only changing certain regions of the image. So you can basically brush over the parts that you want to change, and the rest of the image will remain the same, and then you can describe with a prompt how you want to modify the image.
(Joel Beasley at 00:09:05) Dude, that is so cool. That is so cool. How did you get involved in all of this?
(Anastasis at 00:09:10) I've been interested in AI for quite some time. I would say from the days of high school, I first read a book about neural networks. Neural networks at that time were a technique that had fallen out of favor. People were thinking they had no more promise or potential, and they were trying other techniques for machine learning. But there was something really compelling about this model that's biologically inspired, kind of inspired by how the human brain works and can learn based on data, rather than being explicitly taught specific aspects of the world. They can learn from data, gain knowledge in that way.
(Anastasis at 00:09:58) So I've always been interested in this topic. I've also had an art practice for a long time. I've been very interested in making artwork primarily with technology. So I essentially had these two separate careers where I was working as an engineer at startups, primarily infrastructure and back-end work. And at the same time, I was making art on the side.
(Anastasis at 00:10:26) And at some point, I decided I wanted to entirely focus on making art. And so I went to this art program that's at the university focused on the intersection of art and technology, and that's where I met my co-founders. And we started exploring the potential of machine learning and AI in the creative domain. And really, those techniques were just starting to become powerful. I think that was around 2016, '17.
(Anastasis at 00:10:57) You had those early methods that produced lower-resolution images that weren't quite production-ready, but they were so interesting, even from an experimental point of view. So we started making small projects together. We started making tools for filmmakers, for designers, for architects to be able to use those techniques because, especially at that point, there was a really high barrier to entry to actually be able to use them. You'd spend hours to actually get things to work, to install the dependencies or set up your environment before you could actually do interesting stuff with those techniques. But we saw that once artists had access to those tools, they could make some really interesting work with them.
(Anastasis at 00:11:40) One of my favorite examples from those early times of collaborating with my co-founders was, there was this model that NVIDIA had released based on—that was trained on self-driving car footage. So basically, it would take a general layout of a scene that consisted entirely of street views. So it would be things like cars, pedestrians, traffic signs. And based on that very high-level layout, it would generate a photorealistic version of that layout. So it would generate artificial street views.
(Anastasis at 00:12:20) So it's a very utilitarian model. It's not something that was meant for a creative use case. It was just more as part of the research that was being done for self-driving, for improving self-driving cars. But we decided to take it in a different direction where we built this drawing tool that used that model, and essentially, you could draw your own version of a street scene. You could draw where the pedestrians should be or where the cars should be or where other elements of the scene should be, and then you could generate an image based on that layout.
(Anastasis at 00:12:57) And we had this installation where people, artists could go and try it out. And we saw people making all kinds of surreal imagery from that very narrow model. So they would make giant pedestrians and tiny cars, or the rain consisting of traffic signs were falling on the ground. So that was a first indication of you can make something that's—you can take a model that's traditionally not meant to be for the creative domain or creative use cases, and you can find an interesting angle to it, and you can put it in the hands of artists. And they can find ways in which they can really express an interesting vision from this very not super creative model.
(Joel Beasley at 00:13:50) That is so interesting, man. Raining street signs. That makes me think about prompting. Right? We use a lot of GPT-type tools.
(Joel Beasley at 00:14:02) I've played a little bit with some of the image generation tools as well. And what I've quickly learned—and I lean on the creative side of things—and what I've learned is that prompting is absolutely a skill. The results you're gonna get are directly connected to how good you are at prompting. Obviously, you've got a lot of experience prompting. Can you tell me, what makes good prompts?
(Anastasis at 00:14:28) Yeah. So prompting, as you mentioned, is a skill itself. I think a general misconception of how those tools work is that there's no learning curve, that you can jump in and make amazing, beautiful things out of the spot. That's not the case. There is a skill that you need to develop in how you're able to control those models.
(Anastasis at 00:14:50) I think broadly, one general principle I would say is just being as descriptive as possible, and then trying new keywords, being very gradual with your edits on the prompt. So you can try adding one keyword and then seeing what the effect of that keyword would be on the final output. And then you can collect over time the keywords that really match the style that you want to create. So there's no hard and fast rules around that. It really depends on the style that you want to get to. Those models are very versatile and very capable, and they can generate from photorealistic outputs to more stylistic things that look like 3D renders, 2D, pixel art.
(Anastasis at 00:15:35) It's really bound by your imagination, and it's just a matter of just exploring and iterating and carefully tuning, adding new words or new phrases and seeing what the effect of those is to the final output and then going from there.
(Joel Beasley at 00:15:51) Do you think that there will be an ability? Because my strength was business software applications. I'm a software engineer, and that's where my strength was. So never any of this type of AI model generation. That's always just been hobby interest, me playing with the tools I see.
(Joel Beasley at 00:16:07) But I'm curious, will we get to a point where in the realistic future or the nearby future where I can run a prompt with a GPT, or let's just say Runway since that's your company. I can run a prompt on Runway, and it generates something, and then you say, okay, well, I want to incrementally add a keyword or two. But will we get to the point where I could converse with the model and ask it why it did something and it be able to explain it to me? Or is that beyond the scope of this technology?
(Anastasis at 00:16:40) That's definitely the direction that things are taking. So we're moving from this one-round conversation where you can describe one thing and then get an output to something that's more conversational and more like a creative assistant where you can kind of continue. And there's a few problems to solve along the way. One is a memory system where the assistant can remember what you've asked in the past, like how you what kind of style of outputs you're trying to get, like one of the most common edits. But that's definitely achievable, and I think we're going to get to that point. A big focus for us has been really on getting the outputs to be as high fidelity and high quality as possible.
(Anastasis at 00:17:25) But once that's solved, I think there's a lot of potential interesting use cases around how you actually iterate on the outputs and how you can do more precise edits. Like right now, in order to do these precise edits, you probably need to go back to a more traditional video editing program to just, if you want to make sure that the timing is fully right, if you want to really match, let's say, sound to the file output. But in the future, you can imagine that essentially you can prompt for the interface that you need, the interface that really responds to how you like to perform those edits, how you like to create. And then you can really build a personalized version of a creative software that's really attuned to your own style and your own workflow. So I don't think we're that far from getting there.
(Anastasis at 00:18:16) It's just a matter of solving what's being called multimodal problems. And multimodal meaning building models that understand both language and image and video. And once we have that in place, then you can have a system that not only can generate things for info, but that can also critique them and basically give you feedback and say, "I generated something, but maybe the pace is not quite right," and then kind of do a second run and further iterate with you. So that's definitely in the near future.
(Joel Beasley at 00:18:54) Dude, that is so cool. Do you think that we're—well, I'm going to stop asking the question that way. Do you think that we get to the point—because eventually, we all end up inside the computers anyways. But have you seen any projects where people are using Neuralink-type technology, brain scan reads to think of images and generate it, basically painting your imagination? Have you seen anything like that?
(Anastasis at 00:19:22) There's been some work recently into being able to reconstruct images from, I believe, fMRI scans. So essentially, the setup is that subjects would think of a specific image and a specific subject, and then using kind of the signals extracted from fMRI scans, you'll be able to use an image recognition model to reconstruct what the subject was thinking about. This is still very early-stage research. I don't expect that to become a product that you can use anytime soon. But I do think eventually we're going to get to the point where this thought-to-image or thought-to-video might become a possibility. It's not really the domain of research that we're actively focused on, but I've seen some research that seems really promising on that front.
(Joel Beasley at 00:20:13) I'm sure some governments would like that. Yeah, because that would actually show some really interesting truths about our universe, right? If I think of a red box and mentally imagine that and you scan my brain, and then you think of the same image and scan your brain, if the model could come to the conclusion that we're both seeing the same object, that would be very interesting information, would it not?
(Anastasis at 00:20:40) You mean from the perspective of almost getting to a consensus for how a thought translates into kind of visual imagination?
(Joel Beasley at 00:20:49) Yeah. Like, this is the mental representation of a red box to humans, and it's firing similarly amongst all of us. That'd be cool. That'd be weird.
(Joel Beasley at 00:21:03) What other—are you seeing any, you know, obviously, you're in this industry, and so you're always seeing the latest cutting-edge stuff and the experiments that people are running and you're keeping up with that. What have you seen that's given you some pause? You've seen and you're like, "Wow, that's actually pretty interesting or out there."
(Anastasis at 00:21:22) From a perspective of something that's impressive or something that's really compelling? Something I'm very excited about is coding assistants and language models applied to software engineering. So with systems like Copilot, and there's kind of an emerging set of tooling for that purpose. I really believe that programming in the future will look a lot more like you have an idea of a system design in your head, and then you translate it into an implemented version of it as quickly as possible, and you can spend time iterating on the actual design of the system versus spending time iterating on the actual code and having to understand the kind of deal with the idiosyncrasies of the programming language that you're using. Rather, it's going to become much more being able to describe in natural language the kind of system, the kind of application that you want to build, and then get an initial version of it and then further iterate from there. So I'm very excited about tools like Copilot that really speed up the kind of engineering time and let you get from an idea to a final output way faster.
(Joel Beasley at 00:22:40) Josh, do you—who was the guy we talked to the other day about Copilot that I disagreed with?
(Intro Narrator at 00:22:45) Oh, that was Gal Saraf.
(Joel Beasley at 00:22:47) Gal. I love you, Gal, by the way. But yeah, we had a great conversation because I firmly believe that it is massive value add, that it's here today, that you could inject it into production workflows and save crazy time and energy and effort. And then there are some people like Gal who had some good points, but he just doesn't see it as being there currently. And everybody has their own take on it, right? But for me, I've used it, and while it's not—I guess I'll back up for a second and maybe you can help me understand this thought, but that's why we explore and we can edit and everything. When people respond to me about this—because I'm bullish on it, I like it—they typically, the biggest area of resistance I get is that because it's not perfect today, it's not there. Like, because it's not perfect and because I can't sit down and just tell GPT to write me a world-class book that's competitive with Harry Potter in one prompt phrase, that the technology isn't there yet. We almost imagine it as only existing in its most perfect ideal state, but that's not how anything ever works. How everything always works is it comes out, it's super rough, it can barely walk, you got your baby, it can't even feed itself, and then it grows up and becomes a fully functioning person. So I just notice that some people respond like that, and then other people respond being able to see the potential right from the beginning and get real excited. When people are learning about your technology and you get to see the comments online and you get to see people writing different articles, I'm sure there's some haters that say, like, "Oh, you know, Runway is not there yet," or it's—and then there's people that love it and think it's unbelievable. How do you process and deal and sort of talk with yourself internally about the technology? Because you can see, you have the vision. You can see where it's going to go, and you know where it's at today, and you know where it came from. So how do you handle all of that?
(Anastasis at 00:24:44) I think there's two parts to this. The first part is that it's better to look at the trajectory versus just looking at, taking one moment in time and seeing what the technology is. Like, something that we saw even in kind of 2018, 2019 when Runway was just starting out, the results were not there at all. Like, we have some videos that we tried to generate at that point that looked extremely rough, very low res. It was very difficult to imagine that we'd get to the results that you get today, but you could see that the resolution was doubling every year. You could see that fidelity was doubling every year. So there was no fundamental reason why that couldn't continue. And at any given point, people could point at the flaws of the technology and say, like, "It's never going to get there." But then we're seeing every year the feedback that we're seeing from people about the flaws of the outputs of the images was becoming outdated very quickly. And so every year, when you see that pattern again and again, at some point you kind of assume that things are going to get better, and you learn how to respond to kind of those comments, which is saying, "Just look at the trajectory. Look at where we were a few years ago. Look at where we are today. Assume, extrapolate from this where we're going to be in a few years." And then it's going to be clear that those limitations are not fundamental limitations, but rather things that we just need more engineering time, more maybe scaling those models to be able to solve. The other side of it is that at Runway, we philosophically are not envisioning a future where you type in a simple prompt and then you get back a ready-made, like, final form that has everything is perfect and everything expresses your vision from the prompt perfectly. We see it's always going to be an iterative process because you don't quite—when you're starting a new project, you create a project, you really don't know how, what you want to create fully. You have a very rough idea, an outline or specific aspects that you want to make sure are in the final output. But the other details are not yet determined until you actually sit down and you actually start creating things. And so the way we see Runway is it will just allow you to iterate and explore ideas faster, but it's not going to come up with ideas for you. We don't really see our tool being used for kind of coming up with a vision of what you want to create. It's more like a multiplier of your ability to turn your vision into final outputs. So because of that, we never—I don't want us to get to a state where those models take part in actually coming up with ideas. I think it's really empowering artists, empowering creators to create more, to create faster, to create, to get their vision into a final output in a much more kind of detailed, high-fidelity way than ever before. Yeah, I don't know if this answers your question.
(Joel Beasley at 00:28:09) Yeah, it does. I did want to touch on—I know as we're wrapping up here, we got, you know, five minutes left or so because you have a hard stop. I did see a bullet point in the prep document that said Runway got The Late Show, The Late Show on TV, their edits down from five hours to five minutes. And I was hoping you could just tell me a little bit about that.
(Anastasis at 00:28:29) Absolutely. So Runway is being used in a variety of creative teams, media companies, ad agencies, filmmakers. One of our favorite examples was the VFX team behind The Late Show using Runway for a lot of the sketches, comedy sketches in The Late Show. And as you know, those episodes are being made within the same day, and so the time at which they can make them is very—they need to work on very accelerated timelines. So they use specifically our rotoscoping tool. It's called Green Screen, and it allows you to very quickly separate foreground and background in a video. And this is a very fundamental piece of a lot of effects work because you might want to replace the background from the subject. You might want to bring the subject to a different setting. It's a very common aspect in a lot of those kind of comedy skits. And they use Runway and the Green Screen tool in order to really speed up that process, which traditionally, if you take more traditional kind of rotoscoping software, it's a very tedious process. So you need to go through every frame of the video, kind of manually annotate what's the foreground, what's background, and it can really be cost and time prohibitive in terms of executing a given project.
(Joel Beasley at 00:29:52) See, that is amazing. And for anybody who's listening, you can go check it out. They have great examples, beautiful website, excellent branding. I always like when I see companies that have really good branding. Very visual examples all over your website of how people are using Runway in video production workflows and all of that. But, man, we did it. We made a podcast, man. How do you feel?
(Anastasis at 00:30:13) Feeling good. Thank you for the chat.
(Joel Beasley at 00:30:16) Thank you so much for listening. And if you found this episode useful, please share it with a friend or colleague who you think would get value from it. And if you have topics that you would like to hear discussed on the podcast, either add me on LinkedIn or send me an email, [email protected]. Every time I get an email or LinkedIn message, it absolutely makes my day and inspires me to keep going.