Episode 272 ·
David Turek - CTO at CATALOG
Today we are talking to David Turek, the CTO at CATALOG. And we discuss the incredible innovation of DNA Storage and Computing, how to encode, decode and search DNA based data and the impact advertising single figures of merit like “gigahertz” and “core’s” has on the market.
All of this, right here, right now, on the Modern CTO Podcast!

About David:
David comes to Catalog from IBM where he held numerous executive positions in High Performance Computing and emerging technologies. He was the development executive for the IBM SP program which produced the first commercially successful massively parallel system; he started IBM’s Linux Cluster business; launched an early offering in Cloud computing called Deep Computing Capacity on Demand; produced the Roadrunner system, the world’s first petascale computer; and was responsible for IBM’s exascale strategy which led to the deployment of the Summit and Sierra systems at Oak Ridge and Lawrence Livermore National Laboratories respectively. He has been invited to testify to Congress on numerous occasions regarding the future of computing in the US and has helped establish technical collaborations with universities, businesses, and government agencies around the world. In his free time David enjoys going for walks in the woods with his dog, Huckleberry.
About CATALOG:
The world will generate 160 zettabytes of data in 2025. That’s more bytes than there are stars in the observable universe. Conventional storage media like flash-drives and hard-drives do not have the longevity, data density, or cost efficiency to meet the global demand. CATALOG is building the world’s first DNA-based platform for massive digital data storage and computation.
Transcript
(Joel Beasley at 00:00:00) Hello, my friends. Today we are talking to David, the CTO at CATALOG, and we discuss the incredible innovation of DNA storage and computing, how to encode, decode, and search DNA-based data using chemistry, and the impact advertising single figures of merit like gigahertz and cores has on the market. All of this right here, right now on the Modern CTO Podcast. Here we go.
(Joel Beasley at 00:00:31) This is the Modern CTO Podcast. So when you first saw this company, what was your initial impression?
(David at 00:00:49) Well, curiosity for sure. I mean, I hadn't been keeping up with DNA as a technology. My role in IBM was in many cases at the forefront of technology, but not comprehensively to the point where I would have studied this. And as I dug more into it, I saw that it really represented a confluence of a number of technologies. So you have hardware, chemistry, software, all coming together to create a new approach to certain problems in computing.
(David at 00:01:23) On the surface, people look at it as an accommodation to storing data, so naturally you talk about encoding data into DNA and so on. But I think the real promise down the road will be how we use DNA for computing in concert with data stored in DNA as well, creating almost a completely unique kind of platform here. So it was the opportunity to operate at the cutting edge and to really see if there was an opportunity to transform the industry. The industry in this case being the electronics industry, which we've all grown accustomed to with respect to matters of data storage, archiving, computing, and so on.
(Joel Beasley at 00:02:04) So you're at IBM for how long? Twenty years you said?
(David at 00:02:07) Oh no, much longer than that. I was responsible for our supercomputing business. And so a lot of the number one supercomputers in the world were done on my watch, going back into the 1990s and right up until recent days when we, in 2018, deployed the Summit and Sierra systems at Oak Ridge and Livermore, which at the time were the number one and two systems in the world. I guess they're now two and three courtesy of the Japanese innovations.
(David at 00:02:37) And in between that, of course, there was a lot of other innovation with respect to file systems. I did three file systems. I launched the IBM project in grid computing, Linux and IBM. I was one of the founding members of that and a number of other things as well. So it was pretty eclectic.
(David at 00:02:57) And one of the other things that attracted me about CATALOG was we've been kicking around the ideas of what the future of computing looks like. So Moore's law is coming to an end, and in fact, there are diminishing returns based on every turn of the crank of technology. And that's just one of the constraints. I mean, there are constraints that manifest themselves in speed of spinning disk and how you structure memory. And then, of course, there are energy constraints.
(David at 00:03:25) There are topological problems with enormous core counts. There are networking issues. All these things are sort of a pox on our head simultaneously. So that is background. I began looking at how we thought computing would evolve, and so it really made real the notion of the best tool for the job at hand.
(David at 00:03:50) And what I mean by that is the following. If you think about computing, oftentimes people think of it in terms of applications. It's the improper way to think about it. You should really think about it in terms of a workflow and all the component pieces that constitute a workflow. So I'll give you an example in the oil and gas industry, which has been a heavy user of high-performance computing for decades. You could look at seismic processing or reservoir modeling and say that's what computing is in the oil and gas industry. But when you look at seismic processing, you begin to realize that the algorithmic centricity of that statement, whether it's elastic waveform inversion or Kirchhoff methods or reverse time migration, etc., turn out to be just a very, very small fraction of the entirety of the computing problem that governs the process of understanding seismic processing.
(David at 00:04:49) And most of it's really wrapped around data: data acquisition, factoring data, iterating on it, calling it, curating it. So I did an experiment some years ago where I looked at a seismic processing workflow with a large oil and gas company. And all the time was being focused on improving the computational element associated with waveform inversion. And I finally calculated using the notorious back-of-the-envelope method that if we were able to solve that problem instantaneously, we would have reduced the total amount of compute time in the seismic processing process by maybe 3%. But if we could figure out a way to sort data faster, we could reduce the amount of computing by 50%.
(David at 00:05:43) So it's important, therefore, that you define the domain of application—application here in the sense of the technology that you're deploying—in terms of workflows as opposed to computer applications per se. Computer applications are too easy to sort of narrowly constrain your point of view about the problems that you're trying to solve. You know, so somebody will come along and say, "Well, you know, I ran this particular algorithm through NVIDIA GPUs and I got a 20x speedup." Okay, fine. What did that mean to the overall execution of the workflow? "Oh, half of 1%." So we need to be comprehensive in our understanding of what's going on here. And I thought that as we were looking at the encoding problem of data into DNA, we began to see opportunities to deploy certain kinds of computing models against this data as well, giving the possibility of looking at hybrid kinds of systems deployed in the marketplace.
(David at 00:06:45) You preserve von Neumann systems. It's what you do payroll on. It's what you do a lot of different things on. You carve out a little niche for maybe the application of quantum. You do certain maybe exotic things with AI-based systems.
(David at 00:06:59) And then there are these biological systems, of which DNA is one, or biologically inspired systems. Neuromorphic computing is an example of that. And you put all these things into your bag of tricks and you allocate and deploy them as necessary in the context of the problem at hand. And that's the way we think about it at CATALOG. It's not a displacement of a particular style of computing.
(David at 00:07:23) It's an adjunct, it's an amplification, it's an improvement over certain domains where we can provide some unique value.
(Joel Beasley at 00:07:31) What did you say, neuromorphic computing? What is that?
(David at 00:07:35) It's just sort of a design of computing that's inspired by brain function, for example. So you might look at the way circuits operate in the brain and you might take lessons from that and encode that in silicon and create a neuromorphic chip, which would apply certain neurological principles, if you will, to do computing. And a number of companies have played with that. IBM for sure. Other companies as well for some number of years. And it's just another tool in the bag of tricks or the toolbag that you have.
(David at 00:08:08) So when we think about the future of computing, it's this theme of amalgamation of approaches that need to be brought to bear to work on problems at hand. So I not only focus on optimizing my seismic processing problem from an algorithmic perspective, but I work on all the data issues, which turn out to be substantially more important than the algorithmic piece and all points in between. And I think that's the mindset we have at CATALOG is to be expansive in our consideration of what these workflows look like and then understand where our technology might play a role.
(Joel Beasley at 00:08:47) And would neuromorphic computing extend to software? Like if you, I believe graph databases, if you've ever seen a visualization of those, I believe they were inspired by how the brain stores information with neurons and synapses.
(David at 00:09:01) Well yeah, and in fact, it turns out that the way we encode data at CATALOG is actually using a graph theoretic construct. So our encoding of data effectively builds a graph. So when we say, as we say, that we can encode a terabit of data a day, what it means is we've created a graph that at the terminus of every branch through the graph, there's a representation of a bit of data, which sum up to a terabyte of data, where every one of those pathways through the graph is actually represented by a concatenation of small DNA pieces, which in turn give rise to a longer piece of DNA. And all these long pieces of DNA that represent an entire branch through the graph are unique from each other. So they give a unique representation of an address in a bit stream of a terabyte of data.
(David at 00:10:00) So the DNA molecule that represents bit one is completely different from the molecule that represents bit, you know, 10,000,025, which in turn is completely different from the one that represents, I don't know, the 350 gigabit bit of data. So it's these graph concepts that have kind of informed the way we concatenate DNA molecules to represent data with DNA.
(Joel Beasley at 00:10:30) Can you actually encode a terabyte of data per day currently?
(David at 00:10:34) Terabit of data. Terabit. Yep. So let me explain. So everybody who works in the domain of encoding data into DNA makes claims.
(David at 00:10:46) What we did was—and by the way, these claims are based on paper and pencil calculations, wet lab chemistry, bench chemistry, things like that, and then extrapolations from that to what would be possible. What we did is we actually built a machine to automate the entirety of the process of writing data into DNA. And this machine, as I said at the outset, represents this confluence of software, hardware, and chemistry embodied in a physical device that will deploy pieces of DNA that are predesigned by us. Think of them as Lego building blocks, where each building block is the same length, but each one has a different color.
(David at 00:11:28) And we've got lots of different colors in play. And so when you think about going back to this graph theoretic concept, you know, every segment of the graph is a piece of Lego with a unique color, and then that branches off to two more things. Each of those represents another deployment of a Lego building block. Your Lego building block is a small snippet of synthetic DNA, and on and on it goes. Our machine will actually effectively deploy these Lego building blocks in something that we call a reaction spot, which then gets chemically ignited, if you will, to react and produce this longer string of DNA or a long string of multicolored Legos.
(David at 00:12:14) And it's unique from everything else that we do. Our machine is capable of doing a terabit a day. We haven't run it at that speed yet because there's costs involved and so on, and we're still debugging elements of it as well. But the simple fact that we have a physical device lets us gain insight into a myriad of different kinds of issues that if you're constrained to working out these issues in bench chemistry, you'll never witness. You know, how reliable is the data? How fast is the media passing through the machine that can capture the data that you're encoding? How error prone is it? All these kinds of things are informed by the fact that we have a physical device. And, you know, we've written things like the contents of Wikipedia, which is about 14 gigabytes. We have executed a project for a California-based media company to encode parts of a movie into it and so on, just to demonstrate capability, but to also inform us about the operational parameters of the system.
(David at 00:13:21) So yeah, we can do these things. But more importantly, our experience has told us how we would modify the system to actually take it up to a much, much greater rate of data encoding than what the current design provides. So we think we can get to probably 10,000 times greater speed than what we have today by having sort of understood what the limitations are, what the issues are we're trying to achieve those levels of speed. It's not just chemistry. There's engineering, there's software, and there's hardware design as well.
(David at 00:13:57) And all those things are part of what we've done. A video of that, by the way, is on our website.
(Joel Beasley at 00:14:02) It's amazing. It's amazing. Yeah, you actually watch the machine print the DNA and they explain the Lego blocks a little bit in there. And yeah, they're really great videos. I saw one and they take a trip to, I think they took a trip to Europe because they had a team in two locations and then they were doing stuff, and I think that was the video about them encoding Wikipedia.
(David at 00:14:23) Right.
(Joel Beasley at 00:14:24) It was a good video. Yeah. I have a question. You gave me a thousand questions. Okay, because I'm curious, so please bear with me. I'm like a kid in the candy shop right now. Okay, so we talk a lot about encoding. We talked a lot—or you just said something, you know, we hope to achieve, you know, 10,000 times faster—but some earlier you said something along the lines of Moore's law. A lot of people understand Moore's law, right? That the computing speed doubles. And you said it was kind of coming to an end, hitting some constraints, right? But then we're—but then new technologies are emerging. Do you think enough new technology will just spark out of the ethos and come in to keep Moore's law going, or do you think Moore's law is done?
(David at 00:15:11) Well, I think we should constrain Moore's law to the domain from which it emanated, which is this notion, if you will, of circuit density and how that improves over the course of time. And so to laypeople, the characterization of Moore's law has been by virtue of declarations of technology generations. You're at 14 nanometers, you're at 12 nanometers, you're at seven nanometers.
(David at 00:15:40) Guess what, everybody? There's no such thing as zero nanometer technology. It has to come to an end. And one of the things that we're seeing when I talk about slowdown is the effort required, including the money required to get to the next level of technology, is growing up quite dramatically. And the benefits that accrue as a result of where we are now in the 10, seven nanometer range, what have you, are beginning to get marginalized.
(David at 00:16:08) So the big speedups that you saw maybe 15 or 20 years ago when people were operating at, you know, 45 nanometers or something like that and made a jump down, they're not seen anymore. A new generation of technology now may net you 6%, 7%, 8% in terms of performance. But even that you take with a grain of salt because that doesn't address effective performance. So it typically talks to, you know, max, a synthetic benchmark that gives you some characterization of speed that you're meant to think is sort of universal, but which is in fact not universal. So there are a lot of applications that see very, very small improvements.
(David at 00:16:54) The industry has known about this impending doom scenario, if you will, for at least 20 years. And so that's why you'll note that I think I could be plus or minus off by a year—around 2000 or 2001 was the last commercial you saw on TV, PC commercial, where somebody said, "Buy a new PC because it's now running at X gigahertz and that's intrinsically good." The trend at that time was, "Well boy, by the time we get to 2020, the gigahertz will be so great you're going to need a nuclear power plant to power and cool your PC," which of course was never going to happen. So you have sort of this reductio ad absurdum argument that was there in front of everybody.
(David at 00:17:38) So post that era of singular fixation on gigahertz, if you will, the industry moved en masse to multi-coreism. So now it wasn't about how fast a single processor went. It was about how many cores were on a chip. We've got two cores, we've got four cores, we've got eight cores. And now, you know, laypeople have equated more cores to being intrinsically better than fewer cores.
(David at 00:18:06) And the problem, of course, is that as you do more and more cores, you have this sort of topological problem of how you put them all together to feed them with data. And so just as in the early part of the 2000s, people speculated about the fact that in 20 years you would need a nuclear power plant to cool a PC, now the issue is, "Well, how do I design a network inside of a chip to connect, I don't know, 10,000 cores or something like that?" It's just absurd.
(David at 00:18:36) But these are the kinds of band-aids that come along. And I say band-aids with great respect for band-aids. You know, not everything comes along and solves a problem forever. You make these incremental improvements through time. And listen, if you get 10, 15, 20 years out of an idea, that's pretty good, because by then new ideas will come along. You get another 10 or 15 or 20 years.
(David at 00:19:02) So yeah, things are operating now in a domain of progressively more problematic constraints, which is a wonderful circumstance. And I say that because it's only in an era of constraint that you see real innovation spark forward, because otherwise, too easy. Think about this. In the early days of PCs, you had maybe 64 kilobytes of memory on your system, and people thought, wow, that's a lot. And then it turned out producing more memory got to be pretty cheap, pretty much commodity-like.
(David at 00:19:40) So you went to megabytes of memory and you went to gigabytes of memory. And you didn't have to think very hard about it. You just put more memory in the system. Gave rise to a generation of pretty sloppy programming, because probably the best programmers in the world were the people who came to maturity in the fifties and early sixties where everything was hyper-constrained. And as a result, their programming skills made up the problematic elements of that.
(David at 00:20:06) When you start just throwing around huge megabytes or gigabytes of memory, you could be sloppy in your programming because, well, memory was cheap. So now when you think about Moore's law and the constraints on that, the manifestation of constraint, the marginal improvements that you see through time, despite the ever-increasing cost to get there, will spark new ideas. They'll come and, and by the way, there's one evolving in front of our eyes right now. The idea of accelerators is an idea that was around since the beginning of time in computing. And it never really grabbed hold, whether it was FPGAs or DSPs or whatever idea came forth or specialized custom design accelerators. They never gained traction because by the time it took to get them commercialized, Moore's law improved so much that the advantage provided by an accelerator didn't matter.
(David at 00:21:01) Right? Be much more efficient to simply go after conventional microprocessor design. But what's happened in the last five years is that the onset of a decline of benefit from Moore's law has sparked huge interest in accelerators, sparking huge growth in Nvidia, in AMD, and perhaps others along the way. You know, the RISC-V community, for example, is beginning to come of age, if you will, and probably through this dimension of acceleration. So that's a manifestation of innovation in response to the constraints that everybody sees with Moore's law.
(David at 00:21:40) There will be more like that as well.
(Joel Beasley at 00:21:43) Thanks for clearing that up. I didn't realize that it was just constrained to that one specific. So many people—I mean, I'm sure at one time I did when I read it because at some point I looked it up on Wikipedia—but the way marketing works and conferences and everything, the lines always get so blurred. But I was curious, so a couple weeks ago, I was talking with Thomas Hazel, and he is the founder of this company called ChaosSearch. And we were having this conversation about data and compression.
(Joel Beasley at 00:22:10) And what he had done—and he's very intelligent as an engineer—but they had made, I believe it was a compression algorithm that took data, made it smaller, but then kept the ability to search on top of it. And when they were looking for how to apply it into the marketplace, what they found was companies were spending enormous amounts of money on storing log data to then have—they would store it in their S3 and then an analytics tool would plug into their S3 and read off of it. And so companies are spending millions of dollars a year, because it needed all that raw data. But they were able to compress the data, then store in S3, which just became a business model of cost savings on their log data. And so that company is called ChaosSearch. I thought it was fascinating. But then as you're talking and you're discussing about how you're building this DNA storage concept, I'm curious.
(Joel Beasley at 00:23:02) You've got this awesome technology, just like ChaosSearch has had their awesome innovation and technology. But where's the market-ready place that you're plugging it in?
(David at 00:23:12) Well, so this is a deep question in the sense that any innovation is subject to certain fundamental barriers to acceptance. One is a natural human inclination to simply go along with the way we've always done things. Right? And it's a very natural human kind of response because, listen, if you spent twenty years, thirty years, whatever the length of time is, honing your craft on a particular technology, building your career around it, and then suddenly somebody comes along and says, you know, all that stuff you knew about electronics and physics, it's all going to be supplanted by biology, which you know nothing about. And you're going to have to start over from scratch because the ideas don't apply directly. That's a huge threat.
(David at 00:24:03) Right? And so one of the problems any innovative company has to deal with is how to produce technology in a way that's not scary. And by not scary, I mean something that somebody who is really dramatically vested in the state of the art, the way it exists before the innovation, can look at it and say, yes, I can work with it. I can use it, and it won't threaten me from a career perspective or what have you. This is not an insignificant issue.
(David at 00:24:33) And one has to think about this very carefully from a human perspective to find ways to overcome the intrinsic reluctance of people to embrace new technologies that are out of their domain of understanding, if you will. So that's point number one. Point number two, when we look at customers, we have to be very careful in terms of not being glib about how we define markets. And what I mean by that is the following. It's easy for somebody to come along and say, well, the market for high-performance computing is this, or the market for electric cars is this, or the market for bananas is this, whatever the case might be.
(David at 00:25:15) But it turns out the more you dig down into that sort of amorphous categorization of market, you find out markets are capable of being subsetted quite dramatically. So let's go back to the oil and gas industry for a moment. I won't even begin to guess how many oil and gas companies there are in the world, but there's one thing that's true. There are leaders and there are followers in that industry based on capability, attitude towards risk, financial strength, etcetera, etcetera. So if you look at ExxonMobil or British Petroleum, or one of these major vertically-integrated companies, their attitude towards technology is going to be a lot different than some small national company that doesn't have nearly the same amount of resources.
(David at 00:26:06) Right? And these companies are highly sophisticated. The ones I mentioned, and there are others like them as well—I don't mean to exclude anybody, but we're not gonna talk about every company in every industry. They have deep skills.
(David at 00:26:18) They have knowledge. They understand risk. They have the wherewithal to look at emerging technologies, which they do all the time. You know, as a person who's been in the computing industry for a long time, it was a rare case where I went into a company like that and surprised somebody with a new technology. They had already been thinking about it, looking at it and so on.
(David at 00:26:41) So you have to divide markets into those that are willing to take risks—a couple of examples like that, because they have the wherewithal to accept risk—and those that are not, so focused on leaders and not followers. And then even within those categories, you can auger down more deeply into where they are in their current planning, other constraints that might apply. There might be, for example, political constraints that affect some of the national companies. It might just be behavioral issues that come along. So the question of where we would deploy our technology is actually an exercise in augering deeply into nominal markets where you think it might apply and deriving a refined perspective of requirements that get progressively closer and closer to our ability to accommodate those requirements. So I would not say to you, for example, that the prospect of encoding DNA is something that will eliminate the tape industry. That will never happen. People are used to it.
(David at 00:27:48) It's good technology. It's dense. It's got a lot of interesting prospects. But there are companies we're working with now, commercial enterprises, in which we've engaged in collaborative efforts that are looking beyond tape, not as a displacement, but as an augmentation to their tape infrastructure. Just like people look at quantum, not as a displacement for Von Neumann architectures, but as an augmentation to it.
(David at 00:28:12) Or they look at AI, not as a displacement for everything you do in software, but as an augmentation to it. So it's this idea of where you can take our technology, augment existing approaches, and provide a client with enhanced benefit. So, example. We—and I'll make comments both in terms of storage and compute. With respect to storage, you now hear the emergence of a new kind of vernacular in the storage community. And people talk about "write once, read never."
(David at 00:28:46) Right? And you say, well, how does that make sense? Well, we already have a hierarchy of archival approaches and so on. But then there's always this idea that, well, what if everything else goes wrong? Right?
(David at 00:28:59) We've got to have a real deep backup to make sure that we in the face of disaster recovery can come back online again. So you do what? You encode data in potentially a different medium like DNA, and you sock it away somewhere and you throw away the key for maybe ten, twenty, a hundred years, five hundred years, whatever the case might be. Because as we know from Jurassic Park kinds of representations—but more real-world—you've seen the recovery of DNA from insects in amber and dinosaurs and people frozen in the Alps for twenty thousand years, etcetera. You can keep this stuff around for a long time if you take the proper precaution.
(David at 00:29:40) And the other nice thing is DNA is not gonna change over time in the sense that, you know, I can read a DNA molecule that was produced a million years ago if it's intact. You can't read a tape cartridge that was produced ten years ago with modern hardware. You might be missing device drivers, operating system support, this and that. You might have the cartridge. The media may not have decayed, but your chance of reading is close to zero.
(David at 00:30:09) So people come along and say, well, how do I have an immutable technology that can be preserved forever with effectively no energy footprint that under dire circumstances, I can kind of reconstitute and rebuild my databases or what have you? DNA fills a role like that. Right? Not the only role, but that's an accessible example I think most people would understand. Film archive.
(David at 00:30:33) You know, there are films that were produced in the 1890s. Is that media solid? No. That's why you have people working on transforming it to new media to try to preserve it and so on and so forth. But the electronic media that they produce it on also becomes obsolete, as I alluded to, pretty quickly.
(David at 00:30:54) So you get yourself into this never-ending cycle of forever upgrading the storage of what was produced maybe a hundred years ago, as you try to manage this obsolescence of technology. We're sort of immune to that concept. We can put stuff into DNA in a thousand years from now, because of the genomics industry and everything else. There will be machines that'll read that molecule quite well. And if we give you the recipe for how to decode it, no problem.
(David at 00:31:23) Right? So that's an example in the storage side. It's not the only example, but it's sort of an insightful example that will help people understand where this is going. On the compute side, there are a variety of different things that we can do depending on how you encode your data and how you operate on it. So I, at the outset, said that we encode data in kind of a tree structure.
(David at 00:31:47) Tree structures have implicitly attached to them this notion of branching. It's obvious. Otherwise, it wouldn't be a tree structure. But from a computational perspective, you can take the idea of branching through data and actually do a lot of very interesting operations on top of the data that's been encoded in DNA. So the nature and the way by which we've encoded the data actually opens it up to be operated on in the same kind of fashion.
(David at 00:32:18) Perhaps more accessible is the concept of search, right? So I can invoke chemical processes to expose data to massive parallel search. Because if I take a single molecule of DNA and you said to me, that's nice, but I like a trillion of those. I can do that for you really quickly. I can create as many as you want cheaply, quickly. No problem.
(David at 00:32:41) Harder to do in electronic media. But I can take a trillion molecules and they can all be different and they can represent search targets in a database. Right? Also encoded in DNA. And I don't have to leave the DNA world.
(David at 00:32:58) I can actually infuse my file of encoded DNA and attack it also with DNA molecules that are meant to, let's say, find an anomalous event or something like that, and find that really quickly in a fixed amount of time. What I mean by fixed amount of time is it's independent of how big that search file is, which is not the case in electronic media. If I say find a piece of data in a megabyte file, you say fine. If I say find a piece of data in an exabyte file, you're gonna say that's gonna take longer. And if I say find a piece of data in a yottabyte file, you're gonna say, not in your lifetime, you know, because of the way search is done electronically.
(David at 00:33:44) In chemistry, that's different. If I've got a file of data encoded in DNA and I wanna make it a million times bigger, you know, a billion times bigger, a trillion times bigger, I'm gonna find that missing or anomalous piece of data in the same amount of time. It's effectively gonna be unaffected by the volume of data we're searching. And part of the reason for that is, what does my file look like? It's just a bunch of DNA in a liquid.
(David at 00:34:10) Right? It's not linearly structured on tape. It's not structured in some, well, let's say, more complicated way in NVMe or something like that. I don't have to search a physical device. I just need to find a molecule in a pool of molecules.
(David at 00:34:30) And there are a lot of ways I can do that in a fixed amount of time.
(Joel Beasley at 00:34:34) You're blowing my mind right now, David. All I can think about as you're talking is that we're giant computers, like, as people. Right? Because our chemicals in our body, the way they work, it's like a bunch of autonomous processes that are just firing off. If, you know, you get attacked by a virus, you don't consciously respond. Your body's system responds and instantly finds all the virus and attaches to it and then edits code, and then your body can then look for and defend that virus in the future.
(Joel Beasley at 00:35:04) And, you know, I had talked to, I think her name was Darlene, but she was this awesome CTO at a large company that did organic biotech type stuff. And she was telling me, I asked her this question, I was like, when will we have wings? And she's like, we're a publicly traded company. Because I was wondering when we're gonna have wings. Because the way that they were describing the advancements in technology, it's not sexy. It's not popular. It's not in the magazines or on the TVs.
(Joel Beasley at 00:35:30) But she was describing the revolution happening in biological computing and the things that they're able to build and manipulate and understand, both through modeling it and then physically interacting with it, is like the computing revolution was. It's the next thing.
(David at 00:35:56) Yeah. There's a tremendous universe of possibilities here. We deal in synthetic biology. So we build these DNA molecules essentially from scratch. So they're not harvested from animal, plant, human or anything like that.
(David at 00:36:14) And of course, we construct them in such a way that they're non-biologically active as well. And what we try to do is leverage the behavior of DNA molecules and the way they're structured to help us solve these sort of data/computational problems. And it's impossible to say, but one would expect there to be tremendous amount of crossover from, for example, the work you just quoted with that publicly traded company into what we're doing and what we're doing into other companies as well, as people begin to discover alternative ways to make use of what we're doing in completely orthogonal kinds of domains, if you will. So if I have a machine that can string together building blocks of DNA to create a longer DNA molecule, what else can someone use that machine for? Well, maybe they wanna create biologically active molecules as a therapy for some disease or something like that.
(David at 00:37:16) And it's isolating the technology that we've invented, developed, et cetera, and maybe applying in a different domain as well. So it's a dramatically rich area right now because you always see these Cambrian-like explosions of innovation. I use the Cambrian period as sort of the metaphor for innovation explosion, because that was when life on earth just sort of blossomed out. It always seems like there's this tipping point, if you will, where you reach things because, well, you know, we talked earlier about the ascendancy of accelerators. A tipping point was reached with respect to the benefits that would accrue from the deployment of accelerators versus the singular reliance on Moore's Law, and suddenly a whole new world opened up.
(David at 00:38:07) That kind of stuff is happening in molecular biology. Now things are getting cheaper. Things are getting smaller. Things are getting faster. Our chemistry operates at the picolitre scale, right?
(David at 00:38:19) You know, nano scale, you know, one billionth. Pico scale, that's 10 to the minus nine. Pico is 10 to the minus 12. So that's a thousand times smaller than what people talk about when they talk about nanotechnology. And then there's this talk about getting to femto level chemistry, which is 10 to the minus 15. So a million times smaller than nanotechnology are the domains that the biological industry is working towards.
(David at 00:38:51) Right? Because there's a tremendous advantage in just about anything you do in technology with making things smaller. Cost, obviously. Speed. And then there's sometimes new properties that emerge when you go that small. I would also say one other thing, and I made this comment about technologies sort of migrating through orthogonal or adjacent spaces. There's another kind of migration as well, and that is this coalescence of technology from different spaces into a new thing completely. So why not contemplate the merging of electronics with biology in a single device? And it's an area that we're focused on now.
(David at 00:39:34) Frankly, we're looking at the application of microfluidics and putting everything on a chip. So now you have a silicon substrate, if you will, and the chemistry that we're talking about operating within the confines of that electronic substrate, using microfluidic principles and so on to miniaturize the kind of chemistry going on to make it a closed system so that you don't have to worry about hiring chemists to encode your data into DNA and things like that. And it becomes more of a natural deployment of technology. I would expect this to be a growing trend in the computing industry over the next decade, two decades, et cetera, where you'll see these pockets of innovation pop up. You know, it was neuromorphic or quantum or biological or whatever. And over the course of time, these things will start to be merged together, amalgamated together in interesting ways, trying to leverage the best principles from each of those domains to create something completely new and advantageous to the problems that people are trying to solve.
(David at 00:40:42) So, yeah, there's a really rich atmosphere percolating through the industry right now that's generating just this amazing set of ideas that, you know, you just look around and say, I wonder how I could leverage that. And that's exactly where we want to be. So we've reached a tipping point with some technologies. We've seen the manifestation of constraints in other technologies. And all these things conspire to really percolate and give rise to a tremendous acceleration of innovation.
(David at 00:41:16) And I think that's what we'll see.
(Joel Beasley at 00:41:18) Yeah. I was talking with Robert Sutor, who is quantum computing at IBM, I believe. I don't know, Jacob could correct me. So he had helped correct me and give me some understanding of the quantum space. One thing that he shared with me was I went into it with my programming history, right, as a software developer. And I was like, you know, when am I going to be able to do this but run it on a quantum workload? And he explained to me that these computers are useful for certain problems and then problems that don't even exist today. But one of the areas I saw emerging in my own research was its use in biology. Its potential use in biology because of the models, the ability to create more rich, complex models.
(Joel Beasley at 00:42:06) And it kind of seems like quantum computing would be useful in the biology world. And you were just describing how these technologies would emerge and then they would all kind of come together. But I feel like it's something that we can't see, but they're closely related.
(David at 00:42:25) Well, the application of quantum in the biological world, and maybe better said the chemical world, because at its core, all biology is kind of chemistry—it's a collection of isolated application domains in that area. And the way to think about it is people have been using computing in chemistry problems for a long, long time, but there are some problems where the conventional application of Von Neumann approaches simply don't give you good results at all. And people have calculated that a quantum approach would be much, much better. It echoes again what I said earlier about the right tool for the right problem.
(David at 00:43:08) So you don't want to just categorically say, well, we're going to apply quantum to all problems in chemistry. That makes no sense. So we have to look at these application domains and really refine them quite carefully. It's just like in the DNA examples I gave about search. Well, we're not going to displace Google overnight with respect to search, but there are categories of problems within the search domain, discovery of anomalies for fraud, discovery of rare events in maybe astronomy or high energy physics, something like that, where you're not attacking that problem with Google kinds of approaches, but something like the chemistry embodied in DNA would be appropriate to do something like that.
(David at 00:43:51) So not only do we have to refine markets very precisely, but even within those markets, we have to look at the application domains and really pick the right ones for the tools that we have available to us.
(Joel Beasley at 00:44:07) Now, if I wanted to take a picture, right, like took a selfie of us right now or something, we take that picture, we put it into that robotic arm machine that I saw in the video. It encodes it into this DNA, which if I remember correctly, kind of looked like a white powder, at least that was the visualization in like a little tube. Okay. So now we have it there and I know there's already machines to decode the data. Because I remember in high school there was this big push to encode the human genome and decode it or whatever.
(Joel Beasley at 00:44:36) So I know that there's machines to decode it to some degree. But what is it like? It's an algorithm. So if I have the key, I can decode it and then get my data back. Like, can you walk me through the—I have a picture, I have a piece of data. I'm putting into it. And how do I get it back out?
(David at 00:44:51) Right. So there are gene sequencers that have been inspired by the Human Genome effort and biology since then, of course, from companies like Illumina, Oxford Nanopore. There are a few others as well. We actually partnered with Oxford Nanopore to decode these DNA molecules. And what it does is it'll examine the molecule that you present it, and it will give you a readout of what the base pairs are, et cetera, that constitute the makeup of that molecule.
(David at 00:45:25) And then using algorithms embedded in software, it'll convert that under the schemes that we've provided back into the stream of zeros and ones to characterize the input that we provided. Those machines work in different ways, different companies, different technologies, and they put an emphasis on accuracy, but not so much on speed. So one of the obstacles that we have to work on in our industry, of course, is to get those machines to run a lot faster. In other words, we can write data a lot faster than we can read it. Now that's not a problem in the example I gave about write once, read never, because if you're trying to reread that data that you encoded into DNA, circumstances happen where you're going to be patient and you're going to be careful about reading that data back out.
(David at 00:46:18) But we'd like to see innovations occur in that part of the market that would push speed along as well. And I think as more and more companies embrace this notion of using DNA as a data storage mechanism, you'll see those companies—i.e., the companies that produce these reading devices—put a lot more emphasis on trying to make them go faster, et cetera. So it's a reversal of the process that we provided. It takes an input from the, in this case, the desiccated DNA molecules that you refer to as a white powder. We, by the way, could also keep it in liquid form and present to these devices, and they will come back and produce output that you can reread, knowing your encoding scheme, into the stream of zeros and ones that you start with at the beginning, with high accuracy, by the way.
(Joel Beasley at 00:47:11) Is searching on it possible? Like you mentioned earlier that you can make things bigger, you can manipulate it after—like, once you have data in a chemistry format, you can then, you know, enlarge it, duplicate it, search it. Is there any real world examples of that happening today?
(David at 00:47:32) Well, we've done it in our own tests. So for example, in the example I mentioned earlier about working with a West Coast media company, we encoded parts of a feature film and then decoded it out the back end. So it's not a problem. And by the way, it's a logical question to ask. Same question you would ask to people in the quantum area, which is how do you handle error recovery?
(David at 00:48:05) Right? So there have been decades of experience put into electronic media with error recovery schemes. So in our case, it's actually pretty straightforward. We take data that's initially presented to us in digital form and we'll apply error correction codes and things like that to the data. And all that means is we're producing more DNA than just the raw data per se.
(David at 00:48:29) So there are error correction codes and so on. It's just more molecules we're creating that represents the error correction. And by virtue of doing that and understanding how to do that, the output that we produce, we can characterize the quality of the output in terms of frequency to bit errors and so on, just as you do in the electronics industry. And we think we're quite comfortable thinking that we can get to tape level quality, which is like one bit error in every 10 to the 19 bits. So that's a pretty rare event in its own right.
(David at 00:49:04) But you do that by imposing error correction schemes on the chemistry of what you're doing. And we know how to do that today. We do that today as a matter of fact.
(Joel Beasley at 00:49:15) So when you're searching on—when you talk about searching on the data just to help me get a visualization in my mind, you know, I'd see a software interface and I run a search and I get a result back. But is yours more of, like, a chemistry set thing?
(David at 00:49:32) Right. So a couple of different ways. If you know what you're searching for, right, you can build a molecule, a DNA molecule that effectively incorporates the definitive elements that characterizes your search target. Okay? Now if that already exists in the DNA file that we've created, there are means through chemistry by which this new molecule that we've created can be used to help isolate the presence of that molecule in the existing database.
(David at 00:50:06) Right? So that's a matter of chemistry. The other case is you don't know what you're searching for. All you know is you're searching for something anomalous to what you would ordinarily expect. And so a little more involved and there's a little more chemistry involved for that.
(David at 00:50:22) But that's also a possibility for how you do that. We haven't done that yet. We know how to do that. It involves a little more chemistry that we have to develop and embed in the process that we have. But these are the two kinds of fundamental issues.
(David at 00:50:38) If you know what you're searching for, it's fine. You can represent that as a vector of attributes. You can put that into a DNA molecule. You can submit that into your file of DNA, if you will—file here being a beaker of DNA—and you can isolate the fact of whether that molecule is present or not present. How do you do that?
(David at 00:51:02) Well, you can take the molecule you've created and you create a million or a billion copies of it. Right? And then you can expose that to the file that you have and you can put in markers of fluoresce, the molecules that you're looking for. And if they all connect, you're fine, if you will. So, yeah, there are ways to do that.
(Joel Beasley at 00:51:24) That is so cool. See, it's starting to click for me. So you'd build like, if it were, let's say, a letter written by me. Right? So you got a letter. It's coded. It's in this jar or beaker, and you want to find my name, which is down in the signature line. So you would create a molecule that's looking for the text "Joel Beasley" essentially. And then you would just amplify that, like clone it out a lot, and then put it into the data source, like that beaker, and you would have some detection and then you would see the glow and you're like, there's the glow. Right? That's where Joel is inside of this.
(Joel Beasley at 00:52:04) And there may even be multiple Joels because you say you make more data than you need sometimes. Right? So it can be glowing in multiple spots. That's so interesting. That is fascinating. And then it's like the—so all of this is happening in the lab. So you're proving like these first principles on it, like really low level stuff because later it'll get to the point where you have like a computer controlling interface system that will then go run that process somewhere. Right?
(David at 00:52:37) Yeah. Automation is quite critical.
(Joel Beasley at 00:52:40) Yeah.
(David at 00:52:40) So our process right now is heavily automated on the write side, but it's not so automated on the read side because, again, we're leveraging other technologies from other companies. And we have this intermediate step where we have to essentially prepare the output from the writing process as input to the reading process. So that for us is still, you know, conventional chemistry done by conventional chemists. But everything we do here is a target for automation, miniaturization, acceleration, all those things. So when you have—the way I think about is the following way.
(David at 00:53:19) When you look at the emergence of innovation, and let's say data science as an example. And, you know, let's think about data science over the last five to ten years. As people grew comfortable with the idea of data science, what they discovered was that the cost of actually employing a data scientist got quite far out of hand because there's so few people trained that way. And the problem had a second dimension, which was those trained people pretty much congregated in urban areas. So if you're in New York or London or Paris or Moscow or wherever, you're okay. You could find somebody trained that way.
(David at 00:53:59) But if you happen to be in, oh, I don't know, a small town in Ecuador or, better yet, a small town in maybe Kansas, and you wanted to do something like this, well, you had no resource to call on. People weren't there, you know, training wasn't there, et cetera. And I think when there is a geographic dislocation of skill concentration, it begs for automation. And so what you've seen over the course of time is a lot of these things that were at the beginning very bespoke from a data science perspective are progressively becoming more and more automated. And that's a way to kind of diffuse the capabilities, the ideas, and so on, embodied in data science to as much of the population as possible.
(David at 00:54:48) You know, you don't want to dictate that innovation is only available to people who are located within the five boroughs in New York City. And if you go beyond that, you're out of luck. No, you automate. So automation is the correlate to innovation to promulgate and propagate technology as broadly as possible. It's another reason why we invented the machine that we did.
(David at 00:55:13) We wanted to begin to assess what was automatable, anticipating a future where when we became commercially viable, we could deploy it anywhere. You know, it wasn't, "Oh, jeez. We'd love to do business with you, but you're not in LA. You're out of luck." Right? No, we don't care. You're in Chula Vista, California? No problem. Or you're in Bend, Oregon or someplace like that? No problem.
(David at 00:55:40) So automation is a critical adjunct to the whole idea of innovation. And it's something that the composition of staff at CATALOG embraces because we've got expertise in all disciplines: data science, chemistry, biology, engineering, computer science, all under one roof. It's a small set of people, but we're fairly eclectic in our backgrounds. And we can bring that all together. And it's this diversity of point of view from domain-specific expertise that also is another critical ingredient to stimulate innovation.
(David at 00:56:20) Right? So we talked about constraint. We talked about tipping points of technologies when they would become viable. But I think it's also the creation of a corporate enterprise that has sufficient diversity of thought, measured in terms of diversity of background and expertise, that really becomes a catalyst to take advantage of those other phenomena and figure a way that's novel in the context of the constraints or other impediments that are in place. So we go out of our way to really look at diversity of people and also capacity to grow intellectually, because what you know today and the reason why we hire you today may not be immediately useful to us, but maybe two years down the road it will be.
(David at 00:57:07) By the way, the demonstrated intellectual capacity that we observe today means that you can probably learn everything that we're doing today, so we're comfortable teaching you. But at the end, you know, we want people who are quite comfortable bouncing between software and algorithms and engineering and chemistry and biology. And it's actually one of our competitive strengths.
(Joel Beasley at 00:57:29) I love it. Sometimes when I get real excited about this type of stuff, my brain overclocks and I'm just like, all right, I've got to keep this simple.
(David at 00:57:40) Water-cooled. You can handle that.
(Joel Beasley at 00:57:42) I know, right? No, but there's so much I want to say. I really like the way you described the type of people over there. I feel as if we would get along really well.
(Joel Beasley at 00:57:55) I like exploring different ideas. That's the reason why the podcast works so well for me as my current career position, because I get to go from talking about the farthest ends of one spectrum to another. And I love when I get to meet people like you because you're very intelligent and you're a great talker, and that makes my job really easy. So thank you. But man, we did it.
(Joel Beasley at 00:58:23) We made a podcast. How do you feel?
(David at 00:58:25) Oh, it's fine. That was a good conversation. I think one of the things we'd like to explore at an appropriate time is to examine this from an international perspective and to understand what's going on worldwide. You know, we talk as if we know everything here, but of course there are initiatives going on all over the world and it's important to understand those. That's point number one.
(David at 00:58:51) Point number two, and you may have gotten this from Bob Sutor as well in your conversation on quantum. IBM, of course, was very prominent in terms of articulating this idea of quantum volume, which maybe you talked about in your call with him. But the notion of trying to escape single figures of merit to describe a new technology. So for example, it's easy enough for me to say, "terabit a day." And you latch onto that and say, well, boy, that's a figure of merit.
(David at 00:59:21) Or in the quantum space, it would be, you know, how many qubits do you have? And those single figures of merit lack nuance, and they can become very dangerous. In the supercomputing space, where it all came down to the magnitude of your LINPACK run, the consequence of that over fifteen or twenty years was people started building machines to solve a benchmark as opposed to building machines to solve real-world problems. So you have these oddball designs that did really well on LINPACK and maybe not so well on other things. So one of the themes that CATALOG wants to really engage on in the coming year is this notion of nuance with respect to how one should evaluate progress in this domain of using DNA for computing and storage, and to underscore the themes that are embodied in it.
(David at 01:00:19) We touched on some of them today, by the way. We talked about error correction, but that's a precursor to the notion of reliability. There are other themes. There's consistency. There's energy consumption.
(David at 01:00:32) There's predictability. All the kinds of things that people take for granted with respect to data storage. You know, the example I give to many people is this. So here is a cell phone. Right?
(David at 01:00:47) You probably have one, right? And you probably back it up, right? Because, well, you know, maybe you lose it. Maybe, you know, something bad happens, or maybe you just buy a new cell phone. It's pretty easy just to download an image off iCloud or something like that for your new phone. You're up and running in ten minutes. Here's a question for you. How many times have you investigated the backup that you have in the cloud for accuracy of your phone numbers?
(Joel Beasley at 01:01:15) Well, I don't believe that there's any tools for me to really do it without just doing it. But it's very few. When I transition from one device to another, it works or it doesn't, you know?
(David at 01:01:30) Right. And so by and large, you trust the process. You trust the technology. And you probably also calibrate it by virtue of the fact that, well, you know, if something goes wrong, it won't be the end of the world. I've got another copy somewhere else or I've got phone numbers written down or whatever the case might be.
(David at 01:01:49) But suppose when we were copying your phone we said, you know, I may need to recover this in a thousand years. You know? Would you trust me to say that what you just stored will be there in a thousand years? So this issue of trust with respect to the invocation of process has been something that has been kind of burned into our minds to not worry about. Right?
(David at 01:02:16) So I'm copying a file on my computer. If I look over there and I see the file name, I'm assuming that all the contents I asked to be copied are there, right? I see the file name. I'm probably not even running a checksum against it or something like that. So there's a confidence that accrues to the deployment of process in the electronics industry that needs to be manifest in what I'm doing. Because my guess is man in the street is going to say, wait, what? My data is now in that test tube? You know, how do I know that my data is really there?
(David at 01:02:51) Prove that to me. Well, okay. Do you want me to read it all back to you and put it in electronic form? Is that what it takes? So there are issues like that that are floating around in the perimeter that haven't yet been explored deeply.
(David at 01:03:04) Because when I say to you I've written a terabit of data a day, you're hearing the words, but your mind is sort of filtering what I'm saying through all the filters that have built up over decades in our industry to say, well, I know what that means. It means the data is written and I can read it back. And the question that should be asked is, is that what that means? Right? So there are things like that that I think during the course of the coming year we want to have conversations about, because it's too easy to be glib in this space.
(David at 01:03:40) Right? It's too easy for people to make declarations. It's just like in the quantum space when D-Wave says, I got 3,000 qubits. What are you guys? You're at 63, you're at 75, or whatever.
(David at 01:03:51) So therefore, I'm better than you. Guess what? Your 3,000 qubits can't do much because you don't have provisions for networking or error correction or any of that. So qubits is a figure of merit. It's kind of faulty, which is why I kind of like the quantum volume approach.
(David at 01:04:07) It brings in nuance by virtue of saying value is represented in a multidimensional way. It's not convenient for the lay press. You know, they just want to say, oh, here's the Top 500. Here's number one. We don't have to worry about anything else.
(David at 01:04:20) It's winners and losers. But more technically, people, especially consumers of technology, should be asking different kinds of questions, right, and not assume because I make statements in the DNA world that that's easily translatable to the electronic experience that everybody's had. I live in the DNA world in a substantially more stochastic environment than people do in the digital world, right? Because I'm using chemistry.
(David at 01:04:50) Are all the molecules there? I'm not quite sure. Right? That kind of stuff. So there's a different set of conversations that need to be had, and we'd probably like to talk to you more about that sometime next year.
(David at 01:05:04) We have to manage the hype cycle. That's a fundamental issue. Our ambition is to be commercial in a legitimate way and not simply run the hype cycle up and get acquired by somebody for an overblown—
(Joel Beasley at 01:05:16) Well, I read all about you guys. It was pretty clear in the text. I have been following this company, I think, three years now. Like, I really have. I've been super—you can ask the founders. I forget his name.
(Joel Beasley at 01:05:29) But I messaged him on LinkedIn a couple years ago and I was like, what you're doing is the future. This is so amazing.
(David at 01:05:34) Right. Right. And it's important for us to be really clear on this point. And I'll give you one more example to close. If you look at the fastest supercomputers in the world and everybody beats their chest and says, I've got the fastest supercomputer in the world—
(David at 01:05:52) And right now it's Japan followed by the U.S. And there are rumors that there are even bigger systems in China. Doesn't matter. The question nobody ever asks is, how efficient are these machines? Because I can tell you the efficiency of those machines lurk around 9%.
(David at 01:06:08) So the rough equivalent in buying an automobile would be something like this. Wow. I want to buy your car because you have a speedometer that goes up to 350 miles an hour. Now I know I can never run it that fast, and I can only run it—I live in a small town in Connecticut with dirt roads and this and that. I can probably only go 35 miles an hour, so 10%.
(David at 01:06:34) But, boy, I really want to buy your car because it says it can go as fast as 350. That is the mindset that drives the high end of supercomputing and has for the last twenty years. It's, you know, that extreme sticker on the window as opposed to what you're really going to do with it in real life. And in our case, it's the same thing. It's not a matter of how fast I can build a machine to write data into DNA.
(David at 01:06:59) It's how fast I can write data into DNA that's reliable, repeatable, consistent. All those kinds of things are important. I mean, any fool can spin up a machine that passes things through it quickly. But what comes out of the back end—is it at all useful? And so that's what I mean about managing the hype cycle and making sure that what we've done is something that's realistic in the context of what clients are expecting to see.
(Joel Beasley at 01:07:27) I love that. That's a good goal. I understand that, you know, you guys are putting out this unbelievably futuristic technology and people are going to want to try to condense it and put it in a corner and figure out what the hype metric is—we can call it a hype metric, or you said something more eloquent earlier. But they're going to try to figure out what that is and then watch it progress, right? And so controlling the conversation and helping people understand the story of what to focus on and how to look at the technology, educating them, awareness around that—I think there's a lot to be done there.
(Joel Beasley at 01:08:05) Yes. A hundred percent.
(David at 01:08:07) Absolutely. Well, listen, happy New Year, and thank you for this.
(Joel Beasley at 01:08:10) Happy New Year. You have a fantastic day. Thank you so much for listening. And if you found this episode useful, please share it with a friend or colleague that you think would get value from it. And if you have topics that you'd like to hear discussed on the podcast, either add me on LinkedIn or send me an email: [email protected].
(Joel Beasley at 01:08:32) Every time I get an email or LinkedIn message, it absolutely makes my day and inspires me to keep going.