Episode 710 ·

The Future of AI and Search with Sean Mullaney, CTO at Algolia

Today we’re talking to Sean Mullaney, CTO at Algolia. We discuss Sean’s entrepreneurial journey that brought him to leadership, why internet search is largely outdated, and how AI is transforming search and bringing us into the future.

All of this right here, right now, on the Modern CTO Podcast! 

For more about Algolia, check out their website: https://www.algolia.com/

Have feedback about the show? Let us know here

Produced by ProSeries Media.

About Sean Mullaney:

Sean Mullaney is Algolia Chief Technology Officer (CTO). He is responsible for scaling Algolia’s best-in-class engineering organizations, developing Algolia’s AI powered Search and Discovery technologies, and growing our API-first solutions globally.

He  joined Algolia from Stripe, where he most recently acted as the Chief Information Officer in Europe and led a global engineering organization overseeing the development and operations of 40+ local payment methods across Europe, Asia, and the Americas that processed hundreds of billions of dollars. 

Prior to his time at Stripe, Mullaney oversaw the development of AI powered discovery experiences, including search, recommendations, personalisation and browse as the VP of Engineering at Zalando, the largest eCommerce fashion retailer in Europe with nearly 50M active customers.  Mullaney also spent more than seven years at Google heading various innovation teams, including three years at the Google research labs in Silicon Valley.   

In addition to serving as CTO of Algolia, Mullaney serves as a Board Member for Manna Drone Delivery, a Venture Advisory Board Member for Elkstone Ventures, is a member of the Market Advisory Group at the European Central Bank advising on the design of the Digital Euro and an active startup advisor and a Sequoia Scout angel investor.   

Mullaney earned a bachelor’s degree in computer science from the University of Cambridge.

About Algolia:

Algolia is the world’s only end-to-end AI search and discovery platform. The company delivers a combination of market-leading keyword and natural language processing via vector search – all uniquely packaged on a single API and supported by hyperscale indexing. The company’s vision, mission and purpose is ‘Powering Discovery. Algolia achieves its vision by enabling more than 17,000 customers including Under Armour, Stripe, Petsmart, Walgreens, Proctor & Gamble, Sony, NBC (Universal), British Telecom and Citibank, to build blazing fast and relevant search and discovery experiences for their in-app users and/or online visitors (using any web, mobile or voice device) – by surfacing the desired content instantly and at scale. Algolia powers 1.75 Trillion search requests a year or more than 30 Billion a week – that’s four times more than the combined volume of Bing, Yahoo, DuckDuckGo, Baidu and Yandex). Algolia is used by one in six online users and more than 5 million developers a month.

At the core of Algolia’s platform is hyper-scalability. During the 2022 Black Friday/Cyber Monday weekend, Algolia handled more than 100,000 search interactions a second without a single failure. One customer, Gymshark, experienced 15,000 queries per second with zero interruption to service. According to a Forrester Consulting study, customers gain a 382% Return on Investment with less than a 6 month payback period when deploying the Algolia Search and Discovery platform.  Founded in 2012 in Paris, the company grew and subsequently moved its headquarters to San Francisco to expand its operations and penetrate the North American market.

Transcript

(Intro Narrator at 00:00:00) Today, we're talking to Sean from Algolia about the future of AI and search. You're listening to Joel Beasley, Modern CTO.

(Joel Beasley at 00:00:13) Dude, this is gonna be a blast. Where are you calling in from, by the way?

(Sean at 00:00:17) I'm in Ireland. I'm in Dublin.

(Joel Beasley at 00:00:19) How did you get to Ireland?

(Sean at 00:00:21) Well, pretty short story. I fell in love with an Irish girl.

(Joel Beasley at 00:00:26) Oh, that's easy enough. Yeah.

(Sean at 00:00:28) Yeah. Exactly. But I'm also an Irish American, so they call me a plastic Paddy here.

(Joel Beasley at 00:00:33) Okay. So were you born in America?

(Sean at 00:00:36) I was born in Virginia, but to a very Irish American household, very proudly Irish. So when I married an Irish girl, it was pretty easy and fun to move to Ireland. Such a great country.

(Joel Beasley at 00:00:47) Oh, that's pretty cool. So I was hoping we could start out by you just telling me a little bit about, you know, what Algolia does and what you do there.

(Sean at 00:00:55) Yeah. So a little introduction. I'm the CTO of Algolia. Algolia is the leader in search as a service. So we've been around for about ten years.

(Sean at 00:01:04) We started off very much as developer first, API first search platform. And when you think about the Internet, there's some foundational pieces of infrastructure that every single site or application needs. Search is definitely a foundational part of that Internet infrastructure. And so if you're starting a project or you've got a company, you can allow us to take care of all of your search needs. We're actually, we've grown pretty large. We're almost 2 trillion requests a year on the platform, which makes us the second biggest search engine in the world behind Google. So one out of six Internet users will interact with us every month on one of the websites. As I said, we're infrastructure, so we sit behind the scenes. So you wouldn't see our name or our logo anywhere. But we power 17,000 different websites.

(Sean at 00:01:47) A lot of the e-commerce infrastructure is powered by us as well. And we're going through, you know, the same type of AI transformation that the rest of the industry is. And obviously, AI is incredibly powerful when it's applied to search. And we're starting to see some of the first really incredible applications of this new generative AI applied to our search product at the moment.

(Joel Beasley at 00:02:10) So would a competitor be Elasticsearch, something like that?

(Sean at 00:02:14) Yeah. So Elastic is an open source product that you can deploy yourself. So if you're a developer and you want to run your own AWS, you can download Elastic and run it. At my previous company, Zalando, which is one of Europe's biggest fashion marketplaces, we built and ran our own Elastic cluster.

(Sean at 00:02:31) But it takes a lot of expertise. It's pretty complex to be able to scale it up and run it yourself, to get all the configurations right. So if you're an extremely advanced user and you actually want to have, I think we had a team of about 20 engineers running our search engine at Zalando. That's definitely one way to go. Algolia is very much more like Stripe, you know, where it's like seven lines of code. You install us. It's super simple to get set up, very powerful. Lots of great dashboards and UXs in the background. So business users and non-tech users can also configure and optimize it.

(Joel Beasley at 00:03:04) That is very cool. Yeah. Because, I said, my background was software engineering, and I haven't done it in about four years since the podcast took off. So I'm not in it every day doing it. But that type of technology, so that you don't have to build search from scratch, was just an amazing advancement that we had in our industry. It was just absolutely beautiful.

(Sean at 00:03:27) Yeah. So before I joined Algolia, I actually ran the European payments engineering team for Stripe.

(Joel Beasley at 00:03:33) Okay.

(Sean at 00:03:33) And so I was really excited joining Algolia because we have very much that same type of developer first ethos, really, really easy to use product, taking all the complexity away from a complex, high performance product. And similarly, I remember when I first got started coding on Algolia, I had something up and running in ten minutes. It was incredible. A very powerful search engine with a great experience in the back end. So it really is the type of product for builders and developers who want to get up and running fast.

(Joel Beasley at 00:04:04) Did you know David Singleton over at Stripe?

(Sean at 00:04:08) I did know David. Yeah. David was a phenomenal CTO, very inspiring, and also an Irishman just like the founders of Stripe. So it was very exciting. Stripe's European headquarters were here in Dublin, so we got to see a lot of the founders. And David, they used to come over and visit us a lot.

(Joel Beasley at 00:04:25) Yeah. I got to do an interview with him a few years back, and I was personally a huge fan of Stripe because for me, in my experience, that was the first time I had ever seen the developers have such a huge influence on the business decision of something like payments.

(Sean at 00:04:41) Yeah. And, you know, Stripe and Algolia have a very similar ethos when it comes to building products. So both are very much a product-led growth mentality, where the easier the product is to use for a startup, the easier it is to onboard, to integrate, to get all the UX components, to have the dashboards right, these are things that typically you would find smaller, fast-moving companies value.

(Sean at 00:05:04) It turns out that very large enterprises value the same thing as well. And so, you know, when we were working with Amazon, they used to say, we actually prefer using Stripe because our engineers can actually finish the project way faster rather than to optimize for the 0.01% of cost that you might save going somewhere else. And so, you know, definitely something I believe in quite strongly, having great, easy to use products helps everyone.

(Joel Beasley at 00:05:27) So you founded a company before as well? Tell me about that.

(Sean at 00:05:31) Yes. Yeah. So when I left university, it was the height of the dot-com boom. And I was obsessed with being an entrepreneur. And so the first ten years of my career, we started three companies. One of them called Masabi is now one of the largest mobile train ticketing companies in the world. But back then, we were building it as a video game mobile video game company. And so I've been through that tough startup. After the dot-com bust, it was very hard to do a startup, particularly, we were doing it out of Europe where there wasn't a lot of accelerators or VC capital or interest in funding tech companies.

(Sean at 00:06:11) I call it my ten-year MBA because you just learn so much when you're hustling and on the ground. There's a certain benefit as well not having access to large amounts of capital because you really have to focus on the product and the customer very, very early. And you really have to do everything that you can to start generating revenue to bootstrap yourself. So, you know, very much early hustling, three bootstrap companies. And I do think it was the greatest part of my education.

(Joel Beasley at 00:06:41) I can say I fully agree. When it's your money and your paycheck and feeding your family that's on the line, you learn real fast.

(Sean at 00:06:47) Yeah.

(Joel Beasley at 00:06:48) How the market works.

(Sean at 00:06:49) Yeah. Absolutely.

(Joel Beasley at 00:06:51) So were you still at Masabi when they made the transition to ticketing?

(Sean at 00:06:57) No. They actually pivoted after I left. But it was, you know, really great to see the company, again, having to figure out the market and having to figure out what's gonna work and what doesn't work. And it was only once they pivoted into train ticketing that they managed to raise venture capital and start to scale up the business. So it took a while. I think we did mobile gaming for a while, but it was about a decade too early before the iPhone and the App Store. And then I think for a while, worked on mobile casinos and marketing and, you know, just hustling to try to find, you know, a product market fit somewhere.

(Joel Beasley at 00:07:32) And then, ultimately, you went from those startups to Stripe, or were there some things in between?

(Sean at 00:07:39) Yeah. So I actually then went to Google. And it was funny. I said to my wife, I'm like, listen, you know, if this startup thing doesn't work out, my backup plan is to go and join Google. And she was kind of laughing going, well, that's some backup plan.

(Sean at 00:07:52) But it turns out that Google at the time also really valued people who were entrepreneurial and had hustle and knew how to build things. So I remember this is 2011. My first couple years at Google were, you know, the place was growing so fast. Everything was on fire and chaotic. And so for someone who likes to build things and is a little entrepreneurial, there was no shortage of problems to solve and places to add impact.

(Joel Beasley at 00:08:19) What type of work did you do there? Search?

(Sean at 00:08:21) No. So I actually started off in the Google office here in Ireland, which was part of the ad sales team. So I got involved in a lot of, actually, I was one of the few people with a computer science background in the office. So I was able to figure out a lot of the big data and how to tap into it to help enable our sales teams. So I kind of became the data nerd, and I started our big data team inside the business organization.

(Sean at 00:08:47) But then I quickly found out that when you get a reputation for being entrepreneurial, that there are some other crazy parts of Google that are even more entrepreneurial. And I got a call about a secret Google X research lab project and was drawn out to go and, we moved to San Francisco and got to work down in the Google X labs with the self-driving cars and Google Glass and the balloons. And I worked on with about three or four MIT graduates and their professor on building the world's first drone delivery business.

(Joel Beasley at 00:09:24) Is that still in existence, or was that proof of concept?

(Sean at 00:09:27) Yeah. So within a couple years, we had the first commercial business up and running on the University of Virginia's campus. And the Wing team, as they're called now, are one of the major Alphabet letters inside Google. So they graduated. And I think they just celebrated their 250,000th drone delivery for real customers. So they're operating in a few sites in the U.S., I think Australia, Finland, few other countries.

(Joel Beasley at 00:09:55) And what's their name?

(Sean at 00:09:57) It's called Google Wing.

(Joel Beasley at 00:09:59) I didn't realize, I've heard this before, right? You obviously, you've seen in pop culture and TV and cinema and stuff, this idea of drone delivery. But I thought it was something that wasn't making a lot of progress. But apparently, it is, and it's commercialized, and it's actually operating in certain places.

(Sean at 00:10:16) Yeah. Now we have to go learn about it.

(Joel Beasley at 00:10:17) Now we have to go learn about it.

(Sean at 00:10:18) Yeah. Absolutely. Yeah. And I actually when I moved back from the States, I bumped into a serial entrepreneur here in Ireland, a guy called Bobby Healy, and he had a business plan for starting a drone delivery company in Ireland. And so I kind of joined his board and have been helping this company called Manna. And they're now delivering 200 deliveries a day in a village just outside of Dublin in Ireland.

(Joel Beasley at 00:10:42) That is so neat. How do you feel as a technologist and a consumer? Things are progressing so fast in so many different verticals. It's my full-time job, Sean, to keep up with this stuff. Every day, I learn of a new company or a new, you know, oh, you know, last week. Oh, here's this new company I've never heard of. And they have 100,000 employees, and I'm like, this is crazy. This happens on a daily or weekly basis. It's just all moving so fast. What do you make of it?

(Sean at 00:11:08) Yeah. Well, you know, a lot of times in the early days, things move very fast from a demonstration perspective. So if you take self-driving cars or even the drone projects, you know, you can show progress very fast in the early days. And even ChatGPT right now, huge progress, very early days. But it actually takes a very long time to get past the safety and value levels that you need to commercially scale something out in mission critical systems.

(Sean at 00:11:37) So, you know, the Waymo cars are on the streets of San Francisco now, and it's been, I'm gonna say, twelve years since they first had their prototypes out. And they're still very limited. Again, with the drone delivery, you know, we've been at this for ten years now. And, you know, we're still only in a small number of cities, and, you know, they're still working through the safety levels. So, a lot of technology, you know, we overestimate how fast it goes in the short term, underestimate its impact in the long term.

(Joel Beasley at 00:12:06) So why did you go from entrepreneur to then working on the teams? Or can you compare and contrast the differences in being an entrepreneur or leading a team at a large company?

(Sean at 00:12:18) Yeah. Absolutely. So what I figured out, having done early stage startups, I was with Google during that really super fast growth phase. And, again, I joined Stripe during huge explosion in growth during COVID. But then, you know, at Google, I was there about seven years. And what I realized is that, you know, in terms of the S-curve, when Google became more mature and there were 100,000 employees, velocity became very slow. And there were a lot of politics, a lot of inertia. And so I just kind of learned that the sweet spot for me is finding a company that is past product market fit and ready for very fast growth.

(Sean at 00:12:34) And being able to put a lot of the structure and processes and things in place to enable a company that's growing very fast is actually the thing that I enjoy the most and I get the most value out of. I find with startups, they're extremely fun as well. But by definition, during the first couple years of a startup, you don't have a product with a lot of customers, and you take a lot of detours. Right? And you can end up in a place where, obviously, failure is pretty high when you're doing a startup.

(Sean at 00:13:20) And so, you know, as much as I love working with startups, advising them, I'm an angel investor as well. I don't know whether I'd go back and do a startup again because I find that phase can be a little bit frustrating when you're trying to figure out what the product is, and you don't have a lot of customers. I love the growth phase because you have a lot of customers that are getting a lot of value out of your product. And every day when you show up to the office, you know what you do is gonna have an impact on a ton of people. And so that's very rewarding to me knowing that the work that you do has a big impact.

(Joel Beasley at 00:13:57) So you like being useful.

(Sean at 00:13:58) I love being useful. And, you know, the S-curve as a company grows. I like the vertical bit. Right? When it gets a little bit too mature and boring and, you know, it slows down, I am not as useful at that stage. And when it's too early, I'm probably not quite as useful either.

(Joel Beasley at 00:14:17) I completely understand. Yes. I can see how that, you basically just picked the best part and said, I'm gonna spend time here.

(Sean at 00:14:26) Yeah. And I also think from a risk-return basis, working in that vertical piece, you get the biggest risk-return opportunity. You know, once a company is already Google sized, there's not a lot of risk anymore and the returns are lower. And when you're in the really super risky stage where, you know, nine out of 10 companies won't make it to A round or something, the risk is, the returns are very high, but the risk is also high. So I like that post product market fit stage where I think the risk-return opportunities are the best.

(Joel Beasley at 00:14:56) I've got questions about search stuff.

(Sean at 00:14:58) Perfect. Let's do it.

(Joel Beasley at 00:14:59) So how do you refer to the Googles and the Bings? What type of search is that? What's the words that you use?

(Sean at 00:15:06) Oh, so that would be consumer search or Internet search. You're searching the Internet.

(Joel Beasley at 00:15:12) Do you participate in that type of search or not?

(Sean at 00:15:16) No. So what we do is we power businesses' search functions. So we're there to sit behind the scenes and enable other companies to build great products that are powered by search. And typically, the data that we search through is kind of private data from companies. So it may be like the product catalog that an ecommerce company is selling on their website, or it might be, you know, in a messaging app, all the messages that are going through, or we power file management and discovery experiences.

(Sean at 00:15:46) We do site search for your site. But no, our goal is not to index the entire Internet and try to make it available for consumers in a B to B, B to C experience. Instead, we offer B to B SaaS software that powers search across the Internet.

(Joel Beasley at 00:16:02) Okay. If you could do me a favor, you reach out to Apple iMessage's team and Evernote and tell them that they need your product. Absolutely. Two things I regularly interact with. And for the life of me, I'm like, with all of my engineering experience and everything, I just cannot get the stuff I need when I need it through those search systems.

(Sean at 00:16:21) Well, here's the craziest thing. You know, there were like dozens of search companies back in the late nineties, right? And the technology that they used was basically, you give us all of the web pages. We'll look at all the words in the web page.

(Sean at 00:16:34) And then when you type in a search with a keyword, we'll just go through all the web pages and figure out which ones use the words the most. Right? So we call this keyword matching. You know, you put words into the search, it just matches the pages. And to be honest, that's how most search experiences are still powered twenty years later.

(Sean at 00:16:54) Like, you go to Google, you go to Apple iMessages, etcetera. It's just looking through all the messages and finding matches for the word. And as a user, it's very frustrating because word matching has a lot of limitations. It really doesn't understand the concept of what you're looking for. It can't handle words that have multiple meanings.

(Sean at 00:17:14) It doesn't handle typos well, typically, and it can't find things if the word isn't there. And so this is still a huge limitation on ecommerce sites. It's still a huge limitation on Google, to be honest. And this is the exciting breakthrough that we're going through right now—moving from this like twenty-year-old keyword technology to using AI to really understand the concept and meaning behind the search terms, and then actually being able to understand documents, web pages, product catalogs, and be able to understand the concepts and meanings behind them regardless of the words that are being used.

(Joel Beasley at 00:17:53) That's pretty interesting. Yeah. Because one of the things, like, Google—they do some of that contextual type search, so it will make inferences and whatnot. And then you'll use that. And so you're training your brain on using those types of Google searches.

(Joel Beasley at 00:18:09) Then when you go to a search and a just a straight matching system, you're not trained on how to search on the straight matching system. You're trained for it to have some understanding of context. But when you sit back and you're watching the consumer space explode with AI and Bing does its thing and Google does its thing with Bard and you're watching this back and forth, is this just like a beautiful situation for you guys in your search world where you're just watching these companies innovate and you're cherry picking what you want and how to integrate it into your product?

(Sean at 00:18:40) Yeah. Well, look. I think there are two different things that you have to think about in terms of search. One is the use case where I actually want to retrieve real items from an index. So, for example, in Google's perspective, you want to retrieve actual web pages.

(Sean at 00:18:55) Right? Because there's a web page you want to go to to get your job done. In the case of an ecommerce site, you really want to be able to retrieve a specific product that is on sale that you can buy today. And so we call this part of it retrieval. Now what people are experiencing with ChatGPT is something called generative AI.

(Sean at 00:19:15) And what generative AI does is it goes and reads all of the web pages. And what it tries to do is it tries to create a new response or a new answer that is going to answer the question you ask it. So the response that it gives you is something that's entirely novel and new, never been done before, but it's not actually a real thing. Right? There's no web page that says what the response is that you can go to.

(Sean at 00:19:41) Or imagine you're at an ecommerce store. Like, if they were to generate a product for you that doesn't exist, it wouldn't be very useful. So I think that's one of the things that we're really interested in—is how we can use both retrieval and leverage new technology called vectors, which is the way of understanding human language and matching these concepts, rather than just matching keywords. We're using this both on retrieval and using generative AI to help assist shoppers, for example, on ecommerce sites to find products that they want through a conversation where a chatbot can guide you through the experience. I think you have to differentiate sometimes when it's AI creating something new for you, like an answer, or whether it's retrieving something that exists in the real world.

(Joel Beasley at 00:20:30) And so this is called—I was using the word context, but this is called vector-based search, essentially.

(Sean at 00:20:37) Yeah. So when you go to like an ecommerce site that's powered by us and you type in a question or you type in what you're looking for, instead of matching the words in that query to the words in the product catalog, we turn everything into a vector. And a vector is just like a mapping of the concepts that you're asking for into this high-dimensional space. And then we can go and search for other products that are mapped as vectors nearby. So we do something called vector search.

(Sean at 00:21:04) And these are the same vectors that ChatGPT and large language models like it and generative AI are using. But we're using these vectors to match items using natural language rather than generate something new from the vectors.

(Joel Beasley at 00:21:22) That's pretty interesting. And so then it will search for its neighbors. That makes complete sense to me. Yeah.

(Sean at 00:21:27) So if you go—I mean, one of the crazy things is that these language models work in a language-agnostic way. So you could go to a web page and you could type in—let's say you're looking for a sweater. You could say the word sweater, pullover, jumper. All of these are words that mean the same thing in English. But you could also ask it in French or in German or in 50 other languages.

(Sean at 00:21:48) And all of those concepts would match to a very similar vector in vector space because the concepts are so similar. And then all of the articles, regardless of their language or regardless of what words are used, if they are like a sweater or a jumper or pullover, they map to the same vector space. And then you can kind of search around for similar items.

(Joel Beasley at 00:22:09) So a vector space isn't language dependent?

(Sean at 00:22:13) It's not language dependent. So all of a sudden, you know, if you're an ecommerce site, you can serve customers in 50 different languages, which as a European over here in Europe, we have 27 languages in the EU. And I know it's pretty tough to get your website to work well with all of them.

(Joel Beasley at 00:22:28) So if it's not language dependent, what's the data representation of a vector?

(Sean at 00:22:35) Yeah. So a vector is actually, typically, like 500 or 1,000 decimal point numbers. So they're very, very large. And this is what the breakthrough has been. So we've known that vectors are a good representation of concepts for about a decade.

(Sean at 00:22:53) But only in the last few years has the size of these language models exploded 10x, 100x to what they were before. And that explosion in size, not only the vector space, but the training data that we can feed it—basically, the whole Internet is fed into these large language models to train, and it can figure out all of the relationships between words and the context that they're used in and be able to group them into various clusters in this vector space. And it's a very powerful way of representing human language and human concepts. And it's as I said, it's language independent.

(Joel Beasley at 00:23:30) All right. This is my processing information phase.

(Sean at 00:23:34) That's crazy. You got

(Joel Beasley at 00:23:35) me thinking. Oh, yeah. Because you're saying that you've got these vectors, which are 500 to 1,000 decimal numbers. They're collected in these neighborhoods or these groups of them. And so if I'm representing a sweater, it would be represented as this vector of these 500 to 1,000 decimal numbers. Right?

(Sean at 00:23:53) Yeah. Absolutely. And nearby would be other concepts that are similar to sweater.

(Joel Beasley at 00:23:58) Okay.

(Sean at 00:23:59) The crazy thing is you can add vectors together as well and subtract them and do math with them so that you can actually combine them with multiple words to make new concepts.

(Joel Beasley at 00:24:09) If I build a language model on my infrastructure with the same training data that you build the language model on your infrastructure, are our vectors the same for a concept?

(Sean at 00:24:21) Yeah. If they're trained on the same data set, then they would be the same. Yeah. And, you know, we're language model agnostic, so we can plug in any of the language models. The actual

(Joel Beasley at 00:24:33) cross language models, they're still going to be the same?

(Sean at 00:24:36) Oh, no. So every language model has its own vector space. But we can plug in like OpenAI's models or Cohere's models or, you know, Google's models. There are like dozens of companies building these foundational language models. And you can use different language models for different use cases.

(Sean at 00:24:53) Some are tuned, for example. Like ecommerce, we have one that's tuned to that. And we can also fine-tune them as well on top of the base ones that are provided.

(Joel Beasley at 00:25:02) How much do we understand about how this works? Is there any magic in it at all, or is everything completely understood point to point, end to end?

(Sean at 00:25:11) So we conceptually know how they work. And you can actually read like the original word-to-vector paper that Google published. I think it was like a decade ago. And there's a very simple algorithm. You can play around with it and see what vectors are generated and how they can be added and subtracted in the training data.

(Sean at 00:25:29) So the concept behind them, the basic concepts, have been fairly well understood and are understandable, and you can try it out yourself. The hard bit is then, hey, how do we get like trillions of web pages all over the Internet? And how do we feed this into a language model that's going to end up having, I don't know, 100 billion parameters and is going to be very, very large to train on. And that's why these language models cost like, I don't know, $10, $20, $30 million to train each time and why companies like ours—and any, you know, we're not building new foundational models from scratch.

(Sean at 00:26:06) They're a part of the tech stack. We can like fine-tune them and refine them based on our own customers' data and the use cases we're looking for. But as a base model, understanding human language is something that's kind of coming commoditized part of the tech stack.

(Joel Beasley at 00:26:24) Okay. So OpenAI's absorbing this cost of this $30 million of computational power and engineering time to build a base model of understanding. And then they're making that base model of understanding available to us through an API or whatnot. And then we can fine-tune that model.

(Joel Beasley at 00:26:46) So we can start adding and subtracting vectors based off of the content we send it. But help me understand this, and I know this is, you know, probably beyond the scope of what you're expecting today. But it's something that I've struggled with and maybe if you don't have the answer, you could just tell me where I could go to get the answer. Yeah. Where I should be curious. I have this trouble of understanding.

(Joel Beasley at 00:27:06) I did some research and I found out that like if you want to boot up your own language model on Amazon, it's going to take 48 gigs of memory and it's going to cost you about $900 a month. It's one of the cheapest ways you could do it. Okay? So my background comes from software engineering and my best experience with servers and things like that are servers that are hosting websites and that type of technology. But what's happening with this model?

(Joel Beasley at 00:27:29) Is that one instance of the model? And if you modify anything in that instance, anyone that's interacting with that instance gets the modification? Or how are they splitting these up to where you can have your variation and I can have my variation without each one of those variations requiring an independent boot-up from the entire thing? How does that work?

(Sean at 00:27:51) Yeah. Well, first of all, there are a lot of language models out there. There are open source models. There are models that have 100 billion parameters, and there are models that have like, I don't know, a million. So depending on what your needs are, you can find a whole selection.

(Sean at 00:28:05) So, you know, you figure out your use case, figure out your speed, how much compute you want to spend, how big a server you want up. Or you could just use APIs and access a hosted version of it. But then you do fine-tuning on top of those models. You don't change the model. You just add a little layer on top of it to fine-tune it to your specific use case.

(Sean at 00:28:24) So if you want to create, I don't know, a lawyer language model, you would kind of like take a base model and then fine-tune it on some legal text. Or in our case, we fine-tune on top of like ecommerce data. So there's definitely like a stack and a lot of— But

(Joel Beasley at 00:28:40) how are they doing it multitenant? How are they doing it in a multitenant way?

(Sean at 00:28:44) Well, the language models are stateless. Right? So there's no memory. Like, when you put something through it, it just gives you an output. The fine-tuning is where you can then update the model to improve it.

(Sean at 00:28:57) But the actual inference, like the serving of those language models, is a stateless function.

(Joel Beasley at 00:29:03) That makes more sense.

(Sean at 00:29:05) And one of the—so the hard thing about vectors and the reason why they've not been scaled out more across the industry is because like a lot of research, they're very powerful and they're easy to demonstrate the value of. But we have spent the last two years trying to figure out how to bring them to Algolia scale, so 2 trillion searches a year and, you know, sub-20, 30 milliseconds latency. So like a lot of these technologies, it's seriously compute intensive. The vectors are extremely large to store in memory. And therefore, you have that kind of like the triangle around affordability, quality, and speed.

(Sean at 00:29:43) Right? Pick two. And so it has been impossible to scale up these large language models to very high-scale QPS, very low latency.

(Joel Beasley at 00:29:54) What's that? What's QPS?

(Sean at 00:29:56) Oh, queries per second.

(Joel Beasley at 00:29:58) Okay.

(Sean at 00:29:58) Or requests per second, and to do it at a reasonable cost. So they haven't been rolled out widely in production yet until now. So at Algolia, we have come up with a new representation for vectors called hashes. And these hashes are about a tenth the size of vectors, so it's a vector compression technology. And we're able to store these hashes in standard database format and search through them at high speed just like you would search through a normal data structure in a database.

(Sean at 00:30:31) Typically, the vector search algorithm requires a special tree-like structure, so all the vectors have to be arranged in a tree or a graph. And in order to search through them, it's very expensive and very slow. So Algolia launched NeuralSearch, which is the hashing technology on our own platform. And we're now able to offer like very high speed, high scale, reasonable cost vector search to all of our customers. And some of the amazing results that we've seen is that when shoppers are able to be understood, not matched with keywords, but truly understood language-wise, we're seeing, you know, 20% plus increases in conversion rates like the first day that we switch this on.

(Joel Beasley at 00:31:19) Oh my goodness. Companies love

(Sean at 00:31:21) Companies love it, and it's amazing because we can AB test it for them ahead of time. So the sale is pretty easy when you can show them the AB test and be like, do you want this or not? But it also just goes to show you that shoppers have high expectations now of wanting to be understood.

(Sean at 00:31:38) There's a long tail of queries on websites that are being underserved right now, where customers come in, they ask for something specific. And by the way, the more specific someone asks for something, the more likely they are to buy it if you can find it. But typically, the keywords, if you put in more than two or three keywords, it's hard to match all the records, and it comes back and it's like, nope, sorry, we didn't find anything.

(Sean at 00:32:01) And so what do they do? They leave your website and they go to a competitor site. But now we can use the power of vectors and natural language understanding and these large language models to say, actually, we can't match those exact keywords, but we really understand the concept you're looking for. And here are a bunch of products that actually satisfy what you're looking for, even though they don't contain the same words.

(Joel Beasley at 00:32:24) Interesting. Do you guys have a Shopify plugin?

(Sean at 00:32:26) We do have a Shopify plugin. Yeah. So we support all the major platforms. We've got a really big integration hub where you can bring data from tons of different connectors into our platform.

(Joel Beasley at 00:32:36) That is so cool. I made a new friend last week who built and sold a marketing company. Now he's working at the parent company that bought it. And one of their core parts of their business is Shopify ecommerce marketing.

(Joel Beasley at 00:32:51) And when I was talking with him last week, my first thought as I'm talking to you is I need to call John, and I need to tell him that this new search technology is existing, and it's on Shopify's platform, because that could help him serve his clients better.

(Sean at 00:33:05) We have a one-click installer in Shopify, so any Shopify merchant can get up and running. And then the other exciting thing is lots of customers love the kind of seven lines of code dropped in, search as a service experience. But what we're increasingly finding is that there are a lot of very large companies who have data that is so big and is so integral to their operations that they just wouldn't be able to send it to a third party like us and be able to host it on our platform. So what we're actually doing now is trying to figure out how do we bring Algolia and the power of this vector and hashing technology to people's existing production databases. So if you've got a Postgres database that's a terabyte and it's running a big production system, we can actually install a plugin on your database and help vectorize and hash your entire database so that you can use SQL to get the power of this without having to hand all of your data over to Algolia or hand it over to Elastic or one of these other search platforms.

(Joel Beasley at 00:34:06) Are you guys writing books on this or at least papers on it?

(Sean at 00:34:09) We are. We're working with some pilot customers that are very large companies, the type of companies that we've always wanted to work with. But because of the size of their data and how critical having it served in their own infrastructure is, we haven't been able to work with them so far. So we kind of love this idea of having two products. One is fully hosted, easy to get set up, launched. Let us run the whole thing for you. Or if you already have an existing production system and you don't want to have to duplicate your data, we can come to you and we can run it inside your existing databases.

(Joel Beasley at 00:34:43) Which is the future.

(Sean at 00:34:44) Yeah.

(Joel Beasley at 00:34:45) Okay. So this vector hashing technology, did this happen within one of your teams, or did it happen somewhere else within the organization?

(Sean at 00:34:56) So we actually have been looking for the last couple of years at vector search. We've had high conviction that vectors are going to be the future of search and AI. We went and talked to pretty much every vector database company on the market. And as we were doing this, we really got conviction that the main problem was not the vector technology, but how to scale it. And so a lot of the vector databases, you get started with them, and you either get a very big bill at the end of the first month, or you find that it hits a wall when you try to get a dataset too big.

(Sean at 00:35:29) And so we actually found a company in Australia called Search.io, and we made a small acquisition last October. They had pioneered this hashing technology, and we were blown away by it when we met the team and saw the results in production that they were able to produce. And so what was very exciting is that within three months of the acquisition, the small team at Search.io and Algolia had partnered up together and been able to release a beta and now GA version of this hashing technology inside Algolia for all of our customers.

(Joel Beasley at 00:36:02) Why did they make it? What problem were they facing?

(Sean at 00:36:06) So the problem that they were facing is the problem everyone faces when they start working with vectors. I remember at Zalando in 2017, we had a huge production Elastic service, and we were trying to get vectors enabled for search using one of these plugins. And as soon as we got vectors added to the database, the size of the database and memory just exploded and slowed down every search so much, and our AWS bills got very large. And we just realized, this is not production ready. We're just not able to use the existing vectors and all of their thousand floating point numbers and very high memory in a real world environment.

(Sean at 00:36:46) And so that's the problem that they saw as well as they were heading down this vector search first avenue trying to build for ecommerce. And that's exactly the problem that we found as we were trying to build it as well. And it's the problem that we saw in the vector database market when we went and met with everyone and tried to understand their technology.

(Joel Beasley at 00:37:06) Have you met the guys over at ChaosSearch?

(Sean at 00:37:09) No, I haven't.

(Joel Beasley at 00:37:10) I don't know if they might be a customer of yours already, but they do something similar to this. They're really smart, very collaborative, really cool human beings. And so whenever I find two cool people, I always like to introduce them even if they're in the same area.

(Sean at 00:37:23) That'd be awesome.

(Joel Beasley at 00:37:24) But yeah, they were pioneering some... they would take all the content from within an organization. If you want to find content that's within your organization, it would search across all of your assets. You'd hook your Dropboxes up to it. You'd hook... and they would help you figure that out is what I believe that they were doing. And they were just really cool people. And so I figured, we've got to let them know your stuff exists because they have their own competitors and you're an infrastructure type technology. So they might be able to leverage your vector compression to excel against their competitors.

(Sean at 00:37:59) Yeah, absolutely. And one of our big customers has actually built a product exactly like that. We're powering enterprise file search and enterprise internet search product for one of the biggest players in the industry who've chosen us to be the search backend.

(Joel Beasley at 00:38:14) Oh, that's so cool.

(Sean at 00:38:15) Yeah. No, we definitely have some pretty large customers and large datasets coming through like that. But I mean, I hear you. The problem is so frustrating. I call it the document swamp. It's like everyone's producing stuff across the company, and that's all somewhere in your Google Drive or somewhere in your Dropbox, or even worse, somewhere in your Slack messages. And it's just so hard to find.

(Joel Beasley at 00:38:38) And each one of them has a different search engine. Right? They each work slightly differently.

(Sean at 00:38:43) Well, this is where ChatGPT... I want to be able to ask questions at work to a chatbot, and it's going to come and take all of the knowledge that's been created across the company and actually kind of create an answer for me and give me the links out to it.

(Joel Beasley at 00:38:56) Do you have that today at Algolia?

(Sean at 00:38:58) Well, we're working on our generative AI products. We're calling it conversational commerce. So we're first looking at the commerce industry. So when you go to a website and you have a problem to solve, like I want to go camping this weekend or I'm looking to run a marathon, and you want to get expert advice from someone, this kind of conversational bot will talk you through the different products, the different brands. We've got a very cool demo with an electronics store where you can ask, what is OLED versus LED? This is my budget. Which brand should I buy? I want to connect it to my Apple TV. Real expert domain questions that you would talk to someone in a store to help you solve. We think this is a very powerful influencer and guide to give people conviction to buy a product that they might have... I don't know whether I should buy this one. Is it really going to work with my configuration, or is it really right for me? So we think that's a pretty exciting area around generative AI.

(Joel Beasley at 00:39:59) There are so many amazing things happening. I can tell OpenAI to talk like a specific person. I can say sound like Elon Musk or sound like the greatest marketer in the world. I can have them talk different personas. The intonation, the way that the individuals talk, that's represented as a vector. Is that how it works or no?

(Sean at 00:40:17) Yeah. So on the generative side, what the model is doing is it's trying to predict the next word to add to a sentence based on the question in the vectors that were provided as the question. And obviously, when you provide context in the question, that context then gets represented in the vectors. And so the first part of generative AI is the encoding. So same... what we do is we encode it into vectors. But then when it becomes decoded and tries to predict the next word, it will use "talk like Elon Musk" or "give me a poem like you're a pirate" or something. It'll use that context from the vectors when it's trying to predict the words to output from the model.

(Joel Beasley at 00:41:01) So someone's intonation of speech can be represented as a vector?

(Sean at 00:41:05) Well, it's... so a vector is a very highly complex thing. Right? And so when you put in the context, the question into ChatGPT, and you're like, answer this question, answer it like Elon Musk. The vector space, a thousand dimensions, can capture Elon Musk and ask this question. It has all of these kind of elements to it. These kind of thousand-dimension vector spaces are pretty much infinite in their complexity where you can place ideas and concepts. So yeah, it is kind of represented and encoded in the vector.

(Joel Beasley at 00:41:40) That is so interesting. You could take the entire way that you think and put it into one of these. That is really interesting.

(Sean at 00:41:50) If you really want to have your mind blown, read up a little bit on how crazy high-dimensional space becomes.

(Joel Beasley at 00:41:57) Just high-dimensional space in AI or in language models?

(Sean at 00:42:02) So we're used to... two-D is two dimensions, like a graph with two dimensions. 3D is the world we live in. But as soon as you start thinking about what is a 500-dimensional space or a thousand-dimensional space, and what does a sphere look like or a rectangle or something, it becomes pretty crazy.

(Joel Beasley at 00:42:21) They have visual representations, I'm assuming, that I could see visually representative what a 500-dimension space looks like?

(Sean at 00:42:28) Well, it would be mapped down into two dimensions. But yes, they could describe some of the characteristics and properties of it. For example, a sphere in 500-dimensional space has almost no volume at all, which is counterintuitive.

(Joel Beasley at 00:42:44) What's the most philosophical, heady type things that you've extracted from working in this industry after watching these models exist and actually using them hands-on, about the nature of our universe? What do you think we'll find out in a hundred years?

(Sean at 00:42:59) Well, first of all, the one thing that this vector space really makes me think about is that all of our human brains have a very similar type of encoding system. We all actually think very much alike, but the human language is how we've been able to communicate brain to brain. Right? So if you and I want to exchange concepts, we can't connect our brains together. We have to have an interface, which is human language, in this case, English. But all these interfaces have been developed simultaneously, all the hundred-something languages of the world. But they're all kind of mapping to the same concepts in our brains. And so the idea of language as a kind of interface for brain-to-brain exchange and communication, and how similar we all are, it's really quite a powerful idea.

(Joel Beasley at 00:43:50) Yeah. I think about that quite often. So before I started the show, I had a split in my path. I said, well, I could either develop... I could get into AI and develop personality-type software. I was imagining that I could give Alexa a personality and I could install a different person... this was seven, eight years ago. And then I was like, or I could do the podcast. And I said, the reason why I chose the podcast as the business is because if I was off by two or three years, I would be broke. If I did the podcast, at least I'd have relationships if that didn't work out. Right? So I risk-mitigated there.

(Sean at 00:44:24) Yeah. And look, timing and luck are everything in technology. And so my first startup was a mobile gaming company. We were ten years too early. It was exactly the right idea. We totally had the intuition about the product. But ten years too early, and ten years is a long time to be wrong before you're right. So yeah, timing is an important part about all of this.

(Joel Beasley at 00:44:47) You were 100% right. Timing is so important. So when I was thinking about that, to your point of what you just said about language being an interface, I was thinking about how when we communicate, what we'll ultimately remember from this conversation in 90 days is maybe the main topic of it, and whether or not we really liked each other, how we felt about the conversation. So this interface is really... I think of words and I was being very introspective and curious about how this stuff works when thinking about creating these personalities for this AI. And I was thinking, my brain's constantly checking with my heart, essentially, of how do I feel about these words that are going to be coming out. And then that goes back and forth until I come up with the good enough words to get out there. And so I was thinking that it's just this abstraction. That's when I realized that this language is just an abstraction and what we really are, we're just constantly exchanging these feelings. Where that breaks down for me when thinking about a way of the world is when you need the information and the detailed words to do hard sciences. And I'm like, that's where that breaks down. So it's almost like behaviorally and communication-wise between each other, there's that aspect of it, the feeling aspect of it. And we use the hard words to get to the feeling aspect of it. But then on the hard sciences, the hard words become very, very important. They have to be very specific.

(Sean at 00:46:03) Yeah. And but that's just highly structured language, mathematical formulas, and scientific stuff. But you're right. We both have a vector space in our head, and in a week's time, we're going to remember the concepts. We're not going to remember the words. The words are really there to exchange concepts between our brains.

(Joel Beasley at 00:46:20) I'm going to write that down.

(Sean at 00:46:21) Yeah. Our vector space is probably very different, by the way. So we don't have the same vector space, but we can use language as a common interface. And by the way, that's why ChatGPT has been so explosive, because this is the first time that humans have a natural interface, their language, that they can use to access AI and for AI to respond with an interface that they understand. Right?

(Sean at 00:46:48) So it's the actual interface of natural language in and out of these models that's the most powerful piece of the technology. It means everyone in the world can access AI now. All you have to do is be able to speak language.

(Joel Beasley at 00:47:00) I wonder if one day we'll find out that the concepts of everything already exists. It's just our ability to discover them. Right? Because they could be all around us. All the answers could be all around us, and it's just our ability to figure out how to discover it.

(Sean at 00:47:13) Well, that's what your nine-year-old's baby is doing, wandering around discovering concepts. I always laughed. I was like, my kids are basically AI models being trained, and it's a lot of trial and error and a lot of experimentation. And then they finally grasp the concepts and store it in their memory and then can build the concepts up on top of each other.

(Joel Beasley at 00:47:35) I have had that same conversation before. And to add to it, it reminds me of in the 1960s when you would see these pictures of the 1960s where you have these rooms that a computer is the size of a room, and you can see the transistor. You could see all the parts. It's very clear what's happening. Now it's all condensed and hyper-super-processed and everything like that.

(Joel Beasley at 00:47:56) But with kids, it's human behavior on that scale. You can see anger and joy, and you can see behaviors as if you were standing in that computer room. But when they're adults, it's almost like it's all compressed. Right? And you have to be an expert in it in order to take apart that device.

(Joel Beasley at 00:48:15) Yeah. So yeah.

(Sean at 00:48:16) We're fine-tuning our models as adults, whereas they're building the foundational model still from scratch.

(Joel Beasley at 00:48:23) That is yes. Well, Sean, this was an absolute pleasure. As I become more intelligent in this space and get better questions, maybe we'll have you back on next year, and we can talk a little bit more about the amazing advancements you guys are making.

(Sean at 00:48:36) Yes. Thank you so much for having me on. It's been a really fun chat, and we'd love to come back again.

(Joel Beasley at 00:48:41) Yeah. Thank you so much for doing it. We made a podcast. How do you feel?

(Sean at 00:48:44) Fantastic. It was great fun.

(Joel Beasley at 00:48:46) Thank you so much for listening. And if you found this episode useful, please share it with a friend or colleague who you think would get value from it. And if you have topics that you would like to hear discussed on the podcast, either add me on LinkedIn or send me an email: [email protected]. Every time I get an email or LinkedIn message, it absolutely makes my day and inspires me to keep going.