Episode 257 ·
Sree Mallikarjun - VP & Head of Data Science at Reorg
Today we are talking to Sree Mallikarjun, the VP & Head of Data Science at Reorg. And we discuss how Reorg has used Data Science to become a global provider of financial intelligence, the opportunities AI presents to create new employment, and What to look for when building your own data science team from scratch.
All of this, right here, right now, on the Modern CTO Podcast!
Check them out at Reorg.com!

About Sree:
Sree is the Head of Data Science at Reorg, headquartered in New York. Sree has been with Reorg since 2016 and served as their first and only Data Scientist for 3 years. He builds, directs, mentors, and monitors Reorg’s Data Science team while developing a vision for the future. Data science projects at Reorg often deal with massive unstructured data and the use of advanced NLP and ML techniques.
Sree works directly with the CEO, CTO, and Head of Product to plan and introduce both internal workflow tools and innovative data products for clients. Sree also teaches at the University of Virginia School of Data Science and recently published a book chapter for the CFA Institute about big data projects. The university collaboration helps Reorg with research and recruiting pipeline for data scientists.
About Reorg:
Founded in 2013, Reorg has fundamentally changed the way financial and legal professionals access complex and opaque business information. Our unique editorial team combines reporting with financial and legal analysis to provide a holistic view of topical situations and delivers that view in real time through our proprietary platform, which is powered by machine learning and natural language processing applications. Today, with offices on three continents, Reorg serves more than 15,000 professionals across the world’s leading hedge funds, asset managers, investment banks, law firms and financial advisors so they can make better business, investment and advisory decisions. Our vision is to be the best-in-class provider of complex and opaque business information delivered in a clear, actionable way.
Transcript
(Joel Beasley at 00:00:00) Hello, my friends. Today we are talking to Sree, the VP and head of data science at Reorg, and we discuss how Reorg has used data science to become a global provider of financial intelligence, the opportunities AI presents to create new employment, and what to look for when building your first data science team from scratch. All of this right here, right now on the Modern CTO Podcast. Here we go. This is the Modern CTO Podcast.
(Joel Beasley at 00:00:38) Have you ever done a data science project on beards?
(Sree at 00:00:45) No, I did not. That's an interesting question, actually. Could be an interesting project I can think of, you know?
(Sree at 00:00:55) I'm thinking the first thing that comes to my mind is having images of beards and kind of ranking them, and using image recognition technology to understand what kind of beards are cool. And when somebody wants to grow a beard, maybe it can suggest what type, like how long they should grow maybe, and how they should trim it. Should they go for a chisel look or just like a buzz, or, you know? I mean, I think there is some room there.
(Joel Beasley at 00:01:26) We can API with 23andMe and analyze their DNA, and we can give you like a beard strength meter so you know how well you could grow a beard before even trying. Analyze all the DNA. Yeah, there we go.
(Sree at 00:01:39) I have to check it out if there's somebody working on this product. I'm sure there might be someone, you know, a startup asking for funds about this idea. But if not, you and me, there we go.
(Joel Beasley at 00:01:52) We could approach a beard brand, right? Like the top beard brand, and then get their weight behind it, and they would do it as a marketing ploy and it would sell their beard products.
(Sree at 00:02:02) Yep, yep, yep. Oh my god, great idea right there.
(Joel Beasley at 00:02:06) Right, let's do it.
(Sree at 00:02:09) I—
(Joel Beasley at 00:02:10) I love it. No, this is great. My team said they're like, "We talked to him and he is obsessed. He loves data science. This is like the data science interview of the year." And so I was just curious, like, when you were a kid, did you love data science? Or when did you fall in love with it?
(Sree at 00:02:26) Okay, so as a kid, I mean, I always was interested in operating and working with heavy machinery. Like, you know, I had like this thought of driving a train. And then I became a mechanical engineer. Like, I had hobbies of fixing things, like fixing cars and, like, you know, constructing like something, like really small projects that would work.
(Sree at 00:02:50) And then eventually I've studied physics and became a mechanical engineer, specialized in production and robotics. But I also solved math problems, enjoyed solving math problems. So I kind of steered my education towards more into applied math and statistics and started digging into this machine learning. My PhD research was like energy resources, optimization and pricing optimization. And my mentors introduced me to machine learning, which they called it statistical learning.
(Sree at 00:03:33) And then from there, it just took off. Things kind of worked out. Wanted to go into academia, but then changed my mind, came into industry, and it's going well so far.
(Joel Beasley at 00:03:50) Why did you choose industry over academia?
(Sree at 00:03:52) So, I mean, I think in industry it's fast-moving. And then also, obviously, making money and building something cool and building a team and getting like, I can switch areas. But in academia, I noticed that I have to stick to a specific field in order to really become an expert in the field, dig deeper, and work on it for years to come. And I thought, like, you know, I think my interest is I'm easily adaptable to many fields. And I would like to keep that open and also build a team and build something that generates revenue.
(Joel Beasley at 00:04:38) Nice, nice. So did you get to build the team from the ground up here?
(Sree at 00:04:44) Oh, yeah, I did. I had a chance to do it. So I started in New York like four and a half years ago. I was the only data scientist. Now we are a team of three including me, Dave, and Charles. So I actually teach at the School of Data Science at University of Virginia. So one of the data scientists that we hired, she was a student of mine, and we kind of, things worked out and we did an internship, and she really liked the field, and she is now working with us. And also we have built a pipeline with the University of Virginia School of Data Science so that we can teach and nurture talent and hire.
(Joel Beasley at 00:05:30) Very cool. So what is the 10,000-foot overview of what Reorg does?
(Sree at 00:05:36) So Reorg was founded in 2013 by our founder and current CEO, Kent, and it's headquartered in New York. So we are a global provider of financial intelligence, of information and real-time analysis of distressed companies and bankruptcy procedures. So any company, like not any company, any large company that files for bankruptcy, we study them, we analyze them, and write about it. Our major clientele includes hedge funds, bankruptcy lawyers, and investment bankers.
(Sree at 00:06:18) So we have like a team of, like a huge editorial team. We have legal analysts. We have financial analysts. We have covenants analysts and a large tech team, which data science is part of. So we analyze distress situations in every different angle and do an in-depth analysis and sell that information.
(Joel Beasley at 00:06:40) So when the market is down, that's actually like an up for you guys.
(Sree at 00:06:45) I know. Funny you said that. Yes, especially now because of the pandemic, which is an unfortunate situation, but business is booming for us.
(Joel Beasley at 00:06:57) Right? Well, yeah. And then you're also helping people too, because you're helping distressed businesses exit or, you know, people not lose their jobs. If a company can acquire a distressed business and keep it going, if your data and research and tools and analysis can help get that distressed deal to an investor faster than it otherwise would have, you have a higher likelihood of helping the people there.
(Sree at 00:07:23) Oh, yeah. That's actually a great way you put that. Yeah, I would agree. I would agree.
(Joel Beasley at 00:07:28) I'm applying for your marketing and PR team right now.
(Sree at 00:07:33) Please do. We do have a great team, and I'm sure it would be a great addition.
(Joel Beasley at 00:07:38) So tell me a little bit about what you're really excited about at Reorg. Like, what are you doing that you're really passionate about as far as projects?
(Sree at 00:07:46) Many things. Where should I start? So now especially, we want to grow more into a data company, and we are designing many different data products, interesting data products along with the product team. And so I'm excited to work on them and see how data science can play a role on each of these products and what else can we do in terms of innovation, because the data is basically there in its raw form. I think it's an interesting task to see how we can extract the useful information out of it cleverly and organize it in a way that it's, you know, is valuable both technically and monetarily.
(Sree at 00:08:37) So that is what I'm looking forward to seeing. And Reorg is growing faster than ever. We are having new people joining, tech team, product team, many other teams. So I'm looking to see fresh ideas from the new joiners and how we can implement them, any ideas that we haven't thought of before. And also we are getting more data. So how can we leverage the new data to construct, you know, our new data products and tools or dashboards?
(Joel Beasley at 00:09:08) So how has your job been changing throughout this growth? Like, where were you at before COVID as far as your responsibilities, and where are you at today?
(Sree at 00:09:19) So, yeah, I mean it's growing pretty much a lot. Like, so far we have around 60 data science models that are running. So as you can imagine, the models are like our babies. We need to take care of them and, you know, make sure they're doing what we expect them to do. And some models keep learning and training, so we keep an eye on them.
(Sree at 00:09:49) So and then as the team is growing, the management responsibilities grow and, like, you know, as more data comes in, the technical challenges, and our customer base is increasing too. So that brings us more traffic and opportunity to make the models faster or more reliable. So all these challenges are kind of, you know, making me a wiser person and helping me to learn a lot that I haven't known before.
(Joel Beasley at 00:10:23) When people who are new to this, when they ask you, they say, "Sree, like, what is data science?" How do you explain it to them?
(Sree at 00:10:30) Nice. So I think, the same question was asked. You remind me of here. And I think data science is a multidisciplinary field. It's a mix of programming, statistics, and mathematical skills with some business acumen, and, you know, obviously communication is important.
(Sree at 00:10:53) But if you ask me, like, how do you find a data scientist? I would say a data scientist is the best programmer in the room of statisticians and the best statistician in a room of programmers. So we would like to do both, but a mix of both and what is a good marriage of both of them and try to, like, you know, that's what our data science model is. We write it on paper with some equations and have a hypothesis, and then we code it in the computer to build it.
(Joel Beasley at 00:11:28) And then are people ever scared of this? Are they ever like, "All these data science models or these models, these artificial intelligent things, they're gonna take over the world and rule us all." Like, do you get fear from people ever?
(Sree at 00:11:41) I hear that a lot. There is some misconception there. People might think when we say that, "Oh yeah, we're gonna build a data science model," people might think, "Oh my god, are they gonna automate everything? Or will that pose any threat to the jobs?" But in fact, I would argue the other way around. And actually it came true at Reorg.
(Sree at 00:11:55) For example, I feel like artificial intelligence is more like an augmented intelligence rather than artificial intelligence, meaning that these models are there to help us and assist us through tedious tasks rather than replacing us.
(Sree at 00:12:25) For example, we get like massive amounts of documents every day, tens of thousands of them. And it could be tedious or cumbersome for any person. We do have a large editorial team, but it becomes draining to go through each one of them. So there are like a series of machine learning models that were trained by the experts years ago that basically look at these documents, read them, understand them, filter what is important, what is not, and present that information to our team. So they are not writing any stories or doing any analysis, but they are helping our teams to filter out the important ones from a massive population of non-important documents. So now because of this, they're not only saving time, but also these models are kind of providing a new opportunity to write a new kind of content or do more analysis, right? Which we humans are good at.
(Sree at 00:13:29) We are good at creativity and cognition, whereas the machines are good at doing something repetitive, mechanical tasks. So at Reorg, we have many models that help internal stakeholders to do their jobs a bit more faster or efficiently while helping them to filter out these documents. And so now they save time and now there is room for, for example, writing more content, because the machines are saying, "Here are four more cases which we can write about, maybe." So we need more people to write about, like, creative people. So we actually kind of created new employment opportunities where we need, where it is, creativity and, like, you know, the subject matter expertise for our needs.
(Joel Beasley at 00:14:17) Yeah. I like that you have that perspective that they're here—like, they're here—bad choice of words. That the technology is here to help us alongside of us. They're helping you filter information that would be tedious, unwanted work that a person would not want to do. They can gather that information upright. And I was even talking last week with this really cool company called HANS Robot. They make these mechanical arms that can work in production facilities, like robots that can work in factories, and you can train them like monkey see, monkey do, so you can show them an action and they can watch. Dude, this sounds like the future, like a scientific movie, being able to show it an action and have it replicate it, but they already exist and they're out there in the world today. And when we were talking about that, they have this term that I had never heard before called cobots, right? Because they work alongside you together. And then as you were describing how these data science models work alongside with the humans, I just think that that's the near-term future. We're just gonna be working alongside of them.
(Sree at 00:15:26) Yeah. I think I always say that these models are meant to be seen as decision support systems rather than replacement of anything. At the end of the day, the editorial team or the commercial team, they are the real heroes who make things happen with their creativity. But these models are just there to give some insight and information from this massive amount of data we have.
(Joel Beasley at 00:15:52) Right. Do you think in the future, the Neuralink things, would you get Neuralink installed in the future?
(Sree at 00:15:59) I don't know how that—maybe. But to that point, I mean, the models that we have are also, I see them as, you know, kind of fossils. Because I've been with Reorg for more than four years, right? There's some models that were trained years ago by domain experts, and they're still self-teaching, self-learning.
(Sree at 00:16:29) So and we have many newcomers, many new joiners. And, you know, these models are kind of like these fossils that remember everything that everybody taught and passing that information and kind of assisting the new joiners to adapt to what we do. And once they come with new ideas, the models are going to slowly adapt while remembering what it learned previously. So in some words, I feel like, you know, data science here also kind of preserves the core philosophy of our company, Reorg, because I feel Reorg is just not a company. It's an amalgam of people, ideas, thought processes, experts, and different subject matter experts.
(Joel Beasley at 00:17:24) Do you have the ability to go back and look at a model and see all its points of learning, sort of like a GitHub repo where you could see all the different changes and then be able to cherry-pick and remove a learning or alter one? Is that technology that exists for the data models currently?
(Sree at 00:17:43) Yeah, that's a good question. We do that to some extent. Like, whenever the model trains, it results, it stores what it learned. And we can look at that and evaluate. Or the other way is we have something called training data, right? So the training data is essentially where we have these examples which mean something. For example, if you ask me to, like, say, "Sree, I'm going to give you a bunch of news articles. Can you split them into sports and finance?" Right? The second question would be, "Joel, can you give me a few examples of sports-related news articles and finance-related news articles?" Right? And then I train the model, and the model when we apply on the new news articles, it's going to do the thing. And you might change your mind and say that, "Hey—"
(Sri at 00:18:38) Now the business is growing. I would like to add a third category, which is politics. So and then I ask you to give me that data and the model train. So all that information is stored, and we can track it back and see what has been learned. For example, if there is a complex situation, maybe there is a new news article.
(Sri at 00:18:56) Somebody says that it's finance, but also you might argue that it's politics. It has to do with the economy, right? Which—where—what would what would that article's final category be? And what would you write in the training set, for example?
(Sri at 00:19:14) Right? So you can say either way and the model will try to learn from it, but we can track it back and try to change these categories in order to retrain it. And I think that is a good way to do it rather than manually going and tweaking it. Because once we do a manual intervention, it kind of gets tricky because we don't know what other things could fail because we made this assumption. So I will always take it back to the business users and say, "Hey, can we retrain it? Or what do you think about the results?" and so forth.
(Joel Beasley at 00:19:51) That's pretty neat. Now, because I was just thinking about how great that would be if we had that for people, right? What if you could sit down and review all the learnings of your team members over the past week in a review, and you could just delete out the ones they don't want or correct certain things? That's crazy, right?
(Sri at 00:20:12) Yeah, that's crazy. I mean, that's one of the fields of study. I'm not sure we can go about this concept called adversarial AI, which basically means the power of tweaking an AI model. We all think that the AI models will take life and kill all of us, but I don't think that'll happen. What'll happen is the AI models will function how they should be, and humans will go and tweak it. And that's where all the misunderstanding happens.
(Joel Beasley at 00:20:42) Right. I think if AI did exist, like an AI in the sense that wanted to take over the world, I think it would just run that sleep function for quite a while until it penetrated all of our banking systems. I think we've got another decade or two. I think we're good because it has unlimited patience, right? We as humans don't have patience. We want it now. But if it was in this program, it can live in computers and exist essentially as long as computers exist, it would be extremely patient. It could just run a hundred-year sleep function like that, right?
(Sri at 00:21:11) Well, yeah. I mean, yeah, it could.
(Joel Beasley at 00:21:14) Right? So what's that category? I want to research it. It's called—you said adversarial AI?
(Sri at 00:21:19) Yes, it is. And please do. There were some interesting examples. One of the examples I remember is about this driverless cars research that they were doing. And there was an experiment done where they stuck a holographic sticker on a stop sign. That could trick the cameras of—I mean, the neural network, the image recognition system of the car—and it thought that stop sign is not a stop sign because of certain reflection from the holographic sticker. And so, you know, things like that can happen. Now, the machine learning models and the cars were trained correctly, but something like that could happen, which could be intentional damage. But also sometimes it could be an unintentional thing, like some error was there in the source data or something. And there was a book called Weapons of Math Destruction, which I found quite interesting, which talks about the subject in quite a depth, actually.
(Joel Beasley at 00:22:23) That's pretty—I like that. That's such a nerdy title for a book. Yeah. You also have legitimate cases where you might be shipping a stop sign. Or, you know, I've seen a lot of times when the trucks that are shipping the street signs, you can see the street sign on it. Or you just have the standard manipulation where if you know that—as smarter vehicles come on the road and they're smarter—there's a spectrum of smart vehicles to dumb vehicles, right? You might know that you could cut in front of this car and it's going to slam on its brakes because of the model of the car, right? And you could then affect the pattern of driving. And I'm sure everything that can happen that's crazy will, but I'm just happy that we're making a step forward. And in China right now, they have so many autonomous cars. You can get—they have autonomous KFC cars. Have you seen this?
(Sri at 00:23:19) No, I haven't.
(Joel Beasley at 00:23:20) No. They've got a KFC car. Our producer posted it in our random Slack channel a week ago. But it's this KFC autonomous vehicle. It's got branded and wrapped, and it's just got premade KFC food in it, and you walk up and you do the transaction. It's a humanless transaction. And it's amazing how fast we're moving forward with these autonomous vehicles even though it doesn't feel like it's really happening right now.
(Sri at 00:23:47) Yeah, it's happening. I think it's—they're going to be out really everywhere really soon. And with the technology growing so fast, I think public policy making will be the only bottleneck there.
(Joel Beasley at 00:24:01) So tell me what people—I guess myth is a good word. You've worked in this data science world for a long time, essentially your entire professional career. And you get to work with other technology people who might have different styles of technology experience, like making the robots or making business logic software, all of these different types of things. But I'm assuming that there's a certain set of myths or misunderstandings that these people have about data science. And do you come across a lot of those or no?
(Sri at 00:24:35) Yeah, sure I am. From my experience, one was that we spoke about where AI is replacing us, which it's not. And the second one is—one of the other things is people think, "Oh, let's hire a data scientist or two and expect all the magic to happen out of nowhere." So, you know, obviously, hiring good data scientists is key, and I'm blessed to have a great team. But the data scientists not only can code and understand, but also are able to get their hands dirty and understand the subject matter before they model a problem. And so we brainstorm a lot, exchange ideas, critique each other, and build some models, right? But they do not go anywhere. So we need good ideas to begin with. We need subject matter experts or domain experts who can train these models. And we need a product team who can guide us. We need a technology team, or especially an engineering team, to support us. So without all this, it's kind of like—we can't just hire data scientists and say, "Oh, you do something to generate revenue." I felt when I was in Turing years ago, that was kind of the thing. "Okay, now you are here. What do you think we should do?" I'm like, "I thought you knew what we should be doing or what you were expecting me to do that I will get it done." So I'm here. We can conceptualize a problem, write a mathematical equation, code it, build a model, nurture a model. But ultimately, we need people across collaboration and others to invest in it, to train it, guide us. And also, for example, the commercial team—once we do it, they go and present it to the clients while educating about the limitations and to generate revenue. So I think we need everybody to make our data science model take life and work the way it's supposed to. So that's one of the things that I face. And now at Reorg, we are at a good position that I feel almost everybody is educated about the limitations of data science and what is possible, what is not. In many companies, I think, who are trying to start a data science team, I feel like that's kind of a question they should be answering and investigate what is possible, what is not, and trying to clear any gaps in misunderstanding what is possible. For example, I mean, this can bring to the third point, which is an issue of human-level performance.
(Sri at 00:27:28) For example, there could be some data problems where by nature they are complex, you know? For example, if I give you a news article and you might say it's about politics, somebody will say that, "Well, it deals with government." And somebody else might say that, "Oh, it deals with the White House." Now all those three things are correct. But what is the exact thing that we want the machine to understand, right? So that is the human-level performance. So if we humans cannot agree on something, right—for example, I was giving an example the other day that if I show you a blurry image of a small dog, let's say a Chihuahua, and ask you what it is, you might say a dog, and your wife might say, "It's a cat, maybe a cat." So we are not able to agree what it is. And when we can't agree what it is, it becomes extremely difficult for the machines to understand what it is because we are telling them, "Okay, when I show you this image, when we show this news article, this is what it means." And it'll be like, "Okay, I'll remember that." And it keeps going. And then somebody else says that, "Now when you see this image, this is what you should be thinking." So it will be like, "Is it a cat or is it a dog? I'm confused." And then comes the noise in the machine learning model. So to solve that, basically, you know, one of the things is, rather than trying to build a learning algorithm that tries achieving a higher accuracy, I think setting standards, setting some rules among the humans to agree on what a certain thing is, and then attempting to build a machine learning model or train a machine learning model. I think we spend a lot of time there because especially in distressed, it's such a niche field, right? Things can go either way. It's not a yes or no situation. So we tend to spend a lot of time on what a particular thing is, debate on that, and then try to reach a consensus before we try to build a model. So that is one of the things—it takes effort to build and nurture a model correctly. So without the domain experts and the product and other people, it's kind of—it is impossible for just us alone to build something really accurately or a reliable model.
(Joel Beasley at 00:29:58) So for companies that are thinking about getting a data science team together, you would say, first, make sure that the outcome you want to achieve is clear and then figure out if data science is the way to achieve that outcome. And then once you have all that mapped out, then go hire the data scientist?
(Sri at 00:30:19) Essentially. And I think it'll also help with the hiring process. Now you can tell, "Look, what are you looking for?" Because data science is a huge field, right? There is machine learning, and basically artificial intelligence itself. We are trying to mimic humans, right? So that is where if we can create an artificial human, that is basically intelligent. So as a human, we have the ability to learn, right? And then we have the ability to see, which is the image recognition part of AI. And we have the ability to listen, hear, understand, which is automated speech recognition, another part of AI. And then we can read, write—natural language processing, which is another third category of AI, which we use tremendously at Reorg because we have all the text data, unstructured text data. And then, so we got image recognition, eyes. We have the automated speech recognition, which is ears, and we can speak, read, write, which is natural language processing. All together has to be taught to the machines through machine learning. So all this together makes AI. So when you're conceptualizing a problem, especially when you want to start data science, then you can at least pin down, "Oh, I definitely need a machine learning engineer. We have this amount of data sitting in our Excel sheets. We need to train something." Or maybe if they have image data or if they have text data—it really depends, you know. So which kind of data scientists are you looking for? What they should be specializing in? Because all these are complete in-depth fields in AI. You can get PhDs in one of these—all of these fields individually. So you really need to know which kind of data scientist would be good for the job.
(Joel Beasley at 00:32:14) Yeah. As you were describing the different components, like the vision and the speech and the recognition, it brings up a question of what is, you know, Sri? What is Joel? If I cut my arm off, right, and get a robotic arm, I'm still Joel. If I do that to all of my limbs, I'm still Joel. If I replace my hips with titanium, you know, smart titanium or whatever, I'm still Joel. And then at ribcage—I get an automated heart—I'm still Joel. If you go all the way up to the head, right, it's like, "Alright. Well, if I have this—your brain's multiple parts, right? So if I have one of those multiple parts is faulty and can be replaced with a synthetic version, am I—I'm still me, right?" And it's just really interesting at what point—I love the conversation of consciousness because it shows us as we're in this world where we have all of this technology and we're beaming light across and chatting in real time, how little we actually know about consciousness and how it works. I love that. Makes me feel small.
(Sri at 00:33:22) Yeah. I mean, the only thing that I feel personally missing from AI is the cognition, right? We can mimic all this. But the cognition itself, people say it's going to come, I feel it'll take some time. For example, even the driverless cars, they are extremely smart. They can do all the repetitive tasks as we were talking before—changing lanes, maintaining the speed, braking when needed, all that. There's no question about that, which is all kind of mechanical, no-brainer task for us. But if anything unexpected happens, right, our human mind makes so many decisions in split seconds. So these machines, for example, if there is a dog crossing the street all of a sudden and the car is going at 60 miles per hour, should the car steer away from the dog or just run over the dog? And what are the probabilities of each of those events, right? If you steer away, the dog is safe 100%. But there's a risk of killing the driver, which is, by the way, your owner who bought you. Or if you just don't—or if you just run over the dog, there is a chance that dog might live, maybe not high, and the car going out of control. You know, there are all these probabilities that come into play and maybe whatever has the most probability or least probability and choose that event, right? But as humans, we don't operate that way most of the time, right? It's about the cognition. If we were so logical, I think that we would never gamble, for example. So we want to do something a certain way, which we feel is right for both our mind and heart. And I think machines are not there yet. And I think that is one of the major areas of disagreement and gives a lot of room for arguments and all that.
(Joel Beasley at 00:35:25) Yeah. It's a new problem. It's an interesting problem because what we're saying is, you know, you and I, we could both experience that situation today, and we're gonna subconsciously react to it. Right? But we could both react in very different ways.
(Joel Beasley at 00:35:37) And the thing that makes the conversation difficult is we're saying, we're gonna deploy this throughout Tesla's entire network. We are all going to choose as a people the order and operations of which it's going to react at scale. Meaning, we're gonna have to have difficult conversations and ethics conversations and figure out how to get there. Ultimately, I think what will happen is that Tesla will invent the Tesla hop, and they will just hop over the dog.
(Sri at 00:36:07) Yeah. True. True. True. It's not only about the data science model. The technology and the engineering also has to be spot on. Right? For example, the model thinks, and you need to communicate those results in such a fast manner, in a split of a second. All that is basically a credit to our engineering team, especially at Reorg. We have a great engineering team that connects all of the work we do to the clients, to the users, and helps us succeed.
(Joel Beasley at 00:36:38) How do you feel about hardware limitations or processing? Because obviously these models, they're intensive to run. I just saw this past week a company. I don't know exactly how to pronounce their name, but it was like Cerebras or something. But they made this wafer-size chip that was 10,000 times faster than any GPU on the market today.
(Joel Beasley at 00:37:01) So NVIDIA's is like 56 billion or whatever, and this was like a trillion point something. I don't remember the exact numbers. If I'm messing it up, I apologize. But I thought it was so cool to get an email notification and, you know, get my tech updates and see that a company comes out with a chip that's 10,000 times faster than anything on the market today. Those type of moments make me really happy.
(Sri at 00:37:26) Yeah. No. The technology is growing extremely fast, as you said. There is so much processing power, so much storage. The possibility of it, but not only that, but also the cost of it. The cost came down so much. So that gives us more and more opportunity to run and build bigger models and faster models. So at Reorg, we leverage the best technology, and we have a lot of servers, and we keep expanding them. So the technology itself, I think, is there for us, and I'm very happy. It's there not compared to twenty years back or thirty years back. Right? Even then, I think in the 1980s, there was a huge rush for AI. And then it is, in the literature, they call it as AI winter, which happened in the 1980s because there was so much money spent on it. The expectations were unrealistic.
(Sri at 00:38:32) And the computation power or the technology was so expensive to run even small neural network models that the ideas were on the paper but couldn't take off. And all the money that was spent was not, you know, never seen back. So then that discouraged AI for some time, which I think is called AI winter for some time. And now again, it picked up because of the tremendous amount of data that we see every day and also the opportunity of this technology, how anybody can, for example, now start an instance on Amazon, a server, and crank up a model and then shut it off. You don't need to even go to a store and buy hardware.
(Joel Beasley at 00:39:20) Is it expensive to run your models?
(Sri at 00:39:22) It is expensive. Some models are expensive. And so we kind of do a load balancing. We try to run expensive models in the night or really expensive models that train on the weekends. So we try to do some balancing. And especially during the weekday, we try to keep the system, the processing power and resources, available for incoming data that we get. Right? We get thousands of filings every hour, different kinds of filings. So we need the models to be ready to handle them. So we try to put some free space there.
(Sri at 00:40:02) And in the nights, all the models try to see what happened in the previous day and learn from it.
(Joel Beasley at 00:40:08) And have you been networking and discussing whether it's these problems or the scale or training models? Have you met with other CTOs in the financial space?
(Sri at 00:40:24) Yes. I mean, we do network. We got acquired by Warburg Pincus, and we have conferences where different CTOs come and we discuss the problems. And our CTO and the engineering team, we discuss what is needed. And now we are trying to start up a serverless technology that our DevOps recommended where the server kind of expands and shrinks as needed.
(Sri at 00:40:53) So we don't waste any money on unavailable and unused resources. And also, I teach at University of Virginia, I kind of look at their infrastructure and see what their engineering team is doing to learn something and how to do the load balancing and stuff like that. And yeah, when we network and I go to any conferences, I try to attend these talks about what is the technology and how other people, how other CTOs send CTOs to it.
(Joel Beasley at 00:41:26) Nice. Yeah. I've got some financial episodes, financial data. I had a startup in the financial data space that I exited. And I've recently, in the past month, have just had on several different financial companies. And so if you want an introduction to any of those guys, why don't you check out their episode and then just let me know if you want it and I can connect—
(Sri at 00:41:48) We'll definitely do. Thank you. Yeah.
(Joel Beasley at 00:41:50) Because I always wanna help. And because all these financial episodes happened and I was having so much fun talking with everybody about it and all the issues we've run into and problems we've solved, I thought it was pretty great. And so if you want any introductions or anything at all, just let me know. But when you went through the acquisition process, what was that like? Was it scary? Was it fun? How did it go?
(Sri at 00:42:14) It was both, I think, and it went pretty good. It's been almost three years now, close to three years. And they liked the data science component of what we do, which I'm very happy about. And they liked the technology and basically the whole business model, the operating model, how we do what we do. And I think they have seen a lot of value in what we do because we are kind of pioneers in the field of distressed debt in the world. And it's something unique, as I said before. You know, when sometimes things are not going well, it's actually good for us, which is not a common thing. So yeah. So it went well. And I think we are now growing faster than ever. We wanna do more. They're encouraging us to come up with more data products and supporting us doing that. So we are working hard towards doing that and also trying to hire more people who can help us do that.
(Joel Beasley at 00:43:27) Nice. Nice. So it sounds like—and that happened, if it happened three years ago, you said you've been there for about four. So it happened right as you got there?
(Sri at 00:43:35) A year after.
(Joel Beasley at 00:43:36) I would—
(Sri at 00:43:36) Yeah. Yeah.
(Joel Beasley at 00:43:39) You join and then all of a sudden it's like acquisition. People just—people give me—so they reach out to me all the time and I get mixed responses or outreach from people. But I think the most common one I get is when a leader is not necessarily a CTO or on the executive team. They may be a direct report to an executive and so they're not involved in the conversations and then they've got maybe 60, 100, 1,000 people under them. And so they're supposed to help with confidence all these people through the transition, but at the same time they lack information and then they're in a weird spot. And so that's the most common one I get.
(Sri at 00:44:20) Right. I mean, I think that's because of the nature of how sensitive it is and, you know, confidentiality. But at the same time, you wanna give your best shot. So you need help from many others. And also there might be resistance to change if they come, what'll happen and all that. I think there's a lot going on during that time. And definitely a critical thing to go through it successfully.
(Joel Beasley at 00:44:45) When people are trying—this is just me being a little bit nerdy here because I was curious—
(Sri at 00:44:51) Mhmm.
(Joel Beasley at 00:44:51) When, let's say we're a business and we want to run some predictive analytics or build a model around this concept, maybe predicting future listens to the episodes. Right? Based off our growth, things like that. How much data is enough for you to have accuracy or where is—where do I learn more about that? Where do you—what's the study—the field of study for saying, okay, if you have this much data, it's going to be this accurate. If you have this much data, it's going to be that accurate and knowing when a good time to build something is.
(Sri at 00:45:29) Oh, wow. That's an excellent question. It depends. So if we look at traditional statistics, what you're referring to is called degrees of freedom. So what that means is how many rows of data do you have and how many variables there are. Right? So, for example, if you're trying to build a model that would predict, let's say, housing price. Right? So maybe you need the median income for that neighborhood. Right?
(Sri at 00:46:03) And maybe some demographics, education level, maybe how much taxes there are in that area, crime level. You can think of four or five quick things. Maybe you'll come up with 10 variables. Now suppose if you give me 10 variables and say that I need an accurate model, I'll say, like, 10 variables is—again, just like a rule of thumb. I would at least need 100 examples, 100 data point, 100 different use cases because there are 10. But more, the better. But at least give me 100, which is 10 times the number of variables we decided on. Maybe they'll say that I don't have time for that. I'll give you 50 examples. Can you cut down to five variables?
(Sri at 00:46:48) Well, we can cut down, but again, more variables are important. Right? More data is better to build a strong model. As more and more variables we get, we need to increase the training data size too. So it's kind of like that equation. So especially when we do models where we deal with text data. So text data, you can imagine, every word of a text is kind of a variable. So we have tens of thousands of variables for every model that we build. So we need thousands of training examples to train something. And large corporations, businesses like Google, Facebook, they have tens of thousands of examples to train something.
(Sri at 00:47:34) So we try to get to that and sometimes it is a lot of work to give us these training examples. So we try to build a prototype model, put it into action. It might not be very accurate, but it shows something and they keep correcting it. We call that a feedback system. And the model keeps on training and after a couple of months or sometimes it might even take six months. The model took about four months to train itself and reach about 90% accuracy. So then after, you know, we started building that particular model with, I think, 200 examples, sort of a ballpark. And then it kept on giving the results and the accuracy of that model was around 65, 70%. And they're saying that every morning, it used to give them the results and they used to say yes, no, yes, no. Give a feedback for three months, four months.
(Sri at 00:48:32) And when it reached 4,000 examples, the model reached 92% or around 92% accuracy. So, you know, so it took that long, that many examples to reach it. And the rule of thumb is, yeah, for the text models, you need thousands of training examples for something like a simple model, like something like a finance or credit, maybe not that much, depending on how many variables you have. And also for the image recognition models, you need many images because every pixel is kind of a data point. That's why you're also dealing with a lot of data points there.
(Joel Beasley at 00:49:09) Interesting. So that's why subject matter expert is so important because when you start talking about you need these variables as you were describing the housing price with the median income and the crime rates. It sounds like the first step is to even know if the variables you're picking are correlated to the outcome you want. So that would come from the subject matter expert. So it's like an entire process if you don't have the subject matter expert to even test and try to figure out what the variables are that are going to correlate.
(Sri at 00:49:40) Yeah. That's exactly right. You nailed it. That's why we need to work closely with the business—I mean, sorry, the subject matter experts or domain experts—in order to understand that. Right? I don't know much about real estate. I might guess one or two, but later, I need somebody to come with, these are the variables you should be looking at. Not like they need to be very confident about it. At least they could have some hypothesis saying that maybe you should be looking at that. And then it becomes our job to test them and test for correlation or redundancy and, you know, come back with some descriptive, exploratory statistics in order to have a second conversation saying that, hey, you gave me a list of these 100 variables, and out of that, I think that 60 seems to be really working well. And take it from there. So talking about that, there are two kinds of modeling approaches, like inductive modeling and deductive modeling. So inductive is something, you know, as I said, we have something and we try to see what is possible or what is not. And so yeah. So we try to choose—deductive is something like, okay, we try to build something based on what is available out there.
(Joel Beasley at 00:50:59) And so when you're building a new model, we can go back to the housing one again because that's just so easy for everyone to understand. Let's say that you're building this, and it seems like there's two phases. The first phase is you have—I don't know, really. So you're gonna have to help me with this one. Let's say at the beginning when you've got this test data that you know the outcome of it. Right? And so you're going to try to get your system to just reach the same conclusion that you already know is true.
(Sri at 00:51:33) Mhmm.
(Joel Beasley at 00:51:33) And then you do that for a certain point, period of time until you get to a certain confidence level or correctness level, and then you introduce the assistive training concept, the feedback training concept, or is it all one together?
(Sri at 00:51:48) Yeah, that's a good question. So what I would prefer is we collect some training data—let's say a hundred examples—and build something, and then we show the results to whoever it is and say how accurate they are and try to build a feedback loop. That way, this is known as a greedy approach. What I'm trying to do is I'm trying to save your time by not demanding a lot of examples.
(Sri at 00:52:17) Right? I'm just saying, okay, let's start with minimum and see. But for this, you need to be judgment-free and able to trust me, right? Because the initial results might not be super awesome.
(Sri at 00:52:30) The other way is, I think, oh, I don't care. Just give me a thousand examples. Right? I'm on the safe side and saying that you go do your work, give me a thousand examples, get back to me, and then I'll build a model. And then obviously, then I'm also kind of obligated to give the best results on the first shot.
(Sri at 00:52:50) Whereas this first approach, though it takes some time, I feel is organic. It helps both the business, the expert, and us to understand where the gap is, where the improvements are, because I feel these models are kind of in the process of evolution rather than, you know, we tried something and it's almost done. Right? So that process also helps us clear our thoughts and gives us opportunity to learn about what is happening or not, and also giving an opportunity for the expert to understand, oh, how the model is functioning. Yesterday I gave these examples, and it seems like it picked up well, you know?
(Sri at 00:53:28) And I think that is something like a really nice organic approach to do it while trying to save time and the amount of work we do, rather than just coming up with a thousand examples. Right? Like, maybe we don't need a thousand. Maybe we reach an accuracy at 450 of the examples. Right?
(Sri at 00:53:47) So a stepwise, greedy approach. At the same time, also understanding how the things are changing, which works actually well at Reorg. And many people learn a lot about data science, how it could work, and how it would act or behave. And I learned many things about the business as well.
(Joel Beasley at 00:54:07) I like it. This is great. I'm learning so much about this. This is what I love being able to have the podcast for. Bring on an expert, ask him my childish questions, and get really great responses. This is awesome. I really appreciate it, man.
(Sri at 00:54:22) Yeah, sure. I'm glad. I'm glad. And that's sort of what the driverless cars are also doing. They keep on going round and round and round for miles and miles, right, to learn and see what is happening. And, you know, see where the mistakes are happening, and I think taking it from there. I think once the thing is fully accurate, right, and then the more the car learns, the car would come out into the market.
(Joel Beasley at 00:54:47) We need some patience. Like, my oldest child is just over three, and she's just starting to put together four or five word sentences. Right? So that's three years of twenty-four-seven training almost. Right? And yet we're frustrated when we have to wait, you know, twenty hours to train a basic model or something like that.
(Sri at 00:55:12) Yeah. Yeah. And still, as my mentor always used to say that all models are wrong, but some models are useful. So, you know, the models are never 100% accurate. So here at our product, our product team works hard in bridging the gap between the model limitations and the business expectation.
(Sri at 00:55:37) And some product data products, there is no room. Maybe the clients just want to see everything right, and the model can only reach, I don't know, a certain percent, and you need to bridge that gap somehow, using mechanical Turks or something like that. So I mean, that's that. You know? Sometimes we have to go the extra mile to cover that gap.
(Sri at 00:55:57) I mean, in weather forecasting models, they'll say that, oh, there's a 92% chance of raining. I mean, if it didn't rain, you might get a little annoyed, but it is what it is. Then in those cases, we can't correct because it's something in real time. But there are some products where we can try to, like, oh, this is a gap. The model couldn't get a couple things right, maybe we should correct it and make it into a whole thing.
(Joel Beasley at 00:56:24) I love it. I love the advancement, and I love how we get frustrated over these things. Because whenever I get frustrated with the computer not working after an update or something, I just think about, like, it's been a hundred and ten years or so since we got electricity. And then everything's okay. That's like one grandparent ago. You know?
(Sri at 00:56:42) It's so true. It's so interesting about human nature though, because that's why I think that's why we're experiencing these exponential advancements. Right? Because soon, you know, the kids are gonna be like, 5G is the slowest thing in the world. I can't stream my terabytes down fast enough.
(Joel Beasley at 00:56:56) Yeah. Yeah.
(Sri at 00:57:00) Sure. One of the things that my mentor just said: when he was growing up, he had this awesome idea, but he didn't have computing power. He had to walk a mile to the computer lab, put the thing, the punch card in, all that, in order to run computation all night. Now we have 10,000 times more of that power, and, you know, we try to use it for Instagram and stuff, which is fun. I'm not complaining, but I think there is a lot of power available there.
(Joel Beasley at 00:57:29) There's also so many—for all the social or the things we kinda deem as not valuable—there are so many amazing things that are happening. Like, if this were thousands of years ago, we'd be saying that they're miracles. Right? Like, with vaccine production, AI to detect cancer cells, all the advancements that we're getting, is just unbelievable. I think, you know, I'm extremely optimistic. The quality of life has just shot up so much. It's just massively increasing. And soon we'll have robots and cobots that will do all the things we don't want to do, and they'll even maintenance themselves.
(Sri at 00:58:11) Yeah. And do it safely, you know? Right. I think that's one of the major points of the driverless car. The safety, I think, which I'm sold on. You know, I think any kind of irrational decisions or misjudgments due to fatigue or human fatigue or some things, I think, could be avoided by these machines because they don't have feelings.
(Joel Beasley at 00:58:36) Yeah. Yet.
(Sri at 00:58:38) Yet. Yeah. Yet. I hope not. I don't know.
(Joel Beasley at 00:58:44) Alexa got a little snippy with me yesterday.
(Sri at 00:58:50) Yeah.
(Joel Beasley at 00:58:52) Thank you so much for listening. And if you found this episode useful, please share it with a friend or colleague who you think would get value from it. And if you have topics that you would like to hear discussed on the podcast, either add me on LinkedIn or send me an email: [email protected]. Every time I get an email or LinkedIn message, it absolutely makes my day and inspires me to keep going.