Episode 368 ·

Steven Davis - Avoiding AI Echo Chambers, & Enabling Data Scientists to Do More

Today we’re talking to Steven Davis, the VP of Engineering at Innodata. And we discuss how Innodata enables data scientists to spend more time on data analysis rather than data extraction. Avoiding bias and echo chambers in AI, and tips for meeting the needs of a diverse team as a manager. 

All of this, right here, right now, on the Modern CTO Podcast!

Check out our other episode with Innodata's Chief Product and Marketing Officer, Rahul Singhal

To learn more about Innodata, check them out at https://innodata.com

About Steven Davis:

Steve is the VP of Engineering at Innodata, where he leads the development of AI/ML platforms and products, with a focus on Data for AI and Applied AI geared towards solving customer needs in innovative, effective, and scalable ways. He has significant experience designing and building complex, distributed, big data processing and analytics systems, including cloud ETL architectures, AI/ML services, and web applications for enterprise search, collaboration, and analysis. As a founder of multiple software startups, he has designed and implemented corporate organizational structures, personnel management procedures, and leadership best practices.

Prior to Innodata and data analytics startups, Steve worked at a Federally Funded Research and Development Center (FFRDC) and directly supported an elite U.S. Special Operations Command innovation and fusion cell. In this capacity, he worked directly with all-source, human terrain, geospatial, and other specialized intelligence analysts to rapidly develop tools and techniques to integrate and exploit large volumes of data.

His hallmark in this role was the ability to flexibly develop new technologies to support warfighters in extreme circumstances, deploying new technologies on a vastly accelerated time schedule. During his time with Special Operations, he deployed to Iraq and Afghanistan and received awards from the Defense Intelligence Agency and the National Geospatial-Intelligence Agency for his efforts to integrate systems, implement novel solutions, and support deployed forces.

Steve received a BS with highest distinction in Systems and Information Engineering, and a Master of Science (MS) in Systems and Information Engineering at the University of Virginia.

About Innodata:

Innodata (NASDAQ: INOD) is the world’s leading data engineering company. Prestigious companies across the globe turn to Innodata for help with their biggest data challenges. By combining advanced ML/AI technologies, a global workforce of 3,500 subject matter experts, and a high-security infrastructure, we’re helping usher in the promise of digital data and ubiquitous AI. Our culture of innovation, quality, and service is present in everything we do. We serve publishers, media & information companies, digital retailers, banks, insurance companies, government agencies and many other industries. We take a technology-first approach, applying the most advanced technologies in innovative ways. Founded in 1988, we comprise a team of 5,000 diverse people in 8 countries who are fiercely dedicated to delivering services and solutions that help the world make better decisions.

Transcript

(Joel Beasley at 00:00:03) Hello, my friends. Today we're talking to Steven, the VP of Engineering at Innodata. And we discuss how Innodata enables data scientists to spend more time on data analysis rather than data extraction, avoiding bias and echo chambers in AI, and tips for meeting the needs of a diverse team as a manager. All of this right here, right now on the Modern CTO Podcast.

(Joel Beasley at 00:00:32) Here we go. This is the Modern CTO Podcast.

(Joel Beasley at 00:00:44) Yeah. Could you tell me a little bit about yourself? How'd you get into technology?

(Steven at 00:00:48) So I have always been into computer-related things, trying to learn how to program on my own as a kid, pre-college. And in college, I went to University of Virginia and focused on systems and information engineering for undergrad and grad school, which is a combination of operations research and kind of management consulting and various other things. And I minored in computer science because I still wanted that really hands-on application experience. And I went into a lot of government work with a company called MITRE, which runs a number of federally funded research and development centers. It's a really interesting company. If you ever want to check out MITRE, they do a lot of really fascinating work for the government, and they're a nonprofit entity. So they're kind of this strategic adviser implementation, all kinds of things for the government. So that was the springboard for learning a lot of things at the same time right out of school. I worked for most of the intel and defense agencies in the DC area, and I got a lot of really interesting experience with data and with coming up with platforms and tools to help analysts, to facilitate taking lots of disparate kinds of data and fusing it together to help answer questions. I became the chief of technology for an intel fusion cell.

(Joel Beasley at 00:02:39) What's the intel fusion cell?

(Steven at 00:02:41) So it's a place where you have lots of different kinds of analysts and lots of different kinds of data come together in one place to work together on a specific kind of problem. So imagine expertise in human intelligence, understanding intelligence reports, working with somebody that is an expert in imagery analysis, and working with experts in other domains in the intel and defense world, and even people like myself, technologists that kind of specialize in creating software and manipulating data and helping facilitate the kinds of analytic products that they need to get out. So I would do things like this range of, if we see somebody click a button 30 times in one day, you know, that's something we need to fix. So we'll write a tool to help them so they only click that button one time a day. So they don't have to focus on that button anymore. They can focus on the hard work. And anything from as simple as that to, I know, maybe 15 or so years ago, we were extracting all of the nouns and verbs from every piece of intelligence that we could get our hands on, and we were extracting billions of records and creating tools to help analysts search and query and identify, you know, those proverbial needles in the haystack.

(Joel Beasley at 00:04:19) Yeah.

(Steven at 00:04:20) So that was a lot of really interesting, varied experience, and it took me to Afghanistan a number of times, building tools out in Afghanistan and traveling around training analysts on how to use those tools and how to take advantage of it. I did the same thing in Iraq for a very short period of time. So that's kind of what set me off in the technology direction.

(Joel Beasley at 00:04:49) That's really interesting. A little while ago, we had on a company called IMO. We had their COO, Ivana, on. And what they do is they take medical terminology, and they have, like, these subject matter experts, like medical doctors, just, like, pouring through textbooks and pulling out these terms. And they put the terminology into datasets and deploy that to hospitals so that their systems can understand the notes that the doctors take during their visits with patients and stuff. And that helps with, like, creating automated billing processes and stuff. But it was really interesting to me learning that there's so much work that has to go into just deploying literal words to, for systems to understand the words. And you talking about having to pull all those nouns and verbs for the government work sounds pretty similar. And it's crazy that I just feel like I went so long not hearing about that at all, and now here we are twice.

(Steven at 00:05:58) Yeah. No, no. I mean, you're absolutely right. It really comes back down to that. And probably the reason why people don't talk about it as much is because it has become kind of the status quo and people are focused on higher level activities. Like, as a data scientist, how do I focus on building my neural network and using that to automatically extract data, which is, you know, the ultimate goal. But I think a lot of times people forget how important that data preparation stage is, which is kind of the building blocks of all of the really fancy AI ML stuff that we do today. Whether it's building on top of that named entity recognition, extracting those nouns, verbs, or, you know, people, places, things, or whatever you're interested in, or something more complex. Those are really all the stakes in the ground. And that's actually, at Innodata, the company I'm a part of now. One of the major platforms that we're building is a data for AI platform that focuses on how you get that training data. So it kind of comes full circle to helping the analysts, you know, 15 years ago, who were spending 80% of their time manipulating data and 20% of their time doing the analysis. That's been kind of a common thread throughout my career, whether it's helping analysts in the intel and defense world or helping automotive companies identify emerging issues, or, and now at Innodata, helping data scientists focus on training their models. So flipping that 80-20 so that they're not spending 80% of their time tagging and identifying individual words so they can train their model in only 20% of their time refining things, but rather, really focus on expediting, facilitating that high quality training data so that you can spend more time on that higher level AI ML that we hear about so much more these days.

(Joel Beasley at 00:08:21) That makes a lot of sense. So what are some major differences that you've encountered? Sounds like you've been trying to solve the same problem while serving the government sector and serving the commercial sector. What have been some major differences between the two?

(Steven at 00:08:38) Yeah. That's a really interesting question because I don't know there aren't significant differences from a technology perspective. There are obviously significant differences in the application and what you're using it for and maybe the types of information. But when it comes down to it, the type of technology that you would use today to analyze an intelligence report is very similar, you know, under the hood to something that you would use to extract relevant information from an insurance claim or an earnings report. A lot of those technologies share a lot of things in common. I think, really, when it comes down to it, the differences are really just about the data. So the availability of the data, the quality of the data, the capacity to turn that into something that is useful for training a more sophisticated model, and to prepare it. So a lot, most data is very unclean and needs filtering. It's inconsistent, and so you need a lot of tools to help with that. So I think, you know, when I think about platforms, for data engineers like myself, I think one of the last things that we think of is design and user experience. And so I think that is something that is going to change, that is changing now, and I think we'll see a much larger focus on design and user experience around that data preparation stage, that data analytics stage, really the whole data lifecycle. I think we're going to see more focus on what it means to users because at the end of the day, you can eventually use machine learning to bootstrap and self-learn, essentially, but you still need humans to kick that off, and you need humans in the loop to manage data drift over time.

(Joel Beasley at 00:10:57) Right.

(Steven at 00:10:57) So, yeah, I think one of the bigger differences kind of going back to your question is in the ways that you might interact with systems might be a little bit different between applications, you know, if that's government versus commercial. But, really, I think that the biggest difference we're going to see in the next few years is a focus on what that means for end users.

(Joel Beasley at 00:11:21) That makes sense. Yeah. So before we get into all the nitty gritty of cleaning data and training AI algorithms, let's take a step back. Because a little while ago, we had on your Chief Product and Marketing Officer, Raul. But for those that didn't listen to that episode, can you give me, like, a brief overview of what Innodata does?

(Steven at 00:11:42) Oh, yeah. Absolutely. Innodata is a really fascinating company. So Innodata has been around for about 30 years now, publicly traded on the NASDAQ. Historically, kind of rooted in publishing, the publishing domain, with things like eBooks. And I think, you know, the company evolved by taking a lot of that kind of data engineering experience and expanding into scientific journals, legal journals, medical journals, and over time, built up this impressive team of subject matter experts. And today, Innodata has something north of 3,500 subject matter experts on staff that are, including doctors and lawyers that help not just in that publishing space anymore, but also in all sorts of other industries and verticals to apply that subject matter expertise to data extraction, data annotation, and various kinds of processing. And in recent years, Innodata has taken significant leaps forward in machine learning and artificial intelligence to augment and supplement the good work that the subject matter experts do. So there's a number of products that Innodata offers in addition to the managed services for data processing, automated platforms for data extraction, extracting data points according to a set taxonomy, for example, in the financial or the legal spaces, and also the data annotation platform, which is one of the key platforms that I focus on in Innodata, which really takes into consideration all the lessons learned over the history of the company and the subject matter experts and how we build our own models, leveraging all that experience and turning it into a customer facing SaaS self-service platform that helps facilitate that data from end to end. So from the raw data to powering an actual machine learning model and applying that to a real problem.

(Joel Beasley at 00:14:20) So what does it look like using the data annotation platform in practice? Like, what's the input and what's the output?

(Steven at 00:14:27) Yeah. It's a good question. So it's actually a number of things, depending on the use case. So the way I talk about the annotation platform is workbenches, workflows, and KPIs, which kind of encompass the ecosystem of things that you need in a more end-to-end self-service platform where somebody even outside of Innodata can sign up and click and start their own process to facilitate getting to that final product. And so when you first fire up the platform, you'll configure a new project, and you'll select a workbench that's based on a use case. So if you think of workbenches as my collection of tools that I might use for a particular kind of problem. If I'm categorizing credit card transactions, then I need a workbench that shows me records, and I need certain tools to help me get through a list of records, something that you might be familiar with seeing in a spreadsheet often, tools to help you get through labeling that in an efficient and high quality way. That might be one kind of workbench. Another kind of workbench might be more geared towards annotating entities in line within a document. So any reference you see to Steve or Adam in a document you might want to highlight in relation to an event and connect all of those things visually.

(Joel Beasley at 00:16:06) So to dumb it down a little bit, you're talking about relating all the records in a spreadsheet to, like, a tool for helping you sort through those. That tool could be something like when you're working in a spreadsheet, you have, like, an autocomplete that you can just drag the cell down, and it'll fill in the next line of data based off of what you've been putting in for each record. Is that a good—

(Steven at 00:16:32) No. That is spot on.

(Joel Beasley at 00:16:39) Cool. Okay.

(Steven at 00:16:39) It's kind of the hidden side of data science. It's really preparing all of this data so that you have a set of clean labels that you have high confidence in so that when you go and train your model to automatically recognize the next credit card transaction and see that even though it's clearly a gas station, to identify whether that's actually a gas purchase or if it's more of a convenience store. Maybe the person bought a candy bar instead. So there's a lot of factors that you might take into account. You can only take those into account if you've had a human take a look at that first and make a decision on how to treat those types of things.

(Joel Beasley at 00:17:18) That makes sense. Yeah. Cool. So what are some of the more interesting use cases you've seen for the platform? Just, like, stuff that's kind of out there.

(Steven at 00:17:27) Yeah. So we do a lot of text-based annotation. There's a lot of expertise in the company around the legal, the healthcare domains, the financial services domains. We've also done a lot of work in image and video annotation, which I think is fascinating. I have built a lot of geospatial analytic platforms in the past, so assigning latitude, longitude coordinates to points on a map. But image annotation is a little bit different from that. And I think image annotation is a fascinating world where it ranges anywhere from looking at a shopping cart where you might have occluded items. So you might have a box of cereal, but sitting on top of it is a jar of peanut butter, and you need to be able to account for that in your image recognition model. So those are really interesting use cases and whether the annotator is tasked to outline the full occluded, you know, hidden box of cereal or just the visible portion. Those are all kinds of things that have different implications for what you're trying to do. And as you can imagine, those kinds of decisions are much more critical when you start thinking about autonomous vehicles and recognizing things on the road reliably.

(Joel Beasley at 00:18:58) Yeah. That makes sense. That's life or death. But that actually makes me think of, like, a really specific use case of AI that I was listening to this other podcast called HPE Tech Talk. It's a tech podcast put out by Hewlett Packard. And they were interviewing someone from the Walt Disney Studio Lab about how they're applying AI in the filmmaking process. Because at the end of creating a movie, they have these quality control experts that go through and check out the entire film frame by frame and look for pixel level anomalies on each individual frame. Like, that's the level of detail they get into on this quality control. And they used to have people that were responsible for finding a single miscolored pixel on a single frame of a movie. But now they're able to train their AI, presumably, with some really high quality image recognition that's able to find these anomalies. And just like you were talking about flipping the 80-20 ratio there, now the quality control experts are able to spend a lot more of their time focusing on what to do about the anomalies. Is it worth fixing? And if it is worth fixing, how do you fix it? Because the movie's done. Right. It's—

(Steven at 00:20:25) That's fascinating. I wonder if that can help resolve the coffee cup in the Game of Thrones issue.

(Joel Beasley at 00:20:35) I don't

(Steven at 00:20:35) know if you're

(Joel Beasley at 00:20:36) I'm not familiar. Can you tell me?

(Steven at 00:20:37) There's a big article. I guess it's been a few years. I have no sense of time these days. But with Game of Thrones, there was apparently, I think it was maybe a Starbucks cup that was caught in one of the scenes, I think, kind of a climactic battle scene. And I don't know if they had to go back and reshoot or if they just left it in for posterity, but yeah, that sounds like a really interesting kind of problem.

(Joel Beasley at 00:21:09) Yeah, it sounds like an anomaly that should have been picked up. Yeah, but how do you train an AI to recognize that that doesn't belong in a medieval atmosphere?

(Steven at 00:21:24) Yeah, I don't know. I guess you'd have to classify objects to such detail that you can recognize the objects in various frames. But I think we're probably a long way away from thinking through that unintended consequence. Yeah, we were able to recognize the coffee cup 100%. That was easy, but we forgot to note that it doesn't belong in a medieval fantastical world.

(Joel Beasley at 00:22:00) Yeah. I guess that level of quality control is still up to the humans today. So have you seen any interesting stories of data labeling gone wrong?

(Steven at 00:22:14) You know, there's nothing in particular that I can think of except just kind of the general theme of bias in data, which I think is important. I think bias has always been a challenge for machine learning models. I think we'll continue to see a lot of focus on that as we get more into things like synthetic data generation.

(Joel Beasley at 00:22:43) Right.

(Steven at 00:22:44) So that's one of the things that we're working on at Innodata. As part of the full ecosystem of data, machine learning models, and the application of those, one of the things that we often see is that sometimes there just isn't an adequate level of data available for training. And so I think synthetic data is a really interesting topic. I think the challenges are going to be how to avoid bias and kind of that echo chamber of error that causes lots of bias issues, which have social economic implications. It also just causes issues with data drift over time. If you're training all of your data on a narrow slice of reality, it's going to eventually drift away from that reality as those errors propagate. So that's kind of the importance of continuing to train data over time. It's never done. You always have to account for new things, even for something like if you're trying to recognize events, geographic, sociopolitical events, or medical issues. You know, pre-COVID-19, a machine learning model might not even know what that is. And today, if you were able to continue training your models in whatever domain, you can keep up with things that are unexpected like that.

(Joel Beasley at 00:24:28) So my base level of synthetic data generation is that it's a solution for when you have an incomplete dataset. Right? So a lot of the time, I feel like bias in AI, as you mentioned, just comes from just not having the data, whether it's due to historical inequities or any reason for just not having the right amount of data. The model ends up coming out biased. But how can you tell if you have an incomplete dataset, especially if the model is functioning as you expect it to, and it seems like it's giving you the right answer? How can you tell that the model is biased when your subconscious view of reality might be biased as well?

(Steven at 00:25:20) Yeah, that's—there's an important role there for data scientists and for metrics around evaluating the models. And I wouldn't profess to be an expert in that domain, but I know one of the things that we focus on with our data extraction machine learning technologies is the not just the confidence of our overall extraction, but at a data point level, how confident is the model that it was able to identify this thing correctly? So you might have an earnings report. You might have a legal contract that's 500 pages, and you want to extract the parties and the dates of interest and all sorts of parameters around that complex contract. And at a data point specific level, you can say, well, I'm 99% sure that I got the date of execution correct, but I'm only 50% sure that I got the interest rate that we agreed on correct, because maybe the model wasn't able to have enough examples, enough variance of those examples, or enough context to be able to pull out that specific data point. So that's one of the things that active learning also helps address. Active learning is really about pushing things in reverse. So having the model tell you—prompt the user to say, I need to know more information about this thing. Tell me more about this. So the example is kind of more of a question prompting process where after some initial set of documents, you train your model on initial set of documents, and then the machine says, tell me about X, Y, Z. And it kind of targets those specific questions because it knows that those are the areas where it is least confident.

(Joel Beasley at 00:27:32) That's smart. Yeah. I feel like that's a hallmark of intelligence in a person, like being smart enough to be able to ask for help. And that seems really cool that you're implementing that in computer intelligence. Yeah, that's just a really cool self-awareness.

(Steven at 00:27:51) I hadn't thought of it in that way outside of, you know, the machines becoming self-aware for better or for worse.

(Joel Beasley at 00:28:02) Sounds like that's a use case for better. So earlier, you mentioned—so we talked about the data annotation platform. You also mentioned something about a data extraction platform. Is that, like, the example you just gave of pouring through a legal document and pulling out the data points? Is that what the data extraction platform is?

(Steven at 00:28:24) Exactly. Yeah, that's exactly right. So we take all of that training data that maybe use the annotation platform to identify. You feed that into our machine learning models. They apply that on a per-document basis, and they give you the source transparency. So they tell you exactly what the value of a data point was, how it was extracted from the document, and it points you to the exact location in the document so that as a user, as an SME, as a subject matter expert, you can review that and verify that it's correct. And it gives you those confidence values so that you can have some level of confidence around the model itself. And then you can feed that back in to the overall process. So I think that's also something that we've been thinking a lot more about at Innodata recently, is how all of these things are really connected and how our vision for this interconnected platform, starting with initial data annotation to training and applying models, feeding that back so that annotators have a starting point. So if you imagine not just extracting the data points for a business purpose, but also extracting the data points to help auto-suggest to the next annotator, hey, this is a thing that we think we should be annotating. Do you agree? And that just makes the life of the annotator easier, facilitates the process, and cuts down the time for the annotators as well. So it's really important, just kind of going back to that user experience. We want to cut down that 80% of data manipulation time for data scientists. We also want to do the same for annotators so that they can be more productive, they can generate more high-quality work, which is also important.

(Joel Beasley at 00:30:34) That's really cool. It sounds like a similar advantage that you're giving annotators that CRMs gave salespeople when they first came around. Like, with the advantage of a CRM that has all the notes of every client in there, you can give multiple salespeople the same set of clients, and they can all know what's going on with each of them. Sounds like you can give multiple annotators the same dataset. I mean, yeah, same dataset. They could all work on the same dataset and all be on the same page, thanks to the suggestions and stuff that you're talking about in there.

(Steven at 00:31:13) That's exactly right. And as part of our SaaS platform, you configure—you start with that data workbench, that annotation workbench, and you configure the workflow. So you say, I've got 30 annotators on my team, but I want every document to be seen by at least two annotators. And then you can configure how an arbitrator might operate on disagreements between the annotators. The annotators can leave comments. If there's something that's a little ambiguous or if they want to justify their rationale behind why they made a decision, they can leave those comments behind for the arbitrator to take into account. We do QA on that, and eventually those results get exported and train incrementally, improve machine learning models that then auto-suggest in the system again. So yeah, I think that's a good analogy.

(Joel Beasley at 00:32:11) And this is making me feel better about the future because I know, like, there's a big concern with AI going forward, as it runs more and more facets of society, is making sure that it's explainable. So you can look at it and know exactly why it's making the decision that it's making. And it sounds like you guys are working on building explainability in at every step of the process, starting with the data, which is super important.

(Steven at 00:32:40) Yeah, that's exactly right. So there's a lot of ways to get at that kind of transparency. And one of the important ways to do that is to have a really good handle on how the model arrived there. And if you know what the data was that went into the model, then you can start to investigate if you start to identify any sources of bias.

(Joel Beasley at 00:33:05) That's really cool. So a little while ago, we had on Mark Messina. He's the COO of a company called Geek Plus. And that was a really fun episode because they do robotics in the logistics industry, and he was literally sitting in his office with a glass wall behind him of a bunch of robots running around carrying boxes, which is pretty cool. But so, obviously, in these warehouses that are running mostly off of robotics, they have, like, robots carrying the packages to the humans that are putting them on, like, a rack where they need to go to be shipped out to the end user of the product. Obviously, a lot of it is run by AI. And I was wondering what the differences are between training datasets for an AI that's running, like, software versus an AI that's running a physical, like, robot.

(Steven at 00:34:07) Oh yeah, that's interesting. So there's a lot of overlap because, especially in the world of image and video annotation. So a lot of that annotation is eventually intended to make its way to things like autonomous vehicles, but the same is true for, you know, Boston Dynamics big dog program, I think they call it. I don't remember what they call it. Spot. Yeah. The Boston Dynamics Spot product and other things like that. I think that's another trend that we'll continue to see is beyond just software web kind of ecosystems of machine learning and AI. We'll start to see that creep more into the physical space, which I think is exciting.

(Joel Beasley at 00:35:01) Yeah. I think that's where it gets really cool because it's tangible. Like, you can have YouTube recommend you a video that's, like, exactly what you want to see, and that's really cool and not something that happened ten years ago. But when you can see a robot just walk up and open a door, that's so cool it's scary. Yeah.

(Steven at 00:35:25) Yeah. And even physical devices in the home, like, the Nest thermostat, I think, is really interesting. The focus on user experience there is really fascinating to me. I, in college, I've been haunted by a book that I read in college my entire life. It's called The Design of Everyday Things, and it is mostly about physical object design. It talks about the design of the paper clip, for example, and other really interesting things you probably never would have thought about. It talks about thermostats, like, in your car. If your car is really hot and you want to crank up the AC, people tend to just put it on max rather than set it at, you know, a reasonable temperature. And the question is, does that actually do anything? Is that actually increasing the cool output more than just sending it to the goal temperature? And even, it haunts me with things as simple as doors. So there are certain doors where whether you have a horizontal or a vertical handle, there's an implication. It's telling you that you either need to push or you need to pull. And a lot of times, those doors are wrong. And so whenever I see somebody, a friend or family member, push a door instead of pulling it, and then they're really embarrassed because they did the wrong thing, I'm just like, it's not your fault. It's not your fault. It's bad design. So I always think about that book, and I always think about it in the context of whatever I'm doing, even with building things like the annotation platform or making it so that an analyst doesn't have to click that button 30 times, that they can just click it that one time. Those are all things that can and should be fixed by user experience.

(Joel Beasley at 00:37:36) That makes a lot of sense, especially in the context of AI. Like, when we're starting to implement AI with driving. I mean, now we're seeing a lot of cities open up full self-driving taxis to be available, which is crazy. But I know the kind of the first use cases for most consumer vehicles are going to be, like, recommendations, where it, like, helps guide you with your lane changes or keep you from drifting out of your lane, and cars that have a camera on the driver that can wake you up if you look drowsy. And it's really important to have that level of detail that, you know, you're not putting the equivalent of a vertical handle that you actually need to push instead of pull into your car system.

(Steven at 00:38:36) Right. Yeah. Absolutely. Absolutely. You want to drive people in the right direction.

(Joel Beasley at 00:38:44) Yeah. Literally. Cool. Well, before we get to wrapping up, I want to ask you some leadership questions. Like, I'm sure you lead teams as a VP over there at Innodata. So how would you describe your personal approach to leadership in your role?

(Steven at 00:39:06) Yeah. That's a really interesting question. So I think, you know, there are a lot of models for how to interact with your team as a leader, and I think different management styles and different leadership styles, understanding those and learning about those from books are valuable. But I think one of the things that is often missed is that it doesn't always work for everyone. If you have direct reports or employees, very likely in an appropriately diverse environment, everybody has different needs. And so there's roles for emotional intelligence as a core leadership principle. There are times when even micromanaging is valuable. There are times when delegation and the opposite of micromanagement is valuable. And I think one of the keys to leadership is identifying those and finding out what the right mix is for your team and for individuals, because I think that is what facilitates strong performance and strong cohesion and, ultimately, happiness at the workplace. I've worked with a lot of teams, and I know some people when I put a one-on-one on the calendar, they're thrilled, and they want that to be repeating, you know, every week. And other people are like, Steve, come on. I just talked to you yesterday. We don't need this one-on-one. Let's skip this. So I think it's important to understand what kind of style works for what kinds of individuals.

(Joel Beasley at 00:41:09) That's really good advice. Thank you.

(Steven at 00:41:11) Yeah. I try to have a mix of all of those things. I try to learn from other leaders, other mentors that I have, and implement things that I see. And even within and among the projects that I work on, I probably take on more than I should. So I kind of have this breadth of interaction with things.

(Steven at 00:41:38) And there are certain things that I do tend to micromanage because I know I have a very strong opinion on this button that I want people to click one time. Right? But I absolutely trust the team to do the other 99% their way, you know. And I think there's an important balance there between being too involved or being involved in the right things and just being a support to the overall team, because I think the leader is really there to help support the team, and not the reverse.

(Joel Beasley at 00:42:20) So earlier, we were talking about bias in AI, and it's, I feel like, pretty common knowledge that the best way to combat bias in AI is just to have a really strong, diverse group of humans working on the AI. So what's your approach to making sure your team is diverse and inclusive at Innodata?

(Steven at 00:42:42) Yeah. That's a really good question. I think, you know, I'm really lucky to work at a company that is as global as Innodata is. We have people on our teams all over the world, not just here in the US, but also in Canada, in India, in the Philippines, Sri Lanka, Israel, Germany, Dublin. We have people all over the world.

(Steven at 00:43:08) So I think that kind of geographical diversity is very valuable. When we approach particular problems, we're doing data annotation as a managed service. That kind of geographical expertise and diversity becomes really important. And having the options to tailor a specific challenge, not just to the expertise—so one important aspect of diversity is that diversity in experience.

(Steven at 00:43:42) So we have all of those doctors and all of those lawyers that are helping contribute their experience, but also locally. So for example, if you're trying to annotate credit card transactions and all of your annotators are on the East Coast, there are certain burger joints that they're not gonna recognize. Right?

(Joel Beasley at 00:44:04) Yeah.

(Steven at 00:44:04) So even that kind of diversity geographically, and on a more kind of local scale, is also really important.

(Joel Beasley at 00:44:14) That's really cool. I remember earlier you mentioned that you have, like, over 3,500 subject matter experts employed there. And having them dispersed geographically around the world, I think, is super valuable because now not only can you bring in a subject matter expert to look at the data of a problem, you can bring in multiple experts on the same subject that have different backgrounds to bring just their different views of the same subject that they've all worked and spent a lot of time on. And—

(Steven at 00:44:49) Exactly. Yeah.

(Joel Beasley at 00:44:50) Yeah. That creates, like, a super expert with all of them.

(Steven at 00:44:54) Yeah. It's incredibly valuable for the expertise as a subject matter expert. It's also incredibly valuable from a developer perspective. So we have hybrid teams that are working in the East Coast, and we have teams that are working directly with teams in India on the same kinds of projects. And all of these teams have their own kinds of experience and their own contributions to how they approach a process.

(Steven at 00:45:29) And I think we try to learn from every individual on every team that we have to help improve and evolve our process over time, whether that's engineers developing a system, whether that's the designers, QA, the DevOps, subject matter expertise, and even project managers and other kinds of leadership.

(Joel Beasley at 00:45:56) So if you could go back and talk to yourself towards the beginning of your career, you're an individual contributor moving into your first management role. What's the advice you give them?

(Steven at 00:46:08) Oh, man. That's a tough one. I hadn't thought about that before. I think knowing that there's no right answer is important. So there are a lot of ways to accomplish a problem.

(Steven at 00:46:24) There's similarly a lot of ways to manage a team or provide leadership. Leadership from within a team is just as important as leadership from an official leadership position. I think everybody has that ability to contribute. And whether you're in either of those kinds of roles, just knowing that there's never really a right answer. There's just some answers are better than others.

(Steven at 00:46:57) And everything is a little bit trial and error, and everything is a learning opportunity. So if you apply the wrong leadership style or if you implement the wrong process and it crashes and burns, it's not the end of the world. That's something that's easy to change and evolve over time.

(Joel Beasley at 00:47:24) Today, as you're leading your teams, how do you implement that attitude towards failure to make sure everyone's open to sharing, talking, and learning about it?

(Steven at 00:47:34) Yeah. So it comes back to a balance of process and where the appropriate level of process comes into play. So I think some teams are more naturally communicative and self-aware and feel like they can do those retrospective meetings and talk about it openly and with a goal towards improving the process over time or improving the product over time. Sometimes more process helps. So facilitating communication is the most important step.

(Steven at 00:48:16) It's usually the first step. Daily scrum meetings just to make sure everybody's on the same page and involved is really important. Retrospectives are really important. I think the way that you run the retrospectives, that the teams maybe run their own retrospectives, are really important—kind of the tone that you set and the expectations around what failure means and how improvement happens is really important. And especially today, now that we're, most of the world is much more remote work from home than usual, those kinds of communication channels are more important than ever.

(Joel Beasley at 00:49:04) Awesome, man. Well, before we wrap up, is there anything else we wanna get out there today? Any shout out for Innodata? You guys hiring?

(Steven at 00:49:13) We're always hiring. So, yeah, I would say if you're interested in the data life cycle, if you're interested in helping reduce that 80% time for data scientists or annotators or anyone in the financial services or legal or health care industries be more effective and more efficient. If you're interested in the design process of taking managed services and turning those into platforms and products that we can put in the hands of people outside of the company and run as a self-service tool. I think if any of that resonates, then definitely check out innodata.com.

(Joel Beasley at 00:50:11) Thank you so much for listening. And if you found this episode useful, please share it with a friend or colleague who you think would get value from it. And if you have topics that you'd like to hear discussed on the podcast, either add me on LinkedIn or send me an email [email protected]. Every time I get an email or LinkedIn message, it absolutely makes my day and inspires me to keep going.