Episode 569 ·

The Ins and Outs of Machine Learning with Soundar Pandidurai, Data Science Practice Head at Hexaware

Today we’re talking to Soundar Pandidurai, Data Science Practice Head at Hexaware; and we discuss the ways in which machine learning is being used at scale; how far machine learning has come since its inception; and why machine learning needs to be made clean.

All of this right here, right now, on the Modern CTO Podcast! 

Check out more of Soundar and Hexaware at https://hexaware.com/!

About Soundar Pandidurai:

A multi-skilled AI/ML solutions architect possessing around 27 years of Industry experience with a good all-round ability to juggle multiple projects and meet deadlines whilst at the same time comprehending complex and independent business processes.

Very capable with an ability to identify and deal with clients’ needs by translating them into appropriate technical solutions.

Experienced in providing motivation, guidance, and an up-to-date consultancy service to both colleagues and clients.

About Hexaware:

Hexaware is an automation-led next-generation service provider delivering excellence in IT, BPS and Consulting services. We are driven by a combination of robust strategies, passionate teams and a global culture rooted in innovation and automation. Hexaware’s digital offerings have helped clients achieve operational excellence and customer delight. Our focus lies on taking a leadership position in helping clients attain customer intimacy as their competitive advantage. We are on a journey of metamorphosing the experiences of the customer’s customers by leveraging our industry-leading delivery and execution model, built around the strategy— ‘Automate Everything®, Cloudify Everything®, Transform Customer Experiences®’. Powering Hexaware’s complex technology solutions and services is the Bottom-Up Disruption, a disruptive crowdsourcing initiative that brings about innovation and improvement to everyday complexities and, ultimately, growing the clients’ business. The digitally empowered, diverse and inclusive workforce of Hexaware represents various nationalities, comprising 24,166 employees, and thoroughly lives the company’s philosophy of ‘customer success, first and always’.

Transcript

(Intro Narrator at 00:00:01) Today we're talking to Soundar, Data Science Practice Head at Hexaware, about machine learning at scale. You're listening to the Modern CTO podcast.

(Joel Beasley at 00:00:14) So tell me a little bit about what's your specialty?

(Soundar at 00:00:17) Data science is my specialty. In fact, I started with SPC—it's called Statistical Process Control. Core manufacturing industries, they love it. And I'm sure you've heard about Six Sigma, Nine Sigma, Twelve Sigma, right, and multiple millions of tries. So that is the kind of quality and perfection you can achieve through those techniques. I started my career with that, then slowly moved on, and now I'm with a writing company almost doing the same thing and reducing the error through statistics.

(Joel Beasley at 00:00:52) Nice. And so that lends itself to machine learning, right? What do you get to do with it?

(Soundar at 00:00:57) Right, and let's say it's machine learning. Right? Data science is another word. People call it with a lot of names—machine learning, data science, artificial intelligence. Right? If we make the machine learn what we do or what we can't do, that is, simply put, machine learning. And basically, it helps us—I mean, multiple human brains together—and bring up the insights which are not known earlier to the business and takes out unknown insights which are critical to the business. So that is the power of the machine learning and the statistical algorithms can bring up. I'm really passionate to work with those statistical algorithms and bring up useful insights to the business so that it is beneficial to the business and to my organization as well.

(Joel Beasley at 00:01:48) So it's more of a consulting-type engagement where you work with the business to find insights by using machine learning.

(Soundar at 00:01:55) Right, right. And my portfolio starts with consulting, and I touch on most all the phases, including project implementation.

(Joel Beasley at 00:02:03) What are some of the cooler projects that you've gotten to work on?

(Soundar at 00:02:06) Right, and very niche area. So you might be aware of this particular area, which is semiconductor manufacturing, right? And sometime back in the COVID, everybody was looking at, "Oh, okay, what happened to the semiconductors or IC chips?" Some blockades. Cars are not getting delivered. Scooters are not getting delivered. Correct? And we deal with such companies as well, and we are building machine learning models for them as well. And one particular use case I'm going to talk about in detail. So that industry is very niche and very, very sensitive, you know. Right? You know about the data sensitivity—normal human beings, they care for their own data—but these are all high-cost and highly sensitive. Effective, and no other person should know. And can you imagine one person who is working in that particular company fully on-premise and no cloud technology? You can imagine, without cloud technology, building a machine learning model at scale is always difficult. We have done it. We tasted success and proved that we can do it—multiple machine learning models without cloud technologies then.

(Joel Beasley at 00:03:22) Why is it so difficult to do it at scale?

(Soundar at 00:03:25) Because when we talk about cloud technologies, right, all the hyperscalers—right, like Google, Azure, and AWS—they have the bells and whistles, right? When you start a new machine learning model, you can pick up, say, a template and build a model easily. You can use drag-and-drop features and build a model easily. You can scale it because the technology will allow you to just include multiple nodes. One node you use, you can just use multiple nodes, just drag and drop multiple nodes and add it. That means scaling. I think that is quite possible, easy, because of the services they have, the APIs they have. The ML model at scale is relatively easy when it comes to hyperscalers. But when it comes to on-premise, those things are not there, right? And the drag-and-drop features won't be there, and the APIs and services, right? Suddenly, if I want to multiply multiple nodes, I don't think that's quite possible, you see. And we did it on-premise. That's a specialty of Hexaware, I would say, in this machine learning area.

(Joel Beasley at 00:04:39) So that's one of the ways you specialize, is with scaling on-prem models?

(Soundar at 00:04:43) Right, right. And once anybody specialized scaling ML model on-premise, it is easy. It is a cakewalk when it comes to, you know, using a hyperscaler like Azure, AWS, because things are going to be relatively easy. Many services are just pull it and drop it.

(Joel Beasley at 00:05:01) Is there a benefit in cost to run these models on-prem versus in the cloud?

(Soundar at 00:05:08) More than comparing the cost, it's a business criticality, right, and not allowing my data to be seen to the outside world or making my data vulnerable to keeping it in the cloud. That's how our clients are seeing, I mean, on-premise. We do have clients where we have implemented cloud technologies, ML at scale, but that data is not that critical. But they allow the model to be posted on the cloud where the ML at scale is possible.

(Joel Beasley at 00:05:44) So when we talk about "at scale," maybe you can correct me. Is it relative to each company, or is there a specific threshold where there's a certain amount of data you're processing that means "at scale," or there's a certain part in your growth process that means you are distinctly at scale? Or is it just more of, like, a general word to describe increasing the load on the process?

(Soundar at 00:06:09) Okay, so let me talk from the evolution of the ML, right? And ML was a paper, a research paper, and not able to implement. Then came the computing power, and people started using computing power and started building the model. Deep learning models were still in the papers. Now we are able to use deep learning models because of the computing power, right? And now it's a reality that everybody started using it. So that's the evolution, right? And my point is, ML at scale was there from the beginning. Now it is prominent because of its benefit. So the feasibility—now it is prominent. I can give you maybe a simple example. Let us take customer churn prediction. Very famous, well-known, proven problem statement—customer churn prediction. Say, for an insurance company, what this model does: this model predicts, "Oh, you have one million customers. Out of one million customers, 15,000 customers are going to cut the relationship." That is the prediction. Then their marketing team can act upon it. But before even they cut the relationship, the model will say, "These are all the potential customers. They might leave the relationship." That's a good one, right? Assume that it is giving more than 90% accuracy. That's fantastic. Same customers of that same company—automobile insurance customers, life insurance customers—they may churn in a different way. Yeah, right? If it is statistically proven, essentially, we should have two different models.

(Joel Beasley at 00:07:48) Absolutely.

(Soundar at 00:07:49) One individual model, machine learning model, has become two now. One insurance company not just having just two products—they have five products, six products, right? Automobile insurance, life insurance, health insurance, right? A lot of other property insurance. Say, for example, five different products. That means five different ML models. So one model became two. Two has become five. Now think of the geography. The people—health insurance in the UK—they may act differently than the Euro people. Geographically, that geo and this geo may behave differently. Then split it. Now one can understand the ML at scale is not a concept. It's a need, and it is possible. It is evident, prominent. It is feasible because of the technologies that we have around us.

(Joel Beasley at 00:08:48) No, that makes sense, because it's only going to grow more and more. I know you do the consulting part, but I saw that you had, I think it's a tool called AMAZE. So you make tools, but you also do consulting as well?

(Soundar at 00:09:01) Right, right. And the machine learning model, right? And we use hyperscalers, totally automated, right? Machine learning model building itself is automated. And hyperscalers—you know about all hyperscalers? Yes. So that is a cool way of building models. The point is, when we use such tools and platform, it is easy and quick to build a model and see the output. And instead of building one model in three months, same use case, we build 12 times the same model. That means 12 different iterations, same model, so that you change some idea, you build it again, change another idea, give another input, give another idea, iterate, right? People normally take three months and four months to build one single model. So same problem, but use multiple inputs, multiple ideas, iterate. That is a possibility now, which is happening.

(Joel Beasley at 00:10:02) What's the most difficult thing about machine learning?

(Soundar at 00:10:07) It's about the process, right? And of course, it needs statistical background to help us. So that is the critical thing, and people use this drag-and-drop feature to build up a machine learning model. They can still produce—unless otherwise we know the nuances of, say, a drop box. You just pull a drop box. The drop box, say, it does the hypothesis testing. Unless otherwise we know what is the exact hypothesis testing, right? We can see the results and get the inference, but it is a knowledge of statistics that brings data scientists special. That is the most difficult thing. That's the most difficult thing. Statisticians becoming familiarized with the technology and that particular growth—I mean, statistics plus technology plus domain knowledge, or somebody else has to bring in the domain knowledge. So these three things together will make a good machine learning model. If these three things are not aligned properly, whatever be the technology, I don't think there will be a good machine learning model output.

(Joel Beasley at 00:11:18) Yeah, no, I agree. What's the thing that you're most excited about right now with this technology?

(Soundar at 00:11:24) Right, and people talk about ML at scale and at pace and at scale. Can you imagine 210 models for the electronic manufacturing equipment manufacturing company, right? But the time we have taken to build one model—half a day, I mean, full building, testing, validating, pushing it at scale and at pace. If this is not there, 210 models is not possible. It will take years. The technologies which we have around us is really enabling us, right? And I did mention some of the technologies which does automation of ML right from model building to model deployment. For example, DataRobot, DataIQ, H2O.ai, and so on.

(Joel Beasley at 00:12:19) When the AI ultimately—like, Elon Musk, right? He talks about people that don't understand technology are scared about the future of AI and robots and things like that. What are your thoughts? Are you scared, or no?

(Soundar at 00:12:38) I am part of that community, unless otherwise AI is leveraged for good purpose. It might destroy. So that's the reason even we have—so when we put our own thought process, what kind of AI ML model we have to build and give to the world, it should be clean, unfair, unbiased AI and ML model. It should be clean. So that is our thought process, right? That is also a recommendation from Microsoft. One of the predictive maintenance models, one of the predictive maintenance models in the manufacturing company, it pointed out specific supplier's raw material as a faulty one.

(Joel Beasley at 00:13:20) Why?

(Soundar at 00:13:21) Because of some influence. So basically, that particular model was biased.

(Joel Beasley at 00:13:27) How do you make an unbiased model? I mean, isn't the whole point to have some sort of bias to detect a pattern? How do you make a pattern detection machine unbiased?

(Soundar at 00:13:37) Correct. As a data scientist, if the data scientist is super knowledgeable, that person can make the model biased. You can just give wrong input to the model and make the model give the output, the kind of output he wants. That is a fear people have, no? Yeah. So that's happening in the industry, I'm saying, right? And at the humanity level, and that's different. And Elon Musk is talking about that.

(Joel Beasley at 00:14:05) What do you think we should do about it?

(Soundar at 00:14:07) Everybody should be responsible, right? And when I give an output, whatever be the output, my output is always an ML model. I am responsible for my output. That should be clean, fair, responsible, trustworthy.

(Joel Beasley at 00:14:23) Are you teaching your kids how to do this type of technology? Are they interested in it? Are they not interested in it?

(Soundar at 00:14:31) I let my kids decide the future on their own. And looking at me, I think they are not fond of my way of working and career. They are in a totally different area. Oh. They want to do construction, architecturing, landscaping.

(Joel Beasley at 00:14:51) That's very cool. So, I mean, look, I talk to companies all the time that build massive softwares in those spaces, right? But they're interested in, like, building really big buildings, like skyscrapers and architecture.

(Soundar at 00:15:04) Right.

(Joel Beasley at 00:15:05) That's pretty neat.

(Soundar at 00:15:06) Yeah. So my descendants—they're not going to be in IT. Yeah.

(Joel Beasley at 00:15:09) Well, I hate to let them know that they're going to end up being in IT somehow, because the softwares you use to build those buildings are fairly detailed. Yeah.

(Soundar at 00:15:21) Everybody will end up surrounded by IT everywhere.

(Joel Beasley at 00:15:25) Yeah, I know. Tell me a little bit about how people can learn more about Hexaware. Maybe they have machine learning models that they want to grow. They want to do more machine learning within their company and scale it up. How do people learn more?

(Soundar at 00:15:41) Right. We have a community. It's called Data Science Community. And this community comprised of statisticians, visualization specialists, domain specialists, everybody. That's the reason we call it as a community. They talk to each other, right? They talk about the data science problems inside and outside Hexaware, right? Problems and solutions and new techniques, new areas from outside. Anybody who wants to know about Hexaware Data Science Practice, then we do have a nice website that will talk about what we do and how responsible we are in terms of AI and ML.

(Joel Beasley at 00:16:20) And then is that just hexaware.com?

(Soundar at 00:16:22) Right. Very cool. So we'll put a link in the show notes so that people can get there. Do you have any other topics that we want to discuss or calls to action that we want people to do?

(Soundar at 00:16:32) Right. One topic and one call to action, right? It is beyond ML at scale, right? Every trending topic will go outdated. What are the next trending, right? It is beyond ML at scale, and people are interested now, and they will be interested in deploying the ML model, right? And making it available. Making it available to the end user. People are excited. People were excited at the beginning. If they see an ML model in their own laptop, satisfied. But that is not the satisfaction, right? It should reach the end user. End users should use it and provide feedback. They should feel happy about it, right? That is happening at scale, at pace, right? And that is going to be there at least for the next two, three years. And, I mean, without that, that's going to be the basic for any ML, right? It is not just ML model building. It is also ML model deployment, monitoring, and continuously training it, right? That is going to be there. And that is ML model deployment, monitoring, retraining, and governance.

(Joel Beasley at 00:17:44) And do you build software for this, or do you just organize existing software projects?

(Soundar at 00:17:50) We leverage the existing softwares for this, because Azure has its own stack and AWS has its own stack, Google as well. And there are some niche players like DataRobot, DataIQ, H2O.ai. They all have niche softwares.

(Joel Beasley at 00:18:10) Very cool. All right, what didn't we cover today that we want to get out there to the world?

(Soundar at 00:18:15) Right. And this is the same message, and I would like to tell who were belonging to this community, data science or machine learning, AI ML community, be responsible. Yes. Let us produce output that is fair, trustworthy. That is the recommendation and request from my side.

(Joel Beasley at 00:18:38) It's demand from my side. This has been great, Soundar. Thank you so much, man. We made a podcast. How do you feel?

(Soundar at 00:18:47) Feeling great and fantastic.

(Joel Beasley at 00:18:50) Thank you so much for listening. And if you found this episode useful, please share it with a friend or colleague who you think would get value from it. And if you have topics that you'd like to hear discussed on the podcast, either add me on LinkedIn or send me an email, [email protected]. Every time I get an email or LinkedIn message, it absolutely makes my day and inspires me to keep going.