Jon: Welcome back to the Execution Over Hype Podcast. I'm your host, John Turns. Today we're kicking off a special three-part series where I'm sitting down with Devanch, a veteran builder whose work across healthcare, finance, government, legal, and startups has given him a front-row seat to both the promise and the pitfalls of enterprise AI. His deep experience in these high-stakes environments led him to become the co-founder of Iris. Where he and his team are tackling the core problems preventing organizations from trusting autonomous systems at scale. Let's get into it. Here with Devanch from Iris AI. I'd like to introduce myself to you, Dev. ⁓ so my name's John Turns. I am vice president of strategy here at Saison. Do a lot of things in go-to-market, obviously everything from like the the marketing and strategy of the business itself to account management, business development, ⁓ contracts and and finance and and a bunch of other things. ⁓ mainly what we do here at Saison though is again have always helped them on the operational side of things, be able to do systems integration, take legacy systems, integrate into kind of a single pane of glass, create that operational excellence that that organizations are looking for in terms of stretching, you know, budget dollars while leaning into kind of the next gen level of of of tech. So we do a lot in artificial intelligence and automation. We also do some things in virtual reality and augmented reality. So that people in field services and manufacturing can kind of see where things are, a lot of sensor fusion and integration stuff. So we really try to give companies access to the latest in technology without them feeling like they've got to take a bunch of off-the-shelf software or commit to long term licenses. ⁓ so really what we're trying to do here, right, ⁓ on this podcast is the difference between AI hype and real execution. I'm sure that you've seen it yourself. ⁓ a lot of times everyone's like, ⁓ I need agents and I need ⁓ you know, this latest greatest thing because that's what's gonna drive my business forward. And, you know, as you I'm sure you've experienced, ⁓ you know, in in in in a lot of the startups that you've helped and in your business journey, a lot of it is hype. There's not a lot of value. So what we try to do here is try to cut through a lot of that marketing hype to make sure that we're providing the most value for enterprises that are on the fence. on what they should do to implement AI or technology in general and and again provide that real ROI for them. And so ⁓ with that, I would like you to to to introduce yourself and really looking forward to our conversation today. Thank you, John. ⁓ Dev here for your audience. ⁓ essentially I have two major roles. one is obviously Iris, the legal AI startup that we're a part of, ⁓ basically just implementing everything and getting it ⁓ done. And along with that, this is the one that's a little bit more publicly facing. I run one of the world's largest open source AI research communities. It's called the Chocolate Milk Cult, C U L T. And essentially ⁓ we're about a million and a half people every month, ⁓ where we bring in a lot of resources from OpenAI, Anthropic, Nvidia, et cetera. And the goal is how do you open source the next generation of research? ⁓ basically. There's a lot of problems at the frontier of tech where one individual company is not going to solve it because there's just so many areas to look. So I am essentially a clorified middleman there where I just talk to a bunch of people ⁓ who are working on these problems. ⁓ I'm myself working on a lot of these problems. And then we try to pool in resources together to try to solve, like, hey, let's try to solve memory, let's try to solve KV cash stuff. ⁓ I do a lot of work with ⁓ inference. How do you solve, make inference cheaper, more staitable reasoning, et cetera? So those are largely the two personas that I I'm quite known for in the world. ⁓ I think that gives me a lot of insight into what you're talking about earlier with regards to AI hype versus failures, just so many misunderstandings that people often have with AI. So excited to talk. Yeah, absolutely. So ⁓ you know Right into ⁓ again, just to get a little bit more clarity into your background because I do wanna highlight Iris as ⁓ as a very approachable, you know, legal AI. ⁓ what led you to found that? Given your extensive background, what made you go, you know what, I I wanna start, you know, focusing on on legal? ⁓ that's a great question. So I think there's two answers to that. ⁓ one, I've always been interested in legal AI. So even before Iris through, you know, when I was doing my open source research, we would often spotlight things for legal AI. We were talking about developments that could help speed up legal because I'm in India. One of the biggest problems in India is that there is a you know, and this is true worldwide after COVID, ⁓ court cases are backlogged, things are backlogged, things take too long for access to justice. And I just think when when you have technology that can help things go faster, You know, even before generative AI, there was machine learning that could have been useful there. If for nothing else, just navigate searching through cases. Why was there no good way to search through cases and find issues? Like these were some of the things that I was really interested in. and then what happened is I happened to meet ⁓ my now co-founder at a legal AI conference that we were invited to. He was the head of innovation at a legal ⁓ a pretty large ⁓ multinational law firm. And Because of his position, he actually had the ability to try out whatever tools he wanted. His law firm wasn't even particularly looking to save money, yeah, because they were big and rich. So they were just like, buy whatever, make our work hundred X more valuable. ⁓ and when we met, he was quite frustrated with law legal AI tools. He's like, Hey, why am I paying so much more for Harvey? And seems to give me the same answers that Chat GPT does. I why do why the like thirty tools for practice, he was a litigator. ⁓ so we originally just started talking about like what how could he use Chat GPT better or what were some things he could do. And through that journey we very quickly realized that ⁓ we could actually build a much better legal AI too, because a lot of them misunderstood what it meant to have lawyers that ⁓ AI that lawyers can trust. And that was Iris. We launched publicly a year ago and it's been a pretty crazy journey ever since. ⁓ that's fantastic. I mean, again, something that we always talk about. We help a lot of startups write, ⁓ I'll qu get funding, bring their first MVP to market. And the biggest thing that I always say is that if you want to be successful, it it's not just the idea. You really need to be a subject matter expert. So the fact that you have a lawyer and a litigator that was ⁓ a co-founder of yours obviously aligns with that journey in terms of a better chance of success. ⁓ real quick before we kind of move on to implementing AI in the real world and what that looks like, ⁓ what is the main problem that Iris tries to solve ⁓ again versus these these others, ⁓ if you have kind of a a singular differentiator? Trust. Iris tries to basically ensure that whatever output it has can be trusted and worked through with the lawyer. And that's something that fundamentally legal AI just misses. But also we're seeing AI agents for a lot of other industries miss as well, which is why we've seen an interesting expansion which wasn't really intentional into other verticals because like when you're doing anything where Context is important and where people's livelihoods can be on the line and people need to be able to back a decision made by AI or like back like they b have to justify their addresses to somebody that's a much higher authority. just telling having generative AI, do a bunch of agent calls and give them something isn't good enough. They have to know where it comes from. They have to know that whatever it's quoting is a real source, they have to know that. It's not misinterpreting the data it's quoting because that's a big issue. ⁓ and all of those are things that Iris hasn't a hundred percent solved for, but I'm pretty comfortable in saying that every agency system we've tested this against, Iris has done it back then. That's fantastic. Again, something we talk about here on the podcast all the time is that, you know, a lot of times, ⁓ as as the AI is trying to generate a response, there is trade offs for speed and and to your point that inference level. But if I'm quoting a a court case, right, I I need it to be accurate. I can't have it be made up for the final twenty percent, let's say. And so again, ⁓ you know, that that context and trust is is a lot more important than than I think a lot of people think, ⁓ related related to AI and and to your point, the stakes of leveraging AI in things like court cases makes it ⁓ invaluable ⁓ and necessary to trust. All right, so now we're gonna be moving into our first segment, which is moving AI into the real world. ⁓ the theme here is why some AI projects reach production while others remain experiments. And so, Dev, you've worked across healthcare finance, government, some startups. What's the biggest mistake that organizations make when implementing AI in high stakes environments like law, for instance? I think the biggest one is not ⁓ defining what kind of a relationship they want with AI. ⁓ I think where I see a lot of mistakes happen is they don't really if you don't predefine what exactly you want your AI to accomplish versus what should be left in the hands of your humans and ⁓ like what those boundaries are going to be, you will one, you might have very weird unrealistic expectations and two, ⁓ if you're not willing to make those like very hard boundaries. What will happen is your team is going to end up over investing into something. You know, maybe your system only needed an eighty five percent performance to be useful. ⁓ and actually when you get to ninety-five percent, the extra resources you invested into that would ha actually make the ROI not worth it. But you you because you never define that okay, I'm gonna stop at eighty five percent, or once I have eighty five percent, I'll reconsider what I what my goal should be, et cetera, you end up just trying to do too much. It's it's makes a lot more sense to first define exactly how do you ⁓ what do you what what is your user going to do, what is the system going to do, what does the system need to do to ensure the user can do their work better and then just hammer that in completely as opposed to trying to fight five battles at once. I think that makes a lot of sense. I mean, again, here at Cezanne, we always tell our customers to start with the end in mind. ⁓ whether that's an AI project or a traditional, you know, project where we're we're building something from scratch. ⁓ and I think that, you know, to your point on the AI side, a lot of people just assume, well, we have AI, so it's just gonna be this magic bullet that does all these things when they haven't thought through that that really, really important context of what does good enough look like, right? ⁓ and it's the same thing that I think that, you know, It's it's a colloquialism that that happens all the time, which is, you know, don't let ⁓ progress or perfect be get be the enemy of progress. And so again, AI is a continually evolving thing, and I think that that's really good advice for people looking to to implement something ⁓ to make sure that they have the the end in mind. ⁓ you know, when we talk about disease detection, which is something that AI is obviously spearheading quite a bit, as well as legal reasoning with iris. What patterns separate an AI project that's successful, that can actually ship into a production environment versus those that remain stuck? I think again, the one of the reasons you always try to think with the ROI and outcomes in mind is so that you can define your contracts with your system. For instance, with Iris, there is we can have no less than 100% accuracy when it comes to ⁓ being able to tell our users why our reasoning is working the way it is. Because that level of transparency is what a lawyer needs to know when they're brainstorming strategy, when they're thinking about okay, what other outcomes can I be taking? How do I go back and forth with these, etc.? So that kind of work is something that your system basically has to always get right. ⁓ versus there might be some areas where you have some leeway for error, which is I might apply the long stra wrong strategy here, but because I'm always telling my users I'm applying the strategy ahead of time, this is what I'm thinking. with what's very unique with iris is you you can actually steer it mid-influence. So you don't have to just wait for the final output and read through that where you might get skewed, but you're you're always shown exactly what iris is thinking before it generates anything. So you can just tell it then and there, hey, don't apply the strategy, focus on something else and it maps it. And I think that is very specific because anybody that does knowledge work, we started with lawyers, but also with investors and bankers and compliance guys like this is key. ⁓ you w one of their biggest enemies is that you have a thousand page document where nine hundred ninety words are actually really, really good and then there's ten pages, ten words buried in somewhere which is a big, big red flag that invalidates all of your work and because the nine hundred words are so good that you get complacent there. So before that can I can I build the trust before you see the nine hundred words? Can I build the trust in a hundred, two hundred words that you can read immediately before to push? ⁓ for something like disease detection, which is something I did way earlier with machine learning, ⁓ again the point was what are we doing? I can't detect diseases because I'm not God, but can I at least flag risks early on? Can I tell you why I'm flagging those risks early on so that one when you are going to somebody who's diagnosing this or whatever, it makes your life easier. If I'm trying to forecast how's likes how likely you are to be sick, which was another work that we did for hospitals. basically tell them ahead of time who gets who should be released and what are their likely outcomes. Can I tell you what variables am I hinging my calculations on so that when you make that decision you are not going to be sued for malpractice? ⁓ I think that's whatever you do, you have to make sure that if you're in sensitive areas or high value areas, people aren't looking to do better work. Very few people care that much about what they do. People are mostly avo looking to avoid getting sued. So if you can make the ⁓ their decisions easier. If you can t if you can build systems that help them avoid lawsuits, they can rely on your trust and automations. If your trust and automation is I'm gonna make your life better by doing more work and it's got higher accuracy, more often than not, people aren't gonna care because they're they're you're asking them to take liability on their head and that's always a big ⁓ kind of no no. Yeah, well, as a follow-up question, I think that you kind of sort of answered it, but it's really important about complacency, right? And so a lot of times, again, to your point, if nine hundred and ninety out of a thousand words ⁓ of of an output, right, or nine ninety-nine percent ⁓ of a hundred percent output looks so good, can you can you share from your experience a project where something looked really promising, you pushed it to production, but then there was unexpected challenges once it reached and users that you had to then reevaluate or pull back ⁓ that solution? I mean all the time, mate. All the time. ⁓ for instance recently, it's been a few months, we were releasing agent mode because ⁓ what we realized is this is the first time we really decided that why don't we show all of the reasoning that we do to our users and just give them ⁓ the specific layout. Before this, one of the biggest issues we had is we had a research mode, we had a drafting and answering mode. Where research mode would act it would actually go through and build like low hallucination outputs, etc. Go through things and then you could use drafting mode to basically say, okay, using this research port, please put put out whatever you need to. It was a button and we had like a walkthrough for that button. But you'd be shocked how many people never use that button. So This it was with that agenda in mind that people will k people keep cursing us out saying you've had hallucinations. And when we asked them, did you use research mode? They would be like, What is research mode? And that we'd be This is a button right there in that walkthrough that we did. That has a not only grounded citations, but also a citation checker so that you can verify that this thing exists and that this thing is useful to your output. ⁓ but again, people are people. They they just they decided. So we decided let's try out agent mode. Agent mode is going to do everything for you. So it's going to do the research, it's going to do this, it's going to figure out what it needs to do. And this way you can't complain. ⁓ new problem pops up. People are overwhelmed by it. So suddenly we realize that n some people were unhappy with having to choose modes. We took the modes away from agent for agent mode. Other people come back to us and say, Why did you take away my option to take away modes? Now I now I I have to see this AI working through all of the because we do like 60 simulations and we do like, okay, this strategy can apply here, this strategy can do this. So the people are sending us messages like we are so overwhelmed by this. It takes minutes now to do anything. and people are losing their minds. So ⁓ then we had to try rolling that back because we decided we kind of overindexed on this, make every decision for the user. ⁓ but again I think this is where like I'm telling you about why you have to have user contracts, but this is a really good example of how even you might not as even for very simple things, you might not know what the user needs because so many users for us did not know that they needed contracts. ⁓ like they needed to choose what to what kind of answer they were looking for. Are you looking for a draft? Are you looking for legal reasoning? We could do this in the same chat. That was the intention is you just switch AI's focus, but a lot of people didn't even like, you know, to them AI was so scary that they weren't even interested in taking two minutes to see the learn the differences between the modes. and then we tried to change it to let's do everything for the user. And we realized that a lot of users don't like that because they feel disempowered and they feel overwhelmed. What they want is something that they can follow along with cleanly, but still not have to ⁓ like they don't want to have to think about using a fork. They just use the fork. But at the same time, you don't want to give them like a fork that's vibrating at 100 miles an hour because that just overwhelms their senses. So this like, you know, was an iterative process into the current generation of our technology where we've done a really good job, I think, of both showing everything our AI is thinking about and how it's breaking it down, but never overwhelming them with the research it's doing. You can always dig into more if you want to learn it. But it doesn't overwhelm them and it goes back and forth. And that's actually helped a lot with hallucination catching, et cetera, as well, because now suddenly ⁓ they're able to see immediately. Since like only the information can communicate is maybe ten percent of what we used to communicate before, they're able to focus much more importantly and catch everything that they need to. Yeah, I think that that's really important too when you're building software in general. Sometimes the way that users end up having it, you being a technologist. assume one thing, but then it gets into the hands of somebody else who, you know, is not as tech savvy. They're using it completely differently and it has you rethinking the user experience entirely. I think that kind of is a is a great segue to to this last question for this segment, which is what normally proves more difficult when you're putting something in production? Is it building the technology itself? Is it the organizations in which you're deploying it and and trying to get the contextual data that that they have, or is it the people getting people to actually adopt it or use it the way that you you know, that they should? I don't think you can ever tell your user what they should or should not do, ⁓ to be honest, because ultimately your user is what who decides where utility is. ⁓ I think as a builder that's been one of the lessons I've had to learn. ⁓ but definitely the last thing is the most difficult because I think also when you come from a building perspective and this like even in the lawyers in our team because we have a lot of like people who used to be practicing attorneys, the kind of attorney that would switch focus, join a company like Iris, which doesn't have the PR, doesn't have anything like this, like it's not a mainstream name, and then would get involved and start working at a company like ours and use it. is going to be by on average much more experimental, much more willing to try things out, do things. And this means that even when we design our stuff, we often overlook that ninety nine percent of the attorneys are not that crude. ⁓ so just getting like again the button thing is a very small aspect. But yesterday or day before I was in a law firm and I was trying to debug their ⁓ pro ⁓ problems and they come to me and they tell me, hey, this thing doesn't draft very well. ⁓ it's very the answers it gives me is very short and I was surprised by that because drafting was traditionally one of our strongest use cases. and I asked them, Do you what mode do you use? Because now we've gone from agent research and deep research to basically ⁓ quick, ⁓ something standard, thorough. And by the way, it's a slider. That you see in every message you send. So it I I don't know how you can miss that because it's right there in the message you're sending. But somehow ⁓ they had never put two and two together and said this. So they were just using quick mode. And this is why they were extremely unhappy with the output because again, our walkthrough literally tells you this. But if you want very thorough drafting, go to thorough mode. If you want anything heavy duty and reasoning done, you go through Thorough mode. Quick mode is only when you're about to get in front of a judge for five minutes. You've already done all of the strategy, etc. You just want to refresh as you go to quick mode, you're like, hey, what do think about this? What am I learning about here? Can you summarize this? That's the point of quick mode. But again, the user was just not understanding that. Or like we have a really beautiful word, ⁓ kind of a Google Chrome word. thing built into our platform so that you can redline documents and you can ask AI to edit specific sections, et cetera. People don't use it. Not because they don't like it, they don't know it exists. It's every file we create, there's a drafting assistant that right there and there. But they don't want to click open in drafting assistant. They don't realize that this thing exists. So they keep asking me where how do you work in Word when we have it right there? ⁓ so stuff like that is where ⁓ you often find problems. I think at least when you're Dealing with a lot of non-technical guys is user education is a big issue. People aren't again, people aren't willing to just try things. I think they're scared of like breaking something or testing something. So their first instinct is always going to be if this is not immediately in one step obvious, I will just zone out. ⁓ and I think that's more than anything else. I think that's the problem that any ⁓ system has to solve. Is how does it become boring, reliable, and easy to use? Absolutely. But that's regardless of AI, right? I mean, change management and user adoption, I think, continues to be ⁓ some of the hardest challenges to overcome when it comes to any software practice, which which by the way is why, you know, the Facebooks of the world, the apples of the world, the Googles of the World spend hundreds of millions of dollars a year on user research. So ⁓ again, I think I think that kind of goes hand in hand on you know, where the priorities are for those platforms as well. I think before this I worked largely in machine learning. So I never had to deal with this so much because most of my work was being used internally within like by people to figure out decisions or whatever. Or if I had to present it, I would just make the reports myself and people weren't necessarily digging through the details and models and whatnot. So ⁓ like definitely putting out something where strangers use it in ways that like they're not Grand for is one of the most interesting experiences building software, I think. That wraps up part one of our conversation with Devanche on what it really takes to move AI out of the sandbox and into real-world production. But getting an AI system deployed is only half the battle. The bigger challenge is how do you actually trust it when the stakes are sky high? In part two, we're turning our focus to reasoning, accuracy, and trust. Hit that subscribe button and turn on notifications so you don't miss part two. Until next time, I'm John. Ow.