Jon: Welcome back to the Execution Over Hype podcast. I'm your host, John Turns. Today we're stepping into part two of our three-part series with Devanch, co-founder of Iris. After tackling production reality in our last episode, we're turning our focus to reasoning, accuracy, and trust. We're dissecting the critical difference between surface-level task automation and true reasoning infrastructure, mapping out where enterprise teams must draw the line between autonomous execution and human-in-the-loop oversight. And confronting the ultimate question: when an AI system makes a high-stakes blunder, who actually carries the liability? Let's get into it. We're back with Dev from Iris AI, starting our second segment, which is reasoning, accuracy, and trust. So, Dev, legal AI is often marketed as automation, but you talk about reasoning infrastructure. What's the difference and why does it matter? I think that's a great question. ⁓ I don't know if there it makes sense to have a dichotomy between automation and reasoning infrastructure, because I think one is more from a socioeconomic business sense and the other is more from a technical ⁓ business ⁓ technical aspect of it. But essentially automation is anything that would like a lot of legal AI is basically let me replace a lawyer, let me replace a paralegal, let me replace some third party, ⁓ and do XYZ for you, file for you, do something, give you a legal research or whatever. I think that's the wrong approach. ⁓ I think with something as specialized as legal, trust matters. ⁓ verification matters. You know, I pay for tax guy to do my taxes. Not because my taxes are complicated or I'm really interested in saving money because I tell them I don't like doing this, just whatever the fastest is, do it. I don't care. I paid for the tax guy because I didn't want the American government coming after me. If they're gonna come after somebody, come after the tax guy. ⁓ it's his name on my returns. Same the reason people go to lawyers. It's like I can try to fight my own court case, maybe I'll win. But when it comes down to it, I don't want to take that risk. I want to know go to sleep knowing that I did the best I could. And that involves like going to an expert, having them do things. So in that sense, you would not I don't think you want to automate that because I think that adds more burden to people that you don't go to lawyers when you're happy. You're going to lawyers when you're going through something. And if you're trying to automate law, you're basically putting more burden on people who are un are already unhappy. ⁓ and the philosophy around the reasoning infrastructure is what if instead of that we were just making it better for lawyers to serve things, customers easier, faster, better? How do we make sure that whatever they shouldn't have to do, they don't have to do. So that we c lawyers can focus on understanding the customers' needs, empathy, listening to what the judge says, reading between the lines and stitching everything together to tell the AI, this is could work, this is could not work. You figure I like you do the grant work and let me tell you how to run things. I and I think that that's important too when we're talking about, you know, again, the automation in terms of the process, right, versus, you know, the reasoning side of things of saying, hey, I'm replacing this expertise that I'm paying for, right? Because I don't want the personal risk. ⁓ and so that that also I think ⁓ leads into a great second question, which is, you know, how do you build, you've talked about a couple of times now, how do you build AI systems that are accurate and explainable and trusted, especially in these kind of highly expertise, like highly regulated and and and expertise driven industries, right? Like like medical, like like law and legal. I think you have to work backwards and say what makes something untrustable? What ⁓ where if my output if I were to give you an output, where all would you doubt it? And you work backwards from there. You can doubt it in a few ways. You can doubt the way it reasoned through things. You can doubt the fact that it is the things it has are claiming the things that it's claiming. ⁓ you can doubt the very existence of the facts. And those are the three Hallucinations are not a real thing. I think it's I don't like that as a concept because it groups together three very different kinds of information theoretic errors, ⁓ which is the three that we broke down. You know, you can have a f fake case law. What is much worse, much and like almost no legal AI system even talks about this, ⁓ because they don't know how to solve it except for us, is you can point to a real case and have it be a completely useless one for the argument you're making. Because that case, the context there is an agriculture and it just had a few words that your rab or whatever else picked up, says this is a real case, and you're trying to apply it for commercial law. Like that they have no bearing whatsoever. But because people don't realize this stuff, they'll just build these systems, see case law is real, a few words match, we're good to go. But that you cannot have. And finally, at a much larger philosophical level. ⁓ you want to know why, why this and not that. Why are you saying argue this, argue like sue this person versus try to settle? What are the decision points you're considering? ⁓ so like that's the really the meta level that encapsulates this. So if you can make each of these ⁓ if you can address each of these sources aways, you have it. With iris, first off, like everything we claim, you have inline citations. You can check exactly if you've uploaded documents where in the document we're calling. With websites, you can't really we have the quote from the websites or from like legal cases, et cetera. You don't own the website, so you don't directly want to jump there. But you'll directly be able to see in five seconds that this thing exists, this claim exists, it's applicable, this is why it's applicable. So that addresses really the first concern of fake cases and a little bit of the second concern. The larger are about aspect and strategy. again. What we're we think about is our AI simulates a hundred different things at once. It's going through multiple chains of work, it's doing this. How do I simplify this to my user? Give them what the most important operating decision is. So that when they're looking at what the AI is doing, because like on average our system might take two, three, four minutes for a complex query, can go up to ten minutes. So there's no time limit, but I'm saying that's about the distribution range. So you're sitting there for those five minutes. I don't want you to be like Claude where you just zoned out and it does some thinking, bamboozling, whatever, and gives you an answer. Instead, it should tell you, okay, I've reasoned through this, I researched this, I found this, let me this is how I'm updating my case law strategy. When you start to do that, then the user can get involved and the user either they're gonna look at the AI and say, I love this strategy, nothing to do, just draft it. Or they'll come in and they steer RAI. They say, I really don't like the way you're going here. Please focus on this instead. Or I love the way you're doing this. But instead, what if we were to take a more global view? Give me all of the tell me how this is handled across jurisdictions. ⁓ that kind of work is where people really start to build trust because they see the answer being built in front of them. Yeah, I think that those highlights are really, really important too, because you brought up a really great point in terms of how rag is used on on internal documents and and again in the in the in the advent of case law, right? We're talking about, you know, hundreds of thousands of of external documents related to cases. And so that taxonomy and the context of that taxonomy, I think a lot of people don't quite understand whether it's applied to the legal side or others. I don't think they quite understand how AI does its its reasoning. And so to your point, you could reference ⁓ you could reference a case law that's that's tied to to land ownership, let's say, but that has nothing to do with the context of a particularly separate corporate law case, right? It's just getting it's getting the context confused even though the taxonomy is quasi aligned. And so again, these are just some things that ⁓ over time AI and and humans as or us as as developers probably ⁓ will come up with better ways to tackle, but in the interim, ⁓ again, I think it's a real challenge. I think it's great and you hit the nail on the head in terms of building that context of this is how I'm going through it. Don't don't don't if you don't agree with it, you don't agree with it, but at least Iris is showing you step by step how you're building ⁓ that thought process, let's say, for the machine, ⁓ versus ⁓ a line of context and here's everything I did for you. How how can companies decide where They should have AI operate a little bit more. I again I don't like this term independently. you know, yeah, it's sometimes you will peop we'll we will hear people say auto-magically, right, in terms of what AI does. ⁓ but how could should companies decide where they can they can have their AI run a little bit more open in terms of these repeatable tasking, right? We can call it agenic tasks or agenic workflows. ⁓ versus where we still need human in the loop. I think that's a good question. I think generally speaking again, this is the one thing that I don't think as many companies struggle with because you know where you're going to get sued and where you're not going to get sued. If you whatever you think is not your value, whatever if I were to come to you and say, Listen, I've automated this and you say, Okay, cool, this doesn't mess up my job by more than ten, twenty percent, or give that away to AI. whatever you think is left is what you should have the human in the loop for. Fo like people the reason people can trust Iris is not that they can trust the thousand page draft. They they check that and all. But they know that the two hundred page ⁓ the two hundred word sorry ⁓ summary of its strategy and what it's learning, et cetera, where it's going to give you basically everything it's learned and what how it's going to approach things. Once they're able to trust that fully, they can just they can be much more trustworthy with with reg of like, okay, it's probably gonna get the so exact words correct. It's probably going to have read the document correctly and it's not going to make that a mistake because LLMs almost never nowadays make that mistake. You give them a document, they're pretty good at reading that document and say, This is what it says. ⁓ so there like the document reading and fact matching is commoditized. So They're let they're fine letting the LLM automate that. They're still going to check the final output, but they're fine letting the LLM create that document. They're not fine letting the LLM and we push them to be like, don't be like complacent on this reasoning traces. As it's reasoning, don't look the other way. That's your job. Look at that. This is where this entire output is won or lost. So you have a lumen in the loop there. I think that generally is the f is a really good mental model to think about is for me personally I'm not I write good code, I wasn't the greatest coder in the world. I am okay automating all of the coding. ⁓ I don't need to necessarily even review a lot of it. What I do need to review is how are you deciding the data pipeline? How are you designing the how ⁓ models are moving through, like what are the latency requirements you're keeping? Because if I leave it that to the AI, it's going to decide that. gonna build something that will bankrupt my company. ⁓ so that kind of stuff is where you know I've decided that I never I was never superstar coder. I was always a decent coder with a good understanding of good software design principles that really, really knew what AI battles to pick and didn't know how to pick what not to pick. So I've like doubled down on that aspect and I've kind of become even less concerned with a lot of coding, especially when it comes to portals, dashboards, et cetera. That's I never knew how to do that well. I was copying Stack Overflow from day one. So that is completely AI and I don't think I've lost anything there. Well it's a good follow up though in terms of what you're saying, just kind of if I may, when we're talking about responsibility and accountability in these AI systems, to your point, summarizing a document you know, the AI, the L L ⁓ structures, they're very, very good at that. But when we start talking about agenic workflows where all of a sudden it's going from brief strategy to actual building ⁓ contractual documents that we're gonna submit to a court. And to your point earlier, it's those ten words out of a thousand that actually reverse any of the arguments that you're trying to make or open you up to liability that you thought you were you you were ⁓ Closing, right? Who should be held accountable to that? Like is you know, the AI system technically made a mistake. It's only ten words. But when we're doing things at speed, right, that does create real liability for somebody. So that's why again we are like where Those ten words, where w where is that ten word mistake we made? It could be made in drafting or reading the document because none of these systems are perfect. But nine point they have like more than seven, eight decimal price of accuracy for that. They generally don't make those mistakes. Unless of course you have OCR and messy ingestion, et cetera. That's a whole separate class of engineering problems that we had to solve. But assuming that this it has to read this paragraph correctly, you know. With legal, the other thing is you have ten thousand paragraphs. How do you know which one to read correctly? Again, context management, engineering, knowing where to search, whole separate class of problems. For now, I'm just assuming that before the drafting stage there were ten paragraphs that it had to read and why they were relevant and our AI has identified all ten and it's reading. By and large, I ca ⁓ no even I can't get it perfect, but it's close enough to where. I don't think there's a battle to be fought here. similarly with once those ten paragraphs have been drafted and the legal strategy has been drafted and all of this, putting it perfectly in the right with format, etc., tend to do LOMs tend to be fine. They're not great with formatting, especially for law because lawyers have retarded like ways of doing things. But more better than useful enough to wear ninety p times out of a hundred, they've saved a lawyer sometimes. Well let's take the I was gonna say so yeah, let's take the context a little separately, right? Because this is just important and it's come up across AI systems, not just legal. So I don't necessarily want to focus just on Iris, ⁓ because I'm sure that you ha you have your own kind of ⁓ drafted response ⁓ in in terms of what's needed. So, you know, when there's a mistake made, whether it's ⁓ like and it could be Chat GPT, it could be ⁓ Financial, I'm creating a balance sheet. I'm creating, you know, something else ⁓ within AI, and there's a major mistake. And and I'm s as a CFO, I'm submitting that form that has bubbled up from a financial analyst or a director that's that's created that for me. I'm submitting that to let's say the SEC. It has major mistakes in it. Who who should be held accountable? The AI system or ultimately me as a CFO because I didn't double or triple check the work? I think by and large, right now everybody's going to say hold the CFO accountable. That is always going to be the case. But I think even in there, the reason I was going off that example is you want to when you're thinking about breaking it down, ⁓ you really, really want to be ⁓ very In the future, your user does not want to take liability. Even in the present, your user does not want to be take liability. So playing out what the dynamics will be five, ten years from now, there are certain things they will not take liability for. You know, no matter how shitty of a driver you are, nobody takes a liability for a wheel suddenly spinning off or doing things. So there are just some contracts where you can say the car might not be perfect hundred percent of the time, but it's n the engine's not gonna blow up suddenly. And you have to know where what for your system that engine analogy is for so I think over there people will always take liability for whatever you will say I'm charging you a hundred dollars for. ⁓ in the case of legal or finance that's judgment. They nobody will take there. ⁓ probably five, seven years down the line a a form hasn't been filled out the right way, you will start seeing AI companies get sued for that because at that point people will say this is not what our value add is and this is this is where you guys have to be reliable beyond like ninety nine point nine nine nine percent. So I'm guessing like once regulation gets involved and once like audits get involved, you will start to see the split of some things where the AI company is liable, some things where the user is liable. That wraps up part two of our series with Devonch on what it takes to build AI systems that enterprise teams can actually trust in high stakes regulated environments. In our final episode, part three, we're shifting gears from system architecture to raw execution, building through the hype. Devonch opens up the playbook on how he bootstrapped Iris to $1 million in annual recurring revenue in just seven months without relying on market fluff. We'll break down what founders and enterprise leaders need to stop doing immediately, separate real AI capability from pure marketing noise, and give you three concrete steps. your organization can take right now to move beyond AI experimentation and start shipping production grade systems. Make sure to hit that subscribe button, like the video, and turn on notifications so you don't miss the conclusion to this series. Until next time, I'm John. Out.