Vitaly Sharovatov: Welcome folks to another episode of the Beyond Quality Podcast. Hello, hello. I'm Vitaly. I'm developer advocate at Case. Maria, I'm quality engineering leader at Regent. I'm Anupam. I'm head of AI testing at Test Solutions. And ⁓ we are really happy to bring you another episode. Just a reminder to people who have not tuned in before, Beyond Quality is a collaborative research community focused on advancing software quality practices through open discussion and shared knowledge. We have a focus on quality engineering, but we are also very cross-disciplinary. And we will see that in this episode as well, because for this episode, we have a very, very special guest. Actually, a guest that needs no introduction, at least to our circles. for those of you who have been living under a rock, maybe Lisa, a quick word about yourself in your own terms. Well, thank you. It's wonderful to be here. I'm so honored to have been invited. And I'm Lisa Crispin. I live in Northwest Vermont, which for context is close to Montreal and Canada. And I've been, I'm a tester by trade, but I've been lucky to work on agile, cross-functional agile teams for almost three decades, well, almost three decades now. And very lucky to collaborate with Janet Gregory on a number of books and courses and video courses on adult testing and holistic testing. ⁓ And I've mainly worked on full time on teams, but the last few years I've taken a step back and done working as an independent consultant and also doing a lot of unpaid volunteer work. For example, I'm a guide for the Dora community. And so I really enjoy a chance to give back to the community. I also like to mentor new speakers. So people want to get into speaking, get in touch with me. That's wonderful Lisa. I I've read your work and I've had the privilege of applying it directly to my work, especially when it comes to moving teams to ship in smaller batches, continuous delivery and so on and so forth. So it's a real privilege. And ⁓ on this podcast episode, we will be touching upon a Dora metrics and Dora research collective, but with a specific focus on their latest findings on what it means to ⁓ what it means for AI assisted development. But before we go there, I mean, I think at a time like this, where our industry is transforming as we speak, ⁓ like in December or even November, if you had told me last year that most of my code is gonna be AI generated reliably out of the room, but now I'm on the other side of actually having swallowed that pill. So we... There's nothing that we can say with so much certainty on the one hand. But on the other hand, ⁓ there are some of us like you, Lisa, who've seen similar transformations before. Anyhow, I just wanted to get your take on how this paradigm shift feels for you and what it means for testing in general. Well, it definitely reminds me of when test automation kind of came on the scene. Which it wasn't, mean, we had automated test tools back in the 90s, but when the vendors started to get into it and then the open source of the developers started to get into it, ⁓ when Extreme Programming took off around 2000, late 90s, 2000, and they started developing tools, which they open source. it's like, ⁓ we can really automate a lot of this horrible soul crushing manual work that nobody wants to do. And immediately... testers thought they'd be automated out of a job, companies fired their testers, they didn't need testers anymore. So that all happened. This with this GEN.AI revolution is that to the 10th power. mean, people think it's even more magical than automating tests. And it's a great thing to have. We don't wanna do this whole crushing work. There's a lot of benefits it can bring. But I think part of the reason is as a technology, like I started programming in 1982 and people were talking about AI back then. It's like flying cars. We thought it would never get here. And then all of sudden, not only did they get here, but every day I see something new or something better that it's doing. And I think it's hard for us as humans to observe that change. We talk about cognitive debt. when using agents to generate our code. think we have cognitive debt about the whole thing. It's hard to take in. It's definitely, it might take our jobs for now, but eventually as businesses are gonna realize that they need humans. So we'll see. That's so, so, so, so true. I feel this pressure in the industry that testers are not needed. And I totally... I with what you said about automation. I've heard this myth that manual testers, I hate the word manual, but still manual testers are dead. We don't need them. Like from 2000 and 2010, same, same myth reinforced all over again. And yet it's so interesting that now with LLMs being non-deterministic, we kind of need more QA, maybe not testing per se, but more QA in general, like... When do we apply testing? What testing do we apply? Like thinking about it and then doing it of course. And this requires testers and QA folks. It's so strange. not only that, but hey, it's great that you can give your agents the specifications and it will generate the code. Wonderfully if you've given it good skills files and all those things, it will do a great job. Are the specifications right? Are they meeting the customer needs? To me, that's where testers have always added the most value. Because when they're, my experience, I see it every time I work with a client. If there's not a testing specialist on the team, nobody asks their one of questions. Nobody does a risk assessment. They just take those specifications from the product people who I'm sure they've done user research, they're well-meaning. But we build what they say to the specifications and it isn't what the customers want. What did that do for the company? So, mean, I think we know testers are needed all the way around that whole developer life cycle. Who's watching production usage? Who's watching what the customers do? Who's spotting problems that maybe are not really obvious. Maybe it's only happened to a small percentage of customers, but maybe those are your most important customers. So we need people doing the quality thing the whole way around. I would even add that from the organizational... not even a hierarchy, but organizational design perspective. Testers are the ones who are meant to be questioning everything. So it's like you hire people to do testing specifically for them to be questioning things. So you're designing, let's say, maybe not a conflicts, but like a discussion in by just hiring people with a certain role. And so, with AI amplifying the velocity or like the volumes of code being produced, testers won't have enough time to ask questions. That's what worries me a lot. Well, that's really true too. And you know, here's the deal. We can all ask those questions, but most developer job descriptions and skills on their skills matrices on their career ladders don't mention any testing skills or question asking skills. And they're like, okay, we got this deadline. You know, got this trade show coming up, get this feature built, whatever it is. It's not that they don't want good quality, but that's in their mind, that's not what they're being paid to do. Now we can fix that problem very easily by making testing part of their job, putting in their job description, putting the skills in their career ladder matrix, and then helping them learn the skills, which we might need some testing professionals to come in and help us do that. ⁓ So it's a solvable problem, but yeah, as you say, like I mentioned, the cognitive debt, even if the developers are trying to review the code, it's too much, it's too much, they've gotta back up. They've gotta test the agents, they've gotta test that they can trust what the agents are producing. I saw a great video last week that Anthony Marcano did to show how he test drives his agent testing when he's building the agents. so that he can make sure that they're going to, even though they're non-deterministic, that they can be fairly reliable. That's very interesting point you mentioned. I watched this video as well, and even before seeing this video, I practiced testing my skills, not my skills, the agentic skills. But what I found peculiar is that those tests have to be rerun regularly, because skills deteriorate, degrade. The more context you have, the less reliable even the skills are and the plugins are and everything you have in the context is. And so we are thinking now meta, like we need testers to help us refine our agentic skills so that we apply testing practices per se to the skills so that they don't degrade. Cause otherwise at the end of the day, we're just like after a month of getting someone's skill, we end up having no skill at all. That's a good point I hadn't even thought of that. Wow. Yeah. This is like in my work, I produce content a lot and I'm not a native English speaker. So I always use LLM to teach me better English. So I ask, pardon me, I ask it questions like what I'm writing, does it make sense? Is it grammatically correct, et cetera? And it produces output with ⁓ dashes all the time. I'm so fed up of asking it not to. I have a strict Claude MD rule. Do not ever, ever produce a single ⁓ dash. No, after half a day in a session, It still produces. So what I had to do, I had to create a hook, Claude hook, which runs deterministic Python script checking for ⁓ dash and giving back Claude a slap, essentially saying, do not do this. I told you so. That's basically a guardrail, which industry or industries are applying in the big volumes. Like, okay, now we. came actually to a conclusion that, okay, we need both deterministic parts of the code and LLM generated, which is kind of less reliable and more, there is more variation to this produced code. So I think there is a place for both. It's just, we need to decide how to introduce these code generation tools into our processes to improve and to amplify the output. not to degrade it. So we need to find the way how to integrate it. Do you see, I'm wondering if you already see some good examples how companies do it or do you have any suggestions how to introduce AI into existing processes to improve them? I've got to say, it's not my area of expertise because I'm not working on a team hands-on using these things. I'm just watching other friends who show me what they do at work. ⁓ But my push has always been the same and it still applies even more, I think, with AI is we need, you we mentioned having a diverse, people with diverse skills on our team. And I think it's really more important than ever to collaborate. So I have always worked on teams that did a lot of pairing, did a lot of ensemble work, and it lets us prevent a lot of bugs, because if I'm in an ensemble building code for a story, I ask a lot of questions. And they go, ⁓ yeah. And at the end of session, I say, was I helpful, or did I annoy you with questions? it's always, Lisa, you prevented a lot of bugs. So I just think we've got bigger challenges now. We've got challenges we've never had before. And so having people with diverse experience, mindset, ⁓ skills, ideas, we have to be thinking laterally. We gotta be thinking out of the box. And one way to do that is to get people together that can challenge each other and build on each other's ideas. And also more important than ever, you know me, I'm big on using visual models to help us collaborate. So using those visual models to guide our conversations. I think, I mean, not to toot our own horn, but I still think like the tollistic testing model that Janet and I came up with, Janet Gregory and I came up with a few years ago. It just helps people think it's like, ⁓ well, when we're in the planning stage, we need to make sure we do these things. And ⁓ when we deploy to our test environment, we need to make sure we do these things. And ⁓ here's what we want to watch for in production. Here's the data we want to capture so we can watch it in production. You can't just have one person say, we got the one data expert or we've got the one design expert or whatever. They could do a good job on their own, but now they need to get out of their own head and work with other people. And because these tools can do the grunt work if we train them right, then I think we could get even more collaboration. I'm hoping that using these tools, you know, back in the day, I remember back in the early 2000s, I was working for an internet toy company and I was a QA manager. And the company said, oh, well, you know, before we hired you, we bought this Mercury Windrunner, no offense to Mercury Windrunner, but, and so you need to use that. And I'm like, that's a completely inappropriate tool for our, the stack of technical things we're using. But they paid a lot of money for it. So I hired somebody who already knew how to use it. I was like, okay, go do some regression tests, UI regression test scripts for us. And it's that same sort of thing. It's like, the CTO can't decide when we need this tool. We need multiple people thinking, what is the right tool for this team? What language are they developing in? If we want the developers to play, they're not gonna come over and help me with my Mercury Windrunner script. But in that case, we were writing our website in TCL or Tickle, which is a scripting language. that? And it's very good for scripting automated tests too. So I a book, taught myself Tickle. And then when I had a problem with my test script, could, developers were happy to help me. And then they were interested. ⁓ know, automating these kinds of regression tests, that's really a good idea. Maybe we should do more of that and getting them interested. With these wonderful, magical, shiny new tools we have, What an opportunity, you know, come and see this, come and see me generate some test cases and automate them. I'm going to let the agent do that for me because I've built the skill file and everything. Look at this. Wow, that's great. Well, let's work together to build an agent and a skill file to do this other testing activity. So that's my thought. just really, really wanted to highlight what you just said is exactly what I'm researching now. So in the industry, there is a huge marketing like waves, if not tsunamis of messaging saying AI enhances individual work. Like you take a developer and you make him or her 10x developer. In my opinion, and what I'm trying to validate now, this is the incorrect way, let's say, in suboptimal way of using AI. Because when a single developer produces 10,000 lines of code, second developer or a tester, they can't even possibly validate what's been done there or learn what's been done there. Hence this comprehension cognitive debt and what's even worse intent debt where we stop understanding why the things were done. And in my experience of working in my team, we're trying this now where I and my colleague and she's not even an engineer, we pair on building her a product with an agent. And since the agent can produce code ridiculously fast, we're not waiting for the engineer me to code. So we can be discussing the intent. I can work in small little iterations to see what the agent actually produces to be able to comprehend it. And so this agentic setup enables me with a product manager, with a product marketing manager, with anyone else to be collaborating on the business value delivery straight away and testing it straight away as well. This, in my opinion, can be the way how we use LLMs in the future. I will run this, proper research around this. I will try to run this as an experiment in multiple teams and then see if the comprehension and intent that grows or not. Because this is what, this excites me a lot. I think this could be the way. I'm excited about that too. mean, talk about the fastest feedback loop. Absolutely, Yeah. And then that's where we get around the problem of, we use the specifications to generate the code where the specifications what the user wanted. Well, now we still have to hope that the product owner really knows what the users want. yeah, then they're going to see immediately because you know what? We never know what we want until we put our hands on it and use it. Absolutely. And on the topic of collaboration, I had this experience of actually pair programming or rather actually was ensemble programming with three of us. in a session together with an AI agent. Now, one line of thinking goes, when you have an agent that's already a pair, so you don't need pair programming anymore. But on the other hand, what we actually noticed was when we have the AI agent generating code, that is like a bottleneck that gets out of the way of collaboration in terms of it's not constrained by typing speed or it's not constrained by your knowledge of the syntax. Instead, you can already envision a particular structure, a particular architecture, have it play out, look at the constraints, and then decide going forward whether that makes sense or not. I would even argue that it encourages ensemble programming more, because that was my anecdotal experience with this one example. It was a refreshing finding, and I would be surprised if that's not the case going forward. That's such an exciting thought, you know, because like when I'm on SoundBald, I usually just, you know, even if they told me what to type, my programming skills had deteriorated over the years. And so it's just, I felt like such a model, like it's like, y'all, I'll just watch what you're doing, point out things, ask you questions, and I won't actually be a driver. And... And, this way I could be a driver. Yes. And since you're a great tester, the value coming back to the T shape or the collaborate, not even the shape, the collaborative approach is that you are a great tester. Someone else is a good product manager. You all collaborate not only on producing the things, but also at designing how you're gonna be verifying those things. This is the beauty of it, because the edge cases, like the main problem with test cases generation with the UI is that it's sycophantic. It wants to satisfy the code. So if you are collaborating on writing the code, it is this moment where you can design the test cases as well. And this is even not even TDD, not all people can work in TDD, I get it, but even generating tests after the code or in parallel with the code, this works really well. Yeah, this is exciting. Y'all are answering the questions I've had in my head. Like, how do we do this? The only thing that I've seen here, well, at least where I failed, is that first I started in large batches. Like, I sat with Berna, my colleague, and I'm telling her, go get some coffee, I'll go get some tea for myself, and we will wait till the agent completes this. It's going to take 20 minutes. Then I figure out that... Well, I'm not understanding what it's building. So little small batches are even more important now in this setting. And in pretty much any setting, I would argue that even if it is a sole developer producing the code with an agent, if they work in large batches, well, they are like outsourcing the development to some third party company essentially, not even grasping and understanding of what's being done. And that's always the recipe for failure in my opinion. So what I'm trying to say is that maybe AI isn't bad at all. Maybe it's just highlighting the dysfunctions so much like scatter-gather work where you do this, you do this, then we'll converge, right? It never worked well. It's not working now at all. It just stopped working. Same with let's test something after. It never worked well. Now it just can't work at all. Big batches, if they never worked well, now you just can't, because you're gonna end up having something completely different. So I'm very positive on this. Yeah, I mean, your research has reflected what Dora's research has found. We've always known, mean, small batches are an idea from the 1970s. Why aren't more people doing them? But they found that AI is an amplifier. So if you're doing big batches, the problems that come with doing that are going to just be exponentially bigger. And if you are doing small batches, the benefits are gonna be even better because you're gonna really take advantage of that technology. Yeah, I love that. Now there is a hope that finally because of having to serve machines, we'll finally listen to the things that humans have been telling us all along, right? switching gears to Dora research and capabilities, Lisa, maybe you could talk a little bit about what Dora brings to the table, your involvement there. And from there, we can pivot to the new findings that ⁓ go towards answering the capabilities that you need to build an AI ready organization today. Yeah, I joined the DOR community a few years ago with the suggestion of a friend who was actually helping me learn about AI back then. gosh, I think back to what we were doing back then, it seemed magical then, but it's even more amazing now. But anyway, she should join the door community. I'm like, oh, I don't know. That's probably a lot of site reliability engineers and platform engineers. And while I've worked with those people for my whole career, I'm not one, but he convinced me I went to one of the meetings where they do a 10 minute presentation and then a lean coffee on some topic. And to my surprise, all the topics ended up pretty much being about quality and testing. When you talk about continuous delivery and things like that, these are the topics that come out. And I felt very welcomed into that community. and because I have more free time now that I don't work full time, they asked me if I would be a door guide. And so ⁓ just to try to spread the word and get more people in the community, because like you say, we need people with all the different skills involved in coming up with these ideas. ⁓ So that's how I got into it. you know, the, one of the advantages of that is that we get to learn more about how they do the research. help, we can help, they give us the surveys before they send them out to get our feedback, things like that. They send us the reports before they send them out and get our feedback. And so it gives me a little more insight into what they do and understanding. So these days, like for the 2025 report, they not only do their survey, which I don't know, had 5,000 people or something return it. They also do qualitative research. So they call up, people and companies all over the world and they ask them, you what do you do? And so I think that's important because, you know, the people who aren't on the leading edge of using these tools probably aren't even going to fill out the survey. So there's a lot of people that aren't reflected in there. But when you actually talk to people and understand what they do, then you can start to see and hear more about the problems and things like that. And so, you know, they're finding that AI amplifies everything. It makes a lot of sense, right? but I don't think that's what people were expecting to see. Some things get better, ⁓ but in some cases, code quality, they find the code quality suffered. I wonder why that is, and we have to look at that. And so after a lot, they did the report for 2025, and then they kind of zeroed in to build. They already had the Dora core model, which is a capabilities model. And one of the things that sucked me into Dora is it's not focused on throughput, it's not completely focused on throughput and productivity and how do we get more out of, how do get 10x developers? It's about, how, what makes teams high performing? And then they find out, well, it's when people really enjoy their jobs, feel like they can do good work, they have good documentation to understand how their system works and where things are, and they don't burn out, they don't have the friction. Well, how do we get those things? How do we get those outcomes? And of course, they are more productive and the companies do perform better. So then they try to see what correlates with that. And so that's how they come up with these capabilities, which one of things I love is there it's about having a climate for learning what makes it easy for people to learn how to do their best work. And how do we get the fast flow? What capabilities contribute to that? How do we get the fast feedback? Lisa, sorry to interrupt, but one point that I want to dig a little deeper about was what you said in the beginning about how Dora was more applicable to quality engineering and vice versa than you had initially assumed, right? Like, can you talk a little more about this? Because this was a conception that we also shared when we were preparing for this podcast. Oh, well, I mean, when... You know, when I initially joined, was before the AI trade, now all the discussions are about AI, but then they were all about continuous delivery. And so how can we succeed? Well, you have to be confident about what you're delivering when you're trying to deploy to production every day or multiple times a day. And so how do we get there? And so that, you know, it kind of backs up into, well, What testing activities do we need to do? How do we have reliable automated tests? How do we make time for exploratory testing? ⁓ You know, the big bottleneck has always been things like, well, code review or if we have pull requests, then people get stuck and they wait around and what can we do to make that better? Which to me, that's something quality engineers need to get involved in too, because... if we're gonna review the code, we need to really have a good look at it, not just say, know, looks good to me. And so all those discussions, I mean, they always ended up really coming right into my wheelhouse of, you know, how do we get more reliable tests or how can we manage the risks better? How can we understand our users better? Because that user centricity is so important. That's one of the things that research has shown. And I just feel like as a tester, my goal is always to try to understand. the users and to talk to them. Sometimes it's not possible in some business domains, but at least proxies for the users. Like maybe you have to ask a marketing person or a salesperson, what do they think the users want? But it's important for me to do that and bring that to the team. Does that, does that make sense? Definitely. I mean, think what you said about confidence is even more true now about AI than continuous delivery. Right? mean, think we are in the... true confidence providing business as opposed to the overconfidence. I we spell overconfidence and replace it with qualified confidence, hopefully. ⁓ But yeah, when it comes to these capabilities, and let me just read out the seven capabilities that have been identified. And based on what you told me, these are capabilities that correlate with high-performing teams, which are both delivering, I mean, that are both functionally and from a- satisfaction, job satisfaction perspective. I think there are two aspects when it comes to high performance. One is it's delivering business value. It's operationally efficient. On the other hand, the people there are more satisfied to have less burnout risks and so on and so forth. But you could argue that these two can be collapsed to one because there are very many studies proving that happier people working in a high trust environment and psychological safety, they can be more productive. It's not a guarantee, but they can be more productive. given that the processes do not obstruct them from doing so. Absolutely. So when it comes to the seven capabilities that we got from the report, let me just read these out. The first one being clear and communicated AI stance, healthy data ecosystems, AI accessible internal data, strong version control practices, working in small batches, user centric focus, and quality internal platforms. Now I could list these capabilities out and remove AI because only one of the capabilities mentions AI explicitly and they would still be true. Like what has changed from what we knew before and with AI coming in. think very much changed. I mean, if you look at the original model, it's quite similar. I think they tried to maybe simplify it some and break it to fewer capabilities because there's a lot more capabilities on the original core model. And I think they picked the things that are the most important. Working in small batches, we can tell people that to wear blue in the face, but hopefully, again, AI amplifies the problems without so much more. And I think one of the things that kind of I hadn't thought of it until the research came out is that clear and communicated AI stance because companies that are already pretty dysfunctional and people are just floundering around trying to type faster. Maybe some individual developer says, oh, well, this AI, like maybe I'll try to use Claude or maybe I'll try to use Copilot or whatever. Well, maybe they don't know the dangers from a security perspective and nobody's told them, don't use that when you're working on production code. We don't want to leak customer information out to the, you know, so if they're not, or they might be afraid to try it. Well, I don't know if I'm supposed to do that, so I'm not going to do it. Well, will it replace me? If the company comes up with, you know, if the developer experience people at the company come up with, well, let's see, here's some guidelines. Here are things to be careful and watch for. Don't use these tools if you're working on production code or whatever it is. And here, I wouldn't tell them what tools absolutely to use. Hopefully you don't have to do that, but you can say, okay, you can choose from this selection of AI assistants or whatever, and let people have some choice on that. But give them the guidance, tell them when to do what, when to use which tool, and what to watch out for because the problems with security that could take a company down, you know, so. And the strong version control processes relate to that too. It's like we have to be very, it's like when Agile first came out, people were like, ⁓ we're just gonna throw everything out the window and go fast. Like, no, it requires actually more discipline. And now with AI, even more discipline. We really have to be careful. So we need those policies that are clearly communicated and available to everybody to remind people every day. how to avoid all the perils and pitfalls and enjoy the benefits. I could add here an interesting bit. So my marketing colleagues, colleagues from marketing, they want to build their own agents. And there's another research I've just published yesterday where I'm listing engineering principles and stories behind them. Classical engineering principles like least privilege principle. how it applies to agentic work. Like if you're a marketer, if you're working, if you're building an agent, which will do some competitive analysis, something else, you still must have some engineering principles. Otherwise the risk just skyrockets. And I even gave a talk here in Paris at Anthropic Meetup about using RAG. And marketers and sales were like, ⁓ my gosh, tell me about it. This is so interesting. So what I'm saying is that us engineers and testers now can help others. To me, it reminds me of how Excel emerged. And then everyone can do accounting, right? I mean, no one needs accountants anymore. Not, well, okay, they do need them, but for professional tasks, not for everyday accounting. So same thing here. You still need to have some discipline, some engineering principles. Even for my colleagues, it was a surprise that sessions must end before the context collapses. Otherwise you get performance degradation. they weren't educated about that. They didn't know the internals of the LLM. So it took some explanation time, but then everything is better now. That's a really important thing because it's not just this guidance isn't just for the engineering teams. That's really an excellent point. Yeah, that's really resonates because what I see on the projects I'm working right now is that the shift between not having any guidelines and having at least basic guidelines for AI use changes a lot, especially in case of shadow AI and things like that, because people still use it, but they don't know how to apply it. And ⁓ not everyone has the same background as ⁓ specialists in the AI field. And we need to take care of it really, because a lot of bad things. could happen and they have actually happened. And a lot of security leaks and a lot of companies lost money just because their employees didn't know how to use AI or used AI in secret. Yeah, I think there also need to be guidelines and support for people to realize. A lot more is happening every day. Sure, they're more productive. But if you keep working the same eight hours a day, but you're doing 10 times more, that's not going to help your brain. I keep hearing about BrainFry and the AI Vampire and things like that. And it seems to me that companies could be really thoughtful about saying, you know, we need to take more breaks and, you know, we're producing plenty. So take some of your time to take a walk in because we still need to be creative. We still need to be innovative. And if we still have the nose to the grind zone, we're not typing anymore, but we're still producing code and that's all we're doing. We're not going to be able to innovate and we're going to burn out. So those kinds of guidelines would be really important too. I was quite surprised to see that in Dora 2025 AI capabilities in this research that people self-reported Very little burnout. From what I see, I'm a workaholic. I can admit this simply and easily. I'm a workaholic. I love my work. I love doing things for the community. just love it. It was very hard for me to stop in the evening, but now with the AI, it is even way harder for me because everything is like possible. Like if I need a competitive analysis agent, I'll do it in 20, 25 minutes, five iterations, 25 minutes, done. And I feel like I'm doing so much during the day. Why don't I proceed? And it's 9 p.m. I'm like, ⁓ I'm so tired. And what's even more important is that while agent is doing something, even this tiny iteration, can watch another one in another tab. I can do something with another agent in another tab. And context switching is just crazy. ⁓ I have to apply more discipline to my work. And I'm sure I'm not alone in this. I'm sure that people are burning out. And what you said that companies should ask people to take more breaks. I don't think that many companies will do. The market is not good and companies want to say we are 10 times more important, more performant, even though no one can even measure what this performance is. Yeah, but by doing less, we might be even more productive. That's the thing that people can't grasp. That's true. ⁓ Yeah. One other interesting finding for me from the report was the metric on software delivery instability. And this is a metric that has gone up substantially, even as there are other metrics which have increased like ⁓ code quality, team performance, individual effectiveness, all these have increased. But at the same time, know, software delivery instability has increased as well. Could you talk a little bit more about what this metric means and probably hypothesize why this is the case? From what I've heard from the people in the Dora team, it's that, you know, is that, you know, we're letting agents generate the code. We can't possibly keep up with reviewing that code. It seems to work. And then it gets into production. It didn't really, it wasn't right. So there's, it ends up being a lot of rework involved and rework is a big contributor to Burn It Out. Nobody likes to go back and redo the same things. So people, like you say, people individually felt like they were more productive. Right. Because they're seeing, well, it's just the old thing. We counted lines of code, right? That's how you were productive. And they're seeing more things go out to production faster. But we have a whole software development life cycle. You've got to watch those things in production and see what's that what the customers needed. Did it cause errors we didn't expect because we didn't have time to do exploratory testing on that. And so then I don't know how companies handle this. I bring in a new story to fix the problem or whatever they have to do. We're saying, wow, we're doing the same thing we just did. I think that's one aspect to it. Just I think people probably are tempted to be a little miss and laziness can be really good. But we can become maybe too complacent and think, well, I've been reviewing that code every day for a while and it looks pretty good. So I don't think I need to do it. And like you were saying, Vitali, the agents to generate the skills files don't work as well as they used to, but we're not going back to recheck them. bet. mean, I wouldn't even know. conscious of that problem until you said that. I was like, oh yeah, I bet a lot of other people aren't conscious of it either in their drive every day to do more. So I think that those are some aspects of it. I would also add here that another interesting part here is that even if we... So for instance, to test my agents and my skills, I use deterministic checks and non-deterministic checks, a separate LLM as a judge. Okay, it works. Fine. It works to a degree, but it's all right for me. But then Anthropic cuts the intelligence for Opus. And suddenly it's degrading. I can't even detect this degradation in tests, because before I was using Claude with effort medium and it was absolutely fine. Now I have to switch to effort mocks to get the same type of intelligence that it was there before. I understand their business perspective on that. They got a huge increase in load. They want to cut it somehow. They add this, it's not even a degradation. They just put more thinking into smaller models. I'm not sure even how they do it. But to me, this is like a server running my code, but then starting to run it in a weird way. So with servers, the code either runs or it doesn't. And we have monitoring for that. Here, I don't have any monitoring. I just observe. that it's getting more stupid. And this is so strange. Yeah. And the other aspect of, ⁓ what makes us a little devious is because AI generated code tends to be very cogent, even if it's not correct. What I mean by that is it's formatting was going to be really good. It's going to follow all the, you know, all the rules of your linter. It's not going to have typos. It's going to have well-named variables, but at the same time, if there's something wrong, it just slips through the radar. And it's much harder to detect those things because these were the telltale signs that caused your nose to perk up and figure out if there's a problem in the code that has been committed, right? And you can see big problems coming on the horizon. I have a friend whose team developed, they were asked to develop the skills files and agents and all these things so that if you feed the specs, it generates all the testing, runs all the automated tests so that developers could do all that testing themselves. So they worked for six months. They built a really awesome platform for this. And then they were all laid off because we don't need the testers anymore because they've done this and now the developers have it. Well, if you're saying those agents are going to degrade and the skills falls are going to grade and nobody's monitoring that in a few months, they're going to have a big problem, right? Yeah. Yeah, so we'll look to wrap up this conversation now because we are almost out of time. But yeah, it was a great topic conversation where we touched upon so many different topics, I feel like. And I think one real good takeaway for me is the relevance of quality professionals to Dora and the relevance of Dora to quality engineering per se is something that is a refreshing. insight that I had from this conversation. But not only that, I when we look at what makes AI-generated code or AI-augmented teams more high-performant, it's not very different from what we've known all along as to what makes human teams more high-performing. that's nice parallel that we have. When it comes to one of the other success factors is documentation that is AI-readable is most cases documentation that's also human readable. There are nuances, but it's more or less in the same ballpark. Yeah, so any other last words or wishes to the community, Lisa, from your side before we wrap up? I would just say, I know we human, our human brains don't necessarily care about facts and logic. Unfortunately, we evolved a different way, but we need those. And so I would advise everybody to keep watching the research done by Dora, keep watching the research Vitali. is doing and the experiments he's doing. Look for people who are doing, Anthony Marcano, all these people that are doing experiments. Thank you for doing that and leading the way. But we need to watch what these people at the leading edge are doing because we want the fastest feedback for ourselves of what do our practices need to be? What are our leading practices that will help us the best? Definitely. us happy. encourage all of you to check out Vitaliy's research artifact on ⁓ AI assisted coding and what that means for quality engineering. think there is a ton of insights packed there. So definitely check that out as a follow up to this conversation. So we will link all of this good stuff and any other references that we brought up here in the show notes as we usually do. yeah, so thank you for joining us and we look forward to hearing from you or. hearing you tuning in the next episode. I learned a lot. So thank you. I'll be keep watching these episodes. Thank you so much, Lisa, for coming. It was a true pleasure talking to you, as always. ⁓ my, totally my pleasure. Thank you. I'm stopping. And I've stopped. you