SELLEST OSAST
If you have to do your taxes, do you trust the whiz kid from MIT or the accountant with 30 years of experience? Marc Cluet opens this episode with that question, setting the stage for a candid exploration of modern platform engineering and enterprise AI integration.
Drawing from his experience at Canonical, Rackspace, HashiCorp, and leading AI-driven migrations at a major global bank, Marc sits down with Nicky Pike to confront the core realities of modern software infrastructure. Together, they unpack what actually happens when AI is introduced to legacy systems riddled with fragmented standards and accumulated technical debt.
Marc reveals why line-by-line code translation fails and advocates for translating business intent instead. He details his "Michelin-star menu" strategy to balance developer enablement with strict architectural guardrails, while warning of the unavoidable 16-month productivity valley during AI adoption. Rather than accelerating messy code, Marc emphasizes standardizing systems first or knowing when legacy shell scripts are better off retired than migrated.
This episode is an essential guide for engineering leaders navigating the intersection of enterprise AI hype, platform governance, and legacy modernization.
In this episode, you'll learn:
- Why AI acts as an accelerator for existing technical debt rather than a fix for messy architecture
- How to implement a "Michelin-star menu" that balances developer empowerment with curated platform choices
- The realities of the 16-month productivity dip and why translating underlying intent beats line-by-line code migration
Things to listen for:
(00:00) Meet Marc Cluet
(01:34) How platform engineering became everything
(02:53) From gatekeeper to enabler of development
(06:15) The three buckets for every vendor pitch
(10:14) Making developers maintainers too
(12:00) The ten-year contract test for bad ideas
(14:05) Why 95% of AI pilots fail
(16:45) AI hype versus the old cloud hype
(23:02) Modernizing mainframes and COBOL with AI
(29:00) Why AI is not the new Stack Overflow
(34:21) Building the Michelin star menu
(41:00) The first MCP server worth building
(44:01) The mixing board every IT org is tuning
(48:05) Rapid fire questions
(50:00) Predictions: AI ops roles three years out
(52:09) Defining a Coder: The double-edged sword
Resources
Marc Cluet's LinkedIn: https://uk.linkedin.com/in/marccluet
NÄITA MÄRKUSI 🔗
TRANSKRIBEERI 🔗
If you have to do your taxes, will you trust a whiz kid mathematician from MIT or would you trust a chartered accountant with 30 years of experience? Most AIs had the whiz kid from MIT, so they're really good at what they do, but they know the exact details of what you want.
Nicky Pike (00:16):
This is Devolution, bringing development back to speed, back to focus, back to freedom. I'm Nikki Pike. Okay, so everyone's talking about AI ripping through legacy code bases, translating 20 years of technical debt overnight, and modernizing whole systems while you sleep. But here's the part that nobody's asking. What happens when you point that thing in a system that's already a mess? Does AI clean up the wreckage or does it just build a faster, more confident version of that same wreckage? If your shop has 30 different ways to install the same Java VM and you hand that chaos to AI, what exactly do you think it hands you back? My guest today is Marc Cluet. He's a platform engineer who spent his career on the messy end of infrastructure at Canonical, Rackspace, at HashiCorp, and he's currently driving AI-driven migrations inside one of the biggest banks in the world.
(01:08):
He spent this year in the middle of the exact fight that everyone's about to have. What happens when you turn AI loose on decades of technical debt and who's responsible for cleaning house first? Mark, welcome to The Devolution.
Marc Cluet (01:22):
Yeah, thank you so much.
Nicky Pike (01:23):
It's good to have you here. Anything you want to say before we jump into questions, my friend?
Marc Cluet (01:28):
It's an honor to be here. Thank you so much for having me. Yeah, I hope it will be a very good and interesting hour.
Nicky Pike (01:34):
Oh, I'm sure it will be. Well, let's just go ahead and jump in. You have done platform engineering at all of the companies that we've stated above and now inside one of the biggest banks on the planet. What's the thread connecting all of those together and when did platform engineering stop just being a job title and start being Mark's thing?
Marc Cluet (01:52):
So that's a really good question. The thing here is that platform engineering has evolved throughout the years and has gone from having to go to data center with your CD of the latest Linux Distro and having to store everything manually connect all the ports and everything to something that is very automated now. So we had to change the way the profession works. We had to learn how to code like developers if we didn't know. We learned how to automate and grow the amount of service that you dealt individually. What changed in all of this is the volume of things that you're dealing at the same time, the abstraction level as well that you have right now over everything and any kind of mistakes, any kind of issues then become amplified because you're not affecting just one system, you're affecting 20, 100, 200 systems for a whole company in one go, which before you would have on if you touched the main firewall or the main router.
Nicky Pike (02:53):
And I don't know if you've noticed this, but me being an old platform engineering guy myself, I've seen this iteration where we started off as just the people that provided something to help the developers. And then we became the system that developers ran on and then we became not only the bottleneck, but also the enabler of different development systems. What are your opinions on that evolution that we've seen through platform engineering? Because to me it's becoming a bigger role than I think anybody ever actually thought it would be. Everything runs on a platform now.
Marc Cluet (03:21):
Everything runs on a platform. Everything goes through a platform to go to production in an automated fashion. So the same thing that happened back 20, 30 years ago that when everything went fine, nobody knew about us, nobody even remembered as we were there, it's still the same. But at the same time, one of the things that have increased in a way is that with the whole movement of DevOps and taking down all the silos and everything that has happened throughout these last 15 years, the biggest thing now is that developers have a lot more access to things that they wouldn't have access before. And that's inherently a really positive thing. So developers in the past would be like, "Yep, I finished developing. Here you go, ops guy, do your thing and put it in production." Where now they can just click a button, everything happens.
(04:11):
That doesn't mean inherently that developers know exactly how that system works or they have the expertise, and that's a human thing. It's not that DevOps by breaking the silos made a mess. It's basically that people are good at what they're good at. If you ask me to develop Java, I will make a mess. I'm not a Java developer. I will try my hardest. It won't be pretty. Where if you ask me how to develop a deployment system or a cloud system or orchestration system, I'm the person. I will do it. It will work. It will be amazing. With DevOps, some people saw that of a meshing both developers and operations together and working together, and that worked for a while and it works for small companies. When you're going to bigger companies, that still doesn't work because the person that's developing the code might be in a completely different country and time zone than you and you are making sure that the platform is working.
(05:11):
This is the very interesting thing, right? By enabling developers, we have got ourselves in a very developer-like issue, which is as you can see in people like Alan Jacob, the people that were working at Parpet Labs and a lot of people that have been talking and shouting off the rooftop about the infrastructure of code, they all have this problem now that infrastructure as code have become this large code in a way because this code has not been maintained for five, 10 years and it shows its age. It shows that it starts to have cracks, it starts to not work well. And we have created the same issue that we had in the past. The issue is snowflake code. A person that was in the company five years ago wrote this automation. Nobody knows what it does, but it's critical to go to production, so we're not going to touch it.
(06:05):
And that's the kind of thing that you're dealing with now and it's the kind of thing that hopefully with the next iteration of platform engineering and AI, we will be able to solve.
Nicky Pike (06:15):
Well, and let's pick at that thread for a little bit because when you were at Canonical, you said vendors would walk in and they'd want to see their products and their applications baked into the operating system and you'd sort every pitch into one of three buckets. You build it for free, you split the engineering or you tell them that you're on your own. I think a lot of that applies to how we're building platforms today. So walk me through how you made that call and what did it teach you about the fight that we're seeing that platform engineers are having to pick with everyone around them?
Marc Cluet (06:44):
In Canonical, we did this thing and it was not something that I invented. It was something that we did as a practice. Canonical is an organization that the main benefit is towards the public and towards the community. So anything that any vendor would go and say, "Hey, I would like to have this thing running on Ubuntu or running on this or that," then we would look at it and say, "Okay, what is the value of this towards the public?" If it's something that everybody wanted, for example, Docker, everybody wanted Docker. So it's like, "Yeah, of course we'll bring it in. We'll integrate, we'll do whatever we need and it will run." Where there were other requests from vendors that were a bit more, I would say, gray line in the sense that there was something that was very niche or that would only benefit a very small amount of people.
(07:35):
Then we would say, okay, we are taking this code in, but then if you pay for the maintenance of this code, then it's fair because we have limited resources as everybody does, and we need to ensure that we use our resources for what benefits the most. And if it was something that it was halfway through, then we would split the bill. It's a lesson that I took with me and I applied it in many places and I'm applying it where I'm working right now as well in the sense that I'm trying to simplify the ways we are deploying, as you said at the beginning, 30 ways to deploy Java. And you end up with SMS if you let everyone automate or script their deployments because two different persons will always have two different opinions on how to do it. And if you have enough people, then you have a lot of different permutations of that.
(08:27):
And the biggest issue for all of these of course is that you need to then have a support team able to handle all of these plus making sure or hoping that people that we're creating this code don't leave the company because otherwise we start having Snowflakes, we start having issues. So in that sense, one of the things that I've done repeatedly that works fairly well is saying, okay, let's install Java or let's install Nginx. We're going to do it in one way and then you can configure and extend in whichever way you want. So the Nginx module or collection becomes the muscle and then all the variables and all the configuration becomes the brain. And then you apply your brain because you know exactly how you want that to work and we'll bring the muscle. We'll bring exactly how to do it and it will be one single way for everyone.
(09:22):
The good thing about that is that you have to maintain one version, not certain, but also everybody that was contributing to the 30 versions can come and say, actually, there's this really interesting module that we can enable in Nginx that does open telemetry metrics and then we can just start that as a feature. So we end up with something that is feature reach. The features are enable or disabled depending on what the user wants and everybody can make good use of it. Where if someone comes in and say, I really need this exact way because I don't want to change anything, then you say, okay, then you run it, you maintain it and here's all the list of the risks, security audits and everything else that you need to do and destop lucks. So it's that compromising between what works for the most versus this is the way I've done it and I don't want to change.
Nicky Pike (10:14):
I think that makes sense. When we were originally talking, you were talking about canonical and an operating system, but I think a lot of those same virtues and values bleed through into platforms. I think platform teams have to deal with the same thing. And if I heard you right, you're an advocate of developers not only being consumers of the systems, but also help being maintainers of the system. Is that correct?
Marc Cluet (10:33):
Absolutely, because you're a developer, you have certain interests as a developer. And in every single company that if any, including the one I'm in right now, you have developers that are really seriously good at platform engineering and they know better than even our team how to do certain things. So you want to ensure that you enable all these developers and all this engineering capacity to create something of really high quality world-class while at the same time ensuring that all the other people that, to be honest, they really don't care about deployments. They just want to get the codes to production and they really don't care how it gets there. You make sure that you make it easy for them. You make sure that it's like, you know what, you don't need to write a script, you don't need to write all of this, just put something in this JSON file and we'll deal with it.
Nicky Pike (11:26):
And I get that. I think one of the things that I see is, especially in a company your size, you have a lot of developers out there and I think you got to apply some of the logic of what you said because developers can come in and they want something that may be very specific to them if they're also maintainers of this. How do you make that split between no, this is something that's just for you, this is going to make our platform a little bit of a snowflake or no, this is something that we actually do see over our population and would actually either innovate or broaden the use of the platform that we're providing for our company?
Marc Cluet (12:00):
Yeah, so there's a way that it's a bit brutal sometimes that I use, but it's kind of interesting as well. And what I do is that when I normally do the typical thing of, okay, what's the business value of this? Why do you want to do this? Et cetera, et cetera. A lot of developers, including myself as an engineer, it's like I have the brightest idea in the world, but then I realize that that idea doesn't fit and doesn't give any value. Some things I get too excited, but most of the times I'm like, okay, yeah, no, I'm not going to do that. Nobody needs that. But there's other engineers that are a bit more unfiltered in that sense and they go like, no, no, I really need this. I really want this and I already made the code, just need to add it to your code repository and it will just work and it would be amazing.
(12:48):
And I just tell them, it's like, fine, no worries. I'll do that, but can you please just, we'll call HR and you sign for me a contract that says that you cannot leave the company in 10 years. And they look at you as like, what? It's like, well, this code will be around for 10 years, so let's make sure that you also are around for 10 years because otherwise I'm going to put this in production, other people will use it. And one day maybe you have a great offer to go somewhere else or you decide that you want to retire as a lot of IT engineers and be a farmer for some reason. And then who maintains this code because you know how it works. It's amazing, but nobody else knows how to maintain this. So just sign this 10 years, it's fine. I mean, it's your code, it's amazing.
(13:37):
So yeah.
Nicky Pike (13:38):
Give me a percentage. How many people accept that?
Marc Cluet (13:42):
Zero.
Nicky Pike (13:43):
0%. That' kind of what I saw on that.
Marc Cluet (13:44):
0%. Yeah. Because then suddenly the idea is not so worth it.
Nicky Pike (13:50):
Yeah. When you're putting somebody's neck on the line like that, I think you'll see some back off or definitely some reprioritization of what they thought was important.
Marc Cluet (13:58):
Yeah, definitely. And I said, and don't get me wrong, I'm very excited about my own projects, so I fully get it.
Nicky Pike (14:05):
Well, here's the challenge I want to put on the table. So MIT's NANDA initiative tracked over 300 enterprise AI deployments in 2025, Mark, and they found that 95% of them never measured a dollar of business impact. They never delivered a measurable dollar of business impact. 95%. Now your argument is that AI isn't what failing here, it's what's underneath it that's failing. So what actually separates the 5% that work from the 95% that don't? And before we hear your answer, Mark, that question isn't just for you, it's also for the audience as well. If anybody in the audience is set in a room where leadership pointed to some AI pilot as the future and you already knew it was dead in the water, drop a comment and tell us what actually happened. And while you're there, make sure you subscribed because we're about to get into exactly why this keeps happening.
(14:54):
So what separates that 5% from the 95?
Marc Cluet (14:57):
So there's other factors as well here. The main one is that as any new tool, and you see this not only with AI, you see this as well with cloud at the beginning, you see this with all the trends. Everybody suddenly goes, "We need to have this, otherwise the competition will hit us, but we don't know for what, just make it happen." So suddenly then you have a hammer, everything becomes a nail. And that happened with the cloud. If you remember, everybody was suddenly pushing every single workload to the cloud until they got the first multimillion bill from Amazon and they were like, "Maybe this wasn't a great idea." I think the main problem and the main thing is intent because AI, it's an accelerator, it's a great tool, but if you don't know what you're using it for, it's a problem. And then the second problem you have is data quality for the training because there's this thing, and just excuse my French, but this thing that Cory Doctor says about products and time and the initiatification of a product, the same happens with data.
(16:07):
If you send all this data and feed all this data to AI and all of this data is not showing best practices, it's not showing how to do this in a way that is world-class, then you're going to get bad mediocre codes and you're just going to add to the mediocre code that you have. So before doing any kind of AI rework on your code, the really important thing, and I cannot stress this enough, is be very careful about what you feed to the AI, because if you feed garbage, you will get garbage out. You will not get suddenly something magical that solves all your problems.
Nicky Pike (16:45):
Well, and I think that you mentioned cloud. I mean, this is the same kind of shape that we've seen with all of these different trends and the hypes that we've seen on technology through our career. The one thing I do see a little bit different here, Mark, is the fact that if a company said, "Well, we're not going to go to public cloud," okay, that was their business decision. But with AI, we're seeing something a little bit different of if a company says, "We're not going to AI yet," it affects their stock price. We see stock prices drop because of those kind of statements. So I think part of that, and see if you agree with me on this, is the fact that if they're rushing to get to AI in and they're not thinking through exactly how they're going to use it, and that's part of why we might be seeing that 95% failure rate because they're doing it for reasons other than just amplification or increasing a business value.
(17:34):
They're doing it so that they might keep a stock price or they don't have a bad PR release. Any thoughts on that?
Marc Cluet (17:40):
Yeah, that's a really good one as well. And yeah, I fully agree. There's so much pressure to invest in AI. And I think that pressure is in a way, kind of suspicious that it's in a way made as well by all the big corporations that do benefit in AI, the NVIDIA, the Microsoft, Anthropic, OpenAI, et cetera, et cetera. Because the more AI is there, the more need for memory, the more need for CPUs, for GPUs, and it's a self-fulfilling cycle. And I said, I mentioned cloud in that, but for example, and this is going back slightly more, when we had the dot-com boom back in the '90s at the beginning of the 2000s, a lot of companies were like, "If you are a tier one carrier for the internet and you're not investing $1 billion per year in new fiber, you're going to disappear." And everybody went crazy putting new fiber in.
(18:38):
And a lot of them went bankrupt because they overspent, but their investors were like, "You need to invest, or this is a race. If you don't invest, then you become second and nobody wants to be second." And I said, we ended up with a lot of very cheap fiber afterwards because they all went bankrupt. The bills for installing those fibers were not fully paid. And then we actually had a boom in internet bandwidth, which was great for everyone. But all these companies then, yeah, they went down. The other thing about this to have in mind as well, I was in a talk recently from one of the doctors at Stanford University in computing that he said, and he showed with numbers, that if you invest in AI coding, you go through a value of 16 months. So for 16 months, you are committing that the value that you're getting out of your engineering team and the value that you're getting from your code will be less.
(19:37):
And after 16 months, it becomes again what it was. And after that, then it goes up exponentially. But you need to go through those 16 months of we're going to be less effective, we are going to spend the money on this and we're not going to see a benefit.
Nicky Pike (19:52):
And I think that makes sense because there was this promise of AI, right? Everybody got excited about AI translating code function by function, line by line, one language to another. That's what really gathered people's interest. This is going to help us fix our old technical data. It's going to help us fix our code bases into something that's a little bit more manageable. But you have made the point that going by that approach, it mostly doesn't work. So when people do that, when they bring AI in and they try to go line by line, function by function, what do you see actually breaks when people try that approach?
Marc Cluet (20:27):
The comparison that I give with this normally is that if you translate a language like Russian word by word to English, you're going to get very, very difficult English. And the same happens with languages. You have languages that are structurally very different one from the other, and there's some advantages in some of the structures in those languages and how the code is formed around the language capacity. Now the problem is that the first wave of AI tried to do that. I have language A, I'll just translate to language B and profit, everything will work. But it resulted in really difficult code or code that just wouldn't work. So one of the things that you're seeing a lot more now, and I think that's a bit more the natural way, is that AI is translating code in two phases. They are grabbing the original code and translating that into a spec, the requirements of the code, what the code is actually doing instead of saying line by line and moving the speed, printing this thing.
(21:32):
And it's capturing the intent of the code because at the end of the day, developers for many years have been that, have been people that have been capturing the intents and putting that in code language that was deterministic. So with AI now we have these specs, these requirements of the code, and that expresses what the code was doing, and then we can translate that to a new thing. And I've seen some examples of these with old mainframes and COBOL developers and COBOL language. COBOL is very different to Python, for example. So I've seen some of these successful transformations from COBOL to requirements and spec, and then from Spec to Python and Elite Works.
Nicky Pike (22:17):
This is where I think that your insights from the financial industry are huge here because that's a primary use case. How do we get things off of mainframes and onto something more efficient? Everybody's behind this. Mainframes are expensive, so the companies don't want to keep maintaining them. The cloud companies want people to get off mainframes because that compute has to come from somewhere. It's either coming from the cloud or it's going to a vendor, and we're starting to see a skills gap in COBOL and some of the mainframe languages. So that modernization of mainframe to whatever language you want is one of those places where they're seeing AI could make a huge difference. Are you actually seeing that come to reality? And if it is, what is your process to make sure that it works for them?
Marc Cluet (23:02):
So I've seen some of that. I've not done mainframes directly myself because of course I'm a platform engineer, but I've seen some of that examples work well. And considering that nowadays, if you want a COBOL developer, normally they will charge you something up to a million per year because they're not that many anymore. So it becomes a very sought for skill. Then it does actually make a lot of sense to try to just invest in AI and try to transform that code into something that is a lot more maintainable. What I've seen this work with very well is the area that I'm more a specialist for with these deployments. We have a lot of infrastructure as code and a lot of scripts that define deployments in a way that was working well, let's say 10, 15 years ago, but nowadays it's more of a risk than anything else.
(23:51):
And we go back to that position in which deployments become risky, and that's something that we got out of now with the, I would say, the code getting more and more ancient. We're going back to that, to all the deployments being really risky and non-Kubernetes environments, obviously. The part of this are being able to get that deployment and say, what does this actually do? I copy this for it, I'll check this, I disable this in the load balancer. I make sure that the library is there. And you just put that in words that anybody can understand. And then the end you convert that into something new, something based in Yammer or in JSON that is easy to read, that is machine readable, that then you feed to a new deployment platform. And one of the new, let's say, CD platforms that you have available nowadays that will take that and fly with it and make sure that it's robust, that it's risky, that is not risky.
Nicky Pike (24:49):
Well, and so you're talking about technical depth there, and I think there's nobody that's going to be listening to this interview that hasn't seen that script that's sitting out there for 20 years that nobody knows who wrote it, but we're not sure how important it is to our overall function. Server that's sitting in the corner that nobody wants to unplug because we're afraid it'll bring everything down. When you look at things like that, if you find those scripts or those pieces of code that are sitting out there or part of a platform and servers that you don't understand, and you go and point AI at those to try to understand them because whoever wrote them is gone, what do you actually see from that? What happens when you point AI at code that's been sitting out there that nobody understands, but you're afraid of it because it may be load-bearing?
Marc Cluet (25:30):
So there's two things that happen there, and I think it's really important to have in mind that AI, it's still not, and hopefully will not be for some time, a replacement for human critical mind. So you have all this code that was doing something and then it converts it into something that you can understand. But if you don't have the skills or the experience to say, "Hey, actually this doesn't look right. What AI is saying is a bit out there. It's not really what this should be doing," then just going to, again you could going to again accept something that is actually not fully true. So in all these processes, one of the things that is really important is that AI normally will not fully get it right at the beginning because AI is not great at intent. It will assume it will hope that that's the thing.
(26:26):
Statistically that might be the thing, but it's not the way that was actually working because it might not be exactly in the statistical curve. So you need to have the human capacity to look at it like, no, no, actually have a look again at that function or at that piece of code over there because this doesn't actually sound right. And then of course I go, "Oh yeah, you're absolutely right. They got it wrong because that's how they work." And they will always accept that they made a mistake and move on in a very cheery way. But yeah, you still need that capacity and that critical mind to say, yep, this sounds right or actually.
Nicky Pike (27:00):
I think you said something so important that might've been a little bit buried there, and I hope that the audience is listening, especially the younger generation, you still need the judgment to go in and make sure that what AI is telling you is correct. And I think a lot of people that are coming up in the AI era right now, the junior developers, the newer and junior platform engineers, they're worried about your jobs. But one of the things I always tell them is what we're seeing AI automate are those things that you would've spent your first two years doing anyway. It's taken away some of that toil. Where you need to focus is that judgment gap. You've got to be able to understand if the AI is telling you something very confidently that, oh, okay, it's confident, but it's still wrong. So you can't get rid of those base skills.
Marc Cluet (27:43):
Absolutely. And for example, Matt Pocock, who is someone that has done a lot of AI videos, one of the things he repeats continuously is that AI does not replace software engineering, does not replace basics of software architecture, computer architecture, because all of these things is the critical step that you need in order to take the most capacity out of AI. And in that sense, of course, I mean, I've been in the industry for a while, and I assume so have you, you have that capacity. You look at something and go like, "Yep, this actually doesn't look right and it's this thing here." And then a general engineer comes and says, "How do you know that?" It's like, "Well, I've seen that 40 times and it's always the same, so it didn't change." So being able to have that, "Hey, that doesn't look right, or let me look at this, or let me just check this quickly because these sounds a bit out of place." And it's the same mistakes that you see in a different way when junior developers were copying pieces of code that they didn't really fully understand what they were doing.
(28:49):
It's like, "Oh, this solves my problem. I'll just copy this from Stack Overflow and that's it. It will work." And then suddenly half of your production systems are down and you're wondering why.
Nicky Pike (29:00):
So you bring up Stack Overflow and AI has become the new Stack Overflow. However, one of the things that I tell junior engineers or people that I talk to is it's not the same because a lot of people are just taking AI. It's artificial intelligence. They assume it's right and they're just copy and pasting it. The difference between, well, we did that with Stack Overflow. Well, kind of. You went through, you probably had to read three or more or five different things in Stack Overflow. You found all the things were wrong and you learned something from every one of those that was wrong. We're skipping that with AI. They're not getting that, hey, we definitely know it's wrong and going in and looking at it. So I love that statement about Stack Overflow because that's something that comes up all the time.
Marc Cluet (29:42):
Yeah, absolutely. And that's the thing. One of the things that I think AI is great on and that those people don't use it for that is just sit down and tell it to explain you something. It's like, explain me this concept. I want to understand how this works. And even if you have a vague idea of what that AI is really good at doing that and explaining a concept in a really good way. So if you are afraid as a general engineer that the AI will replace, it's like now AI will actually empower you, first of all, but then you will be able to have that critic mind and that capacity to say when AI is doing something wrong when you actually need to ensure that it's not. That non-determinism of AI is something that can actually create a lot of technical debt and problems if you don't truly right.
Nicky Pike (30:28):
Well, and I think with you leading and running some of these AI migration projects, it feels to me like that's a good starting place for junior developers in AI. And the reason I think that is because you've got a working example. The code's already out there. We need to transform it, we need to modernize it, but I know what good is. I can ask AI to explain to me what the existing code's doing and then translate that into modernization. Now I can still have AI help me write that, but if I got that understanding, then I have that judgment call to say, is what AI is doing right? And I think that's a lot different than greenfield applications. Agree or disagree with that?
Marc Cluet (31:05):
I do agree. And with greenfield applications, I would say one of the things with AI is that you can prototype so fast and it's something that I find incredibly useful. Like this afternoon I was prototyping a new Ansible module for division between collections and playbooks in order to have the separation between muscle and brain. Now we're just playing with it locally. And in a couple hours I went through something that would have taken me easily five days before, which is incredible. But I would say AI is great about prototyping, but once you get into something that is or has been already in production or needs to be in production, that's when the stakes get a lot higher and that's when that determinism really needs to be there. And if you look at the guys, for example, of Swamp Club that I follow very closely, they keep saying the same, determinism in AI, you can get as close to it as possible by making sure that the destination is really well framed.
(32:10):
And if you say this is the destination and you cannot get out of this box, this is the exact destination you need to hit. AI is good at doing that once you tell it exactly what to do. Another comparison of these, I was watching a video the other day and someone said, if you have to do your taxes, you have to do your annual taxes, will you trust a whizzkid mathematician from MIT or would you trust a chartered accountant with 30 years of experience? You would trust the accountant because the whizkid from MIT might be incredible at math, but doesn't know back and forth how tax legislation works. And he said, most AIs had the whiskey from MIT, so they're really good at what they do, but they know the exact details of what you want.
Nicky Pike (32:55):
I couldn't agree with that more. And I think that's a huge statement as well is people start to get in trouble and they start to see variations when they leave decisions up to AI. Because if you leave the decision up to AI, it's going to make a different choice every time. But like you said, if you can confine it to the box, this is what you're going to do, you get pretty good results out of that. Where you start seeing drift is where you start saying, oh, go make a decision for me AI. Well, okay, that's going to be different every single time you ask.
Marc Cluet (33:25):
Absolutely. Sometimes it's fun, don't get me wrong, especially when you're bouncing ideas. But yeah, for production systems, definitely not.
Nicky Pike (33:33):
So here's where we're at. The tech isn't lying to us. Our own mess is. Mark just laid out why bolting AI onto 20-year-old shell scripts and inconsistent deployment pipelines is a recipe for garbage at scale. Next, we get into what he calls the Michelin star menu, the exact framework he uses to turn infinite chaos into a curated set of choices that AI can actually work inside of. Plus the new job title he thinks every engineering org is going to need in the next few years. And if you still haven't subscribed, now's the moment. Hit the subscribe button before we get back into this. Stay with us. Well, I definitely think, and it goes back to that greenfield, you got a little bit more latitude there than you do with migrations because the code needs to operate the same, have the same outcomes regardless of what you migrated to.
(34:21):
All right, well let's get into a little bit about how we solve this. Now, one of the things that I absolutely loved is you described yourself that you were trying to build a menu out of your infrastructure and you had this Michelin star version that every install and every deployment pattern should go through so that AI only ever picks from the vet of choices. Walk us through exactly how you build that menu without just recreating a mess in a fancier wrapper here.
Marc Cluet (34:48):
So it's a really complex thing to do because if you do that with an engineering team, it will take years to go through all the iterations of making sure that every single thing that you do is of the best quality. And that's where communities and that's where enterprise subscriptions do, things that have been written directly by the manufacturer, things like that work really well because you have your window of experience and you know exactly how this thing works within your environment, but the whole world has a bigger, broader window of experience. And if you could concentrate all that brain juice in a way and all that capacity, then you would have something that would be world-class and would be able to solve everything. So I'm trying to reproduce that problem by using the community brain in a way inside a big company because that works well.
(35:42):
There's enough people interested in anything from databases to log balances to networking to Java. So you have everything there. So you just need to make it very easy for them to be able to contribute. And you take that from the open source community, it's really easy to contribute to the open source community and you just need to want to do it and to find something that you can actually solve. By reducing the amount of choices, you make those choices better. And again, we're talking about 30 ways to install Java. If we make that one single way to install a JDK, one single way to install Tomcat, that single way will have all the bells, whistles, checks, balances, monitoring, traces, everything that you need. And it will be feature rich. It will be a good experience where you can just enable exactly what you need. And then I equate that to the Michelin star of that code in the sense that it's something when you go to Michelin Star restaurant, you go there for the food, but you go there for the years of experience that that chef had trying to make that Filet Mignon better and spent 10, 15 years of their life obsessed in making sure that that was the best.
(37:03):
So I'm trying to reproduce that in a shorter time, make sure that that's the best that you can get.
Nicky Pike (37:08):
No, I do think that makes sense. I think the way you described it though, there feels like there's a little bit of a paradox. So we're talking about, and we talked about it just earlier, about limiting AI's choices makes it more predictable, not less useful, it just makes it more producible, like only offering them vanilla versus chocolate for every flavor. But there's also this aspect of we can get things, there is an intelligence built into AI that maybe we want to take advantage of. So where do you think the line is between giving AI too many choices and giving it so few that you're actually kind of limiting what AI is going to be able to do for you?
Marc Cluet (37:44):
So I would say it's the same as monolith based of microservices. You need to reduce the problem. So going back to the Tomcat example, there's 50 different flags that you can have for Tomcat and install Tomcat anywhere in the world with the amount of memory, with the libraries that it supports, with the metrics libraries that it can integrate, et cetera, et cetera. And if you tell AI, build me something to automatically install Tomcat that has everything, it will do a bad job. Where if you say, build me something that only installs Tomcat and does nothing else, it will be a lot more deterministic in the output and then say, okay, you've done this. Now add this more thing over there and then you keep that as small micro additions to what it was doing. It's like just don't touch anything else, just add that thing.
(38:42):
The same way that you're doing a sprint. If you go and say this will take three days or three engineers to do, you break it down. You break it down to something a lot smaller. So it's the same concept in DevOps would apply it to AI.
Nicky Pike (38:56):
And does that go back to this, again, developers being maintainers of the platform that you're building as well as consumers that, okay, here's the opinion of what we need AI to do from an installation and deployment standpoint for Tomcat. They get to come back and say, okay, we need this thing and this thing in addition to, rather than just offering them a full menu of Tomcat. You kind of limit the menu, but you add to it as you start developing people's tastes?
Marc Cluet (39:24):
Yeah, exactly. And it's important to have that friction because if you just give everything to everyone immediately, then in a way people will be unhappy because it's like, oh, this does too many things, this does all these other things, but then now I want this other thing here. But it's really unfair that you're not adding these really small thing when you added already 60 more things. So you need to add a bit of friction and that bit of friction is say, okay, I'm giving you something that does the minimum value product. Oh, but it doesn't do the thing I want. Let's talk about the thing you want. What is the value there? What is the business value? How do you make sure that that works in something that is effective for everyone? And then you have that discussion. And recently I have that with some authentication credentials, storage solution.
(40:19):
Different parts of the business had their own and they were like, we want to just port them as they are. And we ended up with something like five different teams wanting to port their solution. And I just got all of them in a virtual room and I said, okay, we are going to all sit here and we're going to talk until this one solution. And whoever that solution is, it has to have the requirements that we all need in a way that makes sense and everybody will contribute to this because this will be the way to do it. And you need to make sure that those conversations happen because otherwise you end up with everything in the kitchen sink.
Nicky Pike (40:58):
Now, the way I'm understanding you talk about this when you make these decisions and hey, here's the output we're coming in, this feels a lot like we're creating multiple MCP servers. Here's the MCP server to do and define that menu that we put in and we're going to call that when we need it. And I believe that you've said that you think that MCP is one of the best inventions since LLMs themselves, but it feels like that's the way to help keep things consistent across your entire population. So on that line, what would be the first internal tool or server that you'd actually stand up to make that real? Not theory, what do we do in reality if somebody's listening to this?
Marc Cluet (41:36):
So I would say, especially if you are doing an AI first organization, the first thing you need to do is to have a central MCP server. And I said MCP, I think is a great invention. I mean, hats off to Anthropic for this one because it's great because MCP servers make you reflect and think because you've seen all the studies and all the different people trying these. If you have an MCP server with 40 tools and you give those 40 tools to every single model, it will do a horrible job because it will be death by choice. It will use most of your context window. It will use your resources just trying to think what's the best tool to use. Where if you give it you have tool A and tool B overhead, then it's a lot more optimized. It indicates part of the issue that we have.
(42:30):
As humans, we have a lot more bandwidth, we have a lot more capacity. Of course we want 50 tools. It's great to have 50 tools, but at the end of the day it might not be the best use of your time. Well,
Nicky Pike (42:39):
And I think that also helps in the platform team of really monitoring and guardrailing what the agent's using from a logging and audit standpoint as well. If you're trying to put everything in one MCP, you're not able to separate those logs out. You have to look at log line by log line to see what the AI's doing. But if you create separate tools through MCP servers and they're doing separate functions, it feels like it makes it a little bit easier for the platform team to really understand one, not only what tool is providing value the most and where are people seeing versus those that are sitting on the fringe, but also to help with that improvement and that iteration because now you're able to take just that section of laws for a Tomcat server or a Kubernetes MCP or whatever that may be. Is that kind of where you're thinking here?
Marc Cluet (43:26):
Yeah, exactly. It's absolutely that. And also said by having these more MCPs in a way, people avoiding what is the most popular thing to use because maybe you thought that it was the best idea in the world to have a tool that read JSON and transformed it into YAML. But if nobody's using it, then hey, it's not worth running. It's not worth maintaining. It gives you a very clear idea then on where do you need to invest? What is the next thing that you need to do? What is the improvement that you need to do on that tool because that's the one that is being popular. So you need to make sure that it's faster, that it's more effective, and that it gives what people need.
Nicky Pike (44:01):
You did this comparison of you called what we're seeing with AI and how we're bringing it in. You compared it to a mixing board where every fader switch that we have is connected by rubber bands. You push one up, something else goes down, vice versa. What do you think the first thing that's going to go wrong for a company that's trying to move too fast on AI without doing this groundwork first? What sliders are they moving that other sliders that they weren't expecting or getting moved around as a consequence?
Marc Cluet (44:30):
Yeah, so that's a theory that I've been explaining for a while in my presentations. And it's just because I'm fairly visual, so I like that representation, especially because I love music as well. So if you turn suddenly the singers mixer to 10, then you won't hear the rest of the band and the singer might be a bit too loud. So you need to make sure that you balance everything right. And especially when you do a transformation, I always see all these sliders being at zero, one because any transformation and any new technology is just accelerating what you're doing. So in a way it's taking you all the way to 10 or 11. The problem with that is that all these knobs in the mixers, they are all connected with rubber bands because what happens when you pull a rubber band too much, it breaks. And that's exactly the same that it happens in any IT organization.
(45:19):
If you suddenly say, "You know what? We have this amazing tool that does CI and testing automatically and you don't have to give it anything, you have massively to use the time for testing, but then that will add pressure to the next step. If you do testing super easy, then building artifacts will be a thing that everybody will be doing 200 times per day." And maybe the artifact server couldn't handle that and he's like, "Oh my God, what's happening here?" So you're creating that tension. And the same happens with cultural transformations and skill transformations. If you accelerate something too much with AI, that pressure will be shown in other parts of the process. And if you look, for example, as stream value mapping, which is something that I really love to do when I look at the process, you can see exactly where the bottlenecks are.
(46:07):
If you accelerate something that is already not a bottleneck, the bottlenecks become bigger. So if something was waiting for three days when you were delivering 10 of them per week and you suddenly do a hundred of them per week, it's not going to wait less than three days. It's possibly going to wait seven or 10 days because the queue becomes bigger, so you create more tension in that coin. So it's really important when you do these transformations and accelerations to look at the whole thing as a holistic system instead of looking at that one thing that can accelerate your process.
Nicky Pike (46:46):
Yeah, I could not agree more. And I think that is one of the consequences that we're starting to see now is really when we're talking about software engineering, AI has been all focused on developer experience and developer output. AI is basically, and we've had this conversation and I've had this conversation with other people, it's made generation basically free at this point. And exactly what you described is happening. Citizen developers is becoming a thing. It moved that bottleneck, and now we're starting to see our CI/CD pipelines get squashed. We're starting to see our outer loop get crushed. So I think what you're saying, it's not even theory that's being proven at this point. Just look at GitHub and some of the issues that they've had with the influx of code that's been coming in.
Marc Cluet (47:27):
Yeah. And also you can see it in the open source community with the amount of AI contributions they're having now because it has made it a lot easier for everyone to provide code, code that they might not even fully understand, but they are so keen to contribute that now reviewing is becoming a real issue.
Nicky Pike (47:44):
It's amazing to see a lot of the open source projects that are cutting off public PRs for exactly that reason, right? Because yeah, it's easy to generate the code, but each one of those PRs is still requiring human effort to go in and look at it maybe two or three hours per PR. Well, we went from three a month to 300 a month. It's no longer sustainable.
Marc Cluet (48:04):
Yeah, absolutely.
Nicky Pike (48:05):
All right, buddy. Well, now I'm going to hit you with some rapid fire questions. Quick ones, just first thing comes to the top of your mind. No need to put a whole lot into it. Open weight models or frontier models for what you're doing?
Marc Cluet (48:17):
Both.
Nicky Pike (48:18):
Why do you say that?
Marc Cluet (48:19):
Frontier models, really high specialists in some areas. Open weight models, they just can do everything nicely, maybe slightly slower.
Nicky Pike (48:28):
Alright. Rust or grow for your next migration?
Marc Cluet (48:32):
Whoa, depends. Everybody's looking at Rust lately and seeing all the benefits. And I love Rust, to be honest, but he's very specialized for something. So I would say Go would be a bit more of an all player.
Nicky Pike (48:44):
Yeah, I do find that interesting. I think we're watching some shifts in frameworks because of AI, right? Rust is being picked because it has really good error handling, stuff that AI can probably see and review on. Doesn't mean it's the best language, but we are now starting to see architectures come out that are being optimized. And I'll use that in air quotes, optimized for AI development and AI operation. So very interesting. That 20-year-old shell script, should you rewrite it or just kill it?
Marc Cluet (49:15):
Kill it with fire.
Nicky Pike (49:16):
Kill it with fire. Scorched earth on the old tech debt. All right. Canonical, Rackspace, or HashiCorp, which was the best years of your career?
Marc Cluet (49:26):
It's all about timing. I would say Rackspace because I went there when cloud was exploding and DevOps was exploding. It was a really nice and interesting time. I got to speak with some of the original creators of Fancier there because they were ex-Rackers. So yeah, it was really cool.
Nicky Pike (49:43):
Excellent. All right. Now the most important one, vanilla or chocolate?
Marc Cluet (49:47):
Vanilla mostly.
Nicky Pike (49:48):
Ah, we disagree on that one. We disagree on that one. All right. Well, so we're coming up to the end of the episode. We like to end this and have you put out some predictions on it. I'm going to ask you two predictions. Try to put a number on this. Three years from now, what do you think the percentage of enterprise platform teams will actually have an AI ops engineer on staff versus still just kind of winging it with the people they've got now?
Marc Cluet (50:14):
I would say in three years we'll have around 50%.
Nicky Pike (50:17):
50% you think will have an AI ops engineer on board?
Marc Cluet (50:21):
Yeah.
Nicky Pike (50:21):
All right. Well, and on that same lines, this is a question that I know a lot of platform engineers are actually asking right now. What do you think the impacts of AI will have on platform teams and some of the traditional platform engineering roles? Are we going to see new roles like AI ops engineer pop up? Are we going to see extensions of the existing roles or do you think that AI is going to become the new platform engineer?
Marc Cluet (50:46):
I think that we'll have both. We'll have new roles. So for example, I can already see someone being an MCP admin. It's a very specialized job. It's something that requires a deep knowledge of AI that not everybody has. It will reduce also some of the jobs. So the SRE level one, the people that do the fast analysis, AI will be able to do that very well and really fast. So I think we'd need less there. So I think as with DevOps and automation, the skillset will go up more to the abstract and more to the critical thinking. Regarding AI in a way, taking these jobs away, I don't think it will. I said, it will just transform what we have and potentially will give us more time to do more fun stuff.
Nicky Pike (51:28):
Okay. Well, and let me do an extension on that because you said both open weights and frontier models. As we start seeing open weights come in either for cost optimization or personalization reasons, do you think that we're going to start seeing more and more ML ops and data scientists that are part of platform teams to help tune those models?
Marc Cluet (51:46):
Absolutely, because I think that one of the things AI will do really well is accelerate that loop of data-driven decisions with platform engineering. Because with DevOps, we have a lot of metrics and a lot of data and we wouldn't be able to use all of it because it was human scale. If we have something that is really good at crunching numbers, then yeah, it's going to optimize that a lot. I
Nicky Pike (52:09):
Think that's something for people to start looking at is how that platform team's going to grow. All right, buddy, last final question, and this is the one that we ask every guest. You've been at Canonical, you've been at Rackspace, HashiCorp. Now you're doing AI migrations in one of the biggest banks in the world. To you, what does it mean to be a coder?
Marc Cluet (52:28):
I would say it's a double-edged sword because there's a lot of people that think that being a coder is about getting the money. It's one of the jobs that is the most sacrificed, the most frustrating that you can have. So unless you really love this, you're in the wrong place. And I would say coders have the ability to be incredibly creative into a field that is incredibly technical and make something beautiful. Thankfully, still now, we cannot kill people. So I mean, it's still pretty guilt free, so it's good.
Nicky Pike (53:03):
Mark, it sounds like you just advocated for the purge, my friend. Well, we still can't do that today, but hey, knock on wood.
Marc Cluet (53:11):
I'm not trying to say rice off the road. If you compare it to a doctor where every decision is life or death critical, we still can get away with a lot. Yeah, we did not do these database right, but hey, nobody died and the business was some money, but it's fine.
Nicky Pike (53:28):
Yep. All right. I get you now. Well, man, I want to thank you for coming on Devolution. I want to thank you for all of your insights. I know that there's a lot of platform team and platform engineers out there that are looking at exactly some of the questions that we talked about. The final question, man, is what can we say? Can we consider you a full-fledged member of the Devolution moving forward?
Marc Cluet (53:49):
Absolutely. I'll be here every time that you have a new episode listening carefully.
Nicky Pike (53:55):
I appreciate that. All right. Well, thank you very much. And man, thank you for being on the show, and I'm sure that me and you will talk again later.
Marc Cluet (54:02):
Yeah. Thank you so much for having me.
Nicky Pike (54:06):
Thank you for listening to Devolution. If you've got something for us to decode, let me know. You can message me, Nikki Pike, on LinkedIn, or join our Discord community and drop it there. And seriously, don't forget to subscribe. You do not want to miss what's next.