Anne: So hello and welcome to Asynchronous and Unreliable, a weekly podcast where we discuss the latest ideas and concepts in tech. I'm your host, Anne Curry, co author of Building Green Software, the Cloud Native Attitude, and author of the science fiction Panopticon series. today I'm going to be talking to Chad Gibson, the co-founder and CEO of NeuralWatt, a technology company that builds the software to make AI infrastructure radically more power efficient, and who eat their own dog food by offering AI hosting via the Neuralwatt Cloud. So, hello and welcome to the podcast, Chad. Thank you so much for being on.
Chad Gibson: Thank you for having me. It's really exciting to be here.
Anne: So I was just I was just saying to Chad that I'm I'm really keen on you being on this because people have been asking me a lot recently about hosting AI in a way that is more sustainable. And I I commonly refer them to neural what. So why don't we start by you talking me through what neuralwatt does. And and how 'cause you actually take lots of different approaches to cutting the energy use of AI, don't you? And it it'd be quite good to talk through all of them.
Chad Gibson: Yeah. Yeah. Yeah. So so basically NeuroWatt provides different levels of technology to help AI operate effectively with less power. one of the general themes we believe is that a lot of AI's kind of growth and optimism can be can be handled by current resources and current energy. So there's several layers of our software, if you will. Like the the lowest layer sits right above the data center GPUs, which are probably the thing that's consuming the most energy and growing the the most radically in terms of power consumption. And we have some algorithms and and efficiency models that basically allow us to keep these GPUs in the most efficient. We optimize for like tokens per joule or tokens per watt. So we can basically modulate the GPUs. That's like the lowest level and probably the most the deep tech of what we do. And then on top of that, there's different ways that we can expand upon that for different cases. We have this product that allows us to deploy GPUs in like stranded and underutilized or flexible power locations. that again is part of the theme of allowing more AI to be done within current resources. And then on the top of it, we do offer, as you mentioned, a inference cloud for folks who want energy efficient inference. And the really cool thing about that is we provide full transparency, like you know exactly how much energy per request, we provide carbon observability and some unique things we actually offer inference by the kilowatt hour. So all those things kind of compound in a way that allows us to just offer more and more inference or token output with less energy.
Anne: Yeah, I mean it's it's it is interesting. It's not you don't just have one story. That's a lot of different stories, where did you start? Presuming you didn't start with all of those ways all at once.
Chad Gibson: Yeah. No, yeah. Our story our we we offer all those things based upon our experience. So the short story, I don't want to get super long-winded about the origin story, but Scott and I have known each other for a long time. We we both were at Microsoft for a long time, and we actually the last team at Microsoft we worked on was a product called Zune, was ancient Microsoft's iPod competitor. And we were both on the Zoom team. We had actually met before that, but but our paths at Microsoft took different Routes basically. But Scott at Microsoft had had invented a lot of the techniques for carbon observability. So answering the questions of you know how much carbon is being emitted by this network of Windows machines, and that led to optimization, that led him to Intel, where he did a lot of data center work to basically minimize emissions based on data center build outs and footprints and hardware. And so through that work, Scott had kind of built some pretty deep domain experience on how to basically modulate chips for an optimization target. So that's where we started. We started NeuroWatts really began at the end of 2024. And that that lowest piece of technology is where we started. And so we built these this energy efficiency model that through observation allows us to understand the most efficient way to run a GPU based on the work that's being asked of it. And so We we built that system in the model and we trained that model based on a lot of simulated work, a lot of simulated inference workloads, and started talking to customers and doing more trials. And that that core tech 2025 was all about that that just core like NeuroWatt in 2025 was just that piece of software. And we did lots of trials. the trials grew in size and energy impact, and we did a trial at the at over the winter with the utility on the East Coast, where the utility wanted to show and kind of prove that a data center could flex its power to help unlock more flexible power to to different facilities. And there's a big theme about time to power for for data center build outs. And and so we did that work and and when we completed that trial, we actually had built an inference stack in service of that. And when we did did the inference stack, we realized We discovered so much that their the energy efficiency opportunities actually expand as you go up the stack. And that also unlocked all these opportunities with stranded inflexible power. So we launched that product in March. And NeuroWatt, that's NeuroWatt Cloud. And and then through NeuroWatt Cloud, we kind of entered into some partnerships with some other customers who want to utilize the power in their facility that is stranded today. And so all these things kind of build upon each other. You know, NeuroWatt Cloud uses all the technology in our in our in our portfolio, if you will. but but you know, we we're we we'd love to talk to customers for at all aspects of this to to really help AI do more with less, basically.
Anne: that's that that's very interesting. see yeah, it's it's unusual that you you've retained the all of the stack there. You you quite often quite often people will take o take their own products and then run with it, like you have with neural what cloud.
Chad Gibson: Yeah.
Anne: But you also offer all of the stuff all the way up for people who essentially will be your competitors.
Chad Gibson: Yeah, that I mean that was we we we thought a lot about that because when we did NeuroWat Cloud, like one alternative path for us to take was we're just gonna be energy efficient inference. And like it simplifies, we have one class of customer, which is people and inference, and all the other things become part of that. But our mission really is to provide energy efficiency to AI. And so that would kind of preclude us from offering the deeper levels of technology to really help us achieve this goal. So we It's a it's it may be a more complicated path because we have different customers at different levels, but like it it helps us achieve more and basically accomplish more and and it's in service of the mission of you know, we're AI optimists and and we we absolutely believe that we can do a lot to address the fear about, you know, energy and power with AI and it involves, you know, employing all these different tactics.
Anne: Yeah. It's it yeah. I mean I I am also an AI optimist and and it I find it unbelievably frustrating that we don't that the that the data centers industry and the AI data center industry is not is not a really key part of the energy transition because it could be helping. The interesting thing
Chad Gibson: Yeah.
Anne: is why it so often isn't. So I'm I'm quite interested in a ask asking you the question. Why is it so often not? When obviously it can be because because your data centres are.
Chad Gibson: Well, it it absolutely can be. And like in most the last few data center conferences I've all been to, it's all about there's been a huge desire to tell that story because there's a lot of pretty awesome things going on with data centers and helping the energy transition and having data centers be, you know, actual assets for the grid, not not things that are gonna hurt the grid. it's a complicated story to tell, and there's a lot of fear that needs to be overcome and like It's when the fear is prevalent, it's it's kind of hard to tell the the the complicated reason for why that fear isn't always justified. but there there are pretty awesome things hap I g I it's interesting, you know, like a good a good example is like why here in Seattle. I live in Seattle and like there's a moratorium done on data centers in the city. And because there's just this belief that data centers are going to, you know, make everyone's power bills go up and you know, here in Washington. Our grid is mostly renewable energy, with with hydro. Seattle actually has a pretty good grid. And so where there's a good grid, an accessible, you know, renewable energy seems like a great place to have data centers. but the fear needs to be overcome. And the fear requires, you know, telling the broader story. And there's a lot of investments on data centers to actually subsidize, you know, power infrastructure, transmission lines. There's a lot of those stories happening. But I think some some work needs to be done to tell the stories. But but I guess to your point, like we also believe that a lot of the data centers are not fully utilized. You know, the way that that energy is provisioned, the way that you that capacity is is allocated. and it it's no it's of no fault to the data centers. It's just people traditionally think that if I'm gonna deploy a bunch of servers, I need to have this fixed amount of energy. And we provide solutions to that allow, you know, more servers per those. capacity and and more of that current capacity to utilize. I don't think it's gonna take away the need to build more data centers, but it's gonna allow us to do more with with what we have, like a lot of that stranded capacity, if you will.
Anne: Yeah. your comment about there being a lot of hydro in Seattle made me cast my mind back. I I lived in Seattle for two months once when I was working on the exchange project. It rained every day, apart from one day with snowed. But yes, plenty of rain in Seattle Your the the neural watts story is is giving me quite a lot of difficulty in knowing what to talk about. because we've got so you've got so many, many layers. But I so I will focus on just one layer, which is very much at the top of your stack, which is not what we've been talking talking about up till now, which but is the question that comes up over and over again because I'm mostly talking to enterprises, users, what kind of stuff. Most of your models that you run are Chinese models. Well, I think they're all open source models, aren't they? In your in your self, in your hosted stuff, in your what cloud. So tell us a little bit about why you made that decision to all open source, mostly Chinese.
Chad Gibson: Yeah, so we actually started, we launched neuro cloud, we had a diversity of models. We had like some astral models, we had the GPT OSS models that open AI, and we actually just recently started hosting the Gemma model that Google creates. So the the shift towards the the Chinese models has really been customer demand. Like like the the open weight models out of China, in terms of like closing the gap with closed weight frontier models, I mean that gap is closing fast. And so we launched NeuroWatt Cloud. I think we launched with eight models and we wanted so for us model architecture is actually a really fascinating story about AI energy efficiency both with like the emergence of MOE models which is a a way of having the the the weights kind of be only partially activated at a moment which is great for efficiency versus dense models where the whole network is being used as well as small medium and large models and small models for many tasks are amazing and hyper efficient. So we launched with an array of models across small medium and large We we we launched with some models from US, like Mistral's European model, and the demand kind of shifted all towards the the frontier large weight models from China. our we we would love, we would love to like part of our goal and the reason we have like these different layers of stack, like we would love anthropic to license our technology to help make, you know, Fable and Opus more. We would love to to open it. So like we would love to do more with. Closed weight models and frontier models. Like, and we would love to license different parts of our stack. So those, so the the fact that we NeuroWatt Cloud uses open weight models is largely because that's what we have access to. but we we it allows us to basically improve our energy efficiency model at the core of our tech and basically provide energy efficient inference. And the more the more inference work we observe, the better it makes our energy efficient model at the at the base of our system. So That serves the purpose of like kind of helping train our model and and make it more energy efficient. But I mean open weight models, part of Neurolog Cloud has grown very quickly since March, largely because that the gap between open weight and closedweight models is absolutely closing. largely because of a lot of the work happening in China right now to expand the frontier. And you know, I use I use Chat GPT anthropic models a ton. And I use the models we host a ton and seeing that gap close has been pretty fascinating to to observe.
Anne: So that's interesting. So so as an observer of that, probably a fairly close observer compared to most people, I would imagine. Well, yeah, you've got loads of data there. as the the gap closes in terms of functionality between the frontier models and the open white models, what's happening to the efficiency? Because in the past, the the open weight models, well especially the Chinese models, had become extraordinarily efficient compared to the to the frontier models.
Chad Gibson: So the thing about the thing about the the closedway models is we don't entirely know. Like, you know, we know we know that, you know, folks like Anthropic and OpenAI and even their partners like Amazon and Microsoft who host their models have have have huge compute and energy footprints, but like in terms of like jewels per token, like that that is data that only they have access to. So we I can't I can't I can't really I can I can theorize, but I don't really know. So it it very well could be that there's some state-of-the-art energy efficiency happening with some of those models that we don't know. With the openway models, a lot of that and it I think part of it is necessity of invention. Like there are limited accelerators and GPUs, and so it pushes a lot of innovation in the software stack in terms of serving efficiency. And if you take a look at like some of the recent innovations, specifically like Deep Seek4 and You know, the recent GLM five. Two launch going to million context, like there is so much innovation on model architectures, on attention mechanisms, on KV cache f formats. each one of these innovations, like the the thing the thing that we've really learned in the last six months, accelerated in last six months, is the real opportunity of efficiency in the software layer. You know, I think a lot of folks view that AI efficiency is really a hardware thing, like you need new accelerators which makes it more efficient, but there is profound opportunity in the full the full stack of software, all the way from you know the the the layer we exist at which is right above the GPUs to the serving layer to the model layer to the GPU kernel layer. There is profound efficiency opportunities and and and that that whole ecosystem of software efficiency is moving so rapidly and even with its rapid speed there's still so much efficiency still available. So That's the most fascinating thing about what we have, what we observe and see. And, you know, these model releases are happening every four to six weeks. and so with each of them we can really assess how much efficiency this this model architecture or this serving pattern has has has assisted. And it it's it's the the one difference between open weight and and some of the closing models is on open weight, you you can you can observe this. You could observe every step of the transition. And and build more optimizations around that as well.
Anne: Yeah, it really does help to be able to look inside Yeah, so so it's it's it's really interesting that you're you're doing all of this stuff. You're you're you seem to be strangely far ahead of everybody else.
Chad Gibson: well, I think that's that's awesome. I mean, I think we've we've been focusing on this problem for a while, and our thesis is that and I think a lot of people have viewed this, like like energy and power limits will be a constraining factor. And so I think that has we've been and we have also been kind of relentlessly focused on inference, which which you know in the early 2025 was a little bit of a, you know, there's a lot of talk about training and training clusters, and training generally requires a lot of Co-location of GPU compute needs to be in like bigger facilities with with all of the of the compute in the same and so that took a lot of the early attention in terms of data center build-outs. That's why you needed a lot of these gigacampuses to have super large train clusters. Inference still benefits from that, but is largely can be a more fragmented approach. But since our beginning we've been just laser focused on inference and Laser focused inference and laser focused on, you know, making the accelerators and and that focus has has has been, you know, in at and early in our on our journey questionable, but now it's it's paid off because we can we can go really deep on the efficiency of inference.
Anne: I mean so so what what real really sticky difficult things have you hit whilst trying whilst doing this? What were you not expecting to be as difficult as it's turned out to be? If you can tell me that.
Chad Gibson: geez. I mean, I don't know if there's one. There's there's so many, so many hurdles. Like
Anne: Yeah.
Chad Gibson: Well, model evolution speed. I mean the the challenges haven't been like big blockers. It's it's the challenges are are temporal blockers, but moving so quickly. So what's the big one? I mean, first inference demand is kind of insatiable right now. And specifically as open weight models are getting stronger and better and cost effective, like the demand is growing. So we're we're having to we had this this wild experience from like mid-April to like the end of May, where we had just started with NeuroWatt Cloud. we had our first wave of growth and a common pattern when you deploy inference is you just throw more capacity at it. Like an easy way to handle more growth is throw more capacity. And in that era we weren't entirely sure of neurowatt cloud, so we weren't confident in deploying more capacity. So we got into this pattern of just manifesting headroom via software efficiency. And so we would like we would do some work. We would evaluate the okay, we could we could generate 30% efficiency in this layer and we'd do that and we'd get 30% headroom. That headroom would be swallowed. It was like Jevin's paradox every three days. You free up some headroom, you make things faster, it's consumed immediately. And so for like four weeks, it was like Groundhog Day. We're like, okay, that we thought that headroom would last us 10 days, it lasted 24 hours. What do we do now? Okay, let's do it again. And all this, while like in the midst of that, there was a new model change. you know, the so that that was one just and it's not really specific, but the other big challenge was when GLM 5.2 came out, it was a lot of the open weight models at the at that moment were all 200,000, 256k context. So they had a max context limit, which puts some bounds on serving efficiency. It it kind of makes it easier to serve in a way. And then when GLM Five. Two came out along with you know, DeepCP4. They they expanded to million context length like a lot of the anthropic and open AI models. That was a challenge. That was a really difficult serving challenge because part of the nature of our cloud is we have capacity fragmented all over the place. And so managing longer contexts makes the fragmentation challenge harder to overcome. So that was actually a really interesting technical challenge for us to overcome. That took us good 10 days. those are two noteworthy things that come to mind.
Anne: So that's interesting because so something we say a lot in building green software is that it's really hard to run data centers. So, you know, the cloud, if especially if you lose you use cloud services rather than lift and shift, you get the benefit of specialist folk looking after really quite complicated systems that that need to be tuned and need be managed like it sounds like models are just are like that only on steroids, on a on a completely different Different level.
Chad Gibson: Yeah, they they c i i you know it's as much the calling patterns of your customers as the model. It's that combination. So our like our cloud, our cloud, our customers are mostly software engineers. And software engineers, I would I would roughly characterize it as like three-quarters of them are software engineers, a quarter of them are building autonomous agents to do things. And so the autonomous agents themselves have their own calling patterns. And so it's the combination of that class of the classes of customers with the models that make things very challenging. Because even across our customer base, and it I'm the customers are doing great supported things. I don't want to say that the customers are doing bad things, but like some calling patterns are way more complicated than others with certain models. And so it's being able to identify it, route it, address it properly, kind of understand how to handle these classes of calling patterns, that is that combined with the different models, becomes very challenging.
Anne: Yeah. Yeah, yeah. Yeah, yeah. I mean you support a lot of models and I I on your website you're gonna support even more in the future.
Chad Gibson: Yeah. Yeah. We our our view on models is is is to go where there's demand and interest. And like our our the the one bias we have is we would love you know usage across a broad array of model sizes. And like we've we've been doing better at on the small end, you know, like the quint the 35 billion parameters or like the Gemma model is 31 billion parameters. Like so we have a small MOE and a small dense model, and those get a lot of great use. The middle The middle we we have not found like we've we've gone through a lot of of I would say middle size models, but but it's really stratified between like the small models and the extra large models.
Anne: Yeah.
Chad Gibson: and the extra large models is generally the hardest to serve the most demand as well.
Anne: So is that because people are still basically using kind of let's go all purpose, kind of over provisioning on model, do you think?
Chad Gibson: I mean I think I think part of it is simplicity. Like I if I take my workflow, there are still workflows that I'm using a big model to do something that is overkill, but I mean it's on me to, you know, delegate to the small model the task that the small model can do. So it's simple just to use one model for for a given workflow.
Anne: Yeah. Yeah.
Chad Gibson: there's things we're doing to help with that, and there's ways that you can make it easier for different tasks to be delegated. but the simplicity. And one, like I think I think the cutting edge is moving so quickly that people are still realizing how to fully exploit, if you will, or maximize these frontier models. And then when the new frontier comes out, it's kind of like you need to relearn again. How can I re-establish like How can I re-establish my workflow now that knowing the model is, you know, 20%, 30% more capable? So the fact that the frontier is pushing while people are still trying to figure out how to maximize these models in their workflow, it's kind of pushes to this point of simplicity. Well, I'm just gonna push everything on the high model for a bit and see what I can learn and evaluate from that. And then right when I get comfortable with that, yet another new model. comes out and sets the bar even higher.
Anne: Yes. And as you say, there's it there's there's so much potential ability to effectively get more out of the software that we haven't really even scratched the surface of yet.
Chad Gibson: Right. Totally.
Anne: so there's really no end in sight, is there?
Chad Gibson: well I don't see the end. I mean I I think I think conceptually things it's an interesting thing to contemplate. Like cause now the new frontier model, the new open way frontier models are coming out with like 2.8 trillion parameters. And and so generally, if you if you if you take a look at the landscape of like first on hardware, so like NVIDIA and AMD and others Like they they basically look at hardware and like these, you know, it's the the the the capacity of a single chip and like the capacity of a single server. And then generally they have like what's their scale out pattern. So an example with like NVIDIA is they have blackwell, let's take Blackwell, like there's a Blackwell chip, and there's like a Blackwell server, like a B three hundred server that has eight of these interconnected, and then they'll have like, you know, a r a Ruben rack or a bare Ruben with like scaling this out to like the biggest compute in one rack. These frontier models are now pushing the point where you need just about two servers or two 8 GPU servers. But theoretically, the parameter counts could increase to a point where you may see a model where you need a full rack to serve it. So conceptually speaking, the I don't think the ceiling on parameter counts is necessarily in sight. But while parameter counts are being pushed on, you're also seeing models get better at smaller parameters as well. So while the frontier is pushing model size, there's amazing innovation happening with making a 300 billion parameter model better and better and better. So it's like there's also innovation happening on making the value of each weight higher. And then NVIDIA is pushing, you know, the increased capability of each chip. Which therefore is increasing the capability of each server, which is increasing the capability of each scale-up unit. So the fact that all these things are still progressing, I don't I don't have a good sense for which of those three things in the next 12 months is going to be a limiting factor. Like they're all I can see 12 months of growth across all three of those. in 24 to 36 months, it's it's it's it's also for me unclear because it's just moving so quickly.
Anne: Yeah, absolutely. Right, and it's and it's obviously it's a bit it's a bit terrifying because you you've got a good data center, you're doing all the things that we said, if only data centers would do this kind of thing in building grid software. you're you're adding flexibility to the grid by by the talking to working with grids to turn things turn dial up dial down all that kind of stuff. You're using hardware that was that was going to waste or excess electricity that was that was not being used as well as it could. You're you're providing data to people so that they can act on it and, you know, true kind of open source, but also, you know, provided providing the telemetry so that people can go, well if I change that, it would it would improve things like you're doing all the things. But there are so many data centers out there which do sound like, you know, they are pretty terrifying. they're hooking up brand new gas-fired power stations. so yeah you can kind of see there's the hashtag not all data centers, but there are data centers. There are some pretty evil data centers out there and they often have
Chad Gibson: Yeah.
Anne: names.
Chad Gibson: Yeah. Yeah, yeah. I mean like on site power is like if you're if you're impatient, the easiest thing to do is I don't want to say easy, it's not easy, but minimizing reliance on the grid is w is is one way to do it. and there are ways to do that where you could, you know, put it put a modular data center where there's a big solar array where there's abundant solar energy with you know, renewable renewable like with batteries and there's some pretty cool stuff happening there, but there's also ways of just hooking up more diesel generators. And yeah, like I think I I think if we can shine a light on some of the more constructive ways of doing it, can address that fear we talked about and that, you know, anti data center sentiment. but oftentimes, you know, the the the those cases you describe are the ones that become the the examples used.
Anne: Yeah. And of course the difficulty with that is that people hear only about the bad stuff and they think there is only the bad stuff available, that there's no choice.
Chad Gibson: Yeah.
Anne: It's like AI or being sustainable or being green or aligning yourself with the energy transition that, you know, it's an either or and people are always going to choose AI if in those in that circumstance. But I think we do need to shine a light on the fact that there is there are other alternatives. You know, you do not have to to use these giant data centers.
Chad Gibson: Yeah, you know, when we look at like the transition to cloud computing, when people first started moving like on premise workloads to, you know, AWS and Azure and Google Cloud Platform, all those platforms kind of out of necessity started evolving, like providing, you know, carbon and carbon observability for cloud workloads. And that's now a a feature of all those cloud platforms. And it's evolving to provide more, you know, more regional based versus doing tricks with carbon offset credits and whatnot. That level of transparency we absolutely believe is gonna come to AI. that the need for the and we're seeing it in Europe already. We're seeing customers in Europe who kind of want a direct correlation between where these workloads are run and how they can correlate it to to what state of energy generation is is correlated with it. And I do believe that will come across worldwide at some point. don't entirely know know when, but other parts of the world. But but the challenge with that in the immediate short term is the the again, it goes back to speed. And right now everyone is kind of racing and rushing to understand how to leverage AI in a way that it's It's driving more pressure to build out capacity and to, you know, try to service these workloads and ever and ever even the suppliers themselves are trying to figure out how can we scale this effectively, how can we help others figure out how to use AI. So it at the speed is I think putting the pressure on all aspects of this.
Anne: Yeah. Yeah. Yes. So, so if you had advice for listeners, watchers of this podcast about what they can do, what should they be doing? What should we what should they be thinking about? What should what do they need to open their minds to the possibility of in AI?
Chad Gibson: I mean, so first of all, there there are ways. I mean, just to plug neurologal, there are ways to, you know, build and use AI in a way where you know, you know, you can correlate it to energy and carbon usage. So if you yourselves wanna know wanna wanna answer the question of how much energy is my is my AI use consuming, we can answer that for you. Two, there there are ways, and you know, I think you know, to the to to their AI service of choice, you know, they can they they now can speak confidently that like, hey, I mean, like customer demand is going to change this as well. You know, if you're a huge anthropic fan, which I am also a huge anthropic fan, like letting anthropic know that we'd love to know how much energy is being consumed by Opus and I would love to have that kind of transparency. Like, that's a meaningful thing. You know, we've proven it's possible, it can be done. We would love to to help anthropic satisfy that need, but that's something as well. And then when it comes to you know the question of data centers and is it gonna increase my power bill, I do hope that like this conversation has surfaced, there are ways of doing it that can be constructive. And not not all data center is the, you know, gonna put gas generators in my backyard. and the more that we can allow capacity to come online that is, you know. making the grid more flexible. That is, you know, consuming this renewable energy we have is a good thing. And in some cases will probably mitigate how much new, you know, turbines and generators we need to create so create elsewhere. So, you know, hopefully that provides a bit more of a of a of a of a view on on the new the nuance, if you will, of of how they're not all created the same. But yeah, I mean I think I think there are options to to to consume AI and be an AI champion and optimist that allows you to actually understand energy consumption.
Anne: that was very useful to for us all to be reminded of that you do have options here. It isn't it isn't AI or the planets. there there are ways of doing it but are not that not not are not as destructive. So and and a big thing, if you if you can't do anything else, just ask your suppliers. Say, look, I care about this. Can you tell start telling me? Can you start doing something about it? Tell me what you what you're doing, what you're thinking, what your plans are, because I do care about this and I will be making purchasing decisions in future based on people's plans and and and hopes for this kind of stuff. So what so that that was very interesting. Is there is there anything else you want to add? You are allowed to you the you're in here, I'm perfectly happy for you to just keep plugging neural what because you offer a service that I think should be offered more widely.
Chad Gibson: Yeah, I mean the thing the thing that is interesting, and this is like one of the challenges we we deal with, like so we when we offer inference, we we price it by the kilowatt hour, which is which is both
Anne: no.
Chad Gibson: a really cool thing and a complicated thing. And so for users who are used to paying by the token. Consuming their AI by energy is a is just a challenging thing. And so it it we've we've had to to make that easier by providing a lot of transparency. The key thing is it provides transparency. You know how much energy you're here. And now I can optimize for that. So what's fascinating is like, how do I optimize for that? Well, there's certain times the day where energy is more abundant for AI. And and you can do this with carbon as well. There's certain times of the day where your grid is emitting more carbon, and there's times a day where it's not. And so this is like this whole new frontier as of optimization opportunities. And so a lot of our current customers are like pushing us, and it's a really fascinating, fun thing to be pushed on. Because as more of these tasks, if you will, you know, imagine a future where, like, if we think of inference, largely inference stay for our cloud at least, is a human piloting something that's calling our service. more and more it's transitioning to where it's a human that created a autonomous agent that is then calling our service. And as we go to the future, like I was at the AI conference a month ago where they're like planning for a future of like, you know, s autonomous swarms of agents, of software agents. And so the cool thing about the agents is they can if you give them data to optimize, they're really great at optimizing. So we can provide transparency of like, hey There's way less carbon being emitted in the grid right now. Hey, energy is really cheap and abundant right now. And so some tasks that aren't require like immediate response, let's put let's let's do those tasks when there's less carbon being emitted. Like the optimization potential becomes pretty, pretty profound. And now if I'm gonna if I'm gonna give a human customer who's doing who's doing software engineering like these fifty variables to optimize against, it's gonna be hard for them to figure out, I have to do bunch code. Maybe I should wait to do my code next hour. But for the part of their workloads that are increasingly becoming more autonomous, like they're just phenomenal optimization potential. So that's that's a really, really fascinating world where we can empower, you know, both humans and autonomous agents alike to optimize. And we can provide
Anne: Mm.
Chad Gibson: all of this transparency data. To you know, leverage the resources when they're abundant, leverage them when they're cheaper, and reduce the pressure on the resources when they're under great strain.
Anne: Mm-hmm.
Chad Gibson: and so that I'm I think that's a really, really exciting proposition. And it kind of goes, it just goes down this path of doing more with the resources we have, better utilizing the resources we have. And the cool thing is it's not just doing things to, you know, minimize consumption, but it also makes it cheaper. And so there's
Anne: Yeah.
Chad Gibson: just the very, very clear benefit of of consuming things when they're cheaper. And and that I mean everyone loves that.
Anne: Yeah. Well they do. They and they well, it's an interesting saying they they love things to be cheaper, but they don't really like to go too much effort to get it to be cheaper. So
Chad Gibson: Yeah. Yeah.
Anne: but but the the barrier to what is difficult is changing because of AI. and therefore
Chad Gibson: Yeah.
Anne: Yeah.
Chad Gibson: Yeah, I agree. I agree. I mean it's it's like I said, it's it's like all these features and capabilities, we're trying to move as quickly as we can. And while everything around us is changing really quickly and it's it's it's it's f it's exciting. I mean it's it's so I think about my career at Microsoft. Like I I I I loved my time at Microsoft. It's a it's an awesome place. I w got to work on some massive, huge, quite impacting products. But the speed at which I'm moving now is kind of one of the reasons I wanted to to to have, you know, a career post-Microsoft and it's so exciting. It's it's just it's it's being being around the technology this quickly, it's it's it's there's new interesting problems emerging like every week.
Anne: Yeah, I mean I guess we all said we wanted to be in the tech industry to solve difficult problems. We had no idea, really did we?
Chad Gibson: I know, I know, I know. They're are are they ever gonna be solved? I don't know. there's gonna be new ones all the time.
Anne: Well, on that happy note, and it is a happy note because we did to go into the tech industry to be to solve difficult problems and now it's it's more difficult and it's faster moving than ever. You know, my thirty years I was there I was at Microsoft just a little bit before you, I think. and then but it was yeah, it's it's it nothing has changed at the at the rate that that we that like the rate we're seeing at the moment. but so thank you very much indeed for being on the podcast. It was a delight, really interesting. and very inspiring as well. There are things that we can do and tricky, tricky problems to solve. So thank you very much indeed for being on the podcast. And thank you to all our watchers and listeners across the globe. and hopefully I will catch you again on a future episode of Asynchronous and Unreliable Podcast. Thank you very much.
Chad Gibson: Thank you for having me. It's really exciting to be here.
Anne: So I was just I was just saying to Chad that I'm I'm really keen on you being on this because people have been asking me a lot recently about hosting AI in a way that is more sustainable. And I I commonly refer them to neural what. So why don't we start by you talking me through what neuralwatt does. And and how 'cause you actually take lots of different approaches to cutting the energy use of AI, don't you? And it it'd be quite good to talk through all of them.
Chad Gibson: Yeah. Yeah. Yeah. So so basically NeuroWatt provides different levels of technology to help AI operate effectively with less power. one of the general themes we believe is that a lot of AI's kind of growth and optimism can be can be handled by current resources and current energy. So there's several layers of our software, if you will. Like the the lowest layer sits right above the data center GPUs, which are probably the thing that's consuming the most energy and growing the the most radically in terms of power consumption. And we have some algorithms and and efficiency models that basically allow us to keep these GPUs in the most efficient. We optimize for like tokens per joule or tokens per watt. So we can basically modulate the GPUs. That's like the lowest level and probably the most the deep tech of what we do. And then on top of that, there's different ways that we can expand upon that for different cases. We have this product that allows us to deploy GPUs in like stranded and underutilized or flexible power locations. that again is part of the theme of allowing more AI to be done within current resources. And then on the top of it, we do offer, as you mentioned, a inference cloud for folks who want energy efficient inference. And the really cool thing about that is we provide full transparency, like you know exactly how much energy per request, we provide carbon observability and some unique things we actually offer inference by the kilowatt hour. So all those things kind of compound in a way that allows us to just offer more and more inference or token output with less energy.
Anne: Yeah, I mean it's it's it is interesting. It's not you don't just have one story. That's a lot of different stories, where did you start? Presuming you didn't start with all of those ways all at once.
Chad Gibson: Yeah. No, yeah. Our story our we we offer all those things based upon our experience. So the short story, I don't want to get super long-winded about the origin story, but Scott and I have known each other for a long time. We we both were at Microsoft for a long time, and we actually the last team at Microsoft we worked on was a product called Zune, was ancient Microsoft's iPod competitor. And we were both on the Zoom team. We had actually met before that, but but our paths at Microsoft took different Routes basically. But Scott at Microsoft had had invented a lot of the techniques for carbon observability. So answering the questions of you know how much carbon is being emitted by this network of Windows machines, and that led to optimization, that led him to Intel, where he did a lot of data center work to basically minimize emissions based on data center build outs and footprints and hardware. And so through that work, Scott had kind of built some pretty deep domain experience on how to basically modulate chips for an optimization target. So that's where we started. We started NeuroWatts really began at the end of 2024. And that that lowest piece of technology is where we started. And so we built these this energy efficiency model that through observation allows us to understand the most efficient way to run a GPU based on the work that's being asked of it. And so We we built that system in the model and we trained that model based on a lot of simulated work, a lot of simulated inference workloads, and started talking to customers and doing more trials. And that that core tech 2025 was all about that that just core like NeuroWatt in 2025 was just that piece of software. And we did lots of trials. the trials grew in size and energy impact, and we did a trial at the at over the winter with the utility on the East Coast, where the utility wanted to show and kind of prove that a data center could flex its power to help unlock more flexible power to to different facilities. And there's a big theme about time to power for for data center build outs. And and so we did that work and and when we completed that trial, we actually had built an inference stack in service of that. And when we did did the inference stack, we realized We discovered so much that their the energy efficiency opportunities actually expand as you go up the stack. And that also unlocked all these opportunities with stranded inflexible power. So we launched that product in March. And NeuroWatt, that's NeuroWatt Cloud. And and then through NeuroWatt Cloud, we kind of entered into some partnerships with some other customers who want to utilize the power in their facility that is stranded today. And so all these things kind of build upon each other. You know, NeuroWatt Cloud uses all the technology in our in our in our portfolio, if you will. but but you know, we we're we we'd love to talk to customers for at all aspects of this to to really help AI do more with less, basically.
Anne: that's that that's very interesting. see yeah, it's it's unusual that you you've retained the all of the stack there. You you quite often quite often people will take o take their own products and then run with it, like you have with neural what cloud.
Chad Gibson: Yeah.
Anne: But you also offer all of the stuff all the way up for people who essentially will be your competitors.
Chad Gibson: Yeah, that I mean that was we we we thought a lot about that because when we did NeuroWat Cloud, like one alternative path for us to take was we're just gonna be energy efficient inference. And like it simplifies, we have one class of customer, which is people and inference, and all the other things become part of that. But our mission really is to provide energy efficiency to AI. And so that would kind of preclude us from offering the deeper levels of technology to really help us achieve this goal. So we It's a it's it may be a more complicated path because we have different customers at different levels, but like it it helps us achieve more and basically accomplish more and and it's in service of the mission of you know, we're AI optimists and and we we absolutely believe that we can do a lot to address the fear about, you know, energy and power with AI and it involves, you know, employing all these different tactics.
Anne: Yeah. It's it yeah. I mean I I am also an AI optimist and and it I find it unbelievably frustrating that we don't that the that the data centers industry and the AI data center industry is not is not a really key part of the energy transition because it could be helping. The interesting thing
Chad Gibson: Yeah.
Anne: is why it so often isn't. So I'm I'm quite interested in a ask asking you the question. Why is it so often not? When obviously it can be because because your data centres are.
Chad Gibson: Well, it it absolutely can be. And like in most the last few data center conferences I've all been to, it's all about there's been a huge desire to tell that story because there's a lot of pretty awesome things going on with data centers and helping the energy transition and having data centers be, you know, actual assets for the grid, not not things that are gonna hurt the grid. it's a complicated story to tell, and there's a lot of fear that needs to be overcome and like It's when the fear is prevalent, it's it's kind of hard to tell the the the complicated reason for why that fear isn't always justified. but there there are pretty awesome things hap I g I it's interesting, you know, like a good a good example is like why here in Seattle. I live in Seattle and like there's a moratorium done on data centers in the city. And because there's just this belief that data centers are going to, you know, make everyone's power bills go up and you know, here in Washington. Our grid is mostly renewable energy, with with hydro. Seattle actually has a pretty good grid. And so where there's a good grid, an accessible, you know, renewable energy seems like a great place to have data centers. but the fear needs to be overcome. And the fear requires, you know, telling the broader story. And there's a lot of investments on data centers to actually subsidize, you know, power infrastructure, transmission lines. There's a lot of those stories happening. But I think some some work needs to be done to tell the stories. But but I guess to your point, like we also believe that a lot of the data centers are not fully utilized. You know, the way that that energy is provisioned, the way that you that capacity is is allocated. and it it's no it's of no fault to the data centers. It's just people traditionally think that if I'm gonna deploy a bunch of servers, I need to have this fixed amount of energy. And we provide solutions to that allow, you know, more servers per those. capacity and and more of that current capacity to utilize. I don't think it's gonna take away the need to build more data centers, but it's gonna allow us to do more with with what we have, like a lot of that stranded capacity, if you will.
Anne: Yeah. your comment about there being a lot of hydro in Seattle made me cast my mind back. I I lived in Seattle for two months once when I was working on the exchange project. It rained every day, apart from one day with snowed. But yes, plenty of rain in Seattle Your the the neural watts story is is giving me quite a lot of difficulty in knowing what to talk about. because we've got so you've got so many, many layers. But I so I will focus on just one layer, which is very much at the top of your stack, which is not what we've been talking talking about up till now, which but is the question that comes up over and over again because I'm mostly talking to enterprises, users, what kind of stuff. Most of your models that you run are Chinese models. Well, I think they're all open source models, aren't they? In your in your self, in your hosted stuff, in your what cloud. So tell us a little bit about why you made that decision to all open source, mostly Chinese.
Chad Gibson: Yeah, so we actually started, we launched neuro cloud, we had a diversity of models. We had like some astral models, we had the GPT OSS models that open AI, and we actually just recently started hosting the Gemma model that Google creates. So the the shift towards the the Chinese models has really been customer demand. Like like the the open weight models out of China, in terms of like closing the gap with closed weight frontier models, I mean that gap is closing fast. And so we launched NeuroWatt Cloud. I think we launched with eight models and we wanted so for us model architecture is actually a really fascinating story about AI energy efficiency both with like the emergence of MOE models which is a a way of having the the the weights kind of be only partially activated at a moment which is great for efficiency versus dense models where the whole network is being used as well as small medium and large models and small models for many tasks are amazing and hyper efficient. So we launched with an array of models across small medium and large We we we launched with some models from US, like Mistral's European model, and the demand kind of shifted all towards the the frontier large weight models from China. our we we would love, we would love to like part of our goal and the reason we have like these different layers of stack, like we would love anthropic to license our technology to help make, you know, Fable and Opus more. We would love to to open it. So like we would love to do more with. Closed weight models and frontier models. Like, and we would love to license different parts of our stack. So those, so the the fact that we NeuroWatt Cloud uses open weight models is largely because that's what we have access to. but we we it allows us to basically improve our energy efficiency model at the core of our tech and basically provide energy efficient inference. And the more the more inference work we observe, the better it makes our energy efficient model at the at the base of our system. So That serves the purpose of like kind of helping train our model and and make it more energy efficient. But I mean open weight models, part of Neurolog Cloud has grown very quickly since March, largely because that the gap between open weight and closedweight models is absolutely closing. largely because of a lot of the work happening in China right now to expand the frontier. And you know, I use I use Chat GPT anthropic models a ton. And I use the models we host a ton and seeing that gap close has been pretty fascinating to to observe.
Anne: So that's interesting. So so as an observer of that, probably a fairly close observer compared to most people, I would imagine. Well, yeah, you've got loads of data there. as the the gap closes in terms of functionality between the frontier models and the open white models, what's happening to the efficiency? Because in the past, the the open weight models, well especially the Chinese models, had become extraordinarily efficient compared to the to the frontier models.
Chad Gibson: So the thing about the thing about the the closedway models is we don't entirely know. Like, you know, we know we know that, you know, folks like Anthropic and OpenAI and even their partners like Amazon and Microsoft who host their models have have have huge compute and energy footprints, but like in terms of like jewels per token, like that that is data that only they have access to. So we I can't I can't I can't really I can I can theorize, but I don't really know. So it it very well could be that there's some state-of-the-art energy efficiency happening with some of those models that we don't know. With the openway models, a lot of that and it I think part of it is necessity of invention. Like there are limited accelerators and GPUs, and so it pushes a lot of innovation in the software stack in terms of serving efficiency. And if you take a look at like some of the recent innovations, specifically like Deep Seek4 and You know, the recent GLM five. Two launch going to million context, like there is so much innovation on model architectures, on attention mechanisms, on KV cache f formats. each one of these innovations, like the the thing the thing that we've really learned in the last six months, accelerated in last six months, is the real opportunity of efficiency in the software layer. You know, I think a lot of folks view that AI efficiency is really a hardware thing, like you need new accelerators which makes it more efficient, but there is profound opportunity in the full the full stack of software, all the way from you know the the the layer we exist at which is right above the GPUs to the serving layer to the model layer to the GPU kernel layer. There is profound efficiency opportunities and and and that that whole ecosystem of software efficiency is moving so rapidly and even with its rapid speed there's still so much efficiency still available. So That's the most fascinating thing about what we have, what we observe and see. And, you know, these model releases are happening every four to six weeks. and so with each of them we can really assess how much efficiency this this model architecture or this serving pattern has has has assisted. And it it's it's the the one difference between open weight and and some of the closing models is on open weight, you you can you can observe this. You could observe every step of the transition. And and build more optimizations around that as well.
Anne: Yeah, it really does help to be able to look inside Yeah, so so it's it's it's really interesting that you're you're doing all of this stuff. You're you're you seem to be strangely far ahead of everybody else.
Chad Gibson: well, I think that's that's awesome. I mean, I think we've we've been focusing on this problem for a while, and our thesis is that and I think a lot of people have viewed this, like like energy and power limits will be a constraining factor. And so I think that has we've been and we have also been kind of relentlessly focused on inference, which which you know in the early 2025 was a little bit of a, you know, there's a lot of talk about training and training clusters, and training generally requires a lot of Co-location of GPU compute needs to be in like bigger facilities with with all of the of the compute in the same and so that took a lot of the early attention in terms of data center build-outs. That's why you needed a lot of these gigacampuses to have super large train clusters. Inference still benefits from that, but is largely can be a more fragmented approach. But since our beginning we've been just laser focused on inference and Laser focused inference and laser focused on, you know, making the accelerators and and that focus has has has been, you know, in at and early in our on our journey questionable, but now it's it's paid off because we can we can go really deep on the efficiency of inference.
Anne: I mean so so what what real really sticky difficult things have you hit whilst trying whilst doing this? What were you not expecting to be as difficult as it's turned out to be? If you can tell me that.
Chad Gibson: geez. I mean, I don't know if there's one. There's there's so many, so many hurdles. Like
Anne: Yeah.
Chad Gibson: Well, model evolution speed. I mean the the challenges haven't been like big blockers. It's it's the challenges are are temporal blockers, but moving so quickly. So what's the big one? I mean, first inference demand is kind of insatiable right now. And specifically as open weight models are getting stronger and better and cost effective, like the demand is growing. So we're we're having to we had this this wild experience from like mid-April to like the end of May, where we had just started with NeuroWatt Cloud. we had our first wave of growth and a common pattern when you deploy inference is you just throw more capacity at it. Like an easy way to handle more growth is throw more capacity. And in that era we weren't entirely sure of neurowatt cloud, so we weren't confident in deploying more capacity. So we got into this pattern of just manifesting headroom via software efficiency. And so we would like we would do some work. We would evaluate the okay, we could we could generate 30% efficiency in this layer and we'd do that and we'd get 30% headroom. That headroom would be swallowed. It was like Jevin's paradox every three days. You free up some headroom, you make things faster, it's consumed immediately. And so for like four weeks, it was like Groundhog Day. We're like, okay, that we thought that headroom would last us 10 days, it lasted 24 hours. What do we do now? Okay, let's do it again. And all this, while like in the midst of that, there was a new model change. you know, the so that that was one just and it's not really specific, but the other big challenge was when GLM 5.2 came out, it was a lot of the open weight models at the at that moment were all 200,000, 256k context. So they had a max context limit, which puts some bounds on serving efficiency. It it kind of makes it easier to serve in a way. And then when GLM Five. Two came out along with you know, DeepCP4. They they expanded to million context length like a lot of the anthropic and open AI models. That was a challenge. That was a really difficult serving challenge because part of the nature of our cloud is we have capacity fragmented all over the place. And so managing longer contexts makes the fragmentation challenge harder to overcome. So that was actually a really interesting technical challenge for us to overcome. That took us good 10 days. those are two noteworthy things that come to mind.
Anne: So that's interesting because so something we say a lot in building green software is that it's really hard to run data centers. So, you know, the cloud, if especially if you lose you use cloud services rather than lift and shift, you get the benefit of specialist folk looking after really quite complicated systems that that need to be tuned and need be managed like it sounds like models are just are like that only on steroids, on a on a completely different Different level.
Chad Gibson: Yeah, they they c i i you know it's as much the calling patterns of your customers as the model. It's that combination. So our like our cloud, our cloud, our customers are mostly software engineers. And software engineers, I would I would roughly characterize it as like three-quarters of them are software engineers, a quarter of them are building autonomous agents to do things. And so the autonomous agents themselves have their own calling patterns. And so it's the combination of that class of the classes of customers with the models that make things very challenging. Because even across our customer base, and it I'm the customers are doing great supported things. I don't want to say that the customers are doing bad things, but like some calling patterns are way more complicated than others with certain models. And so it's being able to identify it, route it, address it properly, kind of understand how to handle these classes of calling patterns, that is that combined with the different models, becomes very challenging.
Anne: Yeah. Yeah, yeah. Yeah, yeah. I mean you support a lot of models and I I on your website you're gonna support even more in the future.
Chad Gibson: Yeah. Yeah. We our our view on models is is is to go where there's demand and interest. And like our our the the one bias we have is we would love you know usage across a broad array of model sizes. And like we've we've been doing better at on the small end, you know, like the quint the 35 billion parameters or like the Gemma model is 31 billion parameters. Like so we have a small MOE and a small dense model, and those get a lot of great use. The middle The middle we we have not found like we've we've gone through a lot of of I would say middle size models, but but it's really stratified between like the small models and the extra large models.
Anne: Yeah.
Chad Gibson: and the extra large models is generally the hardest to serve the most demand as well.
Anne: So is that because people are still basically using kind of let's go all purpose, kind of over provisioning on model, do you think?
Chad Gibson: I mean I think I think part of it is simplicity. Like I if I take my workflow, there are still workflows that I'm using a big model to do something that is overkill, but I mean it's on me to, you know, delegate to the small model the task that the small model can do. So it's simple just to use one model for for a given workflow.
Anne: Yeah. Yeah.
Chad Gibson: there's things we're doing to help with that, and there's ways that you can make it easier for different tasks to be delegated. but the simplicity. And one, like I think I think the cutting edge is moving so quickly that people are still realizing how to fully exploit, if you will, or maximize these frontier models. And then when the new frontier comes out, it's kind of like you need to relearn again. How can I re-establish like How can I re-establish my workflow now that knowing the model is, you know, 20%, 30% more capable? So the fact that the frontier is pushing while people are still trying to figure out how to maximize these models in their workflow, it's kind of pushes to this point of simplicity. Well, I'm just gonna push everything on the high model for a bit and see what I can learn and evaluate from that. And then right when I get comfortable with that, yet another new model. comes out and sets the bar even higher.
Anne: Yes. And as you say, there's it there's there's so much potential ability to effectively get more out of the software that we haven't really even scratched the surface of yet.
Chad Gibson: Right. Totally.
Anne: so there's really no end in sight, is there?
Chad Gibson: well I don't see the end. I mean I I think I think conceptually things it's an interesting thing to contemplate. Like cause now the new frontier model, the new open way frontier models are coming out with like 2.8 trillion parameters. And and so generally, if you if you if you take a look at the landscape of like first on hardware, so like NVIDIA and AMD and others Like they they basically look at hardware and like these, you know, it's the the the the capacity of a single chip and like the capacity of a single server. And then generally they have like what's their scale out pattern. So an example with like NVIDIA is they have blackwell, let's take Blackwell, like there's a Blackwell chip, and there's like a Blackwell server, like a B three hundred server that has eight of these interconnected, and then they'll have like, you know, a r a Ruben rack or a bare Ruben with like scaling this out to like the biggest compute in one rack. These frontier models are now pushing the point where you need just about two servers or two 8 GPU servers. But theoretically, the parameter counts could increase to a point where you may see a model where you need a full rack to serve it. So conceptually speaking, the I don't think the ceiling on parameter counts is necessarily in sight. But while parameter counts are being pushed on, you're also seeing models get better at smaller parameters as well. So while the frontier is pushing model size, there's amazing innovation happening with making a 300 billion parameter model better and better and better. So it's like there's also innovation happening on making the value of each weight higher. And then NVIDIA is pushing, you know, the increased capability of each chip. Which therefore is increasing the capability of each server, which is increasing the capability of each scale-up unit. So the fact that all these things are still progressing, I don't I don't have a good sense for which of those three things in the next 12 months is going to be a limiting factor. Like they're all I can see 12 months of growth across all three of those. in 24 to 36 months, it's it's it's it's also for me unclear because it's just moving so quickly.
Anne: Yeah, absolutely. Right, and it's and it's obviously it's a bit it's a bit terrifying because you you've got a good data center, you're doing all the things that we said, if only data centers would do this kind of thing in building grid software. you're you're adding flexibility to the grid by by the talking to working with grids to turn things turn dial up dial down all that kind of stuff. You're using hardware that was that was going to waste or excess electricity that was that was not being used as well as it could. You're you're providing data to people so that they can act on it and, you know, true kind of open source, but also, you know, provided providing the telemetry so that people can go, well if I change that, it would it would improve things like you're doing all the things. But there are so many data centers out there which do sound like, you know, they are pretty terrifying. they're hooking up brand new gas-fired power stations. so yeah you can kind of see there's the hashtag not all data centers, but there are data centers. There are some pretty evil data centers out there and they often have
Chad Gibson: Yeah.
Anne: names.
Chad Gibson: Yeah. Yeah, yeah. I mean like on site power is like if you're if you're impatient, the easiest thing to do is I don't want to say easy, it's not easy, but minimizing reliance on the grid is w is is one way to do it. and there are ways to do that where you could, you know, put it put a modular data center where there's a big solar array where there's abundant solar energy with you know, renewable renewable like with batteries and there's some pretty cool stuff happening there, but there's also ways of just hooking up more diesel generators. And yeah, like I think I I think if we can shine a light on some of the more constructive ways of doing it, can address that fear we talked about and that, you know, anti data center sentiment. but oftentimes, you know, the the the those cases you describe are the ones that become the the examples used.
Anne: Yeah. And of course the difficulty with that is that people hear only about the bad stuff and they think there is only the bad stuff available, that there's no choice.
Chad Gibson: Yeah.
Anne: It's like AI or being sustainable or being green or aligning yourself with the energy transition that, you know, it's an either or and people are always going to choose AI if in those in that circumstance. But I think we do need to shine a light on the fact that there is there are other alternatives. You know, you do not have to to use these giant data centers.
Chad Gibson: Yeah, you know, when we look at like the transition to cloud computing, when people first started moving like on premise workloads to, you know, AWS and Azure and Google Cloud Platform, all those platforms kind of out of necessity started evolving, like providing, you know, carbon and carbon observability for cloud workloads. And that's now a a feature of all those cloud platforms. And it's evolving to provide more, you know, more regional based versus doing tricks with carbon offset credits and whatnot. That level of transparency we absolutely believe is gonna come to AI. that the need for the and we're seeing it in Europe already. We're seeing customers in Europe who kind of want a direct correlation between where these workloads are run and how they can correlate it to to what state of energy generation is is correlated with it. And I do believe that will come across worldwide at some point. don't entirely know know when, but other parts of the world. But but the challenge with that in the immediate short term is the the again, it goes back to speed. And right now everyone is kind of racing and rushing to understand how to leverage AI in a way that it's It's driving more pressure to build out capacity and to, you know, try to service these workloads and ever and ever even the suppliers themselves are trying to figure out how can we scale this effectively, how can we help others figure out how to use AI. So it at the speed is I think putting the pressure on all aspects of this.
Anne: Yeah. Yeah. Yes. So, so if you had advice for listeners, watchers of this podcast about what they can do, what should they be doing? What should we what should they be thinking about? What should what do they need to open their minds to the possibility of in AI?
Chad Gibson: I mean, so first of all, there there are ways. I mean, just to plug neurologal, there are ways to, you know, build and use AI in a way where you know, you know, you can correlate it to energy and carbon usage. So if you yourselves wanna know wanna wanna answer the question of how much energy is my is my AI use consuming, we can answer that for you. Two, there there are ways, and you know, I think you know, to the to to their AI service of choice, you know, they can they they now can speak confidently that like, hey, I mean, like customer demand is going to change this as well. You know, if you're a huge anthropic fan, which I am also a huge anthropic fan, like letting anthropic know that we'd love to know how much energy is being consumed by Opus and I would love to have that kind of transparency. Like, that's a meaningful thing. You know, we've proven it's possible, it can be done. We would love to to help anthropic satisfy that need, but that's something as well. And then when it comes to you know the question of data centers and is it gonna increase my power bill, I do hope that like this conversation has surfaced, there are ways of doing it that can be constructive. And not not all data center is the, you know, gonna put gas generators in my backyard. and the more that we can allow capacity to come online that is, you know. making the grid more flexible. That is, you know, consuming this renewable energy we have is a good thing. And in some cases will probably mitigate how much new, you know, turbines and generators we need to create so create elsewhere. So, you know, hopefully that provides a bit more of a of a of a of a view on on the new the nuance, if you will, of of how they're not all created the same. But yeah, I mean I think I think there are options to to to consume AI and be an AI champion and optimist that allows you to actually understand energy consumption.
Anne: that was very useful to for us all to be reminded of that you do have options here. It isn't it isn't AI or the planets. there there are ways of doing it but are not that not not are not as destructive. So and and a big thing, if you if you can't do anything else, just ask your suppliers. Say, look, I care about this. Can you tell start telling me? Can you start doing something about it? Tell me what you what you're doing, what you're thinking, what your plans are, because I do care about this and I will be making purchasing decisions in future based on people's plans and and and hopes for this kind of stuff. So what so that that was very interesting. Is there is there anything else you want to add? You are allowed to you the you're in here, I'm perfectly happy for you to just keep plugging neural what because you offer a service that I think should be offered more widely.
Chad Gibson: Yeah, I mean the thing the thing that is interesting, and this is like one of the challenges we we deal with, like so we when we offer inference, we we price it by the kilowatt hour, which is which is both
Anne: no.
Chad Gibson: a really cool thing and a complicated thing. And so for users who are used to paying by the token. Consuming their AI by energy is a is just a challenging thing. And so it it we've we've had to to make that easier by providing a lot of transparency. The key thing is it provides transparency. You know how much energy you're here. And now I can optimize for that. So what's fascinating is like, how do I optimize for that? Well, there's certain times the day where energy is more abundant for AI. And and you can do this with carbon as well. There's certain times of the day where your grid is emitting more carbon, and there's times a day where it's not. And so this is like this whole new frontier as of optimization opportunities. And so a lot of our current customers are like pushing us, and it's a really fascinating, fun thing to be pushed on. Because as more of these tasks, if you will, you know, imagine a future where, like, if we think of inference, largely inference stay for our cloud at least, is a human piloting something that's calling our service. more and more it's transitioning to where it's a human that created a autonomous agent that is then calling our service. And as we go to the future, like I was at the AI conference a month ago where they're like planning for a future of like, you know, s autonomous swarms of agents, of software agents. And so the cool thing about the agents is they can if you give them data to optimize, they're really great at optimizing. So we can provide transparency of like, hey There's way less carbon being emitted in the grid right now. Hey, energy is really cheap and abundant right now. And so some tasks that aren't require like immediate response, let's put let's let's do those tasks when there's less carbon being emitted. Like the optimization potential becomes pretty, pretty profound. And now if I'm gonna if I'm gonna give a human customer who's doing who's doing software engineering like these fifty variables to optimize against, it's gonna be hard for them to figure out, I have to do bunch code. Maybe I should wait to do my code next hour. But for the part of their workloads that are increasingly becoming more autonomous, like they're just phenomenal optimization potential. So that's that's a really, really fascinating world where we can empower, you know, both humans and autonomous agents alike to optimize. And we can provide
Anne: Mm.
Chad Gibson: all of this transparency data. To you know, leverage the resources when they're abundant, leverage them when they're cheaper, and reduce the pressure on the resources when they're under great strain.
Anne: Mm-hmm.
Chad Gibson: and so that I'm I think that's a really, really exciting proposition. And it kind of goes, it just goes down this path of doing more with the resources we have, better utilizing the resources we have. And the cool thing is it's not just doing things to, you know, minimize consumption, but it also makes it cheaper. And so there's
Anne: Yeah.
Chad Gibson: just the very, very clear benefit of of consuming things when they're cheaper. And and that I mean everyone loves that.
Anne: Yeah. Well they do. They and they well, it's an interesting saying they they love things to be cheaper, but they don't really like to go too much effort to get it to be cheaper. So
Chad Gibson: Yeah. Yeah.
Anne: but but the the barrier to what is difficult is changing because of AI. and therefore
Chad Gibson: Yeah.
Anne: Yeah.
Chad Gibson: Yeah, I agree. I agree. I mean it's it's like I said, it's it's like all these features and capabilities, we're trying to move as quickly as we can. And while everything around us is changing really quickly and it's it's it's it's f it's exciting. I mean it's it's so I think about my career at Microsoft. Like I I I I loved my time at Microsoft. It's a it's an awesome place. I w got to work on some massive, huge, quite impacting products. But the speed at which I'm moving now is kind of one of the reasons I wanted to to to have, you know, a career post-Microsoft and it's so exciting. It's it's just it's it's being being around the technology this quickly, it's it's it's there's new interesting problems emerging like every week.
Anne: Yeah, I mean I guess we all said we wanted to be in the tech industry to solve difficult problems. We had no idea, really did we?
Chad Gibson: I know, I know, I know. They're are are they ever gonna be solved? I don't know. there's gonna be new ones all the time.
Anne: Well, on that happy note, and it is a happy note because we did to go into the tech industry to be to solve difficult problems and now it's it's more difficult and it's faster moving than ever. You know, my thirty years I was there I was at Microsoft just a little bit before you, I think. and then but it was yeah, it's it's it nothing has changed at the at the rate that that we that like the rate we're seeing at the moment. but so thank you very much indeed for being on the podcast. It was a delight, really interesting. and very inspiring as well. There are things that we can do and tricky, tricky problems to solve. So thank you very much indeed for being on the podcast. And thank you to all our watchers and listeners across the globe. and hopefully I will catch you again on a future episode of Asynchronous and Unreliable Podcast. Thank you very much.