Anne: So, hello and welcome to episode 30 of Asynchronous and Unreliable, a weekly podcast where we discuss the most interesting, at least to me, that's the whole point of this, interesting to me, ideas and concepts and technology. I'm your host, Anne Curry, co-author of Building Green Software from O'Reilly, and author of the science fiction Panopticon series. And today I'm gonna do something a little bit different. I today I'm on my own, because today we have a sick household where everybody has a terrible cold. This week we were going to record an episode about feedback loops. John's very interested in feedback loops. It's something they are a concept that came up a few weeks ago when I was talking to Yan Ching Chang about how AIs could be taught to be better managers using feedback loops. So I really, really wanted to dive into what feedback loops are and how useful they are. But Sickness has got in the way of that. John is nowhere near able to talk about feedback loops, and I've even I would be struggling with it a little bit. So that does leave us with a rather interesting situation. and it is a situation that both John and I, my co host John Berger and I are both very familiar with and both very interested in, which is handling unreliability. and you'll notice that this podcast is called Asynchronous and Unreliable because we're very interested in handling unreliability. Asynchronicity and unreliability are a key way of making systems more reliable. Ironically enough, unreliability is an approach to handling and producing better reliability. So by making systems more reliable. By getting them to embrace unreliability. So this week's episode, I warn you now, will be a shorter episode with no guests, just me. And it's I'm going to talk about three techniques for managing unreliability in systems and networking and life in general. And apply those techniques to this podcast. We'll use this episode as a worked example of applied reliability techniques. In the real world. So the first thing I'm to do is is set the scene a little bit by talking about what are the the the major failure points in a communication system so that we can kind of we're all coming, we're all on the same page. So in in most kind of communication systems and most systems are at heart, most human systems are at heart communication systems. You've got the supplier of the piece of communication, which in this case is me and my guests or my co-hosts. The network, which we use to supply that information to folk at the other side of it. That's usually a combination of things. That's the internet, that's my podcast host, Riverside Studios. And then the final key component of a communication system is you. the consumer, the listener or the watcher. Now, in episode 26, I had Sam Newman, author of Building Microservices and his forthcoming Building Resilient Systems on the podcast. And he talked about something that he called a socio-technical system. Now a socio-technical system is basically any kind of complex system. And what I've just talked to you about there, a podcast like this one, is a socio-technical system because you've got people, me or my co-hosts, and you as your li as the listener or watcher involved, and our behaviors change. There are a lot, there are technical things that we can't do much about. So the network, you know, if if the internet goes down, then then that's one kind of failure mode we can't do that much about. But as watchers or listeners or hosts, there's loads of things that we can do to change our behavior to make a system More resilient to failure. And I'm going to be talking about some of those today in this short podcast because that is one of the ways that you can react to situations to make your whole system more resilient, which is to apply a graceful downgrade, a gray out, or graceful downgrade. A downgrade of function which provides most of the value. But not all of the value. So in this case, where I'm going to be applying two types of graceful downgrade to this particular podcast episode. One will be a little bit shorter than usual, and the other is it will just be me. So that's two classic ways of making the system simpler and therefore easier to deliver. So as I said, I'm going to talk about, I've talked about the three components that can fail in the system. the supplier, me in this case, the network, the internet, say, and the consumer, you. Maybe you don't have time. Maybe I normally put this podcast out on a Monday, but maybe you're busy this Monday and you can't listen to it to Wednesday or Friday. That's all of us in this system have failure modes, and that's not wrong. you don't have any responsibility to listen to this podcast as it's as it's live recorded, or immediately it goes out. It is my job here, the job of the entire system, to make it easy for you to consume it whenever you are available to do so. So, so we're going to talk about three three key ways that a system is made more reliable. one is agreeing expectations up front. The second is providing graceful downgrade options, and the third is buffering. Now there are more. And there are more different ways of of thinking about that, but the of these things. But I like those three. So we're gonna talk about those three today. And those three are three that I'm applying specifically to make sure that this episode, episode 30, does still go out and hopefully is still interesting and useful. So so one of the one of the the things that we don't often make enough of is that if you want to sit to be a system to be more reliable, it's great to agree. Expo expectations up front and allow people to handle them in advance so they're not they're not shocked and surprised surprised. That is the fundament of of unreliability, really, is that you need to let people know that a system is unreliable and therefore they should be expecting a degree of variation in the supply of that system. So obviously, one of the things I did with with asynchronous and unreliable. Was attempt to state up front that it would be unreliable through its name. And unreliability really means just don't bet that this will be available on a particular day. Now, I have, to a certain extent, undermined my own messaging there, which is actually quite a common failure mode in systems, to undermine your own messaging by being a little bit too reliable. This is my 30th episode, and I have put every week. At least one episode live, which actually not un not unreasonably might set the expectation that there will continue to be an episode every week. And in fact, this week, even though I'm ill and we're all ill and all my plans have gone out the window, I'm still putting live an episode. But I feel that this episode is part of communicating the fact that actually this this podcast is not meant to be a hundred percent reliable. Maybe next week, hopefully, John will feel better. And we'll be able to have the record the episode on feedback loops during this week, and then I'll be able to go live with it next week. But who knows? this is your warning now that this podcast is called async and unreliable for a reason, and we might not go live next week. Hopefully we will, but we might not. so agreement expectations up front absolutely key to producing resilient systems. And it's much underlooked. We communication is unbelievably important. Upfront communication is unbelievably important if you want a system to be resilient. Because in any socio technical system, it's it's it's a bit of a mouthful that, but i y y hopefully you know what I mean. a system that has people involved in it, that people are really useful to have involved in systems because they can change their behavior. Not all people do, and therefore they those people can be a terrible. problem in your systems. But people who can and do and will and know how to adjust their behavior can be an incredibly useful part of your system. They can go, you know, there's no episode this week. Never mind, I'll listen next week and I'll I'll I'll miss it this week and and I'll I'll just catch up next week. Or maybe I'll watch and listen or watch or listen to an old episode that I haven't listened to yet. So there's lots of People, if they know that something's happened and they know that actually the podcast will be back up again next week, you're keeping them informed, they can make decisions all the time to add stability into a system. And really, with every system you design, you want to be able to use the people who use it as part of the stable stabilization of that system, whether that system is a podcast or a family or a business or a product. it's it you know people can be an incredibly useful stabilizing component but they need to know that they're expected to be and they need to know how they can go about doing that. So agreeing expectations up front phenomenally useful. As I've mentioned before and one of the the techniques that I'm using today is To have planned out some graceful downgrades of functionality up front. Now, we don't know these things that I'm going to do today, I'll be trialling a shorter episode and just presenting on my own without a guest. I'll be trialling those as graceful downgrades of function. Now, when I put this episode out, so this is the first time I've done this. So when I put this episode out, we'll find out whether everybody goes, I hate it, and nobody listens to this episode. Which case I'll know that this isn't really a useful. graceful downgrade of the functionality. But it's impossible for me to know that until I try it. So this will be a good experimentation episode. and another so that's that so agreeing expectations up front or at least stating expectations up front so people can hear them going, yeah, do I like that? Do I not like that? Will I continue to listen to this podcast? Will I move to a different one? That is a very useful thing. The next thing that's useful is thinking about graceful downgrades of function. Now, there's there was a class, there are couple of other graceful downgrades of function here that you might even say is not really a it's just different functionality that I haven't used. So it would have been very sensible of me to have had a co-host lined up who could run the podcast when I wasn't available. The thing is, if I'd done that, it would have been John, but he's ill. So that that is the lesson there that any kind of backup you come up with, that might fail as well. So you might all fail, especially as John and I's having a cold is highly correlated. So I should have taken that into into account. the third wait, so we've got one agreeing expectations up front, two graceful downgrades of functionality. Think about that in advance and then be prepared to to to try it out. the a third technique for producing more resilient systems is buffering. It is doing work in advance of need and storing it somewhere so that somebody can take that work when you when you when you aren't available to do it real time. So today I am recording this episode. The day before I really wanted to go live, which is which is not ideal. But if I'd been sensible, perhaps I would have buffered up five or six episodes in advance so that I wouldn't need to do that. But actually, that isn't really true. I've been in the tech industry and in the networking and networking and resilience part of the tech industry for many decades. So I'm very, very familiar with the idea of buffering. So in this case for a podcast, the sensible thing to do. And often you do you if you listen to podcasts over the summer, lots of folk who record podcasts and are going on holiday will record a whole load of episodes in advance. And that was my my plan. I was going to record a whole load of episodes of this podcast in advance. And then if I was ill or on holiday, I could roll out those episodes. And I was thinking, this is the key to resilience. I'll never have to worry about having a cold because I'll always have a whole buffer, a load of back episodes to to to go live with. And I thought, I was really pleased with this. I thought, that's it, you know, I'll use all my years of expertise to make sure that I have a super, even though I've called my podcast asynchronous and unreliable. I will make it unbelievably reliable by applying all of these techniques to it. But in reality, what I discovered was that I have a socio-technical system here, not a not a completely technical one. And one of the key the one of the key important features in this system is me and my brain and my mind. And I realized that there was a significant problem for me with buffering up a whole load of episodes in advance. And it was a and I'm gonna here I'm gonna refer back to another episode earlier in this post podcast series, the one I recorded with Matthew Skelton, one of the co-authors of Team Topologies. Now one of the the The key things that he talks about in Team Topologies is is is a concept that he himself didn't invent, although I think that team topologies was the first time it was applied to a team. but in this case, this is I'm applying that con this concept to an individual, me, which is cognitive load. what I found when I, because I did at the start, I used to maintain a buffer of about four episodes that I'd recorded and not yet gone live with. And what I realized was that while the episodes were were buffered but not live, I they they I held them very strongly in my head. And those these episodes contained loads and loads of information. and the one of the problems I had was A, it was a lot of information to keep in my head, and I really needed my head space for other things. but the other problem was that when I had new guests on. And recorded new episodes. I couldn't help but refer to what was what had been said in these previous episodes. And that wasn't really fair because these episodes were not available to the other to the guests that I had on my podcast to listen to. Now I think it's perfectly reasonable for me to refer back to what's been said in a previous episode that they haven't listened to. That's absolutely fine. I can say, well, that X said Y, and they go, yeah, yeah. And as a as an a listener, you can go, well, actually, I'll go and have a listen to that episode if I haven't already listened to it. But it's not really fair for me to refer to what's been said in an episode that is not yet live. So my guest can't possibly have listened to it, and my listeners can't possibly have listened to it. And that really, really, that was a major issue in the I would say the first month, two months of this podcast. that I was trying to hold that was holding four episodes in my head. And I found it impossible not to refer to some of the stuff if it was relevant in the the next episode that I was having. So I I came to the conclusion that I couldn't really buffer that many episodes. so yes, that was that was an interesting way of learning about the difficulties with socio-technical systems rather than just technical systems. In a technical system, all you do when you buffer things, if you if you're like using a a content delivery network, for example, helps to make the whole internet more resilient by buffering resources. And what it means though is it just saves a copy of that resource. And on a hard drive it might have a copy of a popular TV show that you don't want everybody to have to download all the way from the US or or Korea or wherever it is that the the popular TV episode is coming from. So that Digital assets that copy can be made is maintained on a disk somewhere closer to the user, usually, or somewhere more conveniently reached when the user wants to stream the content, wants to get access to the content. But that there's no cognitive load associated with that. That's say there there was a lesson there that there is a real difference between buffering as applied to a CDN, where you're just taking a digital asset and sticking it on a hard drive somewhere. And it's conveniently just stored there and waiting for someone to ask for it. And a a podcast episode, which I cannot help but hold in my head until at least until it's gone live. But often then, after it's gone live, I mean, I I remember quite a lot of all of the this this episode 30, and I remember quite a lot of all the previous episodes. partly because I'm very good at remembering things, but but partly because I had to listen to them so many times and edit them and Go through it, read all through the transcripts to make sure they're okay and all that kind of stuff. It makes it very hard for me to go, don't say anything about those last four episodes with this new guest, because they can't have listened to it and and my watchers can't have watched it or my listeners can't have listened to it yet. And they can't listen to it now if they if I if they want to go back and refer to it. That was just too hard. So there is a a valid lesson there, which I only learned by doing it and failing to do it, which is. I can't really buffer podcast episodes in the same way as I would buff buffer as a CDN. I am not a CDN. The so the interesting thing to hear to me, and it is interesting, is that so many other podcasts do do it. I mean, over the summer, most of the podcasts that I listen to, they buffer up a whole load of interviews with X and interviews with Y and interviews with Z. and it doesn't seem to cause them any real problems. I I think they probably do that. They maybe do that by recording episodes that are just kind of slightly orthogonal to the stuff that they're talking about the rest of the time. So it's not really so likely to come up. So it might, for example, be a podcast about the life of X, or the career history of Y, that actually is a little bit historical. Maybe the rest of the podcast is quite. topical, but the ones they record are quite historical. And if you do that, then there's maybe you can you can find a way of separating out the thinking about the historical stuff that really is kind of like nice to know, but isn't all that germane to the to the to the live episodes, the topical episodes. And then you can continue on your weekly topical episodes without getting too distracted. But it might just be that they're all by nature They're all presenters. And I'm not a TV presenter or a radio presenter by nature. Maybe if you are a re radio or a TV presenter, you just get really, really good at excluding all the other streams of communication you do so that they don't cross-contaminate one another. And it might be that I'm just no good at doing that. Either I'm temperamentally no good at doing it, or I just haven't learned how to do it yet. But anyway. that that's been my problem with buffering. Because everybody asks, every single because I know a lot of networking people and they all ask me how many weeks I've got buffered up in case I'm ill or in case something goes wrong. And I used to say four weeks, but now I have to say it's like one week because of the issue with being part of a socio technical system. Now, so I did say I'm going to make this a short podcast. You have heard quite a lot of stuff in this podcast. We've talked about reliability, and we've talked about different techniques for making systems more reliable while being built on top of an innally unreliable reliable base. Because humans But I'm unreliable. I'm I'm the key supplier here, but my co-hosts have also turned out to be unreliable. That's not a problem. That's just humans. We are humans. So how do you put yes? So we've been talking about how you build reliability on on top of an essentially unreliable system. me and my co-hosts, the network itself, we I have been remarkably lucky that there hasn't been a major network outage during the time this podcast has has been going live the last six months. You know, there hasn't been the whole internet hasn't gone. Well, the whole internet hasn't gone down through my entire life, and that's is astonishing. We'll do a whole I'll do a whole episode up with John based on we we take that for granted and we absolutely shouldn't. It's been amazing. Every week actually we should we should thank the the internet stars that the internet has lasted for the entire week because there's really no reason to for us to have assumed that it would be as good as it is at this point. si whatever it is 60, 70, 80 years since the first movements in the direction of an internet work were made. and you. you've probably had cold during what during the time while you've been watching this. You probably have been intending to listen to it on a Monday, but you haven't really had time to listen to it on a on a Friday. Maybe you haven't listened to it for weeks and weeks on end and then suddenly thought, maybe I should listen to this again. Or maybe you just started listening to it. And this is your first episode. In which case I'm Do go back and listen to the earlier ones. They are they're topical, but they do so sound the test of time. I do go back, I was I before I did this one, I went back and listened to Sam Newman's one on resilient building resilient systems, and I and I th I found that quite an interesting perspective to get before I recorded this episode. And I haven't even talked about asynchronicity. I've only talked about unreliability and haven't really talked about asynchronicity. and I will talk a little bit about asynchronicity, I think in probably a different podcast because asynchronicity, it's such a complicated and interesting subject. I think we probably are gonna need a whole episode on on the topic of asynchronicity. And I know that Sam Newman is quite keen to come back and talk about asynchronicity. And John's very keen to talk about asynchronicity. So I will save that so possibly for a future episode where I actually do have a guest because we're all healthy. so thank you very much for listening to this short episode. I do hope that you enjoy it because then it will be one of my graceful downgrade modes. If you don't enjoy it, let me know, and then I'll know it can't be one of my graceful downgrade options. So but thank you very much indeed for listening. And now I will go back to lying on the sofa and read your book while I recover from my cold. So thank you very much, and hopefully see you fit and healthy along with John, where we'll talk about feedback loops on next week's episode of Asynchronous and Unreliable. Thank you very much.