Crafting Code Podcast
$ cd episodes/028-technical-health
~/podcast/episodes/028-technical-health $ ls -1a ~/podcast/episodes/028-technical-health $ cat episode-summary.txtThe change of perspective which comes from reframing a problem can often yield fresh insights or improve our outlook. Discussing the state of a technical system in terms of health may be preferable to the oft-used debt metaphor. In this episode, your hosts discuss how this shift in thinking can help us deal with our software issues, and then we share some ideas which have helped us communicate and manage technical health.
~/podcast/episodes/028-technical-health $ cat references.txt- Learned Optimism. Martin Seligman.
- Technical health, technical debt and code rot
- Fitness functions
- Delivery metrics and productivity
- Refactoring, legacy and rewrites
- Architecture decision records
$ cat transcript.txt
[00:00:16] Allan Stewart: Welcome to the Crafting Code Podcast, where we discuss the importance of doing the right thing at the right time with the right tools. I'm Allan Stewart, a software architect, and lately I've been thinking about the intersection section of job enjoyment and fulfillment.
[00:00:31] Dave Adsit: I'm Dave Adsit, VP of engineering. And recently, I've been thinking a lot about building effective teams and engineering cultures for in-office, remote and remote hybrid environments.
[00:00:44] Allan Stewart: This episode topic is technical health. I think it was a couple of years ago that I started reframing the conversation about technical things away from this idea of debt and started talking about it in terms of health. And apparently other people have been having a similar revelation because the ThoughtWorks Tech Radar started tracking a blip for tracking health over debt. So I think it's an interesting interesting mindset shift to start thinking about it away from a negative frame, right? Debt is usually something we think about as bad and health is a more positive framing. Although inevitably when it comes to technical health, we're telling you about how it's not good instead of excellent, but at least it gives you that positive frame of mind that you're thinking
[00:01:43] Dave Adsit: about. Yeah. As we started talking about this originally, it got me thinking a lot about the concepts behind positive psychology and, you know, it's the founder of positive psychology, Martin Seligman, and how he talks about, he's got that book, Learned Helpless or Learned Optimism, actually is the name of the book. And basically the concept there is that you can teach someone to, through consistent negative reinforcement, you can teach someone learned helplessness. Like they could potentially improve their situation, but because you continually beat them down every time they try, they eventually give up and accept the bad state that they're in. On the other hand, you can learn to think in a much more optimistic way around, "hey, I can improve this situation no matter what it is." And he talks about how the most successful successful salespeople are those who can make 99 calls and have, or make a hundred calls and have 99 people hang up on them and then have one person schedule a meeting and then go to a hundred meetings and have 99 people say no. And only one says yes. And they still consider that a win. Most of us get discouraged. We give up well before that. So there actually is something really important in changing the framing around how we consider a concept so that we are looking at it through a much more positive lens. You know, one of the things that is taught in cognitive behavioral therapy is to look at the positive side of opportunities. If you constantly focus on what's negative, it can put you into a negative mindset, which will decrease the possibilities of you succeeding in the space. And so what we want to do is give ourselves, I mean, it kind of sounds silly because I remember back to the Saturday Night Live sketch around daily affirmations or whatever, right? But actually putting ourselves into a positive mindset can actually impact our ability to have the endurance to go the distance to have the success we need to have. And so I think there's something really positive and powerful about switching the metaphor from debt to health.
[00:03:58] Allan Stewart: And as we discussed on a previous episode, you will always have technical debt, right? If you're thinking about it in terms, because it's not actual debt, it's a metaphor, but this concept of in your system, there's always going to be something that could stand improvement. And so if you are taking that negative stance on it, thinking about it as debt, then you will always be in debt. And it's not recoverable, right? You can't get out of it. You can pay down debt. You can take care of certain problems, but you'll always have more debt. And so I can definitely see how that would feed into that downward spiral of just learned helplessness or just acceptance of, well, I don't know that we need to really worry ourselves with tech debt because no matter what we do always, there's more, there's always tech debt. So like, why do we even try?
[00:04:58] Dave Adsit: Yeah. This system is terrible and it will always be terrible no matter how much work we put into it. So we might as well take shortcuts now because it's not like we're making it worse. Right. Even though we know that we actually are probably making the system worse by taking the next shortcut.
[00:05:13] Allan Stewart: Exactly. And the, the debt metaphor doesn't fit in some cases too because with debt it's all dollars, right, or some kind of money, and you can exchange one kind of money for a different kind of money, and so like, it should all just be equivalent, right, but we know that in software it's not, there's different kinds of prob technical issues, different kinds of complexities, and some of them are more expensive, some of them are stickier and and harder to deal with than others.
[00:05:44] Dave Adsit: Yeah, I remember long ago, you and I worked in a system that still had one little, we can mostly decomposed our monolith into microservices. And we had this one little part left and we always called it the "monolith tax," whenever you had to work in that code that nobody knew very well anymore. It was like things that should have been easy were just a little bit harder than expected and took a little bit longer than expected. And that was just the state of it, right? As we moved more and more functionality into the other architecture that we were moving towards, things will get easier over time and we would be able to move more quickly. And so the overall health of the system was increasing, even though it was very hard to work in the part where that we were considering to be our legacy debt. Yeah.
[00:06:33] Allan Stewart: I think that piece of the system, or at least one of those monolith pieces of the system that lasted a long time, actually outlasted the team that technically owned it because it didn't require a lot of maintenance. It generally was working. It wasn't an area of focus. And slowly over time, the team turned over to the point where nobody actually knew how to run that.
[00:06:59] Dave Adsit: It's not that they didn't know. They were physically incapable because everyone on the team had chosen to use MacBook Pros and the legacy system was .NET Framework and they couldn't even work on it anymore. Even if they had known how, they still would not have been capable of doing the work. Right.
[00:07:21] Allan Stewart: So coming back to this idea of health over debt, is it just a better metaphor? I mean, yes, we're still talking in metaphors, but I think it's more than just a rebranding, right? So when we take something like this and we're trying to change our paradigm, it's not that we're just substituting in one word for a different word. So it's not that we're just saying health and debt are interchangeable, but in fact, we're reframing it in a way that gives us new insight into how we can measure the system, what we can do, like what opportunities we have for working on it. With the debt metaphor, it's always about looking for problems to pay off. But with a health metaphor, well, now you can think about increasing your health. There's the possibility of kind of almost a fitness aspect of it, where you're saying, "hey, in order to to be able to accomplish something, are we fit enough, right?" We might be generally healthy, but is that, does that mean that we're ready to run a marathon, that we might want like an elevated level of fitness, or if we want to be really good at a sport, there might be certain exercises that we need to do. Right. And so carrying that metaphor over into the technical space, there may be areas of your code where you're not as healthy as you want to be. And it's not because you're necessarily in bad health, but you've not prepared yourself for something really important that you know is going to happen that relates to your system technically, right? So for example, can you scale to meet the kinds of load that your particular system lives in, right? So maybe you've got periodic seasonal events. Maybe it's around Black Fridays, or maybe it's around the start of the school year or some other event that causes a big disruption. And are you healthy enough to be able to perform for that event? Just like a, you know, just like a athlete, they're going to have to prepare themselves, practice, get themselves fit so that they're ready for a particular event.
[00:09:42] Dave Adsit: Right. And I like to think about it just to reemphasize that point. There's different types of health based on different goals that you have. You can do strength training because you want to be very physically strong. You want to be able to lift a large amount of weight, or you might be doing strength training to enable you to increase endurance by doing more cardio training based on having more strength. Right. And so different types of of health, I mean, you would consider both a bodybuilder and a marathon runner to be healthy, but they're not interchangeable. And our systems can also be healthy in a way that is fit for one purpose, but not for another. So if you have a large number of consistent users, a high number of daily active users, right? Maybe you've built your system in one way. And if you have really really high spikes, you might build it in a different way. So if you have a consistent daily load, you might say, "hey, our system is very healthy by having a fixed number of servers in place at all times. And if the number of servers drops, we know that we're out of our health band and we want to add one back or whatever." And if you have consistently low level of users with huge spikes, then you're going to have to build a system in a different way. You're going to have have to build things like auto scale. You say, "hey, this system has to have a really, really fast response time on scaling up new instances, new nodes to handle traffic when we start to see a traffic spike, because we don't have time, we don't have a long time to respond." And so both of those systems could be healthy in their context, but it is a different context and we need to be be aware of that. Agreed. Also, both of those systems, I mean, regardless of what type of health you're looking at, there is this idea of built-in continuous preventative maintenance and, and effort. No one says, "Hey, like I'm healthy because once I exercised and now I sit on my couch every day, right?" If you want to be a, if you want to be a bodybuilder or a marathon runner, you are going to be lifting weights and running almost every day. I mean, I don't know how to train for a marathon. What I've seen is running almost every day, but different distances every time. That's the extent of my knowledge because I've never trained for that. But I do know that going to the gym and lifting weights every day is not going to prepare you to run a marathon.
[00:12:16] Allan Stewart: Right. It's not in stasis, right? In the debt metaphor, we think about it in terms of, of, "oh, well, we've got this much debt. And like, why would we keep accumulating more debt?" But in a health metaphor, it's more easily understood that, well, there are some things that just naturally decline. And if you don't keep up, you don't do that preventative maintenance, you're not going to be able to do it. And you can't just pay it off all at once either. Right. I think that's the other thing that the debt metaphor is perhaps guilty of having business people think about, "well, can't we just pay that off? Like, can't we just invest in getting rid of it?" And well, yes, you can for a thing, but then they get frustrated later because there's still more and more and more because you can't just go to the gym as, as we learn every year, you can't go to the gym for January along with your New Year's resolution and then expect to still be fit, you know, come October, November. It's more lifestyle change kind of concepts. I've liked to talk recently about technical lifestyle changes where you're really changing how you go about owning and maintaining your code and getting away from those ideas of, "okay, now we're going to do a project, whether it's a feature project or whether it's a technical debt paying off project, right?" You have to change your mindset of what you do regularly in order to do that. And I think that the health metaphor gets people prepared to hear that and understand that concept much better than the debt metaphor does.
[00:14:00] Dave Adsit: Well, I like to think that if you're building a healthy system, you are probably investing every week in cleaning up up bugs or issues that customers have found, users have found. As opposed to, "we're going to let all that stuff pile up. And then once a quarter, we're going to do a big bug bash where we try to resolve them all in one or two days." The same thing happens. And I would say that we as engineers are more guilty of this than the rest. We say, "hey, after this sprint, we need to take six weeks to rewrite this huge component that we made a giant mess on." And we're going to give ourselves a deadline so that by the time we get close to the end of the six-week period, and we've only done half the work, we can make sure that we rush and rush and rush and rush and finish the rest of it super quickly. Filling it with bugs so that we are guaranteed to need to do this again later this year. We never pitch it to the business that way, but that's what is inevitably going to happen if you try to address all of your technical debt all at once. It's the grand rewrite or or system to effect, right? Yep. We're like, "hey, we made such a mess with this first system, but we learned all the problems we could make. So we're going to make a new system to replace it." And now you have two systems to maintain, and now they're trying to keep up with each other. The new one is trying to outpace the old one. And eventually you end up with just a much bigger mess than you ever predicted. So if we have to change our technical lifestyle, what do we need to do to make that happen?
[00:15:33] Allan Stewart: One of the big things in my mind is just really around how do you make this visible? How do you start measuring the health of the system? Once you start, once you figure out a way that you can quantify it, then you can start moving forward and set some goals, right? We see that in the health space, physical, human health, we get these same kinds of things. We're setting goals for being able to do certain activities or we're setting ourselves up with a particular diet because we know that we should eat the fruits and vegetables all the time and not only when we have scurvy.
[00:16:15] Dave Adsit: Right. Turns out I've never had scurvy. I guess I've eaten enough fruits and vegetables throughout my life to avoid it. Thank goodness. It doesn't sound pleasant. No.
[00:16:24] Allan Stewart: So the first, first way of making this visible that, that I have used that I think is interesting is creating your own tech radar. There, there are multiple ways that you can do a tech radar, but if you just take kind of the basic concept from ThoughtWorks, well, then you end up with these concentric circles and you have this idea of of here are things that we like, that we adopt, that we want to do. And there's other things that are on hold, right? They're the stuff that we're trying to get away from. And it's pretty easy in my experience for a group of developers to look over their system and they will very quickly label the things correctly. These are the things that we like and want to continue doing more of. These are the things that we dislike. Like. And from that, you can kind of start getting an immediate gauge of where are you? Are there a lot of things that are on hold that exist in your system still, but you haven't gotten rid of it yet? It also helps with the directionality. You can start to understand, oh, we're trying to get rid of this particular ORM, or we're trying to adopt this new pattern in our code, and having something like a tech radar, or there are some other things like architecture decision records and things like that that can help with that directionality. But the tech radar is also simple. You can throw one together on a whiteboard or a digital whiteboard like Miro or something like that pretty quickly.
[00:18:08] Dave Adsit: Yeah, I really like the tech I've set up several over different companies I've worked at, and they are very helpful in clarifying what is the core that we're doing? What are things that we're investigating? What are things that we are moving away from? That's super critical. Every long-lived system that I've ever coded in has had at least three different architectural styles that people put in place over time. Right, like, uh, at one point we were just doing connecting controllers directly to the database, and then we read Domain-Driven Design, and now we're putting in place a service layer and a rich domain model. And then we got a little bit more in in depth, and now we've decided to have workflows or, um, the, what are they, what's the other name for that, the, like, the, the script, the, a decision script or whatever that says what it is that needs to happen when this endpoint is called. And in this part of the code, we're trying to do like expose all the entities directly to the web, and then we'll compose them in the front end using GraphQL or whatever. But in this other part of the code, we're doing backends for frontends where every page has a specific endpoint that it gets and a specific endpoint that it posts. And which one of those is the current strategy and which one are we moving away from? And that's always a discussion that you have to have ongoing, right? And if you can have a way of recording it, that everybody on the team can go look, it allows everybody to go a little bit faster and be a little bit better aligned. And if somebody disagrees, they can go argue with the tech radar and have a discussion at whatever your high level architecture meetings are or whatever, rather than just going off and doing their own thing and introducing the third, fourth, or fifth way of doing code in the system.
[00:19:55] Allan Stewart: Yeah. And I think as you were talking, I was thinking about how you should probably want to continue to add more. Yeah. More ways of doing things because you are improving, you are understanding your system better. The world is changing around you. So there's probably, it probably makes sense that you want, at least want something that you're working towards and the old thing that you're moving away from. I mean, not for everything always, but that sense of continual evolution that we're getting better at the software. It's getting better, better fit for our, for our company, or business or whatever we're doing. And we're getting better as our craft of writing code. And so you want to have that change, but you just want to limit. Ideally, you want to limit how many of those are in play at any given time, because even if you know the directionality of it, knowing that there are seven old ones that you're trying to get off of to move to the new one. Well, that that tells you right off the bat something pretty significant about the health of your system.
[00:21:02] Dave Adsit: Exactly. Yeah, I would say it's not a very healthy system if there are so many architectural patterns that nobody knows what the right one is to use at any given point. Right. So having this tech radar, making those things visible, actually helps increase the health without even actually having changed the system yet. But there's a bunch of other things we can do to actually start measuring and fixing and improving our system. One of the really popular ones for the last few years are the DORA metrics, which come out of the DevOps Research Association. Is that right? And those are metrics around deployment frequency, frequency of failure rate, the mean time to recovery, and the lead time for changes from when code is committed until it is in production. I think we've probably all worked on systems where you commit code and you create a batch of code that's ready to go out. And then by the time you've said, "okay, this is dev done," whatever that means. It could be a week, it could be two weeks before it's actually deployed into production. I worked on one system where we intentionally chose to have annual releases. That's a long time between when a request for change is made and it's actually seen by the users. And so these DORA metrics are a way for us to understand system health. And they're actually rigorously researched and they correlate very strongly with team success, product success, and business success. And so I would say if it's taking you a year to make a change in your production environment, that's probably not a very healthy system. I guess I can imagine a couple of cases where that might not be true. But I think that in general, if it takes a long time to deploy a change, you are not operating in a very healthy system or you are low in responsiveness health, reactivity health, right?
[00:23:12] Allan Stewart: Yeah. Yeah. Yeah. And it's nice that those DORA metrics are all fairly easy to understand, right? Sometimes there's, there's questions about, "oh, well, where do we measure this?" Right. Where, you know, what is the cycle time versus the tack time versus, you know, how do you put these measurements into play? But there's also a lot of resources to help you work through that and, and be able to measure it. And, and one of the things that I've liked about it too, is that. But even if you don't go all the way to fully automated metrics, some of them are fairly easy to understand. Like you were just talking about deployment frequency. You can get a sense of that. Even if you don't have like an actual measurement, you can know it's like, "well, we usually deploy every day." And if you know that versus, "well, we always deploy monthly or at the end of every sprint, which is probably two weeks, maybe a week." Right? Like you immediately start having an understanding of, "oh, well, this is where we're at." Even if you haven't gone and actually tracked this, these are all of the deployment markers and what is the mean time between them. Another way that you can think about technical health that I just barely thought of, remembered, the idea of fitness functions from evolutionary architecture is an interesting one where you can go in. The idea of a fitness function is basically you're deciding something that you want to measure and a threshold for how good it is. And then the fitness function is ideally something automated, but definitely something that you can measure regularly to let you know where you're at. So measuring things like code coverage on your tests or cyclomatic complexity are examples of something that you might care about. Or there might be some architectural boundaries that you define, like certain pieces of code are allowed to talk to other pieces of code, but we care about the directionality and the boundary between those layers. And if those boundaries are circumvented or, you know, broken, then you can get an alert and, you know, "oh, this is something that has regressed. It is less good than it was before." And on some of them, you can also set it up in kind of a ratcheting fashion. So right now, maybe we have really low code coverage and we want higher code coverage. And so we find out, "oh, where are we at?" And if we go below our current, then that's a problem. But as long as it keeps going up, then we're okay. And we used to be at 20% and now we're at 30%, and we ratchet it up. So now if we drop back down to 20, it'll alert us rather than letting us kind of slip back down into mediocrity.
[00:26:18] Dave Adsit: Yeah, I really like to think about fitness functions as taking your abstract concepts, like your ilities, right? The scalability, reliability, recoverability, whatever, and putting concrete numbers around them, right? So instead of saying, "we want a highly available website," the fitness function will say something like, "we want a website that is available 99% of the time between, or 24 seven, 365." 65. And so now I can measure that for the system. And if I fall below my target threshold, now I know that my health is dropping. I need to put in work to improve the health of my reliability. Usually, if you say we're going to target 99, most people aren't going to be super happy with that because that's a lot of downtime per year, right? People are going to want to target things things like 99.5, 99.9. I mean, I don't think that most of us need to target 99 or five nines, but I would sure be unhappy if AWS decided that S3 buckets only needed two or three nines of reliability. Yeah. Right. And so the, the, I like to think about the fitness functions, taking an abstract, anything abstract, like, like that, like reliability, availability, scalability, and making it concrete and putting a test around it. So we're looking at it on a regular basis and we know how well we're doing against that target.
[00:27:44] Allan Stewart: And the connection to the health is right there in the name. It's a fitness function.
[00:27:49] Dave Adsit: That's right. I know these servers need to lift more weights because they are not handling requests well enough.
[00:27:58] Allan Stewart: I like that too, because it's fitness, like your health, but also fit for purpose, right? Are you actually healthy enough in the way that you need to be healthy for your purpose? Right.
[00:28:12] Dave Adsit: Yeah, it's a huge waste of money to make a system that's an order of magnitude more reliable than necessary. When we talk about increasing the number of nines of reliability, every additional nine is basically, I mean, I'm going to say 10x cost. I don't know exactly how much it's going to cost in every instance, but it is a substantial cost to increase the number of nines of reliability you have. And beyond a certain point, beyond a reasonable point for your user base, it's unnecessary. Necessary. I'm reminded that people like to use the term real-time. We need real-time reporting. And when we as engineers start thinking about real-time, we're like, "oh man, hard real-time. The system has to crash if it's not hard real-time. How am I going to build a hard real-time system?" And then you go back to the user and you ask, "what do you mean by real-time?" And they're like, "well, it has to be accurate within about a 24-hour period. It has to be up to date as of end of day yesterday." Like, "aha, I can build that system. I don't need to go into the hardware to make that system happen." Yep. Right. So understanding and having a good understanding of what people need across the system is really important. And so it would be a waste of time and money to build a system that is hard real time when what somebody really needs is reports that are accurate as of end of business yesterday. So one of the other concepts that we've talked about a lot and found a lot of value in is scorecards. We've used scorecards for security, for general system health, and a variety of things. And I actually really love the health scorecard you've put together for the company you're working at, Allan. I'd love to hear more about it.
[00:29:55] Allan Stewart: Yeah. So a while back, I discovered that having some kind of simple system helped me communicate better with the non-technical folks. And I also found that it's a lot of work to set up some of these metrics, right? So they're very useful. So like we talked about the DORA metrics, they're very useful to get you a sense of where you're at. And you can measure them, but it can be difficult. Fitness functions are a really cool idea that you have to implement, right? There's work to be done there. And depending on the nature of your company, that can be difficult. I've worked at a couple of smaller startups where it's really hard to justify. You know, when a lot of the system is currently on fire and there are problems that are affecting the day-to-day, it's harder to justify, "hey, we're going to go around setting up some automation for these fitness functions." When we know full well, it's just going to report it's really bad. And so not everything has to be quantitative. Qualitative can be enough.
[00:31:09] Dave Adsit: So measuring the mean time between failures is not a great use of time when you're currently experiencing outage, right? If the website is down, it's not the time to go in and like write the code to measure how often the website goes down.
[00:31:25] Allan Stewart: Exactly. So in the scorecard that I'm using at my current place of employment, I broke it up into some some categories that would be recognizable and hopefully somewhat understandable for people outside of the engineering department. So I've got a whole column of just features. And I broke it, I even broke that up into, these are the key features that are important for our business. These are the supporting features that help our business along. Some of those are just even like the checkbox kind of features that, "well, we have to have this because our a competitor has it, even though it's not our distinguishing feature." So I've got a column for those. I've got a column for different components. These are the deployable things. And this is a little bit less obvious for people outside of engineering, but we're still taking it at a pretty high level. So here's our server API. This is the state of our mobile app. This is the state of the web app. People understand those things pretty well. And so that's my second column. And then I've got a third column for non-functional requirements. So things like accessibility, data integrity, maintainability, security, scalability, testability. And in each of those columns, I've listed out the various things and broken it down. So like some of the components, right, for our mobile app, our deployment process, and the SDK that we're using, automated tests and code quality are examples of things that fit under the mobile app. And then we're just using a really simple kind of traffic light style system to share that with people. So yes, we could probably come up with a more quantitative way that would maybe be even closer to reality, but I found that that's not really necessary for it to be useful. Some things are kind of hard to measure, like latency. Um, and some measurements are hard to decide, like at what, how much test coverage do we need before we call this green versus yellow versus red. But what I found is that you can come to a consensus pretty quickly with a a group of developers, especially on these smaller projects where you don't have as much time and money to invest in automating the metrics, it's often small enough that you can get to consensus pretty quickly. And the traffic light is easy to understand. If it's green, that's good. If it's red, that's bad, right? Red means stop in traffic, right? So you're not getting where you want to go. There's a problem, you're you're stuck on the freeway and it's bumper to bumper traffic. Right, stop is is generally not good, uh, and then yellow is kind of the in between. It's a cautionary state where you're telling, it's like, "hey, this is probably affecting us negatively, but not so bad that we couldn't ignore it for a little while." Right? Yeah, when you have reds, then yellows probably
[00:34:40] Dave Adsit: get pushed to the background, you know, to the back of the to-do list. But if you manage to get to a point where you've addressed all of the things that are actually on fire, maybe know all the yellows are the most important thing to focus on to increase your overall system health. And I can honestly see a case for, "hey, this is a very infrequently used component. And the fact that that it's yellow and not super healthy is kind of okay because we don't really care that much about it. It's not used very much. It doesn't make very much revenue. It's not a priority for the business or for the users at this point. So let's just not address it." Let's instead work in an area where we can add more functionality and hopefully do it in a way that's healthy and And overall increases the amount of green we see on our scorecard by adding new, good, well-maintained functionality of whatever sort. I really like the fact that you've reduced your scorecard to a small number of statuses versus having a more granular, like 100 point scale. What percentage healthy is this thing? Well, I mean, we could spend all day arguing about, is it 56% or 57%? And it really doesn't matter or help at all to know that.
[00:36:00] Allan Stewart: And then we also argue at what threshold of percentage does it change color?
[00:36:06] Dave Adsit: Yeah, right. I would say maybe an exception is if you're, okay, so I have used a security scorecard in the past. And a lot of the parts of the security scorecard would be well served by having a red, yellow, green, right? Right. But I do remember the security person I worked with said, "Hey, I have to report this to the board every month or every quarter. And so this one little metric right here, it's percentile or it's percentage, percentage healthy, whatever that means. And so, you know, every quarter, I just tick that up one. So it looks like we're making a little bit of progress." None of the developers complain if it went from 56 to 57 to 58 to 59, because they all know it's It's like kind of okay, but not fantastic. But, you know, he's like, "but I can show progress on it." And that makes certain people feel good about the investment they're making in security. Yeah. Security is one of those things that is a huge space that you also have a lot of potential for tracking health in different ways, right? What's our password security? Well, if our passwords are stored in plain text, it's probably red or double red. Do you have that as a concept? Black hole? You might need to add black for it. It will kill us. But if we're using like, I don't know, an older password algorithm, that might be orange or yellow, right? Right. Like, "it's fine. It's not hardened against GPUs, but we don't really have time to invest in implementing Bcrypt. So we're just going to leave it because nobody's going to get our database and try to crack it with GPUs." They certainly already have rainbow tables for the low, low value target that we are. But I could see, you know, if you're working your way towards higher and higher levels of security, having a security scorecard that talks about all those, the various things like, what are we doing around security training? What are we doing around penetration testing? What are we doing around code security and code analysis? All of these types of things become important to your overall security profile and security health of your system.
[00:38:22] Allan Stewart: Yeah. And like you're saying, it's easy to add other states, right? So, you know, do you need a double red or a black or an orange? Doesn't have to be red, yellow, green.
[00:38:34] Dave Adsit: Like, we have so many reds, we have to create a distinction between red, red and orange red. It's not as bad as the red reds, right?
[00:38:44] Allan Stewart: Right. Right. And, and it goes back to this concept of understanding your audience and, and what makes it useful and what makes it easy. Right. So you could just do a numeric scale, you know, maybe it's just like a simple, like one to five, kind of like the DEF CON style. You just have to decide at what level are you willing to deal with it? Because you don't want to get into that place where it's like, "well, it's a hundred point scale. And now we're arguing about about percentage points or like decimal." Yeah, you don't want an unnecessary precision. This is just like, especially in this this state where you're not automating the measurements, but you're just getting kind of the feel of it. It's nice to be able to just quickly go through and say, "yeah, this is good, this is bad," but in a way that makes that everybody could just understand, right, especially especially the people who aren't engineers, they don't write code. If they can look at the score sheet and say, "oh, there's a lot of red and yellow in there and that makes me uncomfortable." That's where we want to get, right? Or if they look at it and they say, "oh, hey, look at the progress that we've made over the course of a year. We had a bunch of things that were red and yellow before, and now they're green." They feel good. And hopefully there are are consequences that go along with that. As you make the technical health of the system better, if you're doing it right, it should do things like improve your feature velocity. It should do things like improve the scalability and the reliability of your system. And so people will see that. And if they can draw that correlation from "it was more red and yellow before, and now Now it's more green." And also I feel like my customers are happier than they're going to trust in that system more and you can continue working on it.
[00:40:35] Dave Adsit: Well, and one of the things I think about there is the reason why debt was such an appealing metaphor is because business people understand debt, right? You take on debt. Okay. That's bad. Now our revenue is reduced by our debt, right? Like, we didn't make profit because we still have all this debt. Okay. So that's well understood. And debt is a way to basically manipulate the emotions of the rest of the business because they don't understand what we're doing. It's very complicated and very technical and you require specialized training, just like a lot of us don't understand accounting. So if we're using debt to incite fear, that may not be the best way to collaborate with other members of our company, other people within our organization. And so, but I do think that there is a benefit when we start talking about scorecards and people can actually see it and then they can come to their own conclusions around, "hey, it says right here that card processing is red. Why is card processing red? What's going on there? Can we have a discussion about this?" And then maybe I want to invest in making that green or at least yellow, because as opposed to just saying, "hey, we have to do this big thing and you don't get to have a decision on it. And like, we're mandating that you let us take time to fix X, Y, and Z." And then at the end, we might have actually broken X, Y, and Z a little bit. We fixed one thing and broke two others, you know, kind of what happens when you try to do these big refactor rewrite pushes, we forget functionality that people actually needed. But by making things visible with a scorecard, we can actually bring the interested non-engineering, non-developer parties into the conversation so that they can help us make priority decisions. They might say, "hey, look, I know this one app, you've flagged it as red. I don't care. No one uses it. Not going going to invest in it. Or people use it. We've intentionally decided to strategically move away from it. And so leave it red. We're going to replace it with something else, or we're going to deprecate that part of our business or whatever. This part over here that's yellow, super critical to our success. We need to make it green. How do we change this yellow to a green? What do we need to do? What do we need to buy? How much time do you need?" I feel like you can bring people into the conversation better by making things visible than just by being the expert in the cloud or on the top of the tower yelling down orders at the people below.
[00:43:02] Allan Stewart: Yeah. And it comes to mind that you can also improve your scale beyond green, right? So I don't know what that is. It's double green or it's blue or light green and dark green. I don't know, but you can come up with something. And that might also be a way to strategically look at your system and say, hey, we want this part of the system to be really good, right? Not just passable good, but like great, maybe even world-class or like, because this is our differentiating feature and we want to be able to iterate on it so quickly. We want it to work flawlessly. We want it, whatever it is that makes a difference for your business. Well, that could also be a strategic choice because, you know, getting into that, that concept, I think most, we talked earlier about like, are you just generally healthy or are you healthy for a purpose? And so that would be one way to start looking at that. And so they could say, "okay, things that are yellow and red, well, we're going to invest in those because we don't want to be in poor health." But then beyond that, we also want to invest in certain strategic areas that we're going to say, "hey, it's not good enough to be okay or just healthy. We want to be strong in a particular area that fits the need of the business."
[00:44:32] Dave Adsit: Yeah, so I guess all things said and done, I really like the idea of talking about health instead of debt. And maybe it's in addition to debt, but I do think that you can have a more positive conversation, you can make more progress. You can basically turn around the perception that engineers are constantly pessimistic and start being more optimistic and working with your business partners to deliver on the needs of customers, users, etc. cetera. And I like the idea of using health a lot. I think we've both seen success by using this at different companies we've been at. And I would just encourage people to give it a try and see if it doesn't change the way conversations go when you're talking to people across your company.
Copyright © 2026 - Crafting Code Podcast