My Agents Are Their Own Worst Enemies.

• 1,250 views
vlogvloggervloggingmercedesmercedes AMGMercedes AMG GTAMG GTbig techsoftware engineeringsoftware engineercar vlogvlogssoftware developmentsoftware engineersmicrosoftprogrammingtips for developerscareer in techfaangwork vlogdevleaderdev leadernick cosentinoengineering managerleadershipmsftsoftware developercode commutecodecommutecommuteredditreddit storiesreddit storyask redditaskredditaskreddit storiesredditorlinkedin

In this video, I talk about how my AI agents, if left to their own, continue to build too much CI/CD bloat.

📄 Auto-Generated Transcript ▾

Transcript is auto-generated and may contain errors.

Hey folks, we're going to talk about AI development. Um, one of the things I just wanted to chat through today and I don't have like a particular I don't know like a order of this but just like something that's been on my mind is like um I guess continuous integration testing that kind of stuff when it comes to to AI software development. So when I talk about this stuff normally, you know, building with AI, um I talk about approaching things from from two angles. Like even on the the recent live stream I did this past Monday, I was trying to uh describe this where I think about trying to ensure that agents have proper instructions up front so they do things right the first time. The challenging part is that they're non-deterministic, right? So even if you give them instructions, it's not a guarantee they'll follow them.

And uh of course over time when you're having agents run the more and more things that you're you're trying to expect them to do uh the more opportunity there is for that to get missed. And I kind of argue that it's I I mean the way models are right now uh it's kind of like people right if you gave someone one task to do and said go do this very specifically could do it right you give it give them guidance around how you expect it to be done they could probably do it pretty you know uh pretty straightforward not a lot of room for error and uh if you were to give them a bunch of things to do and a bunch of uh expectations around how those things are done more more room for error. But anyway, uh still important to focus on proper instructions and guidance up front.

Uh but then the other side of it is uh sort of like gating or enforcing uh coming from the other direction. So even if an agent does not follow what you say um at least you can stop it from proceeding. And so even before AI, you know, the way that we do this is things like uh unit tests, integration tests. We have things that run in our continuous integration pipelines, right? So classic example here is you have a uh code change, you know, you push it up uh to your to your CI system. So let's say it's GitHub, right? Open a pull request. Um, maybe that starts in draft mode and does some initial checks and then when it's ready to be reviewed, you arm that for review so other people can check it out. At that point in time, perhaps it's running your heavier integration suite, right?

It's no longer in draft. It's worth running um something a little bit heavier. Uh, and at that point, you're checking all of these things that are important, right? like you want to make sure that whatever code is getting integrated is not breaking other behavior meets expectations so on so forth. Now, like I was saying, this is something that we've done for a long time, well before, you know, using AI for software development. And um I think I think that it's you know increasingly important to have things like this in place so that when you have agents contributing to your codebase you have mechanisms to stop them from doing things that you just would otherwise you know be like yelling at someone hey don't you know we can't possibly have code like this it's going to break stuff. Um, so you you have the automation to stop agents, right?

So my sort of recent like it's not it's not even recent I would say like it's still been like uh like months where I've been struggling with this. it's just becoming like more and more apparent to me is like um when I'm letting AI build out, you know, complete repositories, one of the things that seems to inevitably happen is that the the testing, right? the what we need to be checking at integration time continues to bloat so much that um the agility drops dramatically. And so what do I what do I mean by that? Right? Like if you're starting a new project, right? Or even if you're using a you're working with agents in an existing codebase, let's say that in terms of the the testing automation, it's pretty green, right? Pretty fresh. if you wanted to go run um a build, right?

Depending on how much code there is, that could go from being very fast, if it's something new and early, or maybe it's still a little bit longer because it's a more mature project. Maybe you have uh on the order of like, you know, minutes like 5 10 plus minutes just building because the whatever you're trying to compile is is large, right? Package is large. But um say the tests, right? When you don't have many tests, it's going to be nearly instant, right? And if you're doing like what I would argue are proper unit tests, these should be isolated. They should be extremely fast. So even if you had hundreds of tests, they could still be done within like seconds, which is great, right? like you in CI systems, you're probably paying more time to have like some of the test infra set up like getting the right packages and stuff in place even if you're pulling from cash potentially, right?

Like just these steps to get set up to run the tests when you only have a few unit tests. It's maybe that's even more time than the test. But what ends up happening is that like we end up trusting agents more and more, right? um even if you're in the loop as a human, you know, they're they're coding things, they're adding tests alongside them or going nodding along. Great. Like, yes, we want the test coverage. And um what in my experience so far, what continues to happen is that they keep landing more and more. And if you're um being a little bit more handsoff, they might even be um adjusting some of the the CI system to say like, "Hey, look, we have these other types of tests we want to go add." Now, like, you know, we're we're trying to do the right thing, expanding the coverage to make sure we don't regress on these paths.

So, like conceptually, I think it's it's trying to do the right things. is trying to do behavior in software development that I would argue are good things, right? We're we're expanding the scope of what we're doing and like we need better systems in place to catch this stuff. So, like I I do think that's great, but what ends up happening inevitably is that that system continues to grow and grow and grow and you go from having CI that takes maybe a couple of minutes to 10 minutes to now 20 minutes to, you know, 30 to 45 minutes with maybe a bit of flake, right? you get one of those tests that sneaks in that like out of your thousand tests, you got one that's flaky and um it's enough that when it happens, you have to go rerun the whole test suite. And so, you know, that 30 to 45 minutes, you might have now uh situations where agents are rerunning it.

So now it's 60 minutes to an hour and a half of runtime and it keeps going layer on top of this you have multiple agents working in the same repository right what's going to start happening depending on what your your conventions are if you're waiting for you know you must be at the tip of Maine to be able to integrate well you know two things are racing each other each spending 30 to 45 minutes on CI. One lands first, the other one, what happens? Sorry, got to restart. Right? Like, so there's a lot of different patterns that we that we may have. By the way, I'm not saying that these are things you have to do or you must do or they're the right things.

I'm just saying these are types of patterns that exists in our in our CI systems where in a in a previous world we might you know live with this kind of stuff even as a solo developer right I might argue I always want to make sure that what I'm integrating you know I've been running at the tip of Maine because if I'm behind and I go to run the tests maybe that's not safe right I just want to prove that I'm at the tip but I'm the only develop veloper. So like I rarely ever have to worry about like racing, but now if I have agents, I'm no longer just the only developer and I might have two or more agents running, right? So that the game is changing a bit. My point is that I'm noticing that more and more scenarios where I'm using agents on code bases, what ends up happening is that the um the ability to deliver with like high throughput drops off a cliff.

And so I need to get better at a couple of things. One is having um good CI patterns up front, right? So to give you an example, I got to pass this guy. One sec. Um when I was talking about like you know racing and having PRs and stuff uh you know one integrates you have to wait for the next the you know there there's different solutions for this right like you can have um like GitHub has like natively stacked PRs now which is cool. So, you could say, "Hey, look, like I'm I'm an agent and I'm going to be um delivering some work." Like, what I could do to make the the code more manageable for review is like put up a PR in increments, right? Like for this slice of it, for the subsequent work, like really kind trying to group it logically, which is a nice thing to do.

And then I'm going to stack these PRs so that if they look good, we can basically just press the button and we'll merge the whole stack together, right? Like that could be pretty awesome. Um, but like there's there's drawbacks to that too, right? Like if you have a stack of PRs and the way that your CI is set up, you know, the one at second from the bottom out of five, right? You need a code change. Okay? So, you have the agent update the PR. Now, you have to go run CI, but like how much CI do you have to go run for your stack of PRs, right? Um, you can still merge them all together, but you need a green light on all of them. So, it depends on what your your policies are for integrating. And so I need to find uh better patterns and sort of codify them back into this common repo that I use called Genesis, which is where I have a lot of my templates.

So I've been doing this kind of thing, but it's still not enough, I find. Um so for example, uh I I like arming things to automerge, right? If I'm working with agents and I'm like, "Hey, if we have the right CI in place, then like, yeah, hell yeah. Automerge it when you get the green light, right? I don't want you to sit there waiting for me. Merge it." And um like that doesn't actually work with private repositories in GitHub. So like in Genesis, I have a template for essentially like uh you know rolling my own automerge kind of thing. So, this guy flashing his lights at me here. Not going to have a good time with that. Um, I have to do like uh like one pattern that I see a lot, right, is that it's it glaringly obvious when it happens, but like an agent will change some documentation, right?

changes a some dock somewhere and then I'm like, well, why the hell didn't this PR like land right away? And I go check it out. Um, and that PR is running like 45 minutes of continuous integration. Um, it's like for what? You changed a document and it's like, well, yeah, I mean, but it's CI. We had to go run all the tests and build it. It's like, no, you you really didn't have to do that, right? Maybe you have some things that are like that. But you know, if I'm changing a do like an ADR file, some document for a architecture discussion, I don't need to run CI at least in my repos, I should not have to. So doing like test planners, right? Like the like I was saying, as these test suites grow and grow and grow, it doesn't make sense that if I change one spot or a couple of spots that I go run every single test.

It's very very wasteful. So, how do I have, you know, reusable patterns for um for test planners and I I feel like I'm only scratching the surface with this kind of stuff. I need to have more of them because I continue to run into these things. Um it's a it's a bit of a balancing act, too, because there's some types of things that I want to add in, right? like I want to add in uh CI automation that shows uh the test coverage or I want to you know introduce things like mutation testing, right? I I like having things like this because I think that they're helpful. I don't love gating on code coverage, but I like seeing, you know, if I'm changing 10 files, cool. like, can you give me an idea what kind of test coverage I have on those 10 files? I just want to know cuz it gives me some level of confidence if I'm checking out a PR.

Um, right. And it can even be feedback for the agent like, "Hey, look, like this thing ran and like you didn't have high code coverage." Like you can make uh like allow an agent to to go make a decision about that. Yeah. it was only 40% because I don't know like there was a bunch of data transfer objects and I didn't test getters and setters on you know automate um automatic properties that kind of stuff. Um the the point is that at least the information is there and like decisions could be made but when you add more and more stuff like that there's more and more runtime right even if it's small one day like it it can grow. Um, so this is like one set of things that I really need to continue to explore.

And then the other is like uh really more on the AI side, which is like okay, if I'm letting AI do a lot more development and I'm able to be hands off more and more with like exactly how it needs to get done, I need to give it mechanisms to say, hey, look, like you've been working through this work stream and like you know, if you were to reflect on how much time you're spending, 95% of it is in CI just waiting for stuff or you know waiting for stuff and then it's failing and then you're redoing it again. So you go from a 30 minute CI integration with a fail to now it's an hour because you've run it twice to now it's 4 hours because you've had to keep working on failures. The problem is it will just keep like fixing the thing and then submitting it again and waiting.

And I'm I need to like give it better reflection skills to be able to say like, "Hey, look, like that's not okay." You know, maybe maybe one fail or something because it made a mistake. Sure. But if it's like my only way to go check this stuff is to keep submitting it to CI and it keeps breaking and I've spent 8 hours on it. Like absolutely not right. That's just like not an acceptable way to develop. Um like if a human was doing that I would be like dude what do we have to do differently here? Um, so I I need to do a better job um with having my agents armed with ways to to say, "Hey, look, like CI is getting in the way here." And then because I'm getting close to CrossFit, I'm not going to get a chance to talk about

this more, but like one thing I'm I'm curious about, I like would love to entertain is like like this makes me uncomfortable, but this idea that right now like depending on what the product is, if AI makes a mistake, right, like you integrate something that's not perfect, I bet you in many cases right now for for a lot of the stuff I'm building, It's faster to just like fix the issue, right? To catch it after and fix the issue than to like sit there through more and more CI like hours of CI versus like integrate it and go, "Oops, didn't work. Let's fix it." Um, but that world kind of scares me. So curious. I'm I'm wondering if uh we'll see more of like we're a bit more relaxed and we're finding things and we'll fix them fast after. But I don't know. Love to hear your thoughts on this, too.

So, thanks for watching. That's my thoughts right now on some of the AI stuff I'm struggling with. So, see you in the next video. Take care. This

Frequently Asked Questions

These Q&A summaries are AI-generated from the video transcript and may not reflect my exact wording. Watch the video for the full context.

What are the two angles you consider for AI software development to ensure agents behave correctly?
I think about two angles: giving agents proper instructions up front so they do things right the first time, and gating or enforcing the process from the CI side so they can't proceed if they don't meet expectations. Agents are non-deterministic, so despite guidance, they may not follow it. I rely on unit tests and integration tests in CI to catch issues.
Why does CI throughput degrade as AI agents contribute more code?
As agents contribute more, the testing grows and bloat makes agility drop; builds that were fast can become 30-45 minutes or longer, and flaky tests can cause reruns, extending runtime. The runtime can escalate to an hour or more, especially with multiple agents racing each other.
What strategies do you discuss for managing CI and PRs with AI agents (e.g., stacked PRs, automerge, test planners)?
I talk about stacking PRs so I can deliver work in increments and merge the stack when green. There are caveats for private repos, since automerge doesn't always work there. I also mention test planners to avoid running the full suite for small changes like documentation, and I reference Genesis templates to codify these patterns.