In this video, I talk about how I tried using some free LLM providers instead of frontier models. Buckle up.
📄 Auto-Generated Transcript ▾
Transcript is auto-generated and may contain errors.
Hey folks, I'm going to talk about some AI stuff today. Says an hour to work, which is kind of stupid. Not looking forward to that, but we'll see. Sometimes the traffic is uh actually not bad at all. And it's just Google being kind of silly. But morning started off pretty terrible. Um got up ready to go to CrossFit, which already sucks cuz I don't like waking up early, but that's life now, I guess. And uh I was sitting in this car and it was saying that I didn't have coolant and that's kind of what I was dealing with in our SUV right before we went on vacation. So that was pretty dumb. Um so that kind of ruined the morning. Um I didn't have any spare coolant cuz I had used it in our SUV before.
Um, but yeah, I ended up having to miss the gym and stuff this morning and uh by the time I got sorted out, I realized like it was just that the coolant was like right below the the minimum line. So, but a couple couple drops of distilled water in and I have coolant that's coming later today, so we'll get back on track. But kind of a lousy start, but I've been bummed because a couple days ago I finally hit a big problem and uh that big problem is that I ran out of co-pilot credits. Um they've changed a lot of the limits since last month uh in terms of what we are able to get at Microsoft. And so it is 20% of what it was last month.
So, I thought I was doing pretty good this month being a little bit more conservative and uh then my co-pilot sessions were saying that I was uh that I exhausted limits and I thought that it was like maybe like an API issue because that kind of thing happens. Um and then went looking and sure enough, no, like the limit was just uh reduced uh to 20% of what it was. So, um, obviously panic mode for me because I'm going, "Oh no, how do I how do I build software anymore? I can't remember how to program." Um, no, it's more it's more just that I am trying to do a bunch of things and it means that if I have to start coding by hand or looking for alternatives and I I simply can't keep up with uh a lot of the things I was trying to do, which sucks, right?
like I was uh really enjoying being able to build a lot of things in parallel. So I said, "Okay, well couple couple options, right? One, do I do I bite the bullet and invest into some hardware to run AI models?" And I don't feel qualified to make such a purchase right now. Um but I think that was actually a bit of a wakeup call for me. that's like I probably should um research this kind of thing more. I don't want to jump into a purchase because it's not um it's not cheap, that's for sure. But I also think that if I have to start paying uh for the token usage that I'm, you know, finding productive, then paying for that every month is not going to be sustainable. Like it would make a lot of sense to invest in some hardware. Um, but when I started researching, it's kind of like there's there's numbers all over the place.
And so, like if if you haven't looked recently, like memory prices are absolutely insane. Um, you know, like a lot of things in computing and hardware and stuff like that. It's like costs always come down, right? like we find ways to uh make things faster, get more capacity, and so year over year over year things just come down in price, right? You always have some cutting edge stuff that's pushing some limits and you pay a premium for that. Um and then otherwise things come down, but like memory prices are up such an absurd amount and that's because AI, right? So, um I'm not shocked. Uh I'm not shocked that it's up. I'm shocked like just how much it's up. It's like it's bonkers. So, um that's like like one of the huge driving factors for um you know the cost of putting a a build together that's going to support running models.
But like yeah, the numbers are all over the place in terms of benchmarks. So, like when I'm researching things, it's like you might have someone who's able to run a particular model and they're getting like a particular number of tokens per second, but um there's a caveat and like, you know, if you're not willing to do like that specific thing, then um the whole sort of build falls apart. Like it's not it's going to be dramatically different. Um, and it's like a lot of customization uh around like setting up models and things like that, not just the hardware. So, it's kind of like the way I'm looking at this is like are you willing to pay a lot of money? Like it's thousands and thousands of dollars. So, between, you know, in the neighborhood of like $10 to $20,000. Um, which I think is pretty significant considering most I think the most expensive computer I've ever built in my life was like, you know, $6,000 10 years ago and it was awesome.
It was like the coolest thing that I've I've like ever done with a computer. But, um, you know, that like for gaming and that kind of thing and I was overdue to actually have a an up-to-date computer. So, it's a lot of money to go into a computer and um it's kind of like are you you willing to pay that much plus you know um this space is moving so fast and everyone's like trying to you know build uh more tuned hardware for this sort of thing. So, it's normal in in computer hardware that like there's always going to be new stuff coming out. And so, like, do you go put a build like this together and then like within a year it's already like, oh, that was stupid for you to purchase that because it's already so obsolete that like doesn't make sense. Um, so I'm kind of like struggling with that.
Plus, you know, all the examples are uh, you know, how much are you willing to customize the setup for what you're running on it um to squeeze out the performance benchmarks that people are claiming. So, uh, I'm I'm kind of like pumping the brakes on on a purchase, but at the same time, I think that was a bit of a a wakeup call for me that I'm like, I really need to to understand uh the mechanisms better. So that it's not just, oh, I I use models, I build software with them. It's it's a lot more like I understand um what's needed to run them effectively because that's a it's a huge gap for me right now. And it feels like, like I said, a bit of a wakeup call. So I said, "Okay, well, if I got to pump the brakes on that, um what are what are my options for like something that's free?
Right. Let's go the let's go the exact opposite direction because what I don't want to do is based on my usage. I'm like I'm not I simply don't think it's going to be uh feasible for me to pay my my co-pilot token usage the way I want out of pocket. Like let me let me explore more. So, I decided I was going to do a bit of research and see what's available for like um free model usage. Very weird not going to the speed limit. Um and so I came across free LLM API. Uh so basically open uh open AI compatible sort of a API interface that you can use and then you essentially provide it with a lot of different keys for different providers and then on the fly it will try doing routing to the the model for you.
And so in my head I'm going I know that I know that I'm not getting like Frontier model quality because I'm looking at the list of models and any Frontier models that have some type of quota there are like I'm going to be able to use it for 5 minutes before I'm at a quota and like okay that's I understand that but like uh what's going on for for these other models, right? Like I'm very used to using Frontier models and I have not been coding with these other ones. So what's up? Let's see what's going on. Um so I set this thing up. Um wasn't too bad to set up so I can install it with Docker. Have it running pretty quick. Um and then it was a matter of like signing up for a bunch of different accounts that offer um it's like a variety of things, right?
Like some will have like a monthly quota for different uh models. Some will have um more of like a like a window or a rate limit like a token bucket sort of thing that will refresh or a number of requests that you can run. And so free LLM API does all of this for you um to manage it. I guess like you don't really have to worry but you can also tune it a little bit. Um so you can you know pick a provider and then say like don't use the models from this provider. Um, and it does some things like it will adjust if you're being rate limited, that kind of thing. So, um, you know, it's it's stuff that like I certainly don't want to go build that by hand. So, I'm glad someone else has a project like this. I think it's pretty cool concept.
Here's the the traffic. This is uh me sitting in the fast lane. I think it's a maximum price to sit here and the other lanes are moving faster than me because we're basically parked. Excellent. Um, so dumb. And so I get this thing set up and configuring co-pilot to use it was a little bit more work. Still not bad at all. Um, but usually with this kind of thing, it's like I can sit down with co-pilot and say like, "Hey, I have something like basically set yourself up." Like I've I've become too used to doing that for the last 6 months because it's worked so well. Like, hey, new MCP server. Like, I'm not even going to read the docs like co-pilot. Here's the URL. Like, set it up and if you need me to restart the session, I'll do it. This was different, right? because like I don't my co-pilot credits are uh are exhausted.
So I'm basically sitting there with chat GPT being like what do I got to do to to configure this because I've never set up co-pilot to to point to a different provider. it's always just been, you know, use what's built in. And so yeah, really not so bad, you know, um URL endpoint, uh token, and then I think where it was a little funky and not really much about the setup to be honest, just about getting started is that if I wanted to resume a session, um because I was using Frontier models with like a million token context window, it was struggling to like to basically switch over like I you couldn't even use the /model command which is kind of uh odd like it just wouldn't even show up. So I basically had to force the session uh to resume with a model selected and the air quotes model that I'm selecting is basically just like an auto routing model.
So um it's a I don't know how you I don't know what the right term is. It's not a real model. It's like a a logical model if you will. Um, and then so that will go to the free LLM API that I'm running at a out of a dock or container on another machine on my network. And then we're off to the races, right? So, first thing I, you know, I I say something like, you know, what's the status of the the PR that was last being worked on? And, um, and so it submits it. Great. Obviously, like seeing something go through is good start. And then I see it kind of thinking, running some uh you know some GitHub API calls and then it spits out a response to me and the response is in Chinese. Okay. Um I suppose I probably configured some providers that are out of China.
So now I'm automatically like okay, I got to update some instructions to like say always respond in English. I, you know, wasn't anticipating that I'd ever have to really do that, which is kind of funny, but here we are. So, do that. Restart. Um, try it again. So, now I'm getting some English, but I'm getting like I don't like it almost feels like a a combination of things. Responses that aren't totally thought through. Like it's like I can tell that it's attempting like it's like okay I can see these pull requests and here's the status of things but like almost like no effort put in um the the most cursory check and then I also noticed like some of the output was I want to say truncated um as if it ran out of tokens or something on the output but the the output token context was set pretty high So, I don't think that's the case.
It's almost like it just couldn't format its output properly and just kind of like stopped cuz it went on to continue a sentence and had like here's like the list of things colon and then there was nothing. It was basically like it was truncated too perfectly. It just kind of stopped thinking. It just stopped the turn. Um, and so this is a bit of a foreshadowing because um, it wasn't just when this started up. So kind of give it a nudge and I'm like, "Hey, well, what like finish your thought?" And so then it goes off and kind of does the same thing again and comes back and has a bit more of a complete thought. And basically in for context here, the work that it was talking about is essentially already done. And I was kind of just waiting for, you know, green light on a PR that was already submitted, waiting for CI to run and like it had.
I'm just using it as a sort of like a sanity check to be like, is this thing is this thing on, right? And so finally comes back and I'm like, okay, well, we're going to integrate this thing. So I I merge it in and I and then I ask the next question. So, like based on the work that we had scheduled from the conversation, like what's the next PR? And so, I get a bit of like bit more of a a good feeling cuz it goes and looks and it's like, okay, well, this is the next issue. Um, and it even calls out like it's blocking this other one cuz it sees the relations on on GitHub. Okay, cool. Um, and it's actually also like a PR that's already open. So, good good news, right? Uh, it's just that it has merge conflicts. And so, I actually posted this on Twitter.
Um, so I posted this on my main dev leader Twitter account because I have social media accounts for code commute across every platform and dev leader as well. Uh, but I was posting this from my dev leader Twitter giving updates because this was the most outrageous journey for AI software development that I've ever had. And I could, you know, uh, spoiler alert, I could absolutely understand if people were trying to build software like this with AI, they would say this is dead. Like there there is no way that AI would ever build software. I could absolutely understand uh why that's the case. So, I'm not going to belabor the entire story because it's it's actually just like frustrating for me at this point. Like, it it actually makes it makes me very upset how um how frustrating this experience was. It it became a bit of a joke for me because what I was asking it to do is something I could have solved in minutes and like that's not really the point.
like I was my point is that I'm trying to get this thing set up so that I can do agentic software development. Um, and just sort of realizing like the quality of these models is horrific and simply would not sustain even the most basic tasks. So, um, just to give you an idea of some of the the it was doing, um, it at one point to solve the conflicts, um, pushed up a change to the PR that was empty. It decided that was going to be the move, right? Like if we basically remove all the changes, um, that will solve the conflicts. So, it did that, which caused the PR to close. Um, and then so it couldn't figure out how to reopen it. So it was like, okay, I will, um, you know, I'll make a new PR and like we'll we'll kind of do this properly.
So, okay, great. Opens up a new PR and basically it does the same thing like it basically made a new PR that still had um the merge conflicts. So I'm like, well, no, it's not that's not what we're doing. Um, and so there was a lot of back and forth basically, um, where ev every single time it seemed to like have a bit more understanding what it needed to do and then would com like deliver something completely wrong. For example, it would say in the chat like, "Oh, I should not downgrade this package. We need to use the latest one." And then it would push up a change and I would see that it downgraded a package. And I'm like, well, what's going on? Like, how do I even tell this thing to correct itself? Because what it's saying is the right thing and what it's doing is very wrong.
Um, all in all, I think there were four PRs made. Um, in the end, it it simply could not do it. And again, the context here is this is a PR that was done. There was a change that merged in before that caused merge conflicts in a couple of files. And all that had to happen, just to make it super simple, is that it had to pick uh if you're familiar with merge conflicts, uh had to pick ours and theirs. It couldn't pick one or the other. It needed both. And as long as it listed both um pieces of configuration, it's just like configuration sections were added and two changes added sections. They don't actually conflict. They just happen to be added at the same spot. So if you pick both and keep both, everything just works. Um and it needed to do this across like five or so files.
And they're not like it's not like it had to go figure out some crazy logic um in an algorithm was like colliding with other changes or something like no just a few lines of configuration but it simply could not figure out what to do. Um and again like I'm I I'm talking about this from a perspective of like I'm very used to using frontier models. I I don't have to spend time like this. If I did, like I might as well do it myself. If I'm telling the model, like the agent using this model, the basically exactly what to do step by step, it's literally faster for me to just do it. I'm not I'm not gaining anything from doing it. Okay. So, in the end, from trying to go through this thought experiment, I'm not exaggerating when I say this. I'm trying to give you the the framing of like what this PR looked like, right?
It used 110 million tokens. 110 million. Um, this was something that was running in the background the entire day, like 8 hours. And I'm not saying consistently because I'd poke my head back over and look and then be like, "Oh god, this is really dumb and bad. Um, and interact with it again." But like after 8 hours and 110 million tokens, no result. Um, free LLM API, I think, had like a an estimate of like what do they it's called like dollars saved estimate or something. It's like one of their analytics entries and it said like estimated that it saved me $600. So I'm assuming that means that it was the equivalent of $600 in token usage across these providers. So to try and fix merge conflicts in a PR that it could not do spent a full workday was 110 million credits and $600 sorry 110 million tokens and $600ish worth of tokens from these providers.
And I said it's just not feasible. It's it's completely impractical. I would never in my life get anything done. Um, if I could have lost more hair from being frustrated, I would have, but I'm already bald. So, um, that was that was the most disappointing AI software development I've ever seen. Um, the model that was being used was Neatron Ultra, and I guess I'm I from what I can tell that's an Nvidia model. That was the absolute biggest pile of dog I have ever experienced for AI coding. It was embarrassingly terrible. I had like I I have no words for how frustrating that experience was. So, um, where does that leave me? Right. So, I am out of co-pilot credits for the month. I'm assuming the writing on the wall is going to be that every month the quota is reduced more and more. So, it was reduced uh to 20% of what it was last month.
I'm assuming it's going to get cut down even more next month, which means that if this is already bad where I I have, you know, halfway through the month and I was on vacation for the first bit. So, I'm assuming like I don't know, I have roughly 10 days of runway um with my co-pilot credits in a month. What do I do for the rest? Right. Um so, my options that I see are um you know, be willing to pay more on my out of pocket for my co-pilot subscription, but based on the token usage, I I simply don't know. Uh I think it would be different if I you know I had some some formula that said on average for the number of tokens I'm um I'm spending I have you know clear income from that right so when I think about something like
brand ghost for example when we are helping uh with content generation and stuff like that I actually could come up with a number like that but I'm not using my co-pilot credits for that we we we literally really have you know uh hosted models that we pay for and then pass on the cost right for users that's built into the subscription and everything. So that works but for my own personal development like I don't I don't have such a thing right there there is no if I spend X dollars on tokens that means I'm selling product or service for other amount so I don't have that so brand ghost development can continue um brand ghost uh like service offering can continue because it's all outside of my my personal co-pilot, but my other projects that I've been talking about building, like that's what do I do, right?
And I I was just getting excited making videos for you guys saying like, "Hey, look at the the things I'm putting together." Um, so like I said, options are pay more out of pocket for co-pilot, and it economically seems impractical. Um, there's Claude, which uh, okay, I'll spoil it. I ended up going back with a Claude uh Mac subscription, the 200 bucks a month. Um, so I'm going to try that out. I haven't looked back into Claude um subscriptions and pricing and everything in a while because I think if you if you've been around my channel for a little bit, I've been saying like, hey, I have no problem with Claude, right? It's just that I have co-pilot credits. So like I don't need to use Claude. like basically why would I if I have credits um to use and so yeah I was checking out the pricing and everything and I'm still like super confused by it because it's like there's I don't know how it works.
There's like 5 hour windows where you'll you can get limited within there. Then there's a weekly limit, but there's no nothing published on like how many tokens or requests you can make or like I have literally no idea if I'm paying $200 a month and then about to get like slammed within 5 minutes being like, "Oh, you can't do that." So, I'm kind of doing it with some good faith that like I picked the highest tier Claude subscription I can. I know people love using Claude. I I do think the product is good. I'm not, you know, it's never been the product's not good. Um, so I'm kind of like, okay, if I pick the top thing and I know people are, you know, kicking ass with uh with it, then I'm hoping that's sufficient. I don't think what I'm doing is like some, you know, ungodly type of like software development.
I think there's people doing crazier than me. So, this should be okay. So, that's what I've gone with. Um, and then on the side, I think I need to continue to research more about um AI hardware and having some uh something dedicated. Um, so as I'm talking about this, right, like I am inviting people if you're, you know, listening to this, if you have experience with setting some stuff up, would love to hear, you know, if you've set up something, you have local models running, what that looks like. Uh, I'm saying this out loud so I can hold myself accountable, but uh I am super long overdue for um calling up Hassan uh Habib who is a he's a YouTuber in well he's a former Microsoft employee. He's a a YouTuber. His channel's gone absolutely bonkers, but he's like like one of the I mean this genuinely like one of the nicest souls I've ever met.
And I know that sounds kind of weird, but like when you talk with this guy, you you get this feeling that you're like, "This guy is so wholesome and awesome and just cares about people." it it like it doesn't it doesn't even take long before you're like you have this like overwhelming sense of like this guy is just so genuinely nice and I know that he does a lot of work with um you know with with AI and uh and alternative methods for for models that aren't just like you know blasting frontier models. So, I I want to reach out to him. Um, pick his brain a little bit because I I just know that he's doing awesome stuff in this space. Um, and like I was saying, I invite anyone, you know, kind of watching or listening, if you have experience setting up this kind of stuff, comments are open, man.
I'd love to hear from you. Um, and if I end up doing my own custom build or something, uh, I'll try to to film it. um put a video together or something like that or at least keep people up to speed on the build. Like I I think it's kind of an inevitable thing that's going to have to happen. I don't see another another way. Um unless somehow all of the Frontier model providers just make it free, right? Like I don't know. I just don't think that's going to I don't think it's going to get better. Um, I think that maybe in the future if the models, if these sort of lower grade models of today are the equivalent of the frontier models of today, say in a few years or something like that, maybe, you know, development the way that I've been kind of doing it is more sustainable on uh a lot less dollar.
Um, I think that's kind of the my suspicion, right, is like that's kind of the goal with a lot of these AI companies is to make sure that people can be building things using AI in a way that u the costs are brought down dramatically. Kind of like what I was saying with uh with computer hardware, right? Like usually the costs come down over time. I'm I'm hoping that's what we see a lot more of to the point where these uh you know these other models are just a lot I don't know like the price point is better is what I'm getting at but we'll see. So clawed for now, which is really the most unfortunate thing is that uh like obviously aside from just paying out of pocket again is like I I had all of this set up with um with co-pilot in uh across these different sessions and like you know depending on what I want to work on like I just have things organized in co-pilot.
I feel very fortunate that when I was putting a lot of, if you watch my other videos where I talk about common instructions and and things like that, um, one of the things that I like convinced myself to do early and stuck with it was like when I have co-pilot instruction files, I actually have like a a mini converter/compiler kind of thing that that will just take the co-pilot instruction files and then like uh rewrite them for Claude as well. So I use my co-pilot instruction files as the source of truth. These are ones that are targeted based on file paths. And uh this thing rewrites them for Claude. So that and I I tried it, right? Like the first thing I did was went into a repo with Claude and I said, "Hey, can you see my instruction files?" And it told me like, "Yes, I can see these clawed ones.
You know, I can see your co-pilot ones, but I can see these clawed ones mapped and they're, you know, comparable, that kind of thing." So, I'm glad I did that because all of these videos I've been making, I've been telling you guys that I've been trying to do this uh approach where I take my instructions that I'm like, so say I'm working on a particular project. Here's patterns and practices not only in the code itself, but like for how I'm guiding my agents to build better software. I take those instructions and bring them back to a common spot and then sort of reync them across my projects, right? Like every project is inevitably going to have something that's like specific to it. But like just to give you an example, if I'm building innet and I have xunit for my tests or I have tunit for my tests, there's probably going to be some patterns and practices that I like using for xunit that I like using for tunit.
And then there's going to be like general things that I like in my tests. And so one sec, I just got to check this exit. Why is it saying that? No, we're not doing that. I think it I'm in the fast lane. It's moving and I think Google sees the the other lanes stopped right now. So, it's like get off the highway. But anyway, um I've been doing this pattern, right? So, I'm very glad that I have been and uh I'm glad that it syncs with Claude um Claude instructions so that I'm not, you know, haven't done all of this work to try, you know, augmenting my agentic development just to have it kind of reset to switch what terminal I'm using. Um so far, and maybe this is just me being dumb, I'm willing to accept that. Uh, but I think my biggest complaint with Claude right now is like so co-pilot CLI has autopilot and Claude I believe is just called auto mode.
Um, and so I might have the terminology wrong and I might be missing a setting or something, but when I use autopilot with co-pilot, I basically can give it, you know, I could say like, hey, there's 20 GitHub issues and, you know, we've broken them out, decomposed them, and like maybe the specific implementation details aren't there. I don't love doing that, by the way. I find if I'm doing specking out a lot of work, if I tell every, you know, GitHub issue has like specific implementation, I'm like by the time someone gets to that, it's kind of like the further out you plan, the less accurate you can be. So if I'm planning far out and trying to be very specific about code changes, I'm basically asking for that to not work, right? Like I think it's good to have a direction.
You can have suggestions, but if it's like that work's not going to get picked up until 20 PRs from now, to me, it feels stupid to say here's the exact code you need, like no, it's like here's probably a good direction for it and like you should check these other things to see if they've evolved because odds are it's software, it probably has. But with co-pilot on autopilot, I can give it that issue hierarchy and say just go deliver and it will. And I've had things literally running for like, you know, 4 days just like kind of working through a backlog delivering. And I'm not saying that it's perfect. I'm not saying that it doesn't make mistakes or it doesn't uh sometimes do something dumb where like it, you know, CI is broken and it tries like 10 times. I'm not saying that it never does that.
I'm just saying that I it's been really good that I have the luxury that I can give it this issue hierarchy, let it truly work on autopilot until it's complete or it gets blocked at a point where it's like I just can't proceed. Um, and I have found that with auto mode in um, in clawed code, I like tried doing that a little bit even just across a couple of issues and I noticed that it simply just didn't like it kind of stopped super early um, asking for input and I was like hold on like you're you're on auto mode. like I've already I've already told you like you go do these things and uh and there's just it seems to be a lot more handholding and so I'm assuming I'm doing something wrong and I just have to kind of dig around more in
uh in clawed code because I like I know that there are people that are absolute wizards with uh agentic software development and they use claw code and I'm sure they're not sitting there, you know, every step of the way saying yes, go do this. Yes, go do this. It's not even uh like tool approvals. I'm not talking about that. Like it's not saying like give me permission to run this script or to touch this file. Not like that. I just mean like at a higher level. It's like okay well you know the commits are ready and I'm like great. Like where's the PR? And it's like, "Oh, I was waiting for you to tell me I could make a PR." Like, I literally told you several issues to go deliver them, like opening PRs and like I don't like I just I simply don't understand where the gap is.
So, I don't know if there's something more beyond auto mode. Um, like I said, I I'm absolutely willing to accept that this is me just being dumb, right? um 100%. So, I need to sort that out. Um I'm trying an experiment right now with Clawude Code while I'm driving to work. Um so, one of the the games that I'm making that I made a video about before, like to to talk about it, I haven't like shown it, is it's a like a creature evolution game. Um, so lot of uh I'm I'm super excited to to see this kind of come together. So the again quick recap if you're not familiar with what I'm talking about the the game if if you think about it from like a super casual kind of perspective it's like a glorified Tamagotchi if anyone remembers that or like I don't think Furby's the physical thing you had to actually take care of.
I think you could just interact with them, but you're basically a caretaker and you have some of these fuzzy creatures to take care of, right? So, simple little casual caretaker kind of thing if you want to play it like that. But then I want to have a ton of layers to it. So, there's a lot of uh puzzles to solve where you unlock things across across your space station. And then like the the late game is kind of like then you can go explore planets to go get, you know, more DNA for your creatures. Uh and then the like ultimate sort of end game is like you can literally code in the genetics. So, I This is my first attempt at having something putting assets together for artwork that I'm hoping aren't complete dog Come on, buddy. You got to speed up. You got to speed up.
Jeez. Um, so my my previous attempts, like if we're talking way back, my previous attempts with games, I'm terrible at art. So when you hear me talk about like, oh, like, you know, I've been building role playing games for a long time. They never end up having like an actual playable game because I suck at art. I'm so bad at it. And I uh I know that it's one of those things I'd have to spend time at to try and get better. and I just I simply am not interested. Now AI has changed that, right? Cuz anyone can go to chat GPT now and get pictures and stuff made, which is super cool. And that's come such a long way in the past couple of years. Holy crap. Um, but what about 3D modeling and textures? So my initial attempts at some of this back a few months ago, I was like, nope.
Like it's still not there. I was getting super frustrated. I was just trying to get my role playing game to have some like I don't know like AI generated artwork that I could have a movable character and stuff and it was horrific. And so I heard that like with models like Astra now like it's way better. So I'm super excited to try it out. started trying it out and it was I gave it like some really good reference art and I was like this and like models and I was like this is going to be so good and it made something and it was certainly I won't deny it certainly better than what I saw earlier this year for these for these creatures at least different uh game altogether but something about it I'm like this is like I gave you really good reference models like I don't understand why this is so bad.
And what I learned was that it was not actually using the models. Like I gave it entire resource packs. It was taking screenshots of them and trying to build it entirely from scratch. And I was like, "Oh man, like that 100% explains it." Like that's that's the problem. So to give you um you know to explain and I'm very new to this. I'm not trying to talk about you know AI model generation. Sorry model as in 3D models um art assets. I don't want to talk about this like I'm an expert by any means cuz I'm simply not. But to give you a diff uh an idea of how much it improved. I then told it I want you to actually like I literally gave you the model. Don't just use the screenshot. Um, I want you to take the qualities of like this asset which was basically fully rigged model with like textures and stuff of this like sort of creature and I want you to take this other one.
Right. I like the shape and like the head of this one, but I like how the body of this one has like uh what appears to be very, you know, dense fur. Like the fur on this one looks really good. The other one has the right shape of a body, but it's like kind of just smooth. It doesn't have a good texture on it. And it did it. It actually took like the the hair and the fur and how that was all modeled and applied it to the other one. And I was like, "Holy shit." Like we we're actually doing it. Like it's happening. So uh you know I had run out of AI credits uh right at this point when I started getting a bit of momentum.
So my experiment with clawed code that I'm doing today is I said like you know uh spent more time priming it to be like I'm expecting you to go deliver on this road map like tell me what you need right now to clarify and you know I'm going to kick you off on this work. did that um in the morning and so I'm hoping that when I go home I can see that it's progressed on the art assets. I'm very hopeful. I told it, you know, all the PR's needs, screenshots, like basically had to explain to it, I'm not signing off on the art. Like, you are doing this on autopilot. You know, uh I've given you some direction up front, but you're going to go create all of this and not wait for me, but I do want screenshots of the progress on the pull request.
Then what I'm hoping to do is go review this. Inevitably, like just like with code, I'm fully expecting there's going to be some qualities where I'm like, "Oo, like where did where did that come from? Like that's not okay." Um, and then basically, you know, pause and say, "Okay, well, how did this happen? What do we like, you know, if we want to move in this direction, what do we need to do?" um and just kind of put it back on course and say great like let's let's at least see one or two steps in the right direction. Okay, great. Now kind of put you back on autopilot and put it uh delivering things. So I want to do a few art passes like this because um the creature that was in the game as sort of like the base creature you play with was uh was pretty terrible.
Now I'm I was looking at it I'm like this feels like a this feels like a game. um in terms of like what the asset looks like, but obviously like there's always going to be room to improve. So, I wanted the creature done, which is what it's doing today. And I think it's going to start on some of the actual um like scenery because what Astra had put together before was again for like scenery, it made its own assets from scratch instead of using the models that I I told it to use. So, I think I'm hoping that I get home today, the creature is improved in the game and that the um some of the art assets for the the scenery are are improved. I'm not expecting that this is perfect, right? I want to see that we can make progress in a direction that doesn't feel terrible, right?
I want it to look like it's moving in the right direction. Then I'll pause on the art development, get back into the gameplay mechanics and getting that finished out, and then once, you know, there's more playability with the game, I'll see what the art looks like, and if it's time to update that and kind of go back and forth. Whereas before, what I was doing with co-pilot was I try to just have like art direction and art assets, like you go do that while the other agent is building out the entire game. And then I had a a third agent that was actually trying to do optimizations on on uh the integration pipeline because what I was noticing, and I've talked about this in other videos, um is like like with anything, right? If I'm leaving agents for too long, it is trying to do a lot of the right things, but it kind of goes off um the rails a little bit.
And that happens with uh CI, I find. So, it was like, cool, we have to test all these things because tests are important. I do value that. But as time's going on, I'm like, we're getting less and less done. And I'm noticing like to integrate a PR is now like an hour and a half of like build and test. And like that's not okay. Like if I need to go ship an art asset and it's taking an hour and a half to go get a green light on a pull request, like we're doing something wrong. So I had this third agent that was basically just focused on tuning uh CI. So I'm going to drop that back down to one. Like I said, I don't know what my my claw limits are like and I'll I'll play around with that, see how far we get. But um yeah, biggest thing right now is just that clawed autopilot.
I think I'm misunderstanding something or auto mode, I guess it's called. And I'm hoping to kind of dial that in. So anyway, that's my my long drive into work. How do we Where's the time on this thing? I can't see the recording time. I wonder how long this was. Traffic uh was pretty so Oh well, it's life, man. Thanks for watching. Um, I know it's like an AI kind of random development one, but if uh if you got questions you want answered, leave them below in the comments. My mic's about to die apparently. And otherwise, you can go to codemute.com and you can submit stuff anonymously and then I can try making a video to help answer for you. So, take care and I'll see you in the next one.
Frequently Asked Questions
These Q&A summaries are AI-generated from the video transcript and may not reflect my exact wording. Watch the video for the full context.
- Why did you start exploring free LLM options instead of continuing to pay for Co-Pilot credits?
- I ran out of Co-Pilot credits after Microsoft reduced the limits to 20% of last month, and paying for token usage every month didn’t feel sustainable. I decided to explore free model usage and came across Free LLM API, which helps route to different providers automatically.
- What challenges did you experience when using the free LLM API to handle GitHub PR merge conflicts?
- The model started outputting in Chinese, so I had to update instructions to English. It produced outputs that felt not fully thought through and even appeared truncated. It attempted to resolve merge conflicts by removing changes and closing the PR, then opened new PRs that still had conflicts. After about 8 hours and 110 million tokens, it couldn't finish the task.
- What change did you make after Co-Pilot credits ran out, and how are you integrating Claude into your workflow?
- I ended up renewing with Claude, subscribing to the $200 per month plan, to see if it can handle my needs. I haven’t re-evaluated Claude pricing recently, but I picked the top tier and hope it will be enough. I also use a mini converter to map my Co-Pilot instruction files to Claude so I can reuse patterns across both.