Category: Alex Finn

  • 07/10/2026 – Did ChatGPT 5.6 just KILL Fable 5?

    It actually happens. Chat GPT 5.6 is out and I am totally blown away. But is it better than Fable 5? I have a definitive answer for you. In this video, I’ll cover what Chad GBT 5.6 is, the improvements it made, do an indepth breakdown against Fable 5, show you who the clear winner is, and then show you an amazing workflow that will change how you use AI forever. This is the most important video you’ll watch in a very long time. Make sure you lock in and let’s get into it. Days like today are the absolute best. If you’re into technology, congratulations. You are going to have a lot of fun right after this video. This is Chad GBT 5.6. I’ve had early access. I’ve been using it a ton. So, all the people that go in comments go, “Oh, how’d you get a video out so quick? What the hell? What’s going on?” Shut up. I’ve been using it. Now, let’s get into it. Chad GPT 5.6 Six three models, Luna, Soul, Terra. The one we’re going to talk about today is Soul. That’s their biggest model. That’s the one everyone goes, “Oh, how’s it compared to Fable? Is it better than Fable?” That’s the one we’re going to be talking about is Soul. So Soul’s the smartest. Terra is kind the medium. Luna is the efficient one. The chop GPT Mini. So Soul is the one we care. I care about brains. So let’s talk about that. They also have many different thinking levels now. They have like seven different thinking levels. Ultra is the one I’m using. The reason why we can use the smartest Chad GBT model with the highest level thinking is because Chad GBT is so kind when it comes to usage. It gives you a ton of usage. So you don’t have to be scared like you are with Anthropic of using the highest thinking levels. Then two prompts later you have no more usage. Chad GBT you can use the highest thinking levels and still get a ton of usage out of it. They have a fast mode which is in my opinion revolutionary. I’ll tell you, you’re getting Fable level thinking on a really, really fast speed. One challenge I’ve had with Fable over the last couple weeks is it’s just slow. It’s a slow model. But Chad GBT56 Soul Ultra on fast mode, lightning fast. It is the best model in some ways. We’ll get into the fable comparison in just a second. Stick around with me here. We go in depth. Who’s the winner? I hand a trophy out. It is the best model in some ways, but it is without a doubt the best harness. Codeex, which we’ll also talk about that in a second. It’s gone, is the best harness at the best price at the best speed. Some also also some changes announced today. We’ll go into in-depth breakdowns on subsequent videos. Let me know what you want me to do. Most Chai GBT work, some of these other things. Codeex app no more. It’s gone. They combined it with the chat GBT app, which I actually believe is a very smart move. Codeex was the best harness. They’re bringing all those nice features over to the Chat GBT app. It’s allin-one. I didn’t like the separate apps. I’m a fan of this move. Codeex is no more. It is changed on mobile as well. So, if you use codecs on mobile, if you use codecs on iPad, which I highly, highly recommend you do, it’s now called remote sessions. They really want you using this in a very specific way on mobile, which I’ll talk about a little bit in a second as well. I actually think the iPad experience now is like the best experience. I I use the iPad a ton. I’ll talk about that, too. But the iPad experience for chat, I don’t even think they’re like trying to make it good. It actually is very good. I’ll talk about why I think it’s awesome, too. But let’s do the comparison everyone and their mothers is waiting for Chad GBT56 with Fable 5. Which is the winner? Well, I’m going to show you the comparison between the two, what the differences are, and then I’ll tell you the winner once we go through this. We’ll make this quick. I have been using the absolute booty load out of Fable 5 the last week. I don’t even know what booty load means. I just dropped that. As you can see here, I hit extra usage yesterday. I spent $500 of credits in the last 24 hours of Fable 5. I’ve been using the absolute hell out of it. So, I am in depth in the weeds on Fable 5. I know how it works inside and out. How does it compare to 5.6? Well, let’s take a look at this here. Fable 5 is still better for planning and highlevel thinking and understanding of large code bases. It is the best when it comes specifically to that. I’ll give you an example. I have like an entire software factory I have built out inside of codeex and claude code that allows me to kind of autonomously build a lot of software. It’s very complex, in-depth. If you want a video on that, let me know down below. I’ll do a whole video on loops and software factory if you’re interested. But here’s the thing. When I gave that to Fable 5, I said, “Hey, look at my entire software factory and understand how it works.” It understood it pretty much instantly. It was able to pick it up and start doing work in the factory like it was been there. Like it’s been there for years. When I handed the software factory over to Chat GBT 5.6 initially when I got early access, I’m going to be honest, it did not pick it up immediately. It took a while. I had to kind of nudge it in different directions until it finally grasped how it worked. So when it comes to understanding large code bases, being a high-level wise thinker, understanding from like a CEO level what’s going on, I would say Fable 5 is better. But now let’s talk about actual execution, actual building. 5.6 the winner. 5.6 is the winner for many, many different reasons. The code quality is just as good as Fable 5. The code is putting out is just as good, but it is more eager. So, the tables have kind of flipped here where Chad GBT 5.6 is more eager than Fable. So, it will get tasks done. You can give it major major tasks and it will grab it. It will get it all done. It will add a whole bunch of extra features to it. It’ll be fantastic at that. This was claude strength for a long time. That kind of eagerness of getting a ton done was Claude strength for a very long time. Chad GBT has overtaken that. So when it comes to execution, when it comes to actual building, 56 is the winner. Now after this comparison we’re going through, I’m going to actually show you the workflow. Everyone in their mothers is going on Twitter saying, “Oh, use Chad GBT as a sub agent inside of Claude Code.” No, that’s not the best way to do it. I’ll show you the best way to do it in a second. You’ll learn the best workflow once we go through this. Once I declare the winner between Chad GBT 5.6 ICS and Fable 5 workflow coming right after this. Let’s talk about price and this is going to be a dealbreaker for most people. Fable 5 unreasonably expensive. So expensive I actually think this model will die. I think Fable 5 is the smartest AI model, potentially the greatest technology ever made. I think it’s dead in the water. I was shocked when I started going into API pricing and I within a few hours used hundreds of dollars of credits. That’s absolutely insane. No normal person is going to want to pay that price. And you even look at the enterprise right now. Companies are cutting out frontier models and using things like GLM 5.2 cuz costs are getting so high. I don’t know who Fable 5 is for. It’s absolutely genius level. I wasn’t prepared to tell everyone, hey, you need to be paying API pricing for it. And the moment I started paying API pricing for it, I’m like, wait a second. No, no, no, no, no. Done with this. Done with Fable. uh unreasonably expensive. It is so unreasonable in fact that I believe they’re going to push up Opus 5. I think in the next week Opus 5 comes out. I believe I don’t think it’s Fable level. I believe it’s very close to Fable level, but at way more reasonable prices. I just don’t see how Fable fits into the picture with what I’m about to show you with Chad GPT. No one in their right mind, unless you’re just outrageously rich, is going to be spending $500 a day on credits. And this was even like a full day of coding. This was like an afternoon like 1:00 p.m. to 8:00 p.m. of coding. $500. That is crazy. Chad GBT 5.6. It’s more affordable, right? It’s way more affordable than Fable 5. It is still costly. I have never I’m going to be honest with you and I do this all day. I’ve never hit my limits with Chad GBT before. That’s never happened for me. 55 everything before that. I’ve never hit my limits. I hit my limits for the first time yesterday using Chad GBT 5.6. 6. That’s never happened before. So, it is more expensive than the other Chad GBT models. I did hit my 5hour window. I didn’t come close to the weekly window. 5 hour window though, I hit the limit. I started having to use API pricing. I was using Soul at ultra level. So, like if you go super hardcore on their most expensive model, you could hit limits. But at the end of the day, it is way more affordable and reasonable than Fable 5. UI was always Chad GBT’s weakness. It’s now good. It’s good enough at UI now. They they’ve certainly improved it. It’s good enough. It’s not a concern now where you need to jump to Claude every single time. Claude is better at UI still. Here’s where I actually think the winner comes in. And this is why I am going to declare Chad GBT 5.6 the winner. The harness around Chad GBT is so much better. It’s so much better. The computer use is significantly better. The browser use is significantly better. Claude has made jumps here lately, but Chad GBT, the harness, doing things for you, signing up for websites, getting API keys. I’ve never been more confident when giving a model a task that it will do literally whatever it takes to get it done. It will pop open the browser, will control your computer, it will do whatever it needs in order to get that task done. And that is a feeling and emotion that I’ve never had with Claude inside their harness. Which for me leads me to crown Chad GBT 5.6 the winner. Because that feeling of knowing when you give a task to a model, knowing it will get done is an invaluable feeling. Even though there’s some weaknesses when it comes to planning, highlevel thinking, figuring very large code bases out that it doesn’t really match up with Fabelon, the fact that you can give it a task and you know for certainty at the end it will be done for a reasonable price leads me to make Chad GBT 5.6 the winner. The question becomes also little things here. I’ll do a video. Let me know down below. Do you want a full video on Hermes using Chad GBT 5.6 as well? I’ll do that. But this leads me to the next section of this video, which is the workflow. How do you actually use it? We talked about what is better, how it compares to Fable 5, what the new changes are. How do you actually use it? Everyone in their mother’s been telling you if you pay attention to Twitter, if you’re in the Twitter bubble, oh, set up Chad GBT 5.6 as a sub agent inside Claude Code. No, do not do that. I will show you the best way to do this. And before we do that real quick, if you learned anything so far, leave a like down below, subscribe, turn on notifications. All I do is make the absolute stone cold best content about AI on the entire internet. And if you want to work with me live, I do live boot camps every single Friday in the Vibe Coding Academy. It’s the number one community for people building with AI on planet Earth. We’re also doing a $1,000 hackathon right now. It ends in two weeks. if you can still join it. Get in, build something cool, submit it, you can win $1,000. We’re doing it every single month. Join the Vibe Coding Academy link down below for that. Let’s get into the workflow now. So, this is the Codeex app. I’m using Codex app for now cuz I am filming this right before the model comes out. So, I’m still using the Codex app before it’s combined with Chad GBT. But, here’s how you want to use this. Here’s how you want to use Chad GBT 5.6. First, we’ll talk about from the harness level. You want to choose one computer or device you own to be kind of the head development machine. The reason why you want to do this is the remote control capabilities of Chad GPT are incredible. You can control Chad GPT from any device you have, iPad, iPhone, any of your computers, Mac Mini. So what you want to do is choose your kind of main developer machine. This should be a desktop device. I’m using one of my Mac Studios, my Mac Studio 1. I have three of them. Not to flex on you, but kind of a flex. I’m using my first one as my main developer machine. Now, every other device you have will be a node to that. Here’s how you want to think about it. So, you have your main development machine where you’re using Chad GPT or Codeex and then every other device you own will control that. So if you have a desktop and you have a laptop, your laptop controls it, your iPad, your iPhone, everything controls this one place. And the reason why you want to do it this way is so that you have all your code in one place instead of having some features on your laptop, your desktop, some features on your iPhone, your iPad, it all goes to one place. And because the remote control functionality is so good on the Chat GBT and Codeex apps, you can easily control from all your different devices. So, choose a device. Make sure it’s a desktop if you have one. The reason being is you can leave your desktop on 24/7, 365. Laptops, you got to close it. Puts it in sleep mode, different things like that. So, make sure you have kind of a desktop device, a Mac studio, a gaming computer, whatever you got, make that your main computer. Even if you just have like an old dusty laptop in your closet that you can put on 24/7 mode and plug it in, use that. Something that can be up 24/7, 365. So from there, you have your projects and you have your chats inside of each of the projects. I’m going over kind of highlevel tips in here. We can do a deep dive building out a full app end to end on a subsequent video. If you want, let me know down below. But what we’re going to do is you want to kind of have your factory admin chat in each project. This is the highlevel chat where you can be like okay merge the code, create a PR, do this where it’s kind of managing your entire factory and then every other chat from there is going to be building out specific features for that project. So you have your main admin chat where you can go in from any of your device say hey what’s the status what’s uncommitted code we have what PRs do we have to merge things like that it’s kind of your highle admin and then every other chat is for specific features you’re building out now we said earlier Fable 5 is a better planning model and I I stand by that so what you want to do is you want to use Fable 5 for planning. So, if you have a project, you open it up inside Claude Code, you go into Fable 5 mode. I’d go on high. If you go Ultra Code or higher, you spend hundreds of dollars for a single prompt. Just use high. When you’re doing any sort of planning, you’ll come into Clawude Code. You will ask for the plan. Hey, this is what I’m thinking about building. Please build out a very large spec. If you’ve been paying attention to all my videos, you know I talk about linear and how I use linear in my entire process. So if you’re using linear, have it build the spec out in linear. If not, totally fine to say, “Hey, build a markdown file.” Then what you will do is you will take that markdown file, go into your projects, drop it in, and say, “Here’s the plan. Build this out for me.” Now, a lot of people are telling you to use a Chad GBT sub agent inside of Claude Code to do execution. I don’t like that at all because when you do that, you’re missing out on all the amazing features of the Chad GBT harness. The Chad GBT harness does so many things so well for you. So, for instance, I’m building out my app, Henry Intelligent Machines, link down below if you want to get an early preview of that. And I’m having it test some of the action generation inside the app. In Henry, you launch a venture and then it generates a bunch of actions that you need to do to push your business forward. I wanted to test that. Typically, when testing that, you would need to open up a browser, go through the flow yourself, click on things yourself. But all I do is I go in and I say, “Hey, test this out for me. Load it up into a browser, walk through it for me, do it five times, and then build a report on that.” And that is what Chad GBT does. It opens up its own browser and runs all the servers. It gets whatever API keys it needs, it plugs them in, and it runs through everything. So, it said what I tested, actually opened it up, started a venture, did a bunch of different actions, tested everything out. It did all the testing for me, something that would have typically taken me an hour before, took 24 minutes for chat GBT while I did other things. So your workflow is get the planning from claude code, move it over to chat GPT and then you give the actions to chat GPT. Here is the mindset I have. Anything you would manually need to do before when vibe coding, I just tell Chad GBT to do it. Anything it is, it doesn’t matter what it is. If they say, “Hey, sign up for this account.” I go, “No, you do it.” If they say, “Hey, can you grab this API key?” I say, “No, you do it.” If they say, “Hey, can you connect these two things together or move this file over here?” I say, “No, you do it.” Chad GBT literally 100% of the time I have said, “Hey, no, you do it.” It’s figured out how to do it. That’s why this harness is so good. When it comes to computer use, when it comes to browser use, it is absolutely next level. It is bar none not comparable to Claude Code whatsoever. That’s why the workflow is get your plans from Claude Code, move them over to Chad GPT codeex and then you say, “No, you do this. You handle this and it will get done and it will save you tons of time.” Now, while you’re in here, I do sole ultra for everything. It’s the stone cold best. If you’re on the $200 a month plan, you can get away with using fast for a lot of things. Again, I hit my limits using fast here for the first time in my entire life in Chad GBT. So maybe you don’t use fast quite as much, but if you’re on soul ultra on normal speeds, you can use up your $200 plan, like you you’re not going to hit your limits and you’re going to have plenty of usage, which I cannot say for Fable because I used the plan up in like a day and then when I went into API pricing, it was outrageous, out of control. 5.6 is the winner. It is my new daily driver. It is what you need to be using for pretty much everything outside of highle understanding, planning, and some kind of UI tweaks if need be. It’s still pretty good at UI. By the way, for the hardcore Alex Finn fans who know all about the world famous Alex Finn benchmark. Yes, it blew away the benchmarks, this is the first person shooter test. It is absolutely incredible. It has it has crits, dynamic enemies. It has great explosions. It has great waves. The graphics are fantastic. The sounds are amazing. I don’t know if you can hear the sounds or not. The power-ups are incredible. Like, there’s like 10 different dynamic enemies. Uh, it’s incredible. Chad GBT 5.6 6 won the benchmark test. I am really really impressed. And what’s incredible is over the next few weeks, we’re actually expecting GPT6. We’re expecting the next level. The the rumors I am hearing are that it’s coming very soon. They’re already done training it. They’re just doing testing now. So, expect more videos like this very, very soon. If you learn anything at all, likes, subscribes, turn notifications. Again, live boot camp tomorrow. $1,000 hackathon next couple weeks. Join the Vibe Coding Academy link down below. I love you all. It’s an honor to teach you these things. It’s an honor to work with you.

  • 07/11/2026 – It’s never been more important to learn about local AI – full guide

    The world has changed dramatically this week and I am not understating that. It has never been more important for you to get into local AI. Governments, the US government is banning frontier models for everybody. Hardware prices are becoming outrageous. They are no longer affordable for the average person and they may never be affordable again for the average person. Chinese labs are creating models that are on unparalleled levels. the world is transforming right before our eyes. And I believe the most important thing you can be doing to keep up with all this change is educating yourself and getting into local AI. In this video, I’m going to cover why that is, how the world is changing so rapidly, what hardware you need to get into local AI, what models you should be paying attention to, what use cases there are for local AI, why would you use local AI when you can do Chad GPT for $20 a month? I’ll cover it all. you’ll walk away an expert and even maybe set up with a local model by the end of the video. This could change your life. I don’t think it’s ever been more important to pay attention to the video as it is this one. Just as context, I added this computer to my home AI lab this week. I built this 1590 RTX computer. I already have three Mac Studio 512 GB. I already have two Mac Minis. One of the Mac Minis is not shown here. And I have a DGX Spark. Do you need to get all these hardware? Do you need to spend $40,000 like me? No, you do not. In fact, you might be able to do a really great home AI lab setup with the hardware you already have, with the laptop you got in your closet. I’ll cover all that in a second. But why is this so important? Why am I building up this entire home AI lab with all these different computers? Well, the world has changed a lot this week. First of all, Chad GBT56 was announced yesterday, and with it, it was also announced that only a hand selected group of people by the US government will be able to use it. That stinks. It’s a big party and you ain’t in it, right? You’re not able to use the two best models out there right now, which is Fable 5 and Chad GBT56. They’ll probably return within the next couple weeks. We’ll be able to use it, but the hand selected people by anthropic, by open AI, by the government will have a head start and they will have a massive advantage on you. That sucks. We have officially entered the age of hand selected winners, which isn’t good. On top of that, the prices of literally everything are going up. The prices of every piece of hardware is dramatically going up. Apple announced that the first of many price raises just happened. Many of their devices went up 20%, 25%. I mean, you’re not getting Mac minis for less than $1,500 right now. It’s absolutely insane what’s going on. I went to Micro Center. I wanted to build this computer. I wanted to put 128 gigs of RAM in it, which isn’t the craziest thing to do in the entire world. It would have cost me $4,000 extra dollars, dude. I’m already spending $9,000 on that computer. It would have cost me $4,000 extra if I want to do 128 gig. I mean, even the cost of like the PlayStation 5, the Xbox Series X and S are out of control. The Series S is $500 now, which was the cost of the good model when it just came out. I don’t know if you remember video game consoles, they used to drop in price. They used to be like $99 by the end of their cycle. Now they just go up and up and up. We are living in absolutely insane times. Hardware, while very very expensive, is still obtainable. And that’s why I’m making this video right now is because hardware is still obtainable. I predict it won’t be obtainable. I predict in a year it’s going to be absolutely insane prices, way higher than it is now, and it’s already at insane levels. Two years, I don’t think it’ll be obtainable. We are still years away from humanoid robots being a thing. We are still years away from drones fighting in wars. Most people who are in the know are in agreement that every home will have a humanoid robot and humans will no longer fight in wars. It’ll be all drones. We are nowhere close to that. And every humanoid robot, every drone that fights in a war, every automatic self-driving car out there will need absurd amounts of memory. And those products aren’t even being built right now and we have a bottleneck. So what happens when we actually do start producing humanoid robots? What happens when we actually do start producing tens if not hundreds of millions of drones to go and fight in wars? I’ll tell you this much. You ain’t getting memory when that happens. Memory ain’t going to be available to the normal average Joe when that happens. So yes, prices are absurd right now, but no, I don’t even think we’re close to the end. In fact, I don’t think we’re ever going back to the days of before where these kinds of prices were obtainable for people. So, we live in a situation frontier models being taken away. Hardware prices are absurd. What is the solution? To me, the solution is getting into local AI now. Getting the hardware for local AI now. Cuz when the government can start controlling who uses what models, you need to become sovereign. But the issue is is the window for becoming sovereign is becoming very small because the hardware is getting more and more expensive. So I believe we’re in a small window right now where it is obtainable to get local AI. Even though it is very expensive, it’s still obtainable. That’s why I’m making this video right now. What is local AI for those who are really not in the know? And by the way, there’s chapters down below, I think. Feel free to click around. It is AI that runs on your computer. When you send prompts right now to Chad GPT or Claude or whatever you’re using, it goes to the cloud, right? It goes through your internet, goes to a server, goes to a data center. It’s processed in that data center. Then the answer is sent back to you. That is in full control of those companies. Your prompts live on their servers. The answers they give you are determined by the companies. You are fully out of control. When you run local, the models run on your GPUs on your devices, so on your computer. So your prompts, your answers never leave your computer. They don’t even leave your internet. You don’t even need internet. It’s all done locally on your computer. That’s what I mean by becoming sovereign. No one can take it away. It is completely private. So let’s talk about those advantages. Again, as I said, private. No one can read your logs. When you send a prompt to Anthropic or Claude, if they decide uh to sue you or something like that, they can access your logs. They can read your logs. Anyone at those companies, the government, anyone can read everything you’re saying. So those conversations with your AI girlfriend, they ain’t private. Anyone can read it. The government can read it. So privacy, local AI, completely private, unlimited. You’re not paying a tollgate to send your prompts to the cloud. It’s unlimited. It is run on your computer. You’re just paying the cost electricity going into your computer. And it can’t be taken away unless they come into your home guns ablazing and take your Mac Studio, your DGX Spark, take it away from you. They can’t take it away. It’s all yours. At any given time, as we saw the last couple weeks, it can be taken away. They took away Fable 5. They stopped you from being able to use Chad GBT 5.6. You can’t use the best intelligence out there right now. They determine who can use it, and you ain’t one of them. Before we get into what hardware you need, what models you should be using, let’s just address the elephant in the room, the number one thing naysayers say to me. They go, “Alex, but isn’t local models stupider? Isn’t it dumb? Aren’t they slow?” Yeah, you’re you’re 100% correct. They are stupider than your Chad GBT models. You get chadgbt.com. They are stupider than Opus 48. They are slower, sometimes significantly slower, right? But that’s not the point. You’re 100% right about that. But here’s the thing. One, not for long. I don’t think it’ll be stupider and slower for long. In fact, GLM 5.2, the newest local model, obviously from GLM, is just as smart as Opus 48 I’ve tested. It is just as smart as Opus 48. So, they are getting significantly smarter and they are getting faster. They are a bit slow right now, but these companies including Google are working to optimize these models so they fit on your computers. And that’s why the old paradigm of like hardware going extinct and not being relevant 2 to 3 years after you buy it doesn’t really matter anymore. That that kind of cyclicality of hardware doesn’t really exist anymore because these models are becoming more efficient and the hardware from 5 years ago could start running Frontier models soon enough potentially. That’s why I think the excuse a lot of people make, oh, but won’t the hardware not matter in a few years? No, that’s I I I do think it’ll matter in a few years. In fact, I don’t think you’ll be able to buy more hardware in a few years. So, what you got available today is all you can get. So, little bit of expectation setting right now a little bit stubber. Right now, a little bit slower. Even though it’s stupider and slower, there are use cases that only local AI models can do that Chad GBT and Opus can’t do that I’ll go in as well that are massive advantages that I really think you guys should be getting into. So, let’s first talk about the hardware you should be getting. Can I use Mac minis? Can I use Mac Studio? Should I buy a Mac? Should I buy an Nvidia card? Should I buy a DJX Spark? Well, here’s the three categories of hardware for you. Actually, there might be a fourth one here. Let’s talk about the fourth one, too, which is kind of just dusty old computers/cheep computers, like kind of the budget computers. We’ll talk about in a second. Let’s first go into kind of the more expensive ones people get into. You have the Mac Studios. This is the kind of one I started a big trend on back in January when I first started talking about this. Mac Studios have what’s called high unified memory and low bandwidth. What does that mean? Macs, any Mac computer you get, the advantage of the Apple architecture is it’s unified memory. So the RAM and the VRAM are kind of one thing. the RAM and where you load models into. Models need memory are combined, which means you can load much bigger models. You can load I load GLM 5.2 on one of my Mac Studios. It’s a 250 gigabyte model. You need 250 GB of memory. That’s a tremendous amount. So, you can load massive models, but it has low bandwidth. What does that mean, low bandwidth? That means it can’t process a ton of that memory at once. So, they’re really, really slow. So, you buy a Mac, you can run very smart intelligence, just not going to do it too fast. Then you got these AI computers, these kind of AI workstations that are now coming out like the DJX Spark, the AMD Halos, the DJX station. They have kind of medium unified memory. DJX Spark is 128 gigs, which is really good. Max Studio, I have 512 on these, which is obviously a lot better, but they also have medium bandwidth, so they can do it pretty, you know, a good amount faster than the Max out there. So, that’s pretty good. you can run pretty smart models at pretty good speeds. And then you have kind of the powerhouse chips, the really expensive ones, the 5090s, which is what I just built, and the 6000 Pros, which is now a $12,000 GPU. They have lower VRAM. The 6000 is 96 gig, so actually has pretty good VRAM, but the 5090 uh has 32 gigs of VRAM. So lower VRAM, but very, very high bandwidth, meaning it’s lightning quick, very, very fast. You can run pretty good. actually, you know, for 32 gigs, you can run good models now and you can run them very, very fast. Those are kind of your three options. Then you have the fourth category, which is basically everything else. Your Mac minis, your old laptops, you know, the pretty cheap computers out there, anything below $4,000 computers. You can run some smaller models out there, some Gemma 4s, and they are solid models. You’re just not going to give them any sort of powerhouse use cases. They’re not going to write any sort of kind of quality code for you, but they can do some small processing. They can do embeddings, which is improving basically the memory of your agents. They can do some good things. So even if you just have a Mac Mini, you can still take advantage of local models and you should be taking advantage of them. So it’s up to you what kind of functionality you want. I’m going to go into use cases in a second and that’ll might help you decide what hardware you want, but this is where you’re at. Also, for just from a usability perspective, I personally like Mac OS better than all the other operating systems. I love it. I got an iPhone, I got an iPad, they all integrate together. So, that’s why the Mac Studio is kind of my main driver. AI computers like the Spark, they’re running on Linux, right? The usability for the average person is not great. So, you can still get them. They’re just going to be very usable. You’re probably going to want your Hermes agent to run everything. And you have the powerhouse chips. You can run on Microsoft Windows. People like Windows. People are used to Windows. I personally kind of despise Windows, but you have that option as well. I’m going to get into use cases in a second. I’m going to show you some things I’m doing. But last thing I’ll show you is like kind of what else you need. What software supports the hardware. I’d say two things are critical here. I’d say tail scale. Tail scale is a piece of software that allows you to create your own private network amongst all your devices, right? So I have like seven or eight different computers here. I have tail scale on all of them. So they’re all on kind of the same closed network. And no matter where I am in the world, on my Wi-Fi, not on my Wi-Fi, all my computers and the models can talk to each other. This allows you to leverage the local models you’re running on the devices on any device you have. You if you run a really nice local model, your Mac Studio, if you’re on the tail scale on your phone, you can run it from your phone anywhere in the world, use that model locally on your phone, which is really, really nice. And then Hermes or OpenClaw, this is the killer one. This is one you need. This basically takes away all the need to have technical knowledge. I just have my Hermes on my main Mac studio. And because all my devices are on tail scale, my Hermes can control, load, do whatever it needs with the local models. I can go to Hermes. I can say, “Hey, go my Mac Studio, load Quen 36 onto it. It’ll go over to my Mac Studio through Tailscale, load the model, make sure it’s working, do all the work for me.” It’s basically like my my IT guy. I say what I need. I say, “Hey, I want to write this piece of software. Use the model that’s on this computer. Make a route through this computer and use this other computer and this other local model. Do this.” And it’ll just figure it all out and put all the pieces together. Hermes and OpenClaw are absolutely critical for the setup. You need both of these. And the good news is they’re both free. Now, let’s talk about use cases. This is very important. Everyone goes, “Man, why would I use a local model if I can just use chatgbt.com for $20 a month and it’s way smarter?” Well, here’s the thing. You open up a whole new paradigm of use cases when you have always on ambient 247 AI because the models are unlimited. They’re always on. There’s no rate limits on your computer. You unlock a whole new class of use cases because they can run 24/7 and work around the clock. I’ll show you a few things. This is my home AI platform I built out. By the way, if you learned anything so far, leave a like down below, subscribe, turn notifications. Also, I have a community, Vibe Code Academy. Feel free to join that. I do live boot camps like three times a week in there. So, if you want to join, ask me questions. I’ll walk you through setting this up. I’ll handhold you through it. Join down below for the Vibe Code Academy link down below. But this is my home AI setup. I’m building this platform out myself. My Hermes agent built this out for me. But I want to show you some things here. This is saying failed, but it’s actually working. I am doing multiple things around the clock 247. So, first of all, you can see my running models. On my Mac Studio 2, I have GLM 5.2, which is Opus Level Intelligence. It’s a little slow. On my DGX Spark, I have Ornith 1.0 setup. I still need to set up my RTX 5090 computer. Built that just yesterday, but I’ll be running probably another copy of Ornith 1.0 on there. Very cool local model that just came out. And they’re doing many things all at once. One, they’re doing security scans on my codebase. So, I’m building a SAS right now called Henry Intelligent Machines. It is doing aroundthe-clock security scans on my codebase, looking at different APIs, different pieces of code, making sure it’s secure, making sure nothing’s going on. It’s taking a look at my database. It’s looking for anomalies in my database. This is going around the clock 24/7. You can’t do this with cloud models unless you’re filthy rich because you’re paying a toll booth for every API call to the the cloud models because they’re local and they’re unlimited. I can have it ambient. I can have it working at all times around the clock all day. It would cost me one of these computers to do this with Claude Opus 48, but because it’s unlimited local, it works out for me. I have another thing going on right now, too, which is one of my local models is at all times of the day scraping Reddit, looking at tweets, pulling down posts from X, and it’s looking for business opportunities, opportunities to build SAS, opportunities to build services, software, whatever I need. It’s constantly scraping and scanning the internet to find these opportunities. Now, these opportunities are going into my SAS Henry intelligent machines. My SAS is kind of built around fighting opportunities, but I also take a look every once in a while and I look to see what kind of software I can build, right? Like look at all these bit every day, every couple hours, I get a list of opportunities with sources from people’s tweets saying, “Hey, this person asked for this question or has this problem.” And it shows me the pain. It shows me how to solve it and it ranks the opportunities for me. And this is happening again around the clock every 20 minutes. Just watching the web, looking for signal, looking for opportunities. Again, not possible with cloud models because it’d be way too expensive to do this to run it around the clock. But because I’m running local models, which are unlimited, I can do these types of things. So, I basically have a fleet of employees working around the clock 24/7. Never need to eat, sleep, do nothing, never ask for time off, doesn’t need health insurance. And it’s only possible with local models. Those are the unlocked use cases. And this is just now, right? Think of what happens when these models get more efficient. When Google does better research on how to run models more efficiently locally, you’ll be able to run GLM 5.2 at rapid fast speeds, at cloud speed. So you can have frontier intelligence at amazing speeds and have unlimited of it. This is how you escape the control of what’s going on right now. You escape the control of the frontier labs. You escape the control of the government who’s holding everyone back. This is how you escape the I I hate this use the cringe term. It’s like the most cringe guru term I’m playing, but this is how you escape the matrix. This is kind of it. I also say this and this is the most underrated part people don’t talk about. Two things. One, it’s just fun. It’s fun building your own hardware. It’s fun running local models. It’s fun. You’re allowed to have fun. Believe it or not, in this world, you are allowed to have fun. You’re allowed to experiment and tinker. It’s never been more important to experiment and tinker. So, one, you’re allowed to have fun. Two, it’s educational. You’re learning about the most cuttingedge important technology in the history of our species. Literally, that it literally is. And the best way to learn about it is by getting hands-on and doing it and running yourself and seeing what the capabilities are, right? So, it’s one helpful cuz I showed you all this use case that made my life significantly better. And two, it’s fun. You have a lot of fun doing it. It’s fun learning about the hardware. It’s fun running it. It’s fun booting it up. If you’re anything like me, you can, you know, I built this 5090. I’m running local models, but hey, I’m also playing Cyberpunk 2077 in ultra mode, so that’s pretty fun, too, just as an aside. And then three, educational. You’re learning. Education’s important. You’re learning about the most important technology ever. Now, if you set it up like I set it up where your Hermes is controlling everything, they’re tail scaled out the wazoo. You literally can build anything you want. If you just want to build a simple chat interface that looks exactly like chat GBT, you can go to your Hermes and say, “Hey, build me a chat GBT interface.” If you want to give your local model a face and a name so you can talk to it and you have your own buddy and your own pal on your computer cuz hey, it’s kind of fun looking down your computer knowing there’s an AI agent on it. You can do her. Hey, build a face of my local agent. Put on top it whenever it sends me a message back. Make it talk. You can do you can do whatever you want. It’s complete freedom with the most important technology ever made. I think it has never been more important with hardware becoming about to become unobtainable and the government taken away uh frontier models. I don’t think there’s ever been a more important time to get into this technology. I hope this was helpful. Leave a like, subscribe down below if you want. Let me know in the comments you want more local content. You want me to talk about the home AI lab setup? Where where are you most confused? Where can I help out the most? All my videos are based on your feedback. Let me know down below. Hope this was helpful.