It’s never been more important to learn about local AI – full guide

The world has changed dramatically this week and I am not understating that. It has never been more important for you to get into local AI. Governments, the US government is banning frontier models for everybody. Hardware prices are becoming outrageous. They are no longer affordable for the average person and they may never be affordable again for the average person. Chinese labs are creating models that are on unparalleled levels. the world is transforming right before our eyes. And I believe the most important thing you can be doing to keep up with all this change is educating yourself and getting into local AI. In this video, I’m going to cover why that is, how the world is changing so rapidly, what hardware you need to get into local AI, what models you should be paying attention to, what use cases there are for local AI, why would you use local AI when you can do Chad GPT for $20 a month. I’ll cover it all. you’ll walk away an expert and even maybe set up with a local model by the end of the video. This could change your life. I don’t think it’s ever been more important to pay attention to the video as it is this one. Just as context, I added this computer to my home AI lab this week. I built this 1590 RTX computer. I already have three Mac Studio 512 GB. I already have two Mac Minis. One of the Mac minis is not shown here. And I have a DJX Spark. Do you need to get all these hardware? Do you need to spend $40,000 like me? No, you do not. In fact, you might be able to do a really great home AI lab setup with the hardware you already have, with the laptop you got in your closet. I’ll cover all that in a second. But why is this so important? Why am I building up this entire home AI lab with all these different computers? Well, the world has changed a lot this week. First of all, Chad GBT56 was announced yesterday, and with it, it was also announced that only a hand selected group of people by the US government will be able to use it. That stinks. It’s a big party and you ain’t in it, right? You’re not able to use the two best models out there right now, which is Fable 5 and Chad GBT56. They’ll probably return within the next couple weeks will be able to use it. But the hand selected people by Anthropic, by Open AI, by the government will have a head start and they will have a massive advantage on you. That sucks. We have officially entered the age of hand selected winners, which isn’t good. On top of that, the prices of literally everything are going up. The prices of every piece of hardware is dramatically going up. Apple announced that the first of many price raises just happened. Many of their devices went up 20%, 25%. I mean, you’re not getting Mac minis for less than $1,500 right now. It’s absolutely insane what’s going on. I went to MicroEnter. I wanted to build this computer. I wanted to put 128 gigs of RAM in it, which isn’t the craziest thing to do in the entire world. It would have cost me $4,000 extra dollars. Dude, I’m already spending $9,000 on that computer. It would have cost me $4,000 extra if I want to do 128 gig. I mean, even the cost of like the PlayStation 5, the Xbox Series X and S are out of control. The Series S is $500 now, which was the cost of the good model when it just came out. I don’t know if you remember video game consoles, they used to drop in price. They used to be like $99 by the end of their cycle. Now they just go up and up and up. We are living in absolutely insane times. Hardware, while very, very expensive, is still obtainable. And that’s why I’m making this video right now is because hardware is still obtainable. I predict it won’t be obtainable. I predict in a year it’s going to be absolutely insane prices, way higher than it is now, and it’s already at insane levels. Two years, I don’t think it’ll be obtainable. We are still years away from humanoid robots being a thing. We are still years away from drones fighting in wars. Most people who are in the know are in agreement that every home will have a humanoid robot and humans will no longer fight in wars. It’ll be all drones. We are nowhere close to that. And every humanoid robot, every drone that fights in a war, every automatic self-driving car out there will need absurd amounts of memory. And those products aren’t even being built right now. And we have a bottleneck. So, what happens when we actually do start producing humanoid robots? What happens when we actually do start producing tens, if not hundreds of millions of drones to go and fight in wars? I’ll tell you this much, you ain’t getting memory when that happens. Memory ain’t going to be available to the normal average Joe when that happens. So, yes, prices are absurd right now, but no, I don’t even think we’re close to the end. In fact, I don’t think we’re ever going back to the days of before where these kinds of prices were obtainable for people. So we live in a situation frontier models being taken away. Hardware prices are absurd. What is the solution? To me, the solution is getting into local AI now. Getting the hardware for local AI now. Because when the government can start controlling who uses what models, you need to become sovereign. But the issue is is the window for becoming sovereign is becoming very small because the hardware is getting more and more expensive. So, I believe we’re in a small window right now where it is obtainable to get local AI. Even though it is very expensive, it’s still obtainable. That’s why I’m making this video right now. What is local AI for those who are really not in the know? And by the way, there’s chapters down below, I think. Feel free to click around. It is AI that runs on your computer. When you send prompts right now to Chad GBT or Claude or whatever you’re using, it goes to the cloud, right? It goes through your internet, goes to a server, goes to a data center. It’s processed in that data center. Then the answer is sent back to you. That is in full control of those companies. Your prompts live on their servers. The answers they give you are determined by the companies. You are fully out of control. When you run local, the models run on your GPUs on your devices. So on your computer. So your prompts, your answers never leave your computer. They don’t even leave your internet. You don’t even need internet. It’s all done locally on your computer. That’s what I mean by becoming sovereign. No one can take it away. It is completely private. So, let’s talk about those advantages. Again, as I said, private. No one can read your logs. When you send a prompt to Anthropic or Claude, if they decide uh to sue you or something like that, they can access your logs. They can read your logs. Anyone at those companies, the government, anyone can read everything you’re saying. So, those conversations with your AI girlfriend, they ain’t private. Anyone can read it. The government can read it. So, privacy, local AI, completely private, unlimited. You’re not paying a tollgate to send your prompts to the cloud. It’s unlimited. It is run on your computer. You’re just paying the cost electricity going into your computer and it can’t be taken away unless they come into your home guns ablazing and take your Mac Studio, your DGX Spark, take it away from you. They can’t take it away. It’s all yours at any given time. As we saw the last couple weeks, it can be taken away. They took away Fable 5. They stopped you from being able to use Chad GBT 5.6. You can’t use the best intelligence out there right now. They determine who can use it, and you ain’t one of them. Before we get into what hardware you need, what models you should be using, let’s just address the elephant in the room. The number one thing naysayers say to me, they go, “Alex, but isn’t local models stupider? Isn’t it dumb? Aren’t they slow?” Yeah, you’re you’re 100% correct. They are stupider than your Chad GBT models you get at Chad GBT.com. They are stupider than Opus48. They are slower, sometimes significantly slower. Right. But that’s not the point. You’re 100% right about that. But here’s the thing. One, not for long. I don’t think it’ll be stupider and slower for long. In fact, GLM 5.2, the newest local model, obviously from GLM, is just as smart as Opus 48. I’ve tested. It is just as smart as Opus 48. So, they are getting significantly smarter and they are getting faster. They are a bit slow right now, but these companies, including Google, are working to optimize these models so they fit on your computers. And that’s why the old paradigm of like hardware going extinct and not being relevant 2 to 3 years after you buy it doesn’t really matter anymore. That that kind of cyclicality of hardware doesn’t really exist anymore because these models are becoming more efficient. And the hardware from 5 years ago could start running Frontier models soon enough potentially. That’s why I think the excuse a lot of people make, oh, but won’t the hardware not matter in a few years? No, that’s I I I do think it’ll matter in a few years. In fact, I don’t think you’ll be able to buy more hardware in a few years. So what you got available today is all you can get. So little bit of expectation setting right now a little bit stubber right now a little bit slower. Even though it’s stupider and slower, there are use cases that only local AI models can do that Chad GBT and Opus can’t do that I’ll go in as well that are massive advantages that I really think you guys should be getting into. So let’s first talk about the hardware you should be getting. Can I use Mac minis? Can I use Mac Studio? Should I buy a Mac? Should I buy an Nvidia card? Should I buy a DJX Spark? Well, here’s the three categories of hardware for you. Actually, there might be a fourth one here. Let’s talk about the fourth one, too, which is kind of just dusty old computers/cheep computers, like kind of the budget computers. We’ll talk about in a second. Let’s first go into kind of the more expensive ones people get into. You have the Mac Studios. This is the kind of one I started a big trend on back in January when I first started talking about this. Mac Studios have what’s called high unified memory and low bandwidth. What does that mean? Macs, any Mac computer you get, the advantage of the Apple architecture is it’s unified memory. So the RAM and the VRAMm are kind of one thing. The RAM and where you load models into models need memory are combined, which means you can load much bigger models. You can load I load GLM 5.2 on one of my Mac Studios. It’s a 250 gigabyte model. You need 250 GB of memory. That’s a tremendous amount. So, you can load massive models, but it has low bandwidth. What does that mean? Low bandwidth. That means it can’t process a ton of that memory at once. So, they’re really, really slow. So, you buy a Mac, you can run very smart intelligence, just not going to do it too fast. Then you got these AI computers, these kind of AI workstations that are now coming out like the DGX Spark, the AMD Halos, the DGX station. They have kind of medium unified memory. DJX Spark has 128 gigs, which is really good. Max Studio. I have 512 on these, which is obviously a lot better, but they also have medium bandwidth, so they can do it pretty, you know, a good amount faster than the Max out there. So, that’s pretty good. You can run pretty smart models at pretty good speeds. And then you have kind of the powerhouse chips, the really expensive ones, the 5090s, which is what I just built, and the 6000 Pros, which is now a $12,000 GPU. They have lower VRAM. The 6000 has 96 gigs, so actually has pretty good VRAM. But the 5090 uh has 32 gigs of VRAM. So lower VRAM, but very, very high bandwidth, meaning it’s lightning quick, very, very fast. You can run pretty good actually, you know, for 32 gigs, you can run good models now, and you can run them very, very fast. Those are kind of your three options. Then you have the fourth category, which is basically everything else. Your Mac minis, your old laptops, you know, the pretty cheap computers out there. or anything below $4,000 computers, you can run some smaller models out there, some Gemma 4s, and they are solid models. You’re just not going to give them any sort of powerhouse use cases. They’re not going to write any sort of kind of quality code for you, but they can do some small processing. They can do embeddings, which is improving basically the memory of your agents. They can do some good things. So, even if you just have a Mac Mini, you can still take advantage of local models, and you should be taking advantage of them. So, it’s up to you what kind of functionality you want. I’m going to go into use cases in a second, and that’ll might help you decide what hardware you want, but this is where you’re at. Also, for just from a usability perspective, I personally like Mac OS better than all the other operating systems. I love it. I got an iPhone. I got an iPad. They all integrate together. So, that’s why the Mac Studio is kind of my main driver. AI computers like the Spark, they’re running on Linux, right? The usability for the average person is not great. So, you can still get them. They’re just not going to be very usable. you’re probably going to want your Hermes agent to run everything. And you have the powerhouse chips. You can run on Microsoft Windows. People like Windows. People are used to Windows. I personally kind of despise Windows, but you have that option as well. I’m going to get into use cases in a second. I’m going to show you some things I’m doing. But last thing I’ll show you is like kind of what else you need, what software supports the hardware. I’d say two things are critical here. I’d say tail scale. Tail scale is a piece of software that allows you to create your own private network amongst all your devices, right? So, I have like seven or eight different computers here. I have tail scale on all of them. So, they’re all on kind of the same closed network. And no matter where I am in the world, I’m on my Wi-Fi, not on my Wi-Fi, all my computers and the models can talk to each other. This allows you to leverage the local models you’re running on the devices on any device you have. You if you run a really nice local model, your Mac Studio, if you’re on the tail scale on your phone, you can run it from your phone anywhere in the world. Use that model locally on your phone, which is really, really nice. And then Hermes or OpenClaw. This is the killer one. This is one you need. This basically takes away all the need to have technical knowledge. I just have my Hermes on my main Mac studio. And because all my devices are on tail scale, my Hermes can control, load, do whatever it needs with the local models. I can go to Hermes. I can say, “Hey, go on my Mac Studio, load Quen 36 onto it.” It’ll go over to my Mac Studio through Tailscale, load the model, make sure it’s working, do all the work for me. It’s basically like my my IT guy. I say what I need. I say, “Hey, I want to write this piece of software. Use the model that’s on this computer. Make a route through this computer and use this other computer and this other local model do this.” And it’ll just figure it all out and put all the pieces together. Hermes and Open Claw are absolutely critical for the setup. You need both of these. And the good news is they’re both free. Now, let’s talk about use cases. This is very important. And everyone goes, “Man, why would I use a local model if I can just use chatgbt.com for $20 a month and it’s way smarter?” Well, here’s the thing. You open up a whole new paradigm of use cases when you have always on ambient 24/7 AI because the models are unlimited. They’re always on. There’s no rate limits on your computer. you unlock a whole new class of use cases because they can run 24/7 and work around the clock. I’ll show you a few things. This is my home AI platform I built out. By the way, if you learned anything so far, leave a like down below, subscribe, turn notifications. Also, I have a community, Vibe Code Academy. Feel free to join that. I do live boot camps like three times a week in there. So, if you want to join, ask me questions. I’ll walk you through setting this up. I’ll handhold you through it. Join down below for the Vibe Code Academy link down below. But this is my home AI setup. I’m building this platform out myself. My Hermes agent built this out for me. But I want to show you some things here. This is saying failed, but it’s actually working. I am doing multiple things around the clock 247. So, first of all, you can see my running models. On my Mac Studio 2, I have GLM 5.2, which is Opus Level Intelligence. It’s a little slow. On my DGX Spark, I have Orith 1.0 setup. I still need to set up my RTX 5090 computer. Built that just yesterday, but I’ll be running probably another copy of Ornith 1.0 on there. Very cool local model that just came out. And they’re doing many things all at once. One, they’re doing security scans on my codebase. So I’m building a SAS right now called Henry intelligent machines. It is doing around the clock security scans on my codebase, looking at different APIs, different pieces of code, making sure it’s secure, making sure nothing’s going on. It’s taking a look at my database. It’s looking for anomalies in my database. This is going around the clock 24/7. You can’t do this with cloud models unless you’re filthy rich because you’re paying a toll booth for every API call to the the cloud models because they’re local and they’re unlimited. I can have it ambient. I can have it working at all times around the clock all day. It would cost me one of these computers to do this with Claude Opus 48, but because it’s unlimited local, it works out for me. I have another thing going on right now too, which is one of my local models is at all times of the day scraping Reddit, looking at tweets, pulling down posts from X, and it’s looking for business opportunities, opportunities to build SAS, opportunities to build services, software, whatever I need. It’s constantly scraping and scanning the internet to find these opportunities. Now, these opportunities are going into my SAS Henry intelligent machines. My SAS is kind of built around fighting opportunities, but I also take a look every once in a while and I look to see what kind of software I can build, right? Like look at all these bit every day, every couple hours, I get a list of opportunities with sources from people’s tweets saying, “Hey, this person asked for this question or has this problem and it shows me the pain. It shows me how to solve it and it ranks the opportunities for me.” And this is happening again around the clock every 20 minutes just watching the web, looking for signal, looking for opportunities. Again, not possible with cloud models because it’d be way too expensive to do this to run it around the clock. But because I’m running local models, which are unlimited, I can do these types of things. So, I basically have a fleet of employees working around the clock 24/7. Never need to eat, sleep, do nothing, never ask for time off, doesn’t need health insurance, and it’s only possible with local models. Those are the unlocked use cases. And this is just now, right? Think of what happens when these models get more efficient. When Google does better research on how to run miles more efficiently locally, you’ll be able to run GLM 5.2 at rapid fast speeds at cloud speed. So you can have frontier intelligence at amazing speeds and have unlimited of it. This is how you escape the control of what’s going on right now. You escape the control of the frontier labs. You escape the control of the government who’s holding everyone back. This is how you escape the I I hate to use the cringe term. It’s like the most cringe guru term on planet, but this is how you escape the matrix. This is kind of it. I also say this, and this is the most underrated part people don’t talk about. Two things. One, it’s just fun. It’s fun building your own hardware. It’s fun running local models. It’s fun. You’re allowed to have fun. Believe it or not, in this world, you are allowed to have fun. You’re allowed to experiment and tinker. It’s never been more important to experiment and tinker. So, one, you’re allowed to have fun. Two, it’s educational. You’re learning about the most cuttingedge important technology in the history of our species. Literally that it literally is. And the best way to learn about it is by getting hands-on and doing it and running yourself and seeing what the capabilities are. Right? So, it’s one helpful cuz I showed you all this use case that made my life significantly better. And two, it’s fun. You have a lot of fun doing it. It’s fun learning about the hardware. It’s fun running it. It’s fun booting it up. If you’re anything like me, you can, you know, I built this 5090. I’m running local models. Hey, I’m also playing Cyberpunk 2077 in ultra mode. So, that’s pretty fun, too, just as an aside. And then three, educational. You’re learning. Education is important. You’re learning about the most important technology ever. Now, if you set it up like I set it up where your Hermes is controlling everything, they’re tail scaled out the wazoo. You literally can build anything you want. If you just want to build a simple chat interface that looks exactly like chat GBT, you can go to your Hermes and say, “Hey, build me a chat GBT interface.” If you want to give your local model a face and a name so you can talk to it and you have your own buddy and your own pal living on your computer because hey it’s kind of fun looking down your computer knowing there’s an AI agent on it. You can do go Hermes. Hey build a face my local agent put on top it whenever it sends me a message back make it talk. You can do it. You can do whatever you want. It’s complete freedom with the most important technology ever made. I think it has never been more important with hardware becoming about to become unobtainable and the government taking away uh frontier models. I don’t think there’s ever been a more important time to get into this technology. I hope this was helpful. Leave a like, subscribe down below if you want. Let me know in the comments. You want more local content. You want me to talk about the home AI lab setup? Where where are you most confused? Where can I help out the most? All my videos are based on your feedback. Let me know down below. Hope this was helpful. See you around.