Models Mayhem
If the AI "cover model" seems a bit out of reach, it might be time to reach for the local one.
In the mid-to-late ’90s, Import Tuner magazine was a big deal (at least where I lived, a.k.a. the San Francisco Bay Area). That is to say, import racing was a thing. A huge thing, in fact. “Rice rockets” were all the rage (or the object of getting clowned on, if you thought Kerokerokeroppi stickers somehow made your car faster). And this was way before The Fast and the Furious ever hit the scene (and way before the racers we came to know and love somehow became physics-breaking secret agents).
Another big part of this automobile subculture was import models. Essentially, women posing next to cars. This part of the scene blew up like crazy. So much so that one of the main draws of Import Tuner was the models themselves. They graced every cover, along with some tricked-out featured car. Model interviews were a staple, complete with photo spreads.
When it came to models, you had your big-time models, the women who seemed to have all of the coverage. Mention one of their names at an import car meet-up and you’d have a bunch of drivers frantically looking around to see if she was actually there. But then, you also had other models. Women just starting to get their footing, not so well known on the scene. Maybe not even featured in Import Tuner, but in other ’90s magazines like TokyoPop or something.
So, why in the world am I talking about import car models? Well, all of that resonates with me when it comes to AI models. AI is huge right now, whether you’re an enthusiast or a critic (the same split people had over import cars and the models who posed with them). And where the cars had their models, AI has models of its own. ChatGPT, Claude, Gemini, and Grok represent some of the biggest ones, the ones you can mention in a room full of tech heads to start a commotion. But then, you also have local models, the ones you run on your own hardware. Perfectly fine, but not so well known. I touched on them in a previous article, and today I want to elaborate: why they can be beneficial, and a few you might want to take a gander at, should you have the hardware for it.
Your local in and out (AI, not burgers)
One of the main draws of running local models is not being beholden to whatever the big cloud providers are offering. No more “Uh oh, their servers went down!” or “They updated the model and now it’s giving me bad responses!” None of that applies to you. As long as your computer runs and you have a version of the model you like, it behaves the same way tomorrow that it did today.
Now, the big caveat I need to be upfront about is what you’ll need to run something locally. Unless you have the patience for very long wait times, you’ll want hardware with either a good chunk of RAM (32GB at minimum) or a capable GPU, meaning a graphics card with at least 16GB of VRAM (the memory built onto the card itself). Want to sweeten the deal further? A laptop running LPDDR5X, the faster memory found in most newer machines, or a GPU with 24GB of VRAM or more, opens up better and more varied options. System requirements vary by model, but these numbers give you a baseline.
Of course, money talks, and RAMageddon, the memory shortage driving RAM prices up, is pushing computer prices along with it. But if you already have capable hardware, or you can just plain afford it (a comfy income does it, no Mr. Monopoly bags of cash required), you just might be ready to step into the world of local models. And with our baseline specs in mind, let’s dive into three models made by a big three: Google, Alibaba, and OpenAI.
A model named Gemma (Gemma 4, that is)
If Google’s DeepMind can be equated to the model that’s on the cover of the magazines, then Gemma 4 is the model who isn’t appearing there (but is nonetheless very capable, and even has a few college degrees to boot). The nice thing about Gemma (and most local models) is that, depending on what PC hardware you’re running, you’re bound to find a version right for you. Want to run it on a Raspberry Pi or… a phone?? Versions E2B1 and E4B might be up your alley. Got some beefy boy hardware? Then 31B might be the model you want to go with.
Another plus is the fact that Gemma 4 ships under the Apache 2.0 license. Simply put, if you’re using it for commercial work, you are covered (legally speaking). Previous versions had Google’s own homegrown legal terms & conditions, so this is a glow up of sorts (as far as legalese can be capable of glow ups, anyhoo).
If you’re rocking E4B on your 16GB of RAM with no GPU in sight, you’re already off to a great start. Now, you might not be able to argue with it, the way you can with Gemini. But if you lose Wi-Fi in the middle of the day (or in the middle of the night, with some burning question that’s keeping you up), Gemma 4 will be there and ready to rock.
Yass Qwen (Qwen3.6)
When I say Alibaba, you might immediately think of mass-produced goods from overseas. Or you might even think of cheap PC parts via its other resale front, AliExpress. But Alibaba has also thrown its hat into the AI local model ring, via Qwen, now sitting at version 3.6. They’ve got a full-fat version just waiting for your souped-up rig to run, that being the 27B (albeit slower moving).
But the unique one you might want to try is 35B-A3B. Why’s that? It’s a mixture-of-experts version. In as plain English as your boy can muster, that means only a portion of the model wakes up for any given question. Ultimately, this means your GPU isn’t grinding through all 35 billion parameters every time you ask it something. The whole model still has to sit in memory, mind you. But the work being done? Only about 3 billion parameters’ worth. That’s what the A3B is telling you. I think you get the picture.
Whether you have a crazy PC or a crazy-underpowered one, Qwen3.6-35B-A3B (sounding like a dishwasher model) might be the model you give a spin.
And did someone say GPToss? (no, it’s gpt-oss)
Back in the day, if a popular import model showed up in some other media, like a TV show or something, it was totally unexpected (though today, it’s not far-fetched for one to be a radio show host). In a way, OpenAI did the same thing. Well known for its online-only ChatGPT, they’ve also hit the local scene with gpt-oss. They offer a 120B version if you have the bonkers PC specs to run it, but the 20B version will be the one that most run on their humble systems. Also, it has the same Apache 2.0 license as Gemma 4, and performs about on par with OpenAI’s o3-mini on common benchmarks.
What’s also of note is that you cannot run this model via ChatGPT. In fact, you can’t even access it via OpenAI’s API. This is a local-only model. All this, coming from a company that many would say built the AI house on the platform of running from your Chrome tab. Them also having a local model means that one of the biggest cloud-minded folks has taken heed of the importance of running offline.
Your computer’s next top model
So which model should you run? And what about other models out there that I haven’t even talked about? Well, that all depends on you, my good friend and model enthusiast. For one, there’s the computer consideration of specs that I mentioned above. But also, you need to find the one that feels right to you. Start off by downloading either Ollama or LM Studio. These handle the downloading of and interfacing with local models (without any crazy command line madness, unless you’re into that sorta thing). Then, start with the smallest model, poke and prod your way around it, then work your way up until things start feeling more sluggish than me after a brisket and pie binge.
Maybe those smaller, local models will never be quite as well known as their bigger, online model siblings. Maybe they’ll never wow folks by being the “cover models” either. And yes, maybe the hardware requirements will keep away any but the most intrepid or curious AI enthusiasts. But if you’ve either got the existing capable hardware or some cash burning in your wallet, diving into local AI is definitely worth your attention. The fact that you can still hum along even after a data center goes down may be all the reason you need to pull the trigger. And if some of today’s big models fade into obscurity, much like import models of yesteryear? Your local one will still be there, ready to use, possibly even seeing a second renaissance of its own.
Have your tried running a local AI model? Or are you rocking online models only? Let me know in the comments!
Quick note about all those Bs. The B stands for billion, and the number in front of it counts the model’s parameters, which are the dials it turned during training to learn how language works. The more dials, the more capable, and the more of your computer it eats. A 4B model will run on just about anything. A 31B wants a powerful rig. Click on the 1 to go back to article, you awesome person.



