What I've Learned Picking Local Models
TL;DR: The bigger model isn't always the better model, the benchmark won't tell you which one fits your job, and the only test that counts is the one you run on your own work.
If you want to play with a local AI model — one that runs on your hardware, no cloud, no per-token bill — the ecosystem has gotten genuinely good. It’s also gotten confusing in a specific way that costs people time, so here’s what I’ve learned running these things daily.
Dense beat sparse for my work, and I wouldn’t have guessed
Many model families now ship in two shapes: a dense version, where the whole model participates in every answer, and a sparse mixture-of-experts version, where a router activates only some portion of the parameters per token.
Mixture-of-experts is the more sophisticated design and it’s the one with the better story. Bigger total parameter count, comparable speed, better numbers on the headline benchmarks. On paper it’s the obvious pick.
For coding work, on my hardware, the dense version was better. Not marginally — noticeably. More consistent, fewer moments where the answer had the shape of correctness without the substance.
I’m not going to over-theorize why, because I measured one thing on one kind of task and that’s not a finding, it’s an observation. But it’s the observation that stuck with me, because I’d have bet the other way, and the only reason I know is that I tried both on actual work instead of reading about them.
Benchmarks answer someone else’s question
This is the part I’d most like people to internalize.
Public benchmarks are real and useful and I read them. They are also measuring general capability across a broad set of tasks, and you don’t have a broad set of tasks. You have your task — with your file structure, your conventions, your weird domain vocabulary, your specific idea of what a good answer looks like.
The correlation between “ranks highly in general” and “is good at your thing” is positive and much weaker than the leaderboard’s confidence implies. I’ve had models several places down a ranking beat the one above them on my work, repeatedly, in ways that made no sense until I stopped expecting the ranking to be about me.
Run your own comparison. Take three tasks you actually do, run them past three candidates, look at the output. It’s an afternoon, and it beats any amount of reading.
We eventually built a small internal app to do exactly this — feed new local models through our existing workflows and score them against results we already trust, so “should we upgrade?” has a metric-based answer instead of a vibe. That’s overkill for getting started. But the instinct behind it is not.
The bar isn’t “as good as the frontier”
The mistake I see most often is comparing a local model to the best cloud model, finding it worse, and concluding local isn’t ready.
Wrong comparison. The question is whether it’s good enough for the specific job you’d give it — and a great many jobs in a real pipeline are bounded, structured, and repetitive. Classify this. Extract that. Decompose this spec into tasks with this output shape. For work like that, a model I can run in my office is entirely adequate, and it’s adequate at zero marginal cost, with no rate limit, and no data leaving the building.
That last part isn’t a minor consideration. Some work simply shouldn’t go to a third party, and for that work the choice isn’t between local and cloud — it’s between local and not at all.
Practical notes
Your bottleneck is probably memory, not compute. What decides whether a model runs well locally is usually whether it fits comfortably, not raw speed.
Quantization is a real trade, and it’s not free. A more aggressive quantization of a larger model versus a lighter quantization of a smaller one is a genuine choice with genuine consequences. Test it. Don’t assume.
Model names age badly. The specific models I’m running today aren’t the ones I was running when I wrote this, and they won’t be the ones I’m running next quarter. The method survives; the roster doesn’t. Which is exactly why the method is the thing worth writing down.
The projects, experience and opinions here are mine. AI helped me turn my notes and build records into this piece and polished it for Cairoglyphics.ai.