Plan With the Big Model, Build With the Fast One

TL;DR: A one-command change that puts your expensive thinking where thinking actually happens — plus the planning feature I wish I'd found sooner.

Two small things I picked up recently that changed how I work more than their size suggests. Both are about the same idea: the expensive part of building software is deciding what to build, not typing it.

Different models for different halves of the job

Run /model opusplan in the Claude Code CLI and it will use the larger model during plan mode — when you actually need the deep thinking — then automatically switch to the faster one once you start implementing.

That’s the whole tip. It took me an embarrassingly long time to find it, and it maps exactly onto how the work is actually shaped.

Planning is where the hard reasoning lives. What’s the right structure, what breaks if we do it this way, what did I fail to consider, is this even the correct problem. Get that wrong and no amount of implementation quality saves you.

Implementation, once the plan is genuinely good, is much closer to transcription. It’s not trivial — there’s craft in it — but it’s a fundamentally different cognitive job, and it doesn’t need the same horsepower.

Paying frontier prices for the typing while giving the thinking whatever happened to be selected is a straightforwardly bad trade, and it’s the default a lot of people are running without noticing.

Escalating the planning further

There’s also /ultraplan in the CLI, which takes your design discussion and sends it to a heavier planning session on the web.

Same principle, one notch further out. When the design is genuinely hard — the kind where you can feel yourself going in circles — it’s worth escalating the planning specifically rather than pushing forward and hoping the implementation reveals the answer.

It won’t. Implementation never reveals the answer to a design question. It just produces a very detailed, very confident, very expensive version of the wrong thing.

The pattern underneath

Both of these are instances of something I keep rediscovering: match the model to the shape of the task, not to the importance of the project.

The instinct is to use the best available model for anything that matters. But “this project matters” and “this step requires deep reasoning” are unrelated statements. A critical project has plenty of steps that are pure mechanics, and those steps don’t get better with a bigger model — they get slower and more expensive.

I run the same split at a larger scale in my pipeline: local models on my own hardware for planning and decomposition, frontier models for building and adversarial review, and plain deterministic code for anything a function can answer. /model opusplan is that same architecture, expressed as one command, for a single session.

Which is a nice reminder that most good architectural ideas are fractal. If it’s right at the level of a system, it’s usually right at the level of an afternoon.

The projects, experience and opinions here are mine. AI helped me turn my notes and build records into this piece and polished it for Cairoglyphics.ai.