The Bobs: What If the Consultants Were Actually Useful?

TL;DR: The Bobs: A panel of independent models that argue about your design before anyone builds it — and a synthesizer that turns the argument into a spec you can hand to a builder.

In Office Space, The Bobs are outside consultants brought in to review how things are being done and report back. Everyone dreads them. They are, in fairness, correct about a lot of it.

In my pipeline, The Bobs are a panel of independent AI models — frontier and local — that adversarially review a specification before anything gets built. Everyone should dread them slightly. They are, in fairness, correct about a lot of it.

The premise

There’s a secret to using AI that took me a while to properly appreciate: each improvement in the input you give it produces an exponentially better result. Garbage in, garbage out, amplified — because the model will faithfully, confidently, and at high speed build exactly what a vague spec described.

Which means the highest-leverage place to intervene is before the building starts. A defect caught in the spec costs a paragraph. The same defect caught after the build costs the build.

So the review moved upstream. The Bobs don’t review code (although they can). They review the plan.

Why a panel and not a reviewer

One model reviewing a design tends to agree with it. Not from sycophancy — from architecture. It reads the spec, builds a mental model of the intent, and evaluates the spec against the intent it just derived from that same spec. It’s grading a translation against its own translation.

Several models, working independently, don’t share a mental model. Where they agree, you’ve got something solid. Where they disagree — that’s the signal. Disagreement between competent reviewers almost always marks a genuine ambiguity in the spec: two readers derived two different intents, which means the document permits both, which means a builder could pick either.

I get more value from the disagreements than from the findings.

The two-phase loop

It runs as a cycle, and it can run more than one:

Research. Each model on the roster reviews the material independently. No shared context, no visibility into each other’s answers.

Synthesis. A separate model takes all of it — compares findings, identifies where reviewers converged, flags where they contradicted each other, and fact-checks the discrepancies and tries to confirm anything a reviewer marked low-confidence. Because reviewers are also occasionally wrong, and the failure mode of a panel is treating five outputs as five times the truth.

What comes out isn’t a list of comments. It’s a revised, research-grade specification, with the ambiguities resolved and the reasoning recorded. That document is what goes to the builder.

Then, if the work warrants it, around again.

The roster is configurable per cycle, on purpose

Different cycles want different panels. An early architectural pass wants breadth — models that think differently from each other, even at the cost of noise. A late verification pass (or a code review) wants precision.

So the roster is set per cycle rather than globally, and it draws from whatever’s available: hosted models through an aggregator, local models on my own hardware, and the CLI tools I already subscribe to, wrapped as providers. That last route matters more than it sounds — it means a full multi-model panel can run on subscriptions I’m already paying for, without a metered API key attached to every review.

Rubrics are selectable too, because “review this infrastructure change” and “review this land-use analysis” are not the same question.

Where it stands

The next piece is a read-only status view in the pipeline dashboard, so a run in progress is something I can glance at rather than query. Deliberately read-only — the dashboard is Mission Control, not a control panel. If I want to start something, I’ll do it somewhere that isn’t the place I go to see whether things are on fire.

The part that generalizes

You don’t need any of this to get the benefit. Take your next real design document, hand the exact same prompt to two different models with no shared context, and read where they disagree.

That’s the whole idea. Everything else is just making it happen automatically at three in the morning.

The projects, experience and opinions here are mine. AI helped me turn my notes and build records into this piece and polished it for Cairoglyphics.ai.