Sometimes Many Small Agents Equal One Large Agent
TL;DR: AT&T cut AI costs 90% by replacing big agents with a network of small ones. The same logic is why my planning tier runs on models small enough to sit on my desk.
A large enterprise running AI at serious volume was forced to rethink how it orchestrates all of it, and the rethink cut costs dramatically. The shape of it was simple: replace large agents doing everything with a network of smaller agents each doing one thing.
Ninety percent is the kind of number that makes you check the methodology. But the underlying logic is sound, and it matches what I’ve found building at a very much smaller scale.
The instinct that costs you money
The instinct is to use the best available model for everything, because why wouldn’t you. It’s smarter. It makes fewer mistakes. Sorting out which tasks deserve the good model is work, and skipping that work feels like an efficient decision.
It is not, and the reason isn’t only cost.
A frontier model deciding whether a filename matches a pattern is doing something a fifty-line function does perfectly, faster, deterministically, and for free. You’re not buying intelligence there — you’re buying latency and a bill. Worse, you’re buying variance: a deterministic check gives the same answer every time, and a model gives you an answer that’s right almost every time, which is a meaningfully harder thing to build on.
How this shows up in my pipeline
Project Lumbergh is tiered on purpose, and the tiers exist because different work genuinely wants different horsepower.
Planning and decomposition — taking a spec and breaking it into work that can be built in parallel without collisions — runs on local models on my own hardware. That’s a bounded, structured job with a well-defined output shape, and a model I can run in my office does it well.
Writing the actual code, and reviewing it adversarially, goes to the frontier models. That’s where the judgment is, and that’s where the money should go.
The mechanical checks in between — does this task touch the same file as its sibling, is this branch mergeable, did this daemon report in — are not models at all. They’re code. They should be code. Every time I’ve been tempted to ask a model something a function could answer, the function was the right call.
The part nobody tells you
Small models get better at your specific job much faster than they get better in general.
A local model given a tight task, a clear output format, and a couple of examples of what good looks like will hold its own on that task against something enormously more capable. The gap that matters isn’t raw capability — it’s how much of your problem you’ve bothered to specify. Specification is cheap. Frontier tokens, at volume, are not.
The honest caveat
Splitting one agent into five isn’t free. You’ve added coordination, and coordination is where distributed systems go to die. Now you need to care about what happens when one agent finishes and another never starts, or when two of them decide to edit the same thing.
I’ve spent real engineering effort on exactly those seams, and I’d expect anyone doing this to spend the same. But that’s an engineering problem — bounded, solvable, and the same shape as problems our industry has been solving for decades. The alternative is a cost curve that scales linearly with ambition, and that one doesn’t get solved by being clever.
Many small agents, each doing one thing, with real code holding them together. It’s a less impressive-sounding architecture than one enormous model that does everything. It’s also the one that survives contact with a monthly bill.
The projects, experience and opinions here are mine. AI helped me turn my notes and build records into this piece and polished it for Cairoglyphics.ai.