Building the Middle Manager: My AI Pipeline Evolution

TL;DR: Hitting Enter four thousand times a day is how Project Lumbergh started. What I learned teaching an AI pipeline to manage its own work.

I remember the early days of Claude Code vividly. I sat there, hand hovering over the Enter key, hitting it roughly 4,000 times a day. It was exhausting, mechanical, and frankly, a bit silly. I needed a middle manager. Someone to handle the useless stuff and only surface the actual problems to me.

So, I built Project Lumbergh.

It started as a pipeline that takes a spec, runs it through multiple AIs for review, breaks it down into tasks, and sends agents out to build them. But the real story isn’t the architecture—it’s the evolution. It’s the messy, token-burning, scope-creeping journey of teaching a system to manage itself.

Here is what I learned building an AI workflow that actually works.

The “Loop” Problem

One of the first things I noticed was that when you give an AI access to your computer, it thinks a lot. More than you’d expect. Sometimes, Claude Code gets stuck in a loop—trying something, failing, getting no response, and trying again. It burns through tokens and usage limits while sitting idle.

I had to build guardrails. I started asking Claude to review its own skills and configuration files regularly. At one point, I had given him too many tools, and he was struggling to find the right one. By his own recommendation, we restructured the tools. It’s a reminder: even your agents need a clean desk.

The Feedback Loop

The breakthrough came when I stopped trying to write code manually and started writing the process for writing code. I’ve been a career IT person, but on the networking and hardware side, never as a software dev. Writing smart prompts to describe the process was my solution to getting code built (well).

I’d discuss a project with Claude until we had a solid architectural spec. Then, I’d submit that plan to GPT and a local model on my server Sparky in parallel. Claude would arbitrate the feedback, creating a final engineering document. That document would become instructions to Claude Code, which would write the code and send it to Codex for QA.

Codex would point out bugs. Claude Code would fix them. They’d go around and around.

In one instance, they went 10 rounds. Eventually, they worked it out, and I got working code. It wasn’t magic; it was just persistent, multi-source verification.

The “Lumbergh” Stats

Fast forward to April. I asked Claude for stats on the pipeline. The numbers surprised me:

  • 194 features built in 28 days.
  • ~457,000 lines of code and documentation.
  • 1,231 git commits in a month (a normal dev does 100-200).
  • Zero human hand-coding for the bulk of it.

I had built a small team of AIs that build software for me. In one month, they shipped 194 features and wrote about half a million lines of code, working from one-paragraph descriptions I wrote.

But it wasn’t always smooth. I added a “side door” for tightly-spec’d builds that didn’t need deep analysis, like applying patches, sending them straight to QA. I fed the pipeline fixes to fix itself. Each time, it got a little better. The rate of intakes went up. Sometimes we fed non-Lumbergh projects as tests. It was a long bug-fixing phase, but the overall concepts worked.

The New Guard: “Chief” and “Squasher”

Recently, I’ve been spending more time building tools and agents, making them modular so I can swap models easily. I have a new agent, “Chief,” who operates using Anthropic’s Claude, and wakes on a timer. He checks the pipeline, cleans up the .git if needed, and files bugs he finds while troubleshooting. He starts as a Sonnet-level probe—fast and cheap for mechanical health checks. If he finds an issue needing deeper reasoning, he auto-switches to Opus or Fable, then falls back when done. That saves real tokens.

Then there’s “Squasher.” I had been building so fast, logging bugs to a tracker for future fixing, that I became overwhelmed by the ever-increasing amount of bugs. In one week, Squasher squashed 21 bugs. All by himself. Some were simple patches; others were sent to a panel of AIs for independent triage. In August alone, 317 bug fixes shipped through the pipeline, which is a long way from the mess I was in a couple of months ago.

The Cost of Autonomy

I’m using frontier AI models under subscriptions; for example, I use the Max plan from Anthropic for Claude work ($200/mo). I keep an eye on usage. In one session with Claude Fable, the stats were interesting:

  • Total cost: $429.42 (had I paid via API key)
  • Duration: 8h 36m (API)
  • Code changes: 318 lines added, 129 removed

Takeaway: the subscription is worth it. But it’s a reminder that autonomy has a price. You have to watch the meter.

What I’d Do Differently

  1. Don’t be afraid to ask an AI how to build an AI. Use a conversation with your AI of choice to settle the cloud of uncertainty and pick a path forward.
  2. Verify, don’t trust. I use AI models from other frontier labs to perform code analysis and documentation for diversity. It’s helpful to have an independent model that didn’t write the code look at the code. Remember when teachers used to let us grade our own homework? Yeah, neither do I.
  3. Build the bicycle. A former coworker told me that sometimes we get too busy running to stop and get on the bicycle. I stopped running long enough to build a bicycle, and enjoyed it so much that I kept building it into a bicycle factory. Now I can concentrate on feeding it designs of any kind, instead of manually managing multiple models.

The pipeline isn’t perfect, but neither am I. It’s still learning. But it’s working. And for the first time, I’m not just running; I’m riding my bicycle around my factory, building things everywhere.

I’ve been saying this for a while. This piece was pulled together by AI from 44 of my social media posts (10/2025 - 08/2026) and polished for Cairoglyphics.ai.