The QA Loop That Runs Without Me
TL;DR: When Claude Code finishes writing, Codex reviews it, Claude reads the findings and fixes them, and they go around again — up to five rounds before it stops and asks for me.
Here’s what happens now when Claude Code finishes writing a chunk of my software, and I want to describe it plainly because the shape of it is more important than the tools involved.
- Claude Code finishes. It does not declare victory.
- Codex reviews the work and reports findings.
- Claude Code reads those findings, decides what’s real, and makes changes.
- Back to step 2.
- Up to five rounds. Then it halts and asks for me.
That’s it. That’s the loop. I’m not in it until the end.
Why the loop and not just a review
A single review pass is worth a lot — I’ve written elsewhere about how much one reviewer buys you. But a single pass has an obvious gap: findings get reported and then somebody has to act on them, and if that somebody is you, you’ve just moved the bottleneck rather than removed it.
Closing the loop is what makes it useful. The reviewer’s findings go straight back to the builder, the builder responds, and the reviewer looks again at what actually changed. By round three you’re usually looking at genuine disagreements rather than obvious defects, and genuine disagreements are exactly the thing worth a human’s attention.
Why five rounds and then stop
The cap is doing real work, and I’d encourage anyone building something like this to pick a number and enforce it.
Without a cap you get one of two bad outcomes. Either the loop converges and then keeps spinning, burning tokens polishing something that was finished two rounds ago. Or — worse — it doesn’t converge, because the reviewer wants something the builder can’t give it, and they trade increasingly baroque attempts at the same problem until you notice.
Five rounds without convergence isn’t a failure of the loop. It’s the loop successfully telling me there’s a disagreement it can’t resolve, which is precisely when I want to be interrupted. A stall is information.
Builder is never reviewer
The rule underneath all of this: the model that writes the code is never the model that reviews it.
A model reviewing its own work is agreeable in exactly the places you need it to be difficult. It made the choices. It likes the choices. It will find typos and miss the architectural decision it talked itself into forty lines earlier.
Different model, no stake in the design, no memory of the reasoning that produced it. That’s what you’re buying.
If you want this today
You don’t need a pipeline to get most of the value. If you’re running Claude Code, the Codex plugin lets you have Codex do a QA pass on your code before completion — it’s about as close to free as an improvement gets, and I cannot recommend it strongly enough.
Start there. Get used to reading review output and noticing what it catches that you didn’t. The full loop is a natural thing to want once you’ve seen a review land a finding that would have cost you an afternoon.
Which is the honest reason I built it: not because it was clever, but because I got tired of being the courier between two models that were perfectly capable of talking to each other.
The projects, experience and opinions here are mine. AI helped me turn my notes and build records into this piece and polished it for Cairoglyphics.ai.