I Scaled Writing Code. Then I Had to Scale Finding Problems.

TL;DR: Automating the building half of software makes the finding-problems half the bottleneck. Here's what it looks like to point a second system at your own codebase overnight.

There’s a consequence of automating your building that nobody puts in the pitch: you also have to automate your finding.

Once agents are producing work faster than you can personally read it, the rate at which you can discover problems becomes the thing that governs how fast you actually ship. Not the writing. The looking. And the looking had stayed exactly as manual as it had always been — me, reading, one thing at a time.

I’d scaled one half of the job and left the other half where it was. So I went looking for a way to scale it.

What I pointed at it

AntiGravity 2.0, from Google/Gemini. About ten minutes of setup — which is the first thing worth reporting, because I’d budgeted an afternoon.

The job I gave it: map my pipeline’s code, its history, its logs, and its existing bug records, and go find problems. Not just mechanical ones — behavioral ones. Things that compile and pass and are nonetheless wrong. It runs overnight and produces reports with proposed fixes.

Two details of the setup matter more than the tool choice:

It runs on my own hardware. The GPU in Sparky does the work. For a job that means ingesting an entire codebase and its history, that’s the difference between a routine overnight run and a line item.

It should have run read-only. It didn’t. I gave an analysis tool access to my primary working tree, and it checked that tree onto a branch of its own and left it there — which correctly caused every subsequent pipeline merge to refuse, twenty-nine times, until I reset it. That is my mistake, not the tool’s: there was no reason for an analysis tool to have write access, and “no reason to need it” is sufficient reason not to grant it. If it produces a fix, I want the fix to arrive as a proposal that goes through the same review as anything else — not as a change that already happened.

Mechanical versus behavioral

That distinction is the reason this was worth doing at all.

Mechanical problems are the ones tools have caught for years. Wrong type, unreachable branch, resource not released. Valuable, well-solved, and mostly already handled by things you’re running.

Behavioral problems are the interesting ones. This function does what it says and what it says is wrong for how it’s used. These two components each behave correctly and cannot both be correct at once. This value is checked here and assumed there. That class of defect is invisible to a linter, survives a test suite happily, and is exactly what a system that has read your history and logs alongside your code has a shot at seeing.

Code plus history plus logs is a genuinely different input than code alone. History tells you what people kept changing. Logs tell you what actually happens rather than what was intended. A reviewer with all three is reasoning about a living system rather than a snapshot.

The thing I’d tell someone starting this

Point the second system at the same code you’re already confident about.

The instinct is to aim analysis at the scary parts — the module nobody understands, the thing you’ve been avoiding. But you already know those are risky, and finding problems there confirms something you’d have told me for free.

The value is in the parts you’re not worried about. That’s where an outside reviewer earns its keep, because that’s where you’ve stopped looking. Everything I’ve built in the last year keeps teaching me the same lesson in different costumes: the risk isn’t in the thing you’re watching.

I’ll write more when I’ve worked through the results. For now the setup cost was ten minutes, the runs are overnight and free, and I’ve stopped being the only thing standing between my codebase and the problems in it.

The projects, experience and opinions here are mine. AI helped me turn my notes and build records into this piece and polished it for Cairoglyphics.ai.