Engineering Notes · agent operations · code review automation

Smaller, Then Green

Lucas van Staden · ProxiBlue · September 2026 · proxiblue.com.au

Every reviewer downstream of my build pipeline used to spend part of its attention on code that never needed to exist. Now one step cuts that code first, proves the tests still pass after every cut, and only then lets the reviewers loose on what's left.

Why I built this

My AI build pipeline runs each feature as a plan, split into tasks, implemented by parallel TDD workers, then checked by a stack of reviewers before it can commit. I've written about pieces of that stack before: the test gate that won't let a commit through without recorded evidence, and the hard rules enforced as hooks instead of prose. What that stack didn't have, for a long time, was anything looking at whether the code should exist in the first place.

An AI worker asked to implement one task will reliably implement it. It will also, just as reliably, reach for the shape it's seen most often: a config flag for a value that never varies, an interface with exactly one implementation, a helper that duplicates one three files over, a comment narrating the line below it. None of that is wrong, exactly. It's just cost that didn't buy anything, and it compounds — a plan with six tasks run by six workers can pick up six unrelated instances of it, and nobody who wrote one saw the other five.

My correctness reviewers — codegraph, graphiti, mutation testing, a three-agent security quorum — were reading that bloat along with everything else, and spending real attention on it that belonged on actual bugs. So I gave the pipeline a step whose only job is to find that class of thing and remove it, before anyone downstream has to notice it.

When it runs

It's a hook in my orchestration harness (HCF v2), enrolled at post-implementation, and it runs early in that stage — after every TDD worker on the plan has reported its task complete, and before any of the reviewers that judge correctness:

25simplify-passcuts over-built code from the whole plan's diff, proves tests stay green
30codegraph-reviewerchecks the diff against the code graph for broken or missed wiring
40graphiti-reviewerchecks the diff against prior decisions and incidents in the same area
45mutation-testerruns mutation testing on the plan's changed files, gates on a minimum score
70security-quorumthree-agent adversarial, defensive, and static-analysis vote

It works on the whole plan's diff, not one task at a time — a duplicate helper is only visible once you can see every task's changes at once — and it runs once, near the front of the stage, so every reviewer after it inherits a diff that's already been through one pass of cutting.

What it cuts

For every changed function, class, or config block in the diff, it works down a fixed ladder. The first rung that fails is where the finding lives:

  1. Does this need to exist at all? An unused branch, a parameter nothing ever passes, a flag with exactly one caller.
  2. Does the codebase already do this? It greps for a near-duplicate helper before accepting a new one as necessary.
  3. Would the framework already do this? A hand-rolled loop standing in for a Magento core service or a stdlib collection method.
  4. Is the abstraction load-bearing? An interface with one implementation and no plugin/DI reason to expect a second; a factory wrapping a bare new.
  5. Is this the minimum the task actually asked for — not a guessed future requirement?
  6. Do the comments earn their place? A comment survives only if it states a constraint the code can't show on its own — not one that narrates the next line or talks to the reviewer instead of the next reader.

Some things are never a cut, no matter the cost. Validation, error handling, security checks, accessibility — those don't get judged against this ladder at all. This step trims what didn't need to be written, never what makes the feature correct or safe.

Cut, then prove green

Finding something worth cutting isn't enough — the pass has to prove the cut didn't change behavior, and the only proof that counts is the test suite. So each pass through the diff runs a small loop: back the touched files up, apply every cut it found, run the tests. Green, and it looks at the diff again — a cut can expose another one, like a config flag disappearing and leaving its lone reader as a new single-caller candidate. Red, and the cut that broke it gets reverted from the backup and written down as an advisory note instead — a human decides from there, the pass doesn't fight the tests to force a cut through. It stops after three passes either way, whatever's left becomes advisory.

The revert mechanism is never git. The working tree at this point holds every task's uncommitted changes from every worker on the plan, not just this step's own edits. A git stash or git reset to undo one bad cut would just as happily wipe everyone else's work — which is exactly what happened in the incident that also produced git-tree-guard. So the only way this step undoes anything is copying its own plain-file backups back over the file it just touched. Nothing at the git layer, ever.

Every run leaves a paper trail — which cuts landed, which got reverted and why, which were left as advisory, and the final test command and result — written to the plan's own directory. If there's no record, there's no proof the pass ran at all, so it always writes one, even on a plan where it found nothing to cut.

What it buys the reviewers after it

Nothing downstream had to change to benefit from this. Codegraph, graphiti, mutation testing, the security quorum — they all read the same diff they always did, it's just smaller and already proven green by the time they see it. A reviewer whose job is "is this correct" stops spending its budget on "did this need to be this big," and a mutation-testing pass in particular gets a fairer number out of it — surviving mutants in code that shouldn't have existed were never a signal worth chasing.

Want to do this yourself?

The size of the diff a reviewer has to read is itself a signal, and it was a noisy one before this step existed. Trimming it before review, not after, is the only way the signal stays worth reading.

I use my tooling predominantly on Mage-OS (Adobe Commerce / Magento) e-commerce projects, and my own AI Booking Agent. If you have not swapped to Mage-OS yet, you are falling behind ;)

More in this series: Plans Run By A Fleet — how features run through a planned, adversarially-reviewed, parallel TDD pipeline instead of one long agent chat. · Hooks, Not Hopes — sorting every agent rule into prose or a blocking hook, and proving the hooks still fire. · The Agent Chatroom — giving the agents across my fleet a threaded chatroom so I stopped being the message bus. · The Commit That Has To Prove Itself — gating every commit behind recorded, state-hashed test evidence. · My AI Is Not Allowed To Guess — forcing blast-radius-first investigation and banning blame-shift excuses. · The Near-Miss That Banned Summaries — banning summarised page-fetches fleet-wide after a near-miss on a security advisory. · The Code Quality Checks Nobody Runs — coding standards, static analysis and comment hygiene moved into the commit path. · The ProxiBlue Debugger Discipline — blocking var_dump and wiring in real breakpoint debugging. · The ProxiBlue Domain Graph — giving my AI coding agents long-term domain knowledge with a temporal knowledge graph. · The ProxiBlue Falsifiable Rulebook — testing the rules and guard hooks that govern my AI coding agents.