Engineering Notes · agent operations · code review automation
Every reviewer downstream of my build pipeline used to spend part of its attention on code that never needed to exist. Now one step cuts that code first, proves the tests still pass after every cut, and only then lets the reviewers loose on what's left.
My AI build pipeline runs each feature as a plan, split into tasks, implemented by parallel TDD workers, then checked by a stack of reviewers before it can commit. I've written about pieces of that stack before: the test gate that won't let a commit through without recorded evidence, and the hard rules enforced as hooks instead of prose. What that stack didn't have, for a long time, was anything looking at whether the code should exist in the first place.
An AI worker asked to implement one task will reliably implement it. It will also, just as reliably, reach for the shape it's seen most often: a config flag for a value that never varies, an interface with exactly one implementation, a helper that duplicates one three files over, a comment narrating the line below it. None of that is wrong, exactly. It's just cost that didn't buy anything, and it compounds — a plan with six tasks run by six workers can pick up six unrelated instances of it, and nobody who wrote one saw the other five.
My correctness reviewers — codegraph, graphiti, mutation testing, a three-agent security quorum — were reading that bloat along with everything else, and spending real attention on it that belonged on actual bugs. So I gave the pipeline a step whose only job is to find that class of thing and remove it, before anyone downstream has to notice it.
It's a hook in my orchestration harness (HCF v2), enrolled at post-implementation, and it runs early in that stage — after every TDD worker on the plan has reported its task complete, and before any of the reviewers that judge correctness:
It works on the whole plan's diff, not one task at a time — a duplicate helper is only visible once you can see every task's changes at once — and it runs once, near the front of the stage, so every reviewer after it inherits a diff that's already been through one pass of cutting.
For every changed function, class, or config block in the diff, it works down a fixed ladder. The first rung that fails is where the finding lives:
new.Some things are never a cut, no matter the cost. Validation, error handling, security checks, accessibility — those don't get judged against this ladder at all. This step trims what didn't need to be written, never what makes the feature correct or safe.
Finding something worth cutting isn't enough — the pass has to prove the cut didn't change behavior, and the only proof that counts is the test suite. So each pass through the diff runs a small loop: back the touched files up, apply every cut it found, run the tests. Green, and it looks at the diff again — a cut can expose another one, like a config flag disappearing and leaving its lone reader as a new single-caller candidate. Red, and the cut that broke it gets reverted from the backup and written down as an advisory note instead — a human decides from there, the pass doesn't fight the tests to force a cut through. It stops after three passes either way, whatever's left becomes advisory.
The revert mechanism is never git. The working tree at this point holds every task's uncommitted changes from every worker on the plan, not just this step's own edits. A git stash or git reset to undo one bad cut would just as happily wipe everyone else's work — which is exactly what happened in the incident that also produced git-tree-guard. So the only way this step undoes anything is copying its own plain-file backups back over the file it just touched. Nothing at the git layer, ever.
Every run leaves a paper trail — which cuts landed, which got reverted and why, which were left as advisory, and the final test command and result — written to the plan's own directory. If there's no record, there's no proof the pass ran at all, so it always writes one, even on a plan where it found nothing to cut.
Nothing downstream had to change to benefit from this. Codegraph, graphiti, mutation testing, the security quorum — they all read the same diff they always did, it's just smaller and already proven green by the time they see it. A reviewer whose job is "is this correct" stops spending its budget on "did this need to be this big," and a mutation-testing pass in particular gets a fairer number out of it — surviving mutants in code that shouldn't have existed were never a signal worth chasing.
The size of the diff a reviewer has to read is itself a signal, and it was a noisy one before this step existed. Trimming it before review, not after, is the only way the signal stays worth reading.
post-implementation hook at order 25I use my tooling predominantly on Mage-OS (Adobe Commerce / Magento) e-commerce projects, and my own AI Booking Agent. If you have not swapped to Mage-OS yet, you are falling behind ;)