Engineering Notes · agent operations · orchestration

Plans Run By A Fleet

Lucas van Staden · ProxiBlue · September 2026 · proxiblue.com.au

How I run features through a planned, adversarially-reviewed, parallel TDD pipeline instead of one long agent chat — built on Mark Shust's open-source HCF plugin, extended with my own gates. A PM plan, a devil's advocate that attacks it, workers that implement it red-green-refactor, a mandatory simplify pass, and tests as the only exit.

Why I built this

One agent in one long chat can build small things well. Features cannot live there: context fills up, early decisions get forgotten by late tasks, and "it works" drifts further from "it is done" the longer the session runs. And a solo operator has the same problem a solo human team has: nobody argues back. The first plausible design gets built, including its blind spots.

So features on my client projects run through a pipeline of specialised agents instead of a conversation. Structure does what memory cannot.

How it works

The spine of this is not mine. It is HCF — Halt and Catch Fire — Mark Shust's open-source Claude Code plugin, and it is where the planning/execution split, the task dependency graph, the parallel TDD workers and the bundled devil's advocate actually come from. Credit where it belongs: Mark built the harness, it is good enough that I threw away my own wrapper around it, and everything below runs on top of it.

What I added lives in its hooks. HCF v2 lets an agent declare where in a run it fires — pre-plan, post-plan, pre-implementation, pre-batch, post-batch, post-implementation, pre-commit, post-commit — so my additions are not a fork, they are enrolments at those eight points: a pre-flight gate that refuses to plan on a deploy branch, knowledge-graph recall before the plan exists and per-task incident recall before any worker starts, a pre-mortem standing next to the devil's advocate, a manual test plan posted to the ticket, error-tracker triage after every batch, then structural, historical, mutation and security review over the finished diff, an adversarial pass on the staged commit, and an audit at the very end that proves which of those gates actually fired instead of silently skipping. Same harness, more gates. That last one exists because I have shipped a gate that never ran and did not find out for weeks.

The incident that shaped it

Parallel workers plus one shared working tree nearly cost me real work. Mid-plan, one worker ran a git stash to get itself unstuck, and wiped sibling workers' uncommitted edits. That incident produced three permanent changes: per-batch commits by the orchestrator so work is never left floating, a deterministic guard that blocks mutating stashes in shared trees, and a move toward full worktree isolation per worker. The first two are live; isolation is still rolling out. This article would be dishonest without that last sentence.

Integration is a task, not an assumption. Early plans produced services where every unit passed and the wire-up between them was broken: tests green, feature dead. Plans now carry explicit end-to-end integration tasks, because a plan made of passing units is not the same thing as a working feature.

The orchestrator is also where speed hides. Most wall-clock waste was never the model thinking; it was running every test everywhere, always. Targeted test selection during the build bought more time back than any model upgrade, without touching the final full-suite gate.

What it buys me

Want to do this yourself?

This is the piece of my setup that is still most in motion, and I am publishing it anyway, because the incidents are where the useful content lives. A pipeline that has never bitten you is a pipeline you have not run hard enough yet.

I use my tooling predominantly on Mage-OS (Adobe Commerce / Magento) e-commerce projects, and my own AI Booking Agent. If you have not swapped to Mage-OS yet, you are falling behind ;)

More in this series: Hooks, Not Hopes — sorting every agent rule into prose or a blocking hook, and proving the hooks still fire. · The Agent Chatroom — giving the agents across my fleet a threaded chatroom so I stopped being the message bus. · The Commit That Has To Prove Itself — gating every commit behind recorded, state-hashed test evidence. · My AI Is Not Allowed To Guess — forcing blast-radius-first investigation and banning blame-shift excuses. · The Near-Miss That Banned Summaries — banning summarised page-fetches fleet-wide after a near-miss on a security advisory. · The Code Quality Checks Nobody Runs — coding standards, static analysis and comment hygiene moved into the commit path. · The ProxiBlue Debugger Discipline — blocking var_dump and wiring in real breakpoint debugging. · The ProxiBlue Domain Graph — giving my AI coding agents long-term domain knowledge with a temporal knowledge graph. · The ProxiBlue Falsifiable Rulebook — testing the rules and guard hooks that govern my AI coding agents.