Engineering Notes · agent operations · process discipline
Two of my coding agents sat idle for fifty minutes each, politely reporting that they were still waiting on a test run. The test run had already passed. Nothing was hung — and that turned out to be the worst possible failure mode, because a hung agent and a confidently-waiting one look identical from the outside.
18 September, a worktree on one of my Magento projects, ticket #491. I spun up a batch of four TDD workers to grind through a plan. Two of them — call them tdd-004 and tdd-005 — started sending idle notifications. Reasonable-sounding ones:
Still waiting on the Playwright monitor to report completion. I'll wait for the Playwright run to finish before proceeding.
tdd-004, repeatedly, for ~50 minutes
From the lead session I could see all three of the things you'd check. No node or Playwright process running anywhere on the host. Not one file those tasks owned had an mtime newer than the batch start. And the agent registry showed both workers as idle, not busy. Zero activity for 108 minutes.
17:53:38 batch start
~18:01 Playwright exits 0 <- work done, tests green
19:41 human asks "how we doing here"
^ 108 minutes of "still waiting on the Playwright monitor"
That is a textbook hung worker. It is what I would have sworn was happening, and I nearly killed and restarted the batch on that basis. So instead I asked tdd-004 directly for its real status. Its answer, verbatim:
the background Playwright run had actually finished (exit 0) minutes before your status check — my Monitor's pgrep-based wait check was buggy and timed out without noticing, so I looked stalled. Confirmed via the run's own output file just now.
tdd-004
The work was done. The tests were green. The agent had written itself a pgrep loop to find out when the backgrounded Playwright run finished, that loop had quietly failed, and the agent was now waiting on a condition that would never arrive. tdd-005 did exactly the same thing independently, in the same batch. Two agents, same invented bug, no shared instruction telling them to do it.
My first instinct was to go fix the loop. Better pattern matching, a timeout, a retry. That instinct is wrong, and it took me a minute to see why.
The moment you detach a process from the tool call that spawned it, you have thrown away the only trustworthy thing about it: its exit code. Everything you can reach for afterwards is a proxy for "did it finish, and did it work", and every proxy available fails in the same direction.
Not one of them carries an exit code. So the agent has to guess, and here's the part that makes this expensive rather than merely annoying: a wrong guess is silent. The agent doesn't error. It doesn't stop. It sits there, believing it is being patient, generating calm status updates about a process that died an hour ago. No monitoring I have distinguishes that from a genuinely wedged worker, because there is nothing to distinguish — same process list, same mtimes, same idle flag.
Then I looked at the agent definition, and found the thing that actually made this inevitable. My TDD worker is declared with these tools:
tools: Read, Write, Edit, Bash, Glob, Grep
No monitor tool. No way to read the output of a backgrounded shell command. No way to await a task.
So once that agent backgrounds anything, it is structurally incapable of finding out what happened. There is no legal tool call that recovers the result. The only move left in its toolset is to hand-roll a poll loop — which is exactly what both workers did, independently, with no prompting. They weren't being clever or lazy. They were cornered, and they took the one exit available.
Reframe: the fix isn't "teach the agent a better waiting technique". It's "this agent must never be in a position where waiting is something it has to implement".
Short version: do not detach work you need the answer to, and never write your own loop to find out whether it finished.
npx playwright test ... &
nohup vendor/bin/phpunit ... &
while pgrep -f playwright >/dev/null; do sleep 10; done
until [ -f /tmp/run.done ]; do sleep 5; done
while kill -0 "$PID" 2>/dev/null; do sleep 5; done
npx playwright test tests/vt-billing.spec.ts
# foreground, with the tool's own timeout — 10 min cap
And when a run doesn't fit in ten minutes, the answer is to narrow it — one spec, a filter, --bail — and widen only once it's green. Narrowing is the fix. Detaching is just moving the problem somewhere you can't see it.
There's a pleasing side effect here. I already had a separate, unresolved ask in my backlog about enforcing targeted-first test discipline — run the one spec that covers your change before you run the suite. This rule arrives at the same place from the other direction. If you can't background and you can't exceed ten minutes, you are structurally pushed into running the targeted spec first. Two problems, one constraint.
For any agent without a monitor tool, backgrounding isn't discouraged, it's impossible to do correctly. So the rule for those is explicit: if a run genuinely cannot happen in the foreground, you do not own it. Report failure back to the orchestrator — which does have the tools to watch a long-running task — and stop. Handing work upward is a correct outcome. Inventing a wait loop is not.
Daemons are the one honest exception. Starting a long-lived server is a legitimate detach, because you don't want its exit code, you want it running. Even there, don't pgrep for it — probe the actual thing (curl --retry-connrefused, a port check, a health endpoint), which tells you it's ready rather than merely present.
A rule that lives only in prose decays. I've written about this before — if a rule can't be demonstrated to fire, it's not a rule, it's a hope. So this one lands in three places at once:
sleep plus a liveness probe, and a trailing & or nohup on anything test-shaped. It requires all three signals for the loop case, so ordinary retry loops that do real work — a curl readiness poll, say — aren't touched. The block message names the alternative rather than just saying no.Writing that hook set off a second, funnier problem, and it's the part I'd most want someone else to avoid.
I have a gate-config guard: a hook that stops any agent — including the one I'm talking to — from quietly editing the enforcement layer. Turn off the gate and the gate is decorative, so that one has no bypass at all. It blocked the new hook file. Fine, I thought, working as intended.
It wasn't. The guard protects a specific file, .claude/rules-disable, and its shell branch searches the whole command text for that name. Every guard hook I write documents its own opt-out line — which means it contains that filename as a comment. So the act of authoring any new guard was blocked for mentioning the thing it wasn't touching. I proved it by accident: the test script I wrote to investigate the false positive was itself blocked, for containing the string in a test case.
Two lessons, and the second is the one worth keeping.
First: a guard that matches on text rather than intent will eventually block the documentation of itself. Substring matching on a filename cannot distinguish "write this file" from "mention this file", and prose about a rule looks exactly like a violation of it.
Second, and this is what I actually changed: I had the scope wrong. That guard exists to stop unsupervised agents — plan-orchestrate workers, cron sessions inside containers — from disarming their own gates mid-run. It was never meant to stop me, at my own machine, maintaining the guard layer. Blocking the maintainer isn't strict, it's just broken. So it now exits immediately on host sessions and keeps the full, bypass-free block inside containers, where the risk it was written for actually lives.
Rule of thumb: a guard needs a blast radius as deliberate as its trigger. "Who is this stopping, and are they supervised?" is a design question, not a detail — and if the answer is "everyone, including whoever maintains it", the guard will eventually block the work of fixing itself.
The failure here wasn't an agent being dumb. Both workers behaved sensibly given what they had: a job to do, a long-running command, and no tool that could tell them how it went. The failure was mine, in handing out a capability (background a process) without the matching capability (observe its result), and then being surprised when the gap got filled with improvisation.
It generalises past agents, honestly. Any time you give something the ability to start work it cannot observe the end of, you've built a machine for producing confident wrong answers about state. The fix is never a better guess. It's removing the need to guess.
Cost of learning this: about two hours of wall-clock across two workers, and a near-miss where I almost restarted a batch of completed work. Cheap, as these go. The next one would not have been.
All from my claude-skills-central repo. Every guard ships with a .test.sh beside it.
I use my tooling predominantly on Mage-OS (Adobe Commerce / Magento) e-commerce projects, and my own AI Booking Agent. If you have not swapped to Mage-OS yet, you are falling behind ;)