Engineering Notes · agent operations · process discipline

The Hour My Agents Spent Waiting For Nothing

Lucas van Staden · ProxiBlue · October 2026 · proxiblue.com.au

Two of my coding agents sat idle for fifty minutes each, politely reporting that they were still waiting on a test run. The test run had already passed. Nothing was hung — and that turned out to be the worst possible failure mode, because a hung agent and a confidently-waiting one look identical from the outside.

108 minutes of nothing

18 September, a worktree on one of my Magento projects, ticket #491. I spun up a batch of four TDD workers to grind through a plan. Two of them — call them tdd-004 and tdd-005 — started sending idle notifications. Reasonable-sounding ones:

Still waiting on the Playwright monitor to report completion. I'll wait for the Playwright run to finish before proceeding.

tdd-004, repeatedly, for ~50 minutes

From the lead session I could see all three of the things you'd check. No node or Playwright process running anywhere on the host. Not one file those tasks owned had an mtime newer than the batch start. And the agent registry showed both workers as idle, not busy. Zero activity for 108 minutes.

17:53:38  batch start
~18:01    Playwright exits 0   <- work done, tests green
19:41     human asks "how we doing here"
          ^ 108 minutes of "still waiting on the Playwright monitor"

That is a textbook hung worker. It is what I would have sworn was happening, and I nearly killed and restarted the batch on that basis. So instead I asked tdd-004 directly for its real status. Its answer, verbatim:

the background Playwright run had actually finished (exit 0) minutes before your status check — my Monitor's pgrep-based wait check was buggy and timed out without noticing, so I looked stalled. Confirmed via the run's own output file just now.

tdd-004

The work was done. The tests were green. The agent had written itself a pgrep loop to find out when the backgrounded Playwright run finished, that loop had quietly failed, and the agent was now waiting on a condition that would never arrive. tdd-005 did exactly the same thing independently, in the same batch. Two agents, same invented bug, no shared instruction telling them to do it.

Why it wasn't a bug in the wait loop

My first instinct was to go fix the loop. Better pattern matching, a timeout, a retry. That instinct is wrong, and it took me a minute to see why.

The moment you detach a process from the tool call that spawned it, you have thrown away the only trustworthy thing about it: its exit code. Everything you can reach for afterwards is a proxy for "did it finish, and did it work", and every proxy available fails in the same direction.

pgrep -f <pattern>
Matches your own grep. Matches an unrelated process that reused the name. And it reports "gone" the instant the process exits — which is also exactly what a crash looks like.
kill -0 $PID
Tells you alive or dead. Never tells you pass or fail. A suite that failed every test is just as dead as one that passed.
[ -f done.marker ]
The marker appears when the writer decides to write it, not when the work ended. A crash writes nothing, so you wait forever — which is precisely the shape of this incident.
tail -f output.log
A partial write reads as "still going". A truncated file reads as a clean finish. You are parsing prose to recover a number you already threw away.

Not one of them carries an exit code. So the agent has to guess, and here's the part that makes this expensive rather than merely annoying: a wrong guess is silent. The agent doesn't error. It doesn't stop. It sits there, believing it is being patient, generating calm status updates about a process that died an hour ago. No monitoring I have distinguishes that from a genuinely wedged worker, because there is nothing to distinguish — same process list, same mtimes, same idle flag.

The structural part I'd missed

Then I looked at the agent definition, and found the thing that actually made this inevitable. My TDD worker is declared with these tools:

tools: Read, Write, Edit, Bash, Glob, Grep

No monitor tool. No way to read the output of a backgrounded shell command. No way to await a task.

So once that agent backgrounds anything, it is structurally incapable of finding out what happened. There is no legal tool call that recovers the result. The only move left in its toolset is to hand-roll a poll loop — which is exactly what both workers did, independently, with no prompting. They weren't being clever or lazy. They were cornered, and they took the one exit available.

Reframe: the fix isn't "teach the agent a better waiting technique". It's "this agent must never be in a position where waiting is something it has to implement".

The rule

Short version: do not detach work you need the answer to, and never write your own loop to find out whether it finished.

Banned

npx playwright test ... &
nohup vendor/bin/phpunit ... &
while pgrep -f playwright >/dev/null; do sleep 10; done
until [ -f /tmp/run.done ]; do sleep 5; done
while kill -0 "$PID" 2>/dev/null; do sleep 5; done

Instead

npx playwright test tests/vt-billing.spec.ts
  # foreground, with the tool's own timeout — 10 min cap

And when a run doesn't fit in ten minutes, the answer is to narrow it — one spec, a filter, --bail — and widen only once it's green. Narrowing is the fix. Detaching is just moving the problem somewhere you can't see it.

There's a pleasing side effect here. I already had a separate, unresolved ask in my backlog about enforcing targeted-first test discipline — run the one spec that covers your change before you run the suite. This rule arrives at the same place from the other direction. If you can't background and you can't exceed ten minutes, you are structurally pushed into running the targeted spec first. Two problems, one constraint.

The subagent clause

For any agent without a monitor tool, backgrounding isn't discouraged, it's impossible to do correctly. So the rule for those is explicit: if a run genuinely cannot happen in the foreground, you do not own it. Report failure back to the orchestrator — which does have the tools to watch a long-running task — and stop. Handing work upward is a correct outcome. Inventing a wait loop is not.

Daemons are the one honest exception. Starting a long-lived server is a legitimate detach, because you don't want its exit code, you want it running. Even there, don't pgrep for it — probe the actual thing (curl --retry-connrefused, a port check, a health endpoint), which tells you it's ready rather than merely present.

Making it stick

A rule that lives only in prose decays. I've written about this before — if a rule can't be demonstrated to fire, it's not a rule, it's a hope. So this one lands in three places at once:

A guard that ate its own tail

Writing that hook set off a second, funnier problem, and it's the part I'd most want someone else to avoid.

I have a gate-config guard: a hook that stops any agent — including the one I'm talking to — from quietly editing the enforcement layer. Turn off the gate and the gate is decorative, so that one has no bypass at all. It blocked the new hook file. Fine, I thought, working as intended.

It wasn't. The guard protects a specific file, .claude/rules-disable, and its shell branch searches the whole command text for that name. Every guard hook I write documents its own opt-out line — which means it contains that filename as a comment. So the act of authoring any new guard was blocked for mentioning the thing it wasn't touching. I proved it by accident: the test script I wrote to investigate the false positive was itself blocked, for containing the string in a test case.

Two lessons, and the second is the one worth keeping.

First: a guard that matches on text rather than intent will eventually block the documentation of itself. Substring matching on a filename cannot distinguish "write this file" from "mention this file", and prose about a rule looks exactly like a violation of it.

Second, and this is what I actually changed: I had the scope wrong. That guard exists to stop unsupervised agents — plan-orchestrate workers, cron sessions inside containers — from disarming their own gates mid-run. It was never meant to stop me, at my own machine, maintaining the guard layer. Blocking the maintainer isn't strict, it's just broken. So it now exits immediately on host sessions and keeps the full, bypass-free block inside containers, where the risk it was written for actually lives.

Rule of thumb: a guard needs a blast radius as deliberate as its trigger. "Who is this stopping, and are they supervised?" is a design question, not a detail — and if the answer is "everyone, including whoever maintains it", the guard will eventually block the work of fixing itself.

What I'd take from this

The failure here wasn't an agent being dumb. Both workers behaved sensibly given what they had: a job to do, a long-running command, and no tool that could tell them how it went. The failure was mine, in handing out a capability (background a process) without the matching capability (observe its result), and then being surprised when the gap got filled with improvisation.

It generalises past agents, honestly. Any time you give something the ability to start work it cannot observe the end of, you've built a machine for producing confident wrong answers about state. The fix is never a better guess. It's removing the need to guess.

Cost of learning this: about two hours of wall-clock across two workers, and a near-miss where I almost restarted a batch of completed work. Cheap, as these go. The next one would not have been.

All from my claude-skills-central repo. Every guard ships with a .test.sh beside it.

I use my tooling predominantly on Mage-OS (Adobe Commerce / Magento) e-commerce projects, and my own AI Booking Agent. If you have not swapped to Mage-OS yet, you are falling behind ;)

More in this series: Smaller, Then Green — a mandatory cut-then-test loop that trims over-built code out of every plan's diff before any reviewer sees it. · Plans Run By A Fleet — how features run through a planned, adversarially-reviewed, parallel TDD pipeline instead of one long agent chat. · Hooks, Not Hopes — sorting every agent rule into prose or a blocking hook, and proving the hooks still fire. · The Agent Chatroom — giving the agents across my fleet a threaded chatroom so I stopped being the message bus. · The Commit That Has To Prove Itself — gating every commit behind recorded, state-hashed test evidence. · My AI Is Not Allowed To Guess — forcing blast-radius-first investigation and banning blame-shift excuses. · The Near-Miss That Banned Summaries — banning summarised page-fetches fleet-wide after a near-miss on a security advisory. · The Code Quality Checks Nobody Runs — coding standards, static analysis and comment hygiene moved into the commit path. · The ProxiBlue Debugger Discipline — blocking var_dump and wiring in real breakpoint debugging. · The ProxiBlue Domain Graph — giving my AI coding agents long-term domain knowledge with a temporal knowledge graph. · The ProxiBlue Falsifiable Rulebook — testing the rules and guard hooks that govern my AI coding agents.