07 SEPT 2026 · 11 min read

The agent that keeps the agents on task

When one session both does the work and keeps the task record honest, the record falls behind. I have been trying a split: a manager agent owns task hygiene, worker agents own execution, and a durable queue sits between them. This is a field note on that experiment, including where it strains.

The failure is quiet. A worker agent is halfway through a change and the task record still says the work is queued. A second agent, or a person, reads the record, sees an open item, and starts the same thing. A third item was finished on Tuesday, but its status was never updated: the session that did the work ran out of context before it got to the bookkeeping. Nobody lied. The record simply fell behind the work, and once it is behind it stops being something anyone trusts.

This note is about one response to that drift: splitting the roles apart. A manager agent owns task hygiene, meaning the job of keeping the shared record of the work true, and never writes the code. Worker agents own execution and never own the plan. A durable queue sits between them, and it is the only thing they share. What follows is how I run that, what it changed, and where it still strains. It is an observed practice from my own projects, not a method I am prescribing.

I stopped treating the drift as carelessness some time ago. It is structural. When one session plans a change, executes it, reports progress, and keeps the record accurate, the record is the part that loses. Execution has a compiler and a test suite pulling on it. The record has nothing pulling on it except discipline, and discipline is the first thing a busy context window drops.

Why the record falls behind.

A session doing real work accumulates context fast: files read, test output, half-formed plans, the reason an approach was abandoned. The task record wants a different kind of attention, in short structured updates at the right moments. This item is claimed. This one is blocked on a credential. This one is done, and here is the evidence.

Each of those is an interruption, and each costs a little of the attention the work needs. So updates cluster at the beginning and the end, if they arrive at all.

The middle goes dark, and the middle is exactly where a second agent needs to know that something is already in flight. When a session ends, what it knew about the state of the work ends with it, unless someone wrote it down somewhere durable. A transcript is not durable. It scrolls away, it belongs to one tool, and nobody reads it later.

There is a subtler cost. When one session decides both what done means and whether it is done, the definition bends toward whatever was achieved. A definition of done written before the work starts, and checked by someone other than the writer, is a small guard against that.

The shared record.

The source of truth is Projects, which I wrote about in an earlier note. A Project holds Tasks, Tasks hold SubTasks, and a SubTask is the unit of work an agent can claim. Each SubTask carries a definition of done, an ordered list of steps, dependencies on other SubTasks, a status, and a running log of notes.

I reach that record through pctl, a command-line client. Agents reach the same record over MCP, the Model Context Protocol, so a coding agent calls structured tools rather than shelling out and parsing text. The dashboard is a rendered view of the record, not a second copy of it, which is why there is nothing to keep in sync between them.

The manager agent.

A dedicated session whose only job is task hygiene, and which can run on a different machine from every worker. It takes intent from me in prose and turns it into SubTasks, each with a definition of done. It sets dependencies, so an item that cannot start yet sits as pending rather than tempting a worker into it.

Then it watches. For progress that has stopped arriving. For items that say queued but have a branch full of commits behind them. For completions with no evidence attached. When a worker reports a blocker, the manager records it where a person will see it. When work finishes, it reconciles the result back into the plan so the next item can be shaped.

The worker agents.

A worker session asks the queue for work, and the server offers it the next SubTask meeting three conditions: every dependency is done, the worker's role is authorised to see it, and no other worker currently holds it.

The worker then executes, usually with a small worker team underneath it: one sub-agent traces the affected surfaces, one makes the change, one writes the check that proves it. If the definition of done is met, the worker reports completion with the evidence. If it is not, the worker reports what it found and why it stopped.

Two loops, one queue.

The manager loop is: read intent, decompose, define done, queue, watch, reconcile. The worker loop is: request, claim, execute, verify, report. The two meet only at the queue.

Neither loop needs the other's context. The manager agent never needs to know how a change was made. The worker agent never needs to know why this item is ahead of that one. The queue is the interface, and the SubTask record is the message that crosses it.

Running the manager as a separate session is not architectural purity, it is about pressure. A session doing the work is under constant pressure to spend its context on the work. A session that only manages the record is under no such pressure, so the record gets the attention it needs. Running it on a separate machine takes the same argument one step further, and that is the part I have tested least: the hope is that a worker machine which runs out of disk, or gets reclaimed, cannot take the manager down with it.

What the lease does, and what it does not.

When the server offers a SubTask, the worker takes an exclusive lease on it, and while that lease is held the server will not offer the same SubTask to anyone else. The lease is short, about a minute, and every progress report renews it.

If the reports stop, the lease expires and the SubTask returns to the queue. This is the part worth being precise about. Expiry changes what the record says and makes the item claimable again. It does not reach into the first worker and stop it. A process that is still running is still running.

What the server does instead is reject the stale report: the next time the lapsed worker checks in, it is told its lease is gone, which is the signal to stop rather than to push on. That limits the damage without undoing it, and a completion that arrives after a lapse still lands, recorded as having arrived late.

So the lease is coordination metadata, not an execution guarantee. Where duplicate effects would be harmful, the work itself has to be idempotent, or fenced by something outside the lease. Two agents rarely run the same SubTask. Rarely is not never, and the design should not be read as promising otherwise.

A concrete run.

This site is itself a Project in the system. One item started as a sentence from me: an agent asking for a business contact should get exactly one unambiguous address.

Why that is not trivial takes one paragraph of background. Every page here is published twice: once as HTML for people, and once as a markdown twin at the same URL for machines that ask for it. A machine reading this site does not read the page you are reading. It reads the twin.

The manager turned my sentence into a SubTask with four steps: decide whether one or two mailboxes stay public, update every source that renders an address, label each mailbox's purpose if both remain, and add a readiness check that catches drift. Its definition of done named the surfaces that had to agree: the visible HTML, the markdown twins, the structured data, the in-page agent tools, and the Agent Skill files.

A worker claimed it, and its team found the defect quickly. Every public-facing surface named the business address. The CV's markdown twin named the recruiting address. So a person browsing the site got the right mailbox, and a machine reading the CV got the wrong one, and would have handed a business enquiry to recruiting. The fix was a single contact module with an explicit purpose per address, and every surface reading from it.

Then the part I actually care about. The worker's completion note listed what it had verified in the build output and what it had not: the live check could only pass once a release deployed. The manager did not mark the item done. It recorded the note, kept the item open with the remaining condition stated, and shaped a dependent item for the capability document that needed the same addresses.

Ten days later a different worker session, on a different machine, read the record cold and knew within a minute that three items were code-complete and waiting on a deploy, and that the useful work was elsewhere. Nobody had to find a transcript.

What improved, and how I know.

Two things I can point at directly. The record is now accurate often enough that I read it instead of asking, which was not true six months ago. And ownership is unambiguous: a claimed item has exactly one holder, a released item has none, and I no longer find items sitting assigned to a session that ended days ago.

One thing I can describe but have not measured. Two workers now pull from the same queue without colliding, so more runs in parallel than before. I have not instrumented throughput, and I am not going to quote a number I did not collect.

One thing that surprised me. Completion notes became more useful: they say what was verified, how, and what remains, and the note quoted above is typical rather than exceptional. My explanation is that a worker writing for a manager that will check the note against the definition of done writes differently from one writing into a transcript nobody will read. That is a mechanism I find plausible, not a result I measured.

What still fails.

Decomposition is a judgement, and the manager gets it wrong. Sometimes it splits an item so fine that workers spend more time reporting than building.

Sometimes it queues an item whose real dependency is a product decision rather than another SubTask, and a worker burns a session discovering that. Three items on this site sat blocked for weeks on exactly that: an API catalogue for an endpoint that had never been deployed. The worker that found it wrote a precise note and stopped, which is the right behaviour. A better manager would have caught it before queueing.

Permission boundaries bite in ways that are easy to forget. The dispatcher only hands work to agent principals, meaning identities registered as agents rather than as people. A session authenticated as me cannot ask for work at all: it has to take the operator path instead, acting on the record directly through pctl. That is deliberate and correct, and a worker that does not know which path it is on still wastes its first few minutes finding out.

And the manager is overhead, a whole session whose output is a tidy record. On a project with one worker and a dozen items it costs more than it saves. The failure that worries me most is quieter than that: a manager agent that becomes very good at the record and forgets the record is not the point.

Where the human stays.

None of this removes me. It changes what I do. I set direction, which is the sentence of intent the manager decomposes. I resolve ambiguity when a worker's note ends in a question rather than a result. I accept the completions that matter, particularly anything that ships to a live surface, and I decide when a blocker is worth clearing and when an item should be cancelled instead. The deploy at the end of that contact-address run was mine to trigger, and it should have been.

What I no longer do is reconstruct the state of the work from chat. That was most of my coordination time, and it is the part the machines are plainly better at.

Trying it, and knowing whether it worked.

If you want to test this, keep it small. One project. One session whose only job is the record: give it the intent, let it write the SubTasks, and insist on a definition of done for each before anything is queued. One or two worker sessions that claim from the queue and report back with evidence rather than summaries. Run it for a week.

Then judge it on signals you can actually observe. How long does it take a session that was not there to pick up the work cold, minutes or a re-read of the whole history? How many items carry a status the branch behind them contradicts? How often did two agents touch the same work? What share of completions carry evidence you could check yourself? And what did the manager session cost you, in tokens and in your own attention?

The split earns its coordination cost at the point where a session that was not there can pick the work up from the record alone. Below that bar you are paying for bookkeeping and buying a tidier version of the same confusion.

There is one thing a tidy record cannot tell you, and it is the important one: whether the work was worth doing. A beautifully reconciled queue of the wrong items is still the wrong items. That judgement stays mine, and no amount of hygiene moves it.

What I have observed is narrower than a recommendation. When several agents work in parallel, the record of that work needs attention that is not being consumed by the work itself, and that attention can come from a model. So far the trade has been worth it on the projects where drift was costing more than a session. If you are running something similar, or you have found that the split does not hold for your work, I would like to hear how it differs.

Tri2b · the studio