Pauses and restarts
Why a paid model call is never repeated, whichever worker resumes the ticket.
A harness may pause a turn through the ledger (a budget threshold or a repeated tool call). The engine writes the pause ahead, and resumes the harness from its last token, so a paid model call is never repeated. The same token lets a turn continue after a worker dies. The harness saves it after every model reply and every tool result, so a crash repeats at most the call that was in flight.
The broker
The broker provides the queue, the per-ticket lock, the events across processes and the cancel
signals. The afe-redis plugin implements it on Redis.
- a Stream with a consumer group for the queue.
SET NX PXwith owner-checked renewal for the lock.- pub/sub for the notices and the cancels.
- short-lived keys for the worker heartbeats.
The broker never carries an event's content. A worker records the event in the store and only pings the broker. Every process reads the gap back from the store by sequence number, so no event is lost or duplicated.
When a worker dies
A worker that dies loses its lock (it expires), and its unacked job is redelivered to another worker. That worker continues the ticket from its last checkpoint. A node whose result was written ahead, but whose checkpoint was not saved, is not called again.
A job delivered three times without being finished puts the ticket in needs_human, with the
reason "gave up after 3 attempts" and the retry, take over and close actions. retry queues
it again with the attempt count reset. At startup, every worker also queues a recovery job for
every running ticket that no lock holds.
A broker outage does not stop a worker. Its claim loop, heartbeat and cancel subscription log the
error class once and retry, so the process stays live. /readyz reports broker until Redis is
back. A cancel sent during the outage still stops the ticket at its next node boundary.
A resume belongs to its request
A resume job carries the id of the request it answers. The engine reads the ticket's current
request at execution time, and applies the answer only when the id matches. A stale answer is a
no-op, so a duplicate resume never skips a later pause.
A request that offers actions needs an explicit answer among those offered. A resume with no
answer cannot pass it. Each ticket_resumed event records the request id.
Cancel versus pause
ticket.cancel stops the harness turn in progress at once, on whichever worker runs it, and in
afe run --local. The ticket becomes cancelled, the artifacts and ledger rows written so far
are kept, and one ticket_cancelled is emitted.
ticket.pause still stops at the next node boundary. A turn interrupted and later resumed would
otherwise pay its model call twice.
Crash recovery
A worker resuming a running ticket continues it from its last checkpoint. The workers also
queue a recovery job for every running ticket no lock holds at startup. Tickets that wait for a
person (paused_budget, paused_human, needs_human) stay as they are, with their pending
request readable.
A turn that crashed with pauses already answered replays them in order, and resumes the harness from the last token. No paid model call is repeated.
afe run --local never recovers other tickets.
Troubleshooting
A ticket sits in needs_human with the reason "gave up after 3 attempts".
Three different workers claimed the job and none finished it, usually because the node itself
crashes or hangs. Fix the underlying cause first, then retry to queue it again with a fresh
attempt count, or take_over to run the node by hand.
/readyz reports broker as unhealthy
The broker is unreachable. Workers stay live and keep retrying, so no work is lost, but new
tickets cannot be enqueued and cancels are delayed until Redis answers again. Check Redis itself,
not the worker.