Skip to main content

Pauses and restarts

Why a paid model call is never repeated, whichever worker resumes the ticket.

A harness may pause a turn through the ledger (a budget threshold or a repeated tool call). The engine writes the pause ahead, and resumes the harness from its last token, so a paid model call is never repeated. The same token lets a turn continue after a worker dies. The harness saves it after every model reply and every tool result, so a crash repeats at most the call that was in flight.

The broker​

The broker provides the queue, the per-ticket lock, the events across processes and the cancel signals. The afe-redis plugin implements it on Redis.

  • a Stream with a consumer group for the queue.
  • SET NX PX with owner-checked renewal for the lock.
  • pub/sub for the notices and the cancels.
  • short-lived keys for the worker heartbeats.

The broker never carries an event's content. A worker records the event in the store and only pings the broker. Every process reads the gap back from the store by sequence number, so no event is lost or duplicated.

When a worker dies​

A worker that dies loses its lock (it expires), and its unacked job is redelivered to another worker. That worker continues the ticket from its last checkpoint. A node whose result was written ahead, but whose checkpoint was not saved, is not called again.

A job delivered three times without being finished puts the ticket in needs_human, with the reason "gave up after 3 attempts" and the retry, take over and close actions. retry queues it again with the attempt count reset. At startup, every worker also queues a recovery job for every running ticket that no lock holds.

A broker outage does not stop a worker. Its claim loop, heartbeat and cancel subscription log the error class once and retry, so the process stays live. /readyz reports broker until Redis is back. A cancel sent during the outage still stops the ticket at its next node boundary.

A resume belongs to its request​

A resume job carries the id of the request it answers. The engine reads the ticket's current request at execution time, and applies the answer only when the id matches. A stale answer is a no-op, so a duplicate resume never skips a later pause.

A request that offers actions needs an explicit answer among those offered. A resume with no answer cannot pass it. Each ticket_resumed event records the request id.

Cancel versus pause​

ticket.cancel stops the harness turn in progress at once, on whichever worker runs it, and in afe run --local. The ticket becomes cancelled, the artifacts and ledger rows written so far are kept, and one ticket_cancelled is emitted.

ticket.pause still stops at the next node boundary. A turn interrupted and later resumed would otherwise pay its model call twice.

Crash recovery​

A worker resuming a running ticket continues it from its last checkpoint. The workers also queue a recovery job for every running ticket no lock holds at startup. Tickets that wait for a person (paused_budget, paused_human, needs_human) stay as they are, with their pending request readable.

A turn that crashed with pauses already answered replays them in order, and resumes the harness from the last token. No paid model call is repeated.

afe run --local never recovers other tickets.

Troubleshooting​

A ticket sits in needs_human with the reason "gave up after 3 attempts". Three different workers claimed the job and none finished it, usually because the node itself crashes or hangs. Fix the underlying cause first, then retry to queue it again with a fresh attempt count, or take_over to run the node by hand.

/readyz reports broker as unhealthy The broker is unreachable. Workers stay live and keep retrying, so no work is lost, but new tickets cannot be enqueued and cancels are delayed until Redis answers again. Check Redis itself, not the worker.

See also​