Budgets and loop guards
Why the ledger, not a counter in memory, decides when a flow must stop or ask a person.
Budgets
Every model and tool call is recorded in the ticket's ledger: tokens, cost and time. Budgets are checked against the ledger before each model call, including the tokens the next call would add. When one turn fires several tool calls, the call that crosses a threshold runs. The others wait for the answer.
The <budget> block is volatile: the engine appends it after the history, as the last message of the
call, and never writes it into the saved history. The system prompt, the tools and the history before
it keep a byte-identical prefix across the calls of a turn. The provider can cache that prefix, so the
long turn costs less.
The ledger also records the prompt tokens a provider cache served (cached_tokens). The ticket
totals show the cached tokens and the fresh ones (tokens_in - cached_tokens) apart. The cost is
not recomputed: it stays the one the provider reports, where a cache hit is already discounted.
Loop guards
A turn that calls the same tool with the same arguments over and over is a loop. No budget dimension stops it before its own threshold is crossed, and a loop of cheap tool calls may cross none of them. The engine counts consecutive identical tool calls instead.
- from the third call, the tool result carries a note. It asks the model to change strategy, or to deliver the best result it has and say what is still open.
- at the fourth call, the turn pauses for a person with the same choices as a rejected artifact.
retrygives the model another chance and resets the count.take_overstops the node for a person.closecloses the node.
The count is per activation. It resets as soon as the call changes, or a person answers.