Skip to main content

Operations

Use this page to check health endpoints and Prometheus metrics exposed by long-running processes.

Health, readiness and metrics​

afe serve --ops-port N and afe worker --ops-port N serve three endpoints on a port of their own, separate from the API. They carry nothing secret and ask for no token; they bind to 127.0.0.1 unless --ops-host says otherwise.

Endpoint200 when503 when
/livezserve: it answers. A worker: its claim loop turned in the last few seconds, or it is drainingthe claim loop is stuck
/readyzthe store and the broker answer a ping within 2 sa dependency does not answer, named one per line; a worker after SIGTERM (draining)
/metricsalways: the Prometheus text format—

Liveness never calls the store or the broker, so an outage makes processes unready, not restarted. Each process counts what it records itself, so summing the processes counts each fact once:

MetricLabels
afe_tickets_started_total (where the ticket is opened: serve)flow
afe_tickets_finished_total (where it ends: a worker)flow, state
afe_nodes_total, afe_node_duration_seconds (histogram), agent and script nodesflow, type
afe_model_tokens_totalflow, model, direction (in, out)
afe_model_cost_usd_totalflow, model
afe_tool_calls_totalflow
afe_pauses_totalflow, kind
afe_queue_waiting, afe_queue_claimed, afe_workers_live (serve only)—

No label carries a ticket id, an input value or a secret.

Every long-running process serves /livez, /readyz and /metrics on an operations port of its own, away from the API port.

Process/livez/readyz
servethe event loop answersstore and broker reachable, current revision loaded
workerthe claim loop ran in the last few secondsstore and broker reachable; not draining
operator (V7c)the reconcile loop is aliveit holds the leader lease
UI (V7)the process answersserve reachable
  • Liveness never depends on Postgres or Redis: a database outage must not restart every Pod.
  • Metrics are in the Prometheus format, computed from the engine's events and the queue metrics: tickets by flow and state, queue depth, node durations, model tokens and cost. They carry no secret and no ticket id. With the Prometheus operator, the afe operator adds a ServiceMonitor.
  • Traces go to Langfuse or to any OpenTelemetry collector over OTLP.

See also​

  • CLI for --ops-host and --ops-port.
  • API for the application WebSocket.