Operations
Use this page to check health endpoints and Prometheus metrics exposed by long-running processes.
Health, readiness and metrics
afe serve --ops-port N and afe worker --ops-port N serve three endpoints on a port of their own,
separate from the API. They carry nothing secret and ask for no token; they bind to 127.0.0.1
unless --ops-host says otherwise.
| Endpoint | 200 when | 503 when |
|---|---|---|
/livez | serve: it answers. A worker: its claim loop turned in the last few seconds, or it is draining | the claim loop is stuck |
/readyz | the store and the broker answer a ping within 2 s | a dependency does not answer, named one per line; a worker after SIGTERM (draining) |
/metrics | always: the Prometheus text format | — |
Liveness never calls the store or the broker, so an outage makes processes unready, not restarted. Each process counts what it records itself, so summing the processes counts each fact once:
| Metric | Labels |
|---|---|
afe_tickets_started_total (where the ticket is opened: serve) | flow |
afe_tickets_finished_total (where it ends: a worker) | flow, state |
afe_nodes_total, afe_node_duration_seconds (histogram), agent and script nodes | flow, type |
afe_model_tokens_total | flow, model, direction (in, out) |
afe_model_cost_usd_total | flow, model |
afe_tool_calls_total | flow |
afe_pauses_total | flow, kind |
afe_queue_waiting, afe_queue_claimed, afe_workers_live (serve only) | — |
No label carries a ticket id, an input value or a secret.
Every long-running process serves /livez, /readyz and /metrics on an operations port of its
own, away from the API port.
| Process | /livez | /readyz |
|---|---|---|
serve | the event loop answers | store and broker reachable, current revision loaded |
worker | the claim loop ran in the last few seconds | store and broker reachable; not draining |
| operator (V7c) | the reconcile loop is alive | it holds the leader lease |
| UI (V7) | the process answers | serve reachable |
- Liveness never depends on Postgres or Redis: a database outage must not restart every Pod.
- Metrics are in the Prometheus format, computed from the engine's events and the queue metrics:
tickets by flow and state, queue depth, node durations, model tokens and cost. They carry no
secret and no ticket id. With the Prometheus operator, the afe operator adds a
ServiceMonitor. - Traces go to Langfuse or to any OpenTelemetry collector over OTLP.