Skip to main content

Kubernetes

How afe runs on Kubernetes today.

Four milestones have landed. V6d added configuration revisions, Kubernetes-compatible manifests and the operations endpoints. V6e added the image and the kind environment. V7b added ticket Pods, workspace bundles, netlock, the worker's RBAC and the kind tests. V7c added the operator, the AgentFlowEngine custom resource, the CRDs, KEDA and Ingress.

One custom resource describes the installation, and the operator reconciles it. Requirements: FR1, FR6, FR21 and FR25–FR28.

What runs today​

  • The afe-kubernetes plugin (sandboxes/kubernetes) implements the Sandbox port over the kubectl CLI. It uses one ReadWriteOnce PVC and one Pod per ticket in the sandbox namespace. It applies them, waits for Ready, and runs exec/run/stream/put/get. It removes them by the afe.ticket label.
  • deploy/kind/sandbox/ carries the sandbox namespace (Pod Security privileged). It also carries the worker's namespaced Role and RoleBinding (pods, pods/exec and PVCs only, never secrets), and the ingress-only NetworkPolicy as a second layer. The worker annotates its Pod and PVC with the lifecycle markers, so the Role also grants patch and update on both.
  • netlock (docker/netlock) is the init container that drops every packet leaving the Pod but the ones from the proxy's uid, on IPv4 and IPv6. docker/egress-proxy is the in-Pod squid allowlist the sandbox reaches at localhost:3128.
  • The workspace lives in the Pod and is durable. After every node that ran it, the engine checkpoints the workspace and saves an incremental bundle in the store. A lost Pod is recreated and restored from the last checkpoint. See Workspaces.
  • docker compose brings the whole stack up on one machine with make compose-up. The Docker adapter keeps one volume and one container per ticket, the same volume + exec model (see Run the whole stack with Compose).
  • tests/kind (marker kind, make kind-test, not in make ci) proves it on a real cluster (see Deploy on kind). It runs a git + deep flow to done, and deletes a ticket Pod after a checkpoint to check it restores without repeating a completed node. It also checks network isolation with and without an allowlist, and that the Pod and PVC are gone at ticket end.
  • plugins/afe-operator/ is the operator (afe-operator, module afe_plugins.operator). It watches the AgentFlowEngine resources of its namespace, validates the configuration through afe_core, and reconciles the owned objects. Its status carries the conditions and the current revision.
  • deploy/operator/ is the kustomize install: the generated CRDs, the operator Deployment, its ServiceAccount and two namespaced Roles. make operator-install applies it once. What the operator must not own stays in deploy/operator/admin/ (see Install the operator).

The engine removes a ticket's Pod and PVC by the afe.ticket label when the ticket ends (done, failed, cancelled). Opening a ticket's Pod is idempotent: it is reused while its image, limits, env names and egress are unchanged, and recreated on a signature change (V7b R3).

A crash can still leave an orphan behind. The operator's reaper lists the sandbox objects by the afe.ticket and app.kubernetes.io/managed-by: afe labels. It deletes the terminal ones, and the running ones whose heartbeat is older than sandbox.orphanGrace. It keeps a paused Pod, and it reads no store and no API (PRD FR28).

Principles​

  • Scaling comes from the design that already exists. afe serve only serves the API, afe worker processes claim tickets from a Redis queue, and everything a resume needs lives in Postgres. On Kubernetes, more throughput is more worker replicas.
  • The web UI comes with serve. afe serve --ui serves the bundle on the API's own host and port, so there is no separate UI Deployment, image or Service to run.
  • No shared storage. Customers may have ReadWriteMany volumes and still refuse to use them, so nothing depends on them. Each ticket keeps its files in a Pod of its own, on a local ReadWriteOnce volume.
  • Workers keep no files. They talk to the ticket Pod through the Kubernetes exec API. The Git token stays in the worker.
  • The operator manages the lifecycle, not the scaling (V7c). It validates the configuration, deploys the processes, manages or references Postgres and Redis, and reports status.
  • One engine per cluster. Segregated tenants are not a goal for now.

Topology​

kubectl apply ──▶ AgentFlowEngine + Flow, Agent, Schema… (afe.dev) ──watch──▶ afe-operator
│ validates,
│ publishes the
│ revision, deploys
┌─ namespace afe ──────────────────────────────────────────────────────────────▼─────────┐
│ │
│ Ingress (WebSocket + UI) ──▶ Service ──▶ Deployment serve ×N (the UI comes with it) │
│ │ enqueue ▲ events │
│ ▼ │ │
│ Deployment worker ×1..M ◀── KEDA ScaledObject (queue lag) │
│ no volume; the Git token and the model keys stay here, │
│ serve holds no credential │
│ │
└───────────────────────────────────┬────────────────────────────────────────────────────┘
│ Kubernetes API: create Pod and PVC, exec (stdin/stdout)
┌─ namespace afe-sandbox ───────────▼────────────────────────────────────────────────────┐
│ Pod afe-<ticket>-<uid> + PVC ReadWriteOnce (local-path, TopoLVM, a cloud disk) │
│ one per ticket; the worker's RBAC reaches only this namespace │
└────────────────────────────────────────────────────────────────────────────────────────┘
┌─ namespace afe-data ───────────────────────────────────────────────────────────────────┐
│ CloudNativePG Cluster: tickets, ledger, checkpoints, revisions, workspace bundles │
│ Redis StatefulSet (AOF): queue, locks, notices, heartbeats │
└────────────────────────────────────────────────────────────────────────────────────────┘
Langfuse or any OpenTelemetry collector: outside, referenced by URL and a Secret

The data services sit in their own namespace, so that the worker's pods/exec permission never reaches them. The operator creates the serve and worker Deployments, the Services, the config ConfigMap, the Ingress, the ScaledObject and a managed store or broker. It creates none of the namespaces, the worker's RBAC, the envSecret or the data Secrets. Those, and the storage, runtime and ingress classes, stay the admin's manifests (deploy/operator/admin/).

The ticket Pod​

Pod afe-<ticket>-<uid> volume: PVC ReadWriteOnce, deleted with the ticket
├─ init netlock NET_ADMIN, runs once iptables: drop every packet leaving the Pod,
│ except those of uid 1337 (the proxy)
├─ sandbox uid 1000, no capabilities /workspace/repo
│ file tools, local git, /workspace/worktrees/<scope>
│ ws.exec, ws.test, acp agent
└─ proxy uid 1337, only with egress squid with the project's allowlist
no service account token · seccomp RuntimeDefault · RuntimeClass runc (default) or Kata
  • One Pod per ticket, not per scope. The lane worktrees share one clone. CPU and memory limits apply to the whole ticket.
  • The file tools run inside the Pod. On a single machine they run on the host and must confine every path. In the Pod there is nothing to protect, since the sandbox holds no secret.
  • Environment variables by name. An acp agent's model key comes from a Secret in afe-sandbox managed by the operator, referenced with secretKeyRef. The worker never handles the value.

Network without relying on the CNI​

A plain Kubernetes NetworkPolicy is not enough, for two reasons:

  1. It is only an API: the CNI plugin enforces it, and some (flannel, for one) ignore it silently.
  2. It matches addresses and Pods, not host names, so an allowlist such as openrouter.ai cannot be written with it.

The init container netlock therefore sets iptables rules in the Pod's network namespace, the technique Istio uses for its sidecar. Only the proxy's user may open connections.

The sandbox can reach nothing but localhost:3128, where the proxy applies the allowlist. No DNS query leaves the Pod directly. The sandbox has no NET_ADMIN, so it cannot undo the rules. A ticket without egress has no proxy, and nothing leaves at all.

The rules rest on netfilter and -m owner, which gVisor does not implement. Under gVisor, netlock would be a no-op, and the Pod would only look isolated. The manifest therefore refuses a runtimeClassName naming gVisor, and the runtime stays runc (the default) or Kata.

Only the init container needs NET_ADMIN, so the afe-sandbox namespace must allow it (Pod Security level privileged, or an exception in Kyverno or Gatekeeper). Where the CNI enforces policies, the operator adds a second-layer NetworkPolicy. With spec.afe.sandbox.networkPolicy.manage true (the default) it creates afe-sandbox-ingress in the sandbox namespace: ingress only, default deny. With manage false the admin owns the policy and the operator creates none.

Git through bundles​

The Pod never gets the Git token, and never needs the network for Git.

clone worker: git clone with the token (temporary folder) ─▶ git bundle ─▶ exec stdin
─▶ Pod: git clone from the bundle
save Pod: WIP commit ─▶ incremental git bundle of afe/* ─▶ exec stdout ─▶ worker ─▶ store
restore worker: bundle from the store ─▶ exec stdin ─▶ new Pod: git fetch
push Pod: git bundle afe/<ticket>-<uid>/* ─▶ exec stdout ─▶ worker: git push with the token

After each node that writes files, the workspace is committed and its bundle saved in the store. A lost node or Pod therefore does not break the promise of a restart without repeating paid work. The next worker restores the workspace on a new Pod, and goes on from the saved checkpoint.

The ceiling is the repository size: a very large repository is transferred twice at clone time. A cache inside the cluster comes if that becomes a problem.

A ticket from start to end​

client ── ticket.start ──▶ serve ── pins the current revision, enqueues ──▶ Redis
worker ── claims the job, takes the ticket lock (Redis)
── loads the ticket's revision and checkpoint (Postgres)
── opens the ticket Pod, or restores it from the last bundle
── agent node: model calls from the worker; ws.* and ws.exec through exec in the Pod
── the node wrote files ──▶ WIP commit, bundle saved in the store
── events ──▶ Postgres + a Redis notice ──▶ serve ──▶ WebSocket clients
── the ticket ends ──▶ Pod and PVC deleted by the afe.ticket label

a worker dies ▸ its lock expires ▸ another worker claims the job
▸ resumes from the saved token and finds the Pod again

Configuration revisions​

folder (afe serve -c) ─┐
├─▶ validate ─▶ normalize (prompts inlined, Runtime left out)
ConfigMap named by configMapRef ─┘ ─▶ sha256 ─▶ revisions table in the store (immutable)

ticket.start ─▶ ticket.revision = the current one ─▶ the worker loads it by hash ─▶ runs
resume, recovery, the UI, a new run of the same ticket ─▶ always the ticket's own revision
  • A new configuration never changes a running ticket, and it needs no worker restart. Workers read the revision of each ticket from the store.
  • A revision freezes the configuration, not the code. A newer engine image must still read the checkpoints of older tickets. The contract is a window of one persisted format generation (B25), and outside that window the engine fails hard instead of re-reading silently.
  • The operator reads the ConfigMap named by spec.config.configMapRef in the CR's namespace. It writes the keys as files, generates the Runtime from spec.afe and validates the whole set with afe validate. An invalid set publishes nothing and carries the per-field messages in ConfigValidated=False.
  • spec.config.revision is an optional pin. When it matches the computed revision the workloads roll. When it does not, RevisionPinned=False, reason RevisionMismatch, names both hashes, and nothing rolls.

Manifests as custom resources​

V7c. The files you write for a folder are the ones you kubectl apply.

apiVersion: afe.dev/v1alpha1
kind: Schema
metadata:
name: text-request # a DNS-1123 label: lower case, digits and `-`
annotations:
afe.dev/description: What the user asks for
spec:
fields:
topic: string

A Kubernetes object carries a few rules a plain file does not.

RuleWhy
Names and node ids are DNS-1123 labels (at most 63 characters).Kubernetes rejects any other object name.
The description is the afe.dev/description annotation.metadata.description is not a Kubernetes field, and the API server drops it.
A prompt is inline text, or a file reference that only a folder resolves (a ConfigMap on Kubernetes).A custom resource has no files next to it.
In when, - becomes _: text_request.topic == 'k8s'.text-request.topic would read as a subtraction.
Custom resources have full names (agents.afe.dev) and short names (afeagent).Agent, Flow and Plugin are common kind names.

Runtime is not a custom resource: the operator generates it from AgentFlowEngine.

The operator​

One AgentFlowEngine describes the installation. spec.config names the source ConfigMap and an optional revision pin. spec.afe carries the Runtime semantics (store, broker, tracer, sandbox). spec.deployment carries the workload knobs. Unknown keys are refused.

apiVersion: afe.dev/v1alpha1
kind: AgentFlowEngine
metadata:
name: afe
namespace: afe
spec:
config:
configMapRef:
name: afe-source # the ConfigMap holding the flow manifests
revision: 3f2a91c0d4e8 # optional pin; a mismatch freezes the rollout
afe:
store:
external:
secretRef:
name: afe-postgres
key: dsn
# or managed: { instances: 3, storage: 20Gi } # a CloudNativePG Cluster
broker:
external:
secretRef:
name: afe-redis
key: AFE_REDIS_URL
# or managed: { storage: 2Gi } # a Redis StatefulSet, no auth
tracer:
otlp:
endpoint: http://otel-collector:4318
headersSecretRef: { name: otlp-headers }
sandbox:
namespace: afe-sandbox
storage: 2Gi
runtimeClassName: kata # runc (default) or Kata; gVisor is refused
envSecret: afe-sandbox-env
deployment:
image:
repository: ghcr.io/agentflowengine/afe
tag: "0.7.0" # or digest: sha256:...
api:
replicas: 2
apiTokenSecretRef:
name: afe-secrets
key: AFE_API_TOKEN
workers:
replicas: 2
concurrency: 4
# The model key and the Git token both stay worker-only; serve never holds a credential.
envSecretRefs:
- { name: afe-secrets, key: OPENROUTER_API_KEY }
- { name: afe-secrets, key: GITHUB_TOKEN }
autoscaling:
minReplicas: 1
maxReplicas: 20
triggerAuthentication: keda-redis # an admin-provided TriggerAuthentication
ingress:
className: nginx
host: afe.internal
tls:
- secretName: afe-tls
  • managed is opt-in, never a default. A managed store is one CloudNativePG Cluster, and the operator references the -app Secret CloudNativePG writes. A managed broker is a Redis StatefulSet without authentication. An authenticated or external service is external, and its secretRef is projected into the workloads by reference only.
  • KEDA is optional. With workers.autoscaling and KEDA present the operator owns a ScaledObject and writes no replicas on the worker Deployment. Without KEDA the workers run at maxReplicas, and AutoscalingReady=False, reason KEDAAbsent.
  • Orphans are reaped, not adopted. The reaper removes terminal sandbox objects and running ones older than sandbox.orphanGrace (default 30m). It keeps paused objects.
  • What stays admin. The namespaces, the worker's ServiceAccount/Role/RoleBinding, the envSecret, the data Secrets and the storage, runtime and ingress classes. The installer ships them as examples under deploy/operator/admin/.
  • No secret value ever enters the CR. Every secret is a *Ref (name plus optional key), and the operator names it without reading it.
  • serve holds no credential. It runs no harness and resolves the providers without asking for a key. The operator projects no model key and no forge token onto it. workers.envSecretRefs and the envFrom mounts stay worker-only, so the process behind the API Service is never a way to reach a provider or repository credential (AFE-293).

The operator adopts nothing. If a desired object already exists with the right name but without the operator's ownerReference and label, the operator does not overwrite it. It sets Ready=False, reason ResourceConflict, and emits an Event naming the foreign object. The rest of the reconcile continues.

The operator repeats six steps on each reconcile.

1. configuration read configMapRef ─▶ afe validate ─✗─▶ ConfigValidated=False, nothing changes
─✓─▶ publish the revision
2. data a managed CNPG Cluster or the external Secret · a managed Redis or external
3. processes Deployments serve and worker · Services · Ingress · the config ConfigMap
4. scaling a KEDA ScaledObject on the queue lag when KEDA is installed
5. clean-up the sandbox objects the reaper removes, reported in status.orphansRemoved
6. status conditions, the current revision, components, endpoints and the orphan count

What is needed and what is not​

A small number of pieces carry real weight. Several common Kubernetes pieces are not needed at all.

NeededWhy
CloudNativePGserve and the workers connect to the -rw service, with the uri key of the -app Secret. S3 backups matter, since checkpoints, ledger and bundles live there.
A local or ReadWriteOnce storage classOne volume per ticket Pod.
KEDA (optional)Its redis-streams scaler reads the consumer group lag, so the engine exports nothing extra for scaling.
Kata (recommended). gVisor refused.Ticket Pods run code written by a model. netlock needs netfilter, which gVisor does not implement.
Not neededWhy
ReadWriteMany storageWorkspaces live in ticket Pods.
Calico or Ciliumnetlock isolates the Pod. A NetworkPolicy is only a second layer.
CNPG Pooler (PgBouncer)Every process has its own small pool. Migrations take a transaction-level advisory lock, which PgBouncer's transaction mode allows anyway.
Redis Sentinel or a Redis operatorPostgres is the source of truth, and workers queue running tickets again at startup. One StatefulSet with AOF is enough. Valkey works as well.
Argo, Tekton, one Job per ticketThe queue, the ticket lock and the resume token already schedule the work.
A Docker daemonTicket Pods replace containers.

Running locally with Docker only​

Two modes run without a real cluster.

ModeWhat runsFor
Dockerafe run --local, or make compose-up (a Compose profile with serve --ui, a worker, Postgres, Redis and the UI). One volume and one container per ticket.development and demos.
kind or k3dKubernetes inside Docker: the ticket Pods, netlock, the namespaces and the RBAC from deploy/kind/, and the operator from deploy/operator/.the kind tests (make kind-test) and a production-like cluster.

The Docker and Kubernetes sandboxes share the same model: a volume plus exec, with the workspace inside the ticket's container or Pod. In Compose the worker needs the Docker socket, which is root on the host. That is acceptable on a laptop, not in production.

Limits​

  • A CI that needs Docker itself (testcontainers, for one) cannot run in a ticket Pod without privileges.
  • Stdio plugins and stdio MCP servers run inside the worker container: its image must contain their runtimes, or they run as HTTP plugins in their own Deployments.
  • Exposing serve outside the cluster needs the browser authentication decided in V7.

Troubleshooting​

A ticket Pod never reaches Ready. Check the sandbox namespace's Pod Security level allows the init container's NET_ADMIN, and that the node has capacity for the ticket's sandbox.cpus/memory. afe-kubernetes waits for Ready and then raises SandboxGone, which the engine treats as a lost sandbox: it removes the Pod and PVC and requeues the ticket.

A sandbox call reaches a host outside sandbox.egress Check the runtime is not gVisor. netlock needs netfilter and -m owner, which gVisor does not implement, so a ticket on gVisor would only look isolated. The manifest refuses a runtimeClassName naming gVisor for this reason.

A ticket's Pod and PVC are still there after it ended The engine removes them by the afe.ticket label when the ticket reaches done, failed or cancelled. A worker crash between the ticket ending and the removal can leave an orphan. V7b does not scan for them. The V7c operator reconciles the label and removes and reports orphans.

See also​