Kubernetes
How afe runs on Kubernetes today.
Four milestones have landed. V6d added configuration revisions, Kubernetes-compatible manifests and
the operations endpoints. V6e added the image and the kind environment. V7b added ticket Pods,
workspace bundles, netlock, the worker's RBAC and the kind tests. V7c added the operator, the
AgentFlowEngine custom resource, the CRDs, KEDA and Ingress.
One custom resource describes the installation, and the operator reconciles it. Requirements: FR1, FR6, FR21 and FR25–FR28.
What runs today
- The
afe-kubernetesplugin (sandboxes/kubernetes) implements theSandboxport over thekubectlCLI. It uses one ReadWriteOnce PVC and one Pod per ticket in the sandbox namespace. It applies them, waits forReady, and runsexec/run/stream/put/get. It removes them by theafe.ticketlabel. deploy/kind/sandbox/carries the sandbox namespace (Pod Securityprivileged). It also carries the worker's namespaced Role and RoleBinding (pods,pods/execand PVCs only, neversecrets), and the ingress-onlyNetworkPolicyas a second layer. The worker annotates its Pod and PVC with the lifecycle markers, so the Role also grantspatchandupdateon both.netlock(docker/netlock) is the init container that drops every packet leaving the Pod but the ones from the proxy's uid, on IPv4 and IPv6.docker/egress-proxyis the in-Podsquidallowlist the sandbox reaches atlocalhost:3128.- The workspace lives in the Pod and is durable. After every node that ran it, the engine checkpoints the workspace and saves an incremental bundle in the store. A lost Pod is recreated and restored from the last checkpoint. See Workspaces.
docker composebrings the whole stack up on one machine withmake compose-up. The Docker adapter keeps one volume and one container per ticket, the samevolume + execmodel (see Run the whole stack with Compose).tests/kind(markerkind,make kind-test, not inmake ci) proves it on a real cluster (see Deploy on kind). It runs agit+deepflow todone, and deletes a ticket Pod after a checkpoint to check it restores without repeating a completed node. It also checks network isolation with and without an allowlist, and that the Pod and PVC are gone at ticket end.plugins/afe-operator/is the operator (afe-operator, moduleafe_plugins.operator). It watches theAgentFlowEngineresources of its namespace, validates the configuration throughafe_core, and reconciles the owned objects. Its status carries the conditions and the current revision.deploy/operator/is the kustomize install: the generated CRDs, the operator Deployment, its ServiceAccount and two namespaced Roles.make operator-installapplies it once. What the operator must not own stays indeploy/operator/admin/(see Install the operator).
The engine removes a ticket's Pod and PVC by the afe.ticket label when the ticket ends (done,
failed, cancelled). Opening a ticket's Pod is idempotent: it is reused while its image, limits,
env names and egress are unchanged, and recreated on a signature change (V7b R3).
A crash can still leave an orphan behind. The operator's reaper lists the sandbox objects by the
afe.ticket and app.kubernetes.io/managed-by: afe labels. It deletes the terminal ones, and the
running ones whose heartbeat is older than sandbox.orphanGrace. It keeps a paused Pod, and it
reads no store and no API (PRD FR28).
Principles
- Scaling comes from the design that already exists.
afe serveonly serves the API,afe workerprocesses claim tickets from a Redis queue, and everything a resume needs lives in Postgres. On Kubernetes, more throughput is more worker replicas. - The web UI comes with
serve.afe serve --uiserves the bundle on the API's own host and port, so there is no separate UI Deployment, image or Service to run. - No shared storage. Customers may have ReadWriteMany volumes and still refuse to use them, so nothing depends on them. Each ticket keeps its files in a Pod of its own, on a local ReadWriteOnce volume.
- Workers keep no files. They talk to the ticket Pod through the Kubernetes exec API. The Git token stays in the worker.
- The operator manages the lifecycle, not the scaling (V7c). It validates the configuration, deploys the processes, manages or references Postgres and Redis, and reports status.
- One engine per cluster. Segregated tenants are not a goal for now.
Topology
kubectl apply ──▶ AgentFlowEngine + Flow, Agent, Schema… (afe.dev) ──watch──▶ afe-operator
│ validates,
│ publishes the
│ revision, deploys
┌─ namespace afe ──────────────────────────────────────────────────────────────▼─────────┐
│ │
│ Ingress (WebSocket + UI) ──▶ Service ──▶ Deployment serve ×N (the UI comes with it) │
│ │ enqueue ▲ events │
│ ▼ │ │
│ Deployment worker ×1..M ◀── KEDA ScaledObject (queue lag) │
│ no volume; the Git token and the model keys stay here, │
│ serve holds no credential │
│ │
└───────────────────────────────────┬────────────────────────────────────────────────────┘
│ Kubernetes API: create Pod and PVC, exec (stdin/stdout)
┌─ namespace afe-sandbox ───────────▼────────────────────────────────────────────────────┐
│ Pod afe-<ticket>-<uid> + PVC ReadWriteOnce (local-path, TopoLVM, a cloud disk) │
│ one per ticket; the worker's RBAC reaches only this namespace │
└────────────────────────────────────────────────────────────────────────────────────────┘
┌─ namespace afe-data ───────────────────────────────────────────────────────────────────┐
│ CloudNativePG Cluster: tickets, ledger, checkpoints, revisions, workspace bundles │
│ Redis StatefulSet (AOF): queue, locks, notices, heartbeats │
└────────────────────────────────────────────────────────────────────────────────────────┘
Langfuse or any OpenTelemetry collector: outside, referenced by URL and a Secret
The data services sit in their own namespace, so that the worker's pods/exec permission never
reaches them. The operator creates the serve and worker Deployments, the Services, the config
ConfigMap, the Ingress, the ScaledObject and a managed store or broker. It creates none of the
namespaces, the worker's RBAC, the envSecret or the data Secrets. Those, and the storage,
runtime and ingress classes, stay the admin's manifests (deploy/operator/admin/).
The ticket Pod
Pod afe-<ticket>-<uid> volume: PVC ReadWriteOnce, deleted with the ticket
├─ init netlock NET_ADMIN, runs once iptables: drop every packet leaving the Pod,
│ except those of uid 1337 (the proxy)
├─ sandbox uid 1000, no capabilities /workspace/repo
│ file tools, local git, /workspace/worktrees/<scope>
│ ws.exec, ws.test, acp agent
└─ proxy uid 1337, only with egress squid with the project's allowlist
no service account token · seccomp RuntimeDefault · RuntimeClass runc (default) or Kata
- One Pod per ticket, not per scope. The lane worktrees share one clone. CPU and memory limits apply to the whole ticket.
- The file tools run inside the Pod. On a single machine they run on the host and must confine every path. In the Pod there is nothing to protect, since the sandbox holds no secret.
- Environment variables by name. An
acpagent's model key comes from a Secret inafe-sandboxmanaged by the operator, referenced withsecretKeyRef. The worker never handles the value.
Network without relying on the CNI
A plain Kubernetes NetworkPolicy is not enough, for two reasons:
- It is only an API: the CNI plugin enforces it, and some (flannel, for one) ignore it silently.
- It matches addresses and Pods, not host names, so an allowlist such as
openrouter.aicannot be written with it.
The init container netlock therefore sets iptables rules in the Pod's network namespace, the
technique Istio uses for its sidecar. Only the proxy's user may open connections.
The sandbox can reach nothing but localhost:3128, where the proxy applies the allowlist. No
DNS query leaves the Pod directly. The sandbox has no NET_ADMIN, so it cannot undo the rules. A
ticket without egress has no proxy, and nothing leaves at all.
The rules rest on netfilter and -m owner, which gVisor does not implement. Under gVisor,
netlock would be a no-op, and the Pod would only look isolated. The manifest therefore
refuses a runtimeClassName naming gVisor, and the runtime stays runc (the default) or Kata.
Only the init container needs NET_ADMIN, so the afe-sandbox namespace must allow it (Pod
Security level privileged, or an exception in Kyverno or Gatekeeper). Where the CNI enforces
policies, the operator adds a second-layer NetworkPolicy. With
spec.afe.sandbox.networkPolicy.manage true (the default) it creates afe-sandbox-ingress in the
sandbox namespace: ingress only, default deny. With manage false the admin owns the policy and
the operator creates none.
Git through bundles
The Pod never gets the Git token, and never needs the network for Git.
clone worker: git clone with the token (temporary folder) ─▶ git bundle ─▶ exec stdin
─▶ Pod: git clone from the bundle
save Pod: WIP commit ─▶ incremental git bundle of afe/* ─▶ exec stdout ─▶ worker ─▶ store
restore worker: bundle from the store ─▶ exec stdin ─▶ new Pod: git fetch
push Pod: git bundle afe/<ticket>-<uid>/* ─▶ exec stdout ─▶ worker: git push with the token
After each node that writes files, the workspace is committed and its bundle saved in the store. A lost node or Pod therefore does not break the promise of a restart without repeating paid work. The next worker restores the workspace on a new Pod, and goes on from the saved checkpoint.
The ceiling is the repository size: a very large repository is transferred twice at clone time. A cache inside the cluster comes if that becomes a problem.
A ticket from start to end
client ── ticket.start ──▶ serve ── pins the current revision, enqueues ──▶ Redis
worker ── claims the job, takes the ticket lock (Redis)
── loads the ticket's revision and checkpoint (Postgres)
── opens the ticket Pod, or restores it from the last bundle
── agent node: model calls from the worker; ws.* and ws.exec through exec in the Pod
── the node wrote files ──▶ WIP commit, bundle saved in the store
── events ──▶ Postgres + a Redis notice ──▶ serve ──▶ WebSocket clients
── the ticket ends ──▶ Pod and PVC deleted by the afe.ticket label
a worker dies ▸ its lock expires ▸ another worker claims the job
▸ resumes from the saved token and finds the Pod again
Configuration revisions
folder (afe serve -c) ─┐
├─▶ validate ─▶ normalize (prompts inlined, Runtime left out)
ConfigMap named by configMapRef ─┘ ─▶ sha256 ─▶ revisions table in the store (immutable)
ticket.start ─▶ ticket.revision = the current one ─▶ the worker loads it by hash ─▶ runs
resume, recovery, the UI, a new run of the same ticket ─▶ always the ticket's own revision
- A new configuration never changes a running ticket, and it needs no worker restart. Workers read the revision of each ticket from the store.
- A revision freezes the configuration, not the code. A newer engine image must still read the checkpoints of older tickets. The contract is a window of one persisted format generation (B25), and outside that window the engine fails hard instead of re-reading silently.
- The operator reads the ConfigMap named by
spec.config.configMapRefin the CR's namespace. It writes the keys as files, generates theRuntimefromspec.afeand validates the whole set withafe validate. An invalid set publishes nothing and carries the per-field messages inConfigValidated=False. spec.config.revisionis an optional pin. When it matches the computed revision the workloads roll. When it does not,RevisionPinned=False, reasonRevisionMismatch, names both hashes, and nothing rolls.
Manifests as custom resources
V7c. The files you write for a folder are the ones you kubectl apply.
apiVersion: afe.dev/v1alpha1
kind: Schema
metadata:
name: text-request # a DNS-1123 label: lower case, digits and `-`
annotations:
afe.dev/description: What the user asks for
spec:
fields:
topic: string
A Kubernetes object carries a few rules a plain file does not.
| Rule | Why |
|---|---|
| Names and node ids are DNS-1123 labels (at most 63 characters). | Kubernetes rejects any other object name. |
The description is the afe.dev/description annotation. | metadata.description is not a Kubernetes field, and the API server drops it. |
| A prompt is inline text, or a file reference that only a folder resolves (a ConfigMap on Kubernetes). | A custom resource has no files next to it. |
In when, - becomes _: text_request.topic == 'k8s'. | text-request.topic would read as a subtraction. |
Custom resources have full names (agents.afe.dev) and short names (afeagent). | Agent, Flow and Plugin are common kind names. |
Runtime is not a custom resource: the operator generates it from AgentFlowEngine.
The operator
One AgentFlowEngine describes the installation. spec.config names the source ConfigMap and an
optional revision pin. spec.afe carries the Runtime semantics (store, broker, tracer, sandbox).
spec.deployment carries the workload knobs. Unknown keys are refused.
apiVersion: afe.dev/v1alpha1
kind: AgentFlowEngine
metadata:
name: afe
namespace: afe
spec:
config:
configMapRef:
name: afe-source # the ConfigMap holding the flow manifests
revision: 3f2a91c0d4e8 # optional pin; a mismatch freezes the rollout
afe:
store:
external:
secretRef:
name: afe-postgres
key: dsn
# or managed: { instances: 3, storage: 20Gi } # a CloudNativePG Cluster
broker:
external:
secretRef:
name: afe-redis
key: AFE_REDIS_URL
# or managed: { storage: 2Gi } # a Redis StatefulSet, no auth
tracer:
otlp:
endpoint: http://otel-collector:4318
headersSecretRef: { name: otlp-headers }
sandbox:
namespace: afe-sandbox
storage: 2Gi
runtimeClassName: kata # runc (default) or Kata; gVisor is refused
envSecret: afe-sandbox-env
deployment:
image:
repository: ghcr.io/agentflowengine/afe
tag: "0.7.0" # or digest: sha256:...
api:
replicas: 2
apiTokenSecretRef:
name: afe-secrets
key: AFE_API_TOKEN
workers:
replicas: 2
concurrency: 4
# The model key and the Git token both stay worker-only; serve never holds a credential.
envSecretRefs:
- { name: afe-secrets, key: OPENROUTER_API_KEY }
- { name: afe-secrets, key: GITHUB_TOKEN }
autoscaling:
minReplicas: 1
maxReplicas: 20
triggerAuthentication: keda-redis # an admin-provided TriggerAuthentication
ingress:
className: nginx
host: afe.internal
tls:
- secretName: afe-tls
managedis opt-in, never a default. Amanagedstore is one CloudNativePGCluster, and the operator references the-appSecret CloudNativePG writes. Amanagedbroker is a RedisStatefulSetwithout authentication. An authenticated or external service isexternal, and itssecretRefis projected into the workloads by reference only.- KEDA is optional. With
workers.autoscalingand KEDA present the operator owns aScaledObjectand writes noreplicason the worker Deployment. Without KEDA the workers run atmaxReplicas, andAutoscalingReady=False, reasonKEDAAbsent. - Orphans are reaped, not adopted. The reaper removes
terminalsandbox objects andrunningones older thansandbox.orphanGrace(default30m). It keepspausedobjects. - What stays admin. The namespaces, the worker's ServiceAccount/Role/RoleBinding, the
envSecret, the dataSecrets and the storage, runtime and ingress classes. The installer ships them as examples underdeploy/operator/admin/. - No secret value ever enters the CR. Every secret is a
*Ref(nameplus optionalkey), and the operator names it without reading it. serveholds no credential. It runs no harness and resolves the providers without asking for a key. The operator projects no model key and no forge token onto it.workers.envSecretRefsand theenvFrommounts stay worker-only, so the process behind the API Service is never a way to reach a provider or repository credential (AFE-293).
The operator adopts nothing. If a desired object already exists with the right name but without the
operator's ownerReference and label, the operator does not overwrite it. It sets Ready=False,
reason ResourceConflict, and emits an Event naming the foreign object. The rest of the reconcile
continues.
The operator repeats six steps on each reconcile.
1. configuration read configMapRef ─▶ afe validate ─✗─▶ ConfigValidated=False, nothing changes
─✓─▶ publish the revision
2. data a managed CNPG Cluster or the external Secret · a managed Redis or external
3. processes Deployments serve and worker · Services · Ingress · the config ConfigMap
4. scaling a KEDA ScaledObject on the queue lag when KEDA is installed
5. clean-up the sandbox objects the reaper removes, reported in status.orphansRemoved
6. status conditions, the current revision, components, endpoints and the orphan count
What is needed and what is not
A small number of pieces carry real weight. Several common Kubernetes pieces are not needed at all.
| Needed | Why |
|---|---|
| CloudNativePG | serve and the workers connect to the -rw service, with the uri key of the -app Secret. S3 backups matter, since checkpoints, ledger and bundles live there. |
| A local or ReadWriteOnce storage class | One volume per ticket Pod. |
| KEDA (optional) | Its redis-streams scaler reads the consumer group lag, so the engine exports nothing extra for scaling. |
| Kata (recommended). gVisor refused. | Ticket Pods run code written by a model. netlock needs netfilter, which gVisor does not implement. |
| Not needed | Why |
|---|---|
| ReadWriteMany storage | Workspaces live in ticket Pods. |
| Calico or Cilium | netlock isolates the Pod. A NetworkPolicy is only a second layer. |
| CNPG Pooler (PgBouncer) | Every process has its own small pool. Migrations take a transaction-level advisory lock, which PgBouncer's transaction mode allows anyway. |
| Redis Sentinel or a Redis operator | Postgres is the source of truth, and workers queue running tickets again at startup. One StatefulSet with AOF is enough. Valkey works as well. |
| Argo, Tekton, one Job per ticket | The queue, the ticket lock and the resume token already schedule the work. |
| A Docker daemon | Ticket Pods replace containers. |
Running locally with Docker only
Two modes run without a real cluster.
| Mode | What runs | For |
|---|---|---|
| Docker | afe run --local, or make compose-up (a Compose profile with serve --ui, a worker, Postgres, Redis and the UI). One volume and one container per ticket. | development and demos. |
| kind or k3d | Kubernetes inside Docker: the ticket Pods, netlock, the namespaces and the RBAC from deploy/kind/, and the operator from deploy/operator/. | the kind tests (make kind-test) and a production-like cluster. |
The Docker and Kubernetes sandboxes share the same model: a volume plus exec, with the workspace inside the ticket's container or Pod. In Compose the worker needs the Docker socket, which is root on the host. That is acceptable on a laptop, not in production.
Limits
- A CI that needs Docker itself (testcontainers, for one) cannot run in a ticket Pod without privileges.
- Stdio plugins and stdio MCP servers run inside the worker container: its image must contain their runtimes, or they run as HTTP plugins in their own Deployments.
- Exposing
serveoutside the cluster needs the browser authentication decided in V7.
Troubleshooting
A ticket Pod never reaches Ready.
Check the sandbox namespace's Pod Security level allows the init container's NET_ADMIN, and
that the node has capacity for the ticket's sandbox.cpus/memory. afe-kubernetes waits for
Ready and then raises SandboxGone, which the engine treats as a lost sandbox: it removes the
Pod and PVC and requeues the ticket.
A sandbox call reaches a host outside sandbox.egress
Check the runtime is not gVisor. netlock needs netfilter and -m owner, which gVisor does not
implement, so a ticket on gVisor would only look isolated. The manifest refuses a
runtimeClassName naming gVisor for this reason.
A ticket's Pod and PVC are still there after it ended
The engine removes them by the afe.ticket label when the ticket reaches done, failed or
cancelled. A worker crash between the ticket ending and the removal can leave an orphan. V7b
does not scan for them. The V7c operator reconciles the label and removes and reports orphans.