Deploy on kind
Goal: run the engine on a local kind cluster and drive a ticket through it.
Prerequisites
- Docker.
- kind.
kubectl.
make kind-up deploys the whole engine on the cluster: serve, two workers on the dry harness,
Postgres through CloudNativePG and Redis. It also deploys the V7b sandbox pieces: the
afe-sandbox namespace, the worker's RBAC, netlock, and the images afe-sandbox:dev,
afe-netlock:dev and afe-egress-proxy:dev.
Steps
-
Create the cluster and forward its ports:
make kind-up # cluster `afe`, image afe:dev, CloudNativePG, Redis, serve and workersmake kind-forward # API on 127.0.0.1:8765, ops on 9464 (serve) and 9465 (a worker); leave it onmake kind-uptakes the token fromAFE_API_TOKENin.env. Without that variable it keeps the token it generated the first time. The token lives only in the cluster's Secret.make kind-upalso picks the model harness. Its last line says which one it chose. It readskind: workers on the dry harness (fake models)orkind: workers on the live harness (real models). Read it before you run a flow. A live cluster spends real tokens. -
In another terminal, export the token from the Secret and run a ticket:
export AFE_API_TOKEN=$(kubectl --context kind-afe -n afe get secret afe-secrets \-o jsonpath='{.data.AFE_API_TOKEN}' | base64 -d)afe run writer-reviewer -i request.yaml # stops at the approvalafe resume T-0001 approveafe workerscurl -s 127.0.0.1:9465/metrics | grep afe_tickets -
To run on real models, put
OPENROUTER_API_KEYin.env(never committed). A plainmake kind-upthen goes live: it applies thedeploy/kind-liveoverlay and the workers run without--dry. The Secret gets the key. That is the safe default now, so a real key does not stay unused while the workers run fake models.make kind-up LIVE=0forces dry even when the key is present.make kind-up LIVE=1is explicit live. Without the key it stops with a message.- With no key and no
LIVE, the workers stay dry.
Editing
.envand runningmake kind-upagain is enough: it rewrites the Secret and restartsserveand the workers.The running cluster reports its mode too. Query the worker's annotation:
kubectl --context kind-afe -n afe get deploy worker \-o jsonpath='{.metadata.annotations.afe\.dev/model-mode}'# dry or live
Testing the cluster
make kind-test runs the tests against the cluster:
- the V6e tests run a ticket through two workers and check the metrics. They also scale Redis to zero (the Pods turn unready and are not restarted), and queue a ticket with no worker.
- the V7b tests run a
gitanddeepflow todone. They delete the ticket Pod after a checkpoint, and check that a worker restores the workspace on a new Pod without running a completed node again. They also check network isolation: a Pod reaches nothing by default, or only an allowlisted host when the project has one. Last, they check that the ticket's Pod and PVC are gone at the end.
make kind-test expects dry workers, and it is not part of make ci. make kind-down deletes
the cluster.
Every GitHub Actions job of this repository runs on the board's self-hosted runner, because the
GitHub-hosted runners are blocked by the billing. runs-on: ubuntu-latest, or any other
GitHub-hosted runner, is forbidden. The guard in tests/ci/test_self_hosted_runner.py fails
make ci on a hosted job, reads every workflow file .yml and .yaml alike, and refuses a
job-level uses: reusable-workflow call. A nightly run repeats make kind-test on the
runner and reports every red test on one kind-regression issue (.github/workflows/kind.yml). It
never blocks make ci or a merge. Without a self-hosted runner, deploy/systemd/install.sh
installs the same run as a user timer on the operator's machine.
The kind workloads carry a CPU and a memory limit too: serve, the workers, Redis and the
CloudNativePG Postgres. The nightly e2e/integration suites and the kind run share one
concurrency group, so two heavy runs never overlap on the runner.
Runner disk hygiene
A self-hosted runner keeps the Docker build cache, the images, the volumes and the uv cache between
jobs. So a nightly kind run would grow the disk by gigabytes each time. The kind job runs
scripts/ci/runner_hygiene.sh guard 20 before make kind-up. It frees the space a previous run
left behind. It stops the job with a clear message when / has fewer than 20 GB free. A
runner cleanup step runs the same script with clean in always(), after make kind-down, so
the space is freed even when the run fails earlier.
scripts/ci/runner_hygiene.sh has two subcommands. clean caps the build cache at 3 GB, then
prunes the dangling images, the stopped containers, the unused volumes and the uv cache. It prints
df -h / and docker system df before and after. guard <min-free-gb> runs clean first, then
fails when the free space on / is below the threshold. The script is written for the dedicated
runner. docker volume prune frees every unused volume, which is fine when the only containers on
the host are the throwaway ones a kind run creates.
Troubleshooting
LIVE=1 needs OPENROUTER_API_KEY in .env
make kind-up LIVE=1 ran without the key in .env. Add OPENROUTER_API_KEY=... to .env and
run the command again.
The flow finishes instantly, as if the model were fake
The workers are on the dry harness. make kind-up prints its mode as the last line. The worker's
afe.dev/model-mode annotation also reads dry. Put a real OPENROUTER_API_KEY in .env and run
make kind-up again. If you kept LIVE=0, drop it.
make kind-forward shows nothing on 127.0.0.1:8765
The port forward runs in the foreground and must stay open in its own terminal. Check it is still
running, and that no other process holds port 8765, 9464 or 9465.
afe run or afe resume cannot reach the API
AFE_API_TOKEN in your shell does not match the cluster's Secret, or make kind-forward is not
running. Re-export the token with the command in step 2 and confirm the forward is up.