Skip to main content

Deploy on kind

Goal: run the engine on a local kind cluster and drive a ticket through it.

Prerequisites​

  • Docker.
  • kind.
  • kubectl.

make kind-up deploys the whole engine on the cluster: serve, two workers on the dry harness, Postgres through CloudNativePG and Redis. It also deploys the V7b sandbox pieces: the afe-sandbox namespace, the worker's RBAC, netlock, and the images afe-sandbox:dev, afe-netlock:dev and afe-egress-proxy:dev.

Steps​

  1. Create the cluster and forward its ports:

    make kind-up # cluster `afe`, image afe:dev, CloudNativePG, Redis, serve and workers
    make kind-forward # API on 127.0.0.1:8765, ops on 9464 (serve) and 9465 (a worker); leave it on

    make kind-up takes the token from AFE_API_TOKEN in .env. Without that variable it keeps the token it generated the first time. The token lives only in the cluster's Secret.

    make kind-up also picks the model harness. Its last line says which one it chose. It reads kind: workers on the dry harness (fake models) or kind: workers on the live harness (real models). Read it before you run a flow. A live cluster spends real tokens.

  2. In another terminal, export the token from the Secret and run a ticket:

    export AFE_API_TOKEN=$(kubectl --context kind-afe -n afe get secret afe-secrets \
    -o jsonpath='{.data.AFE_API_TOKEN}' | base64 -d)
    afe run writer-reviewer -i request.yaml # stops at the approval
    afe resume T-0001 approve
    afe workers
    curl -s 127.0.0.1:9465/metrics | grep afe_tickets
  3. To run on real models, put OPENROUTER_API_KEY in .env (never committed). A plain make kind-up then goes live: it applies the deploy/kind-live overlay and the workers run without --dry. The Secret gets the key. That is the safe default now, so a real key does not stay unused while the workers run fake models.

    • make kind-up LIVE=0 forces dry even when the key is present.
    • make kind-up LIVE=1 is explicit live. Without the key it stops with a message.
    • With no key and no LIVE, the workers stay dry.

    Editing .env and running make kind-up again is enough: it rewrites the Secret and restarts serve and the workers.

    The running cluster reports its mode too. Query the worker's annotation:

    kubectl --context kind-afe -n afe get deploy worker \
    -o jsonpath='{.metadata.annotations.afe\.dev/model-mode}'
    # dry or live

Testing the cluster​

make kind-test runs the tests against the cluster:

  • the V6e tests run a ticket through two workers and check the metrics. They also scale Redis to zero (the Pods turn unready and are not restarted), and queue a ticket with no worker.
  • the V7b tests run a git and deep flow to done. They delete the ticket Pod after a checkpoint, and check that a worker restores the workspace on a new Pod without running a completed node again. They also check network isolation: a Pod reaches nothing by default, or only an allowlisted host when the project has one. Last, they check that the ticket's Pod and PVC are gone at the end.

make kind-test expects dry workers, and it is not part of make ci. make kind-down deletes the cluster.

Every GitHub Actions job of this repository runs on the board's self-hosted runner, because the GitHub-hosted runners are blocked by the billing. runs-on: ubuntu-latest, or any other GitHub-hosted runner, is forbidden. The guard in tests/ci/test_self_hosted_runner.py fails make ci on a hosted job, reads every workflow file .yml and .yaml alike, and refuses a job-level uses: reusable-workflow call. A nightly run repeats make kind-test on the runner and reports every red test on one kind-regression issue (.github/workflows/kind.yml). It never blocks make ci or a merge. Without a self-hosted runner, deploy/systemd/install.sh installs the same run as a user timer on the operator's machine.

The kind workloads carry a CPU and a memory limit too: serve, the workers, Redis and the CloudNativePG Postgres. The nightly e2e/integration suites and the kind run share one concurrency group, so two heavy runs never overlap on the runner.

Runner disk hygiene​

A self-hosted runner keeps the Docker build cache, the images, the volumes and the uv cache between jobs. So a nightly kind run would grow the disk by gigabytes each time. The kind job runs scripts/ci/runner_hygiene.sh guard 20 before make kind-up. It frees the space a previous run left behind. It stops the job with a clear message when / has fewer than 20 GB free. A runner cleanup step runs the same script with clean in always(), after make kind-down, so the space is freed even when the run fails earlier.

scripts/ci/runner_hygiene.sh has two subcommands. clean caps the build cache at 3 GB, then prunes the dangling images, the stopped containers, the unused volumes and the uv cache. It prints df -h / and docker system df before and after. guard <min-free-gb> runs clean first, then fails when the free space on / is below the threshold. The script is written for the dedicated runner. docker volume prune frees every unused volume, which is fine when the only containers on the host are the throwaway ones a kind run creates.

Troubleshooting​

LIVE=1 needs OPENROUTER_API_KEY in .env make kind-up LIVE=1 ran without the key in .env. Add OPENROUTER_API_KEY=... to .env and run the command again.

The flow finishes instantly, as if the model were fake The workers are on the dry harness. make kind-up prints its mode as the last line. The worker's afe.dev/model-mode annotation also reads dry. Put a real OPENROUTER_API_KEY in .env and run make kind-up again. If you kept LIVE=0, drop it.

make kind-forward shows nothing on 127.0.0.1:8765 The port forward runs in the foreground and must stay open in its own terminal. Check it is still running, and that no other process holds port 8765, 9464 or 9465.

afe run or afe resume cannot reach the API AFE_API_TOKEN in your shell does not match the cluster's Secret, or make kind-forward is not running. Re-export the token with the command in step 2 and confirm the forward is up.

See also​