Skip to main content

Docker sandbox

Goal: run commands in the Docker sandbox and open only the network a project needs.

Prerequisites​

  • Docker, reachable from the machine that runs the worker.
  • A Project manifest with sandbox.image set.

Steps​

  1. Run commands from a deep agent with execute: true, or from a script node. Both use the container the afe-docker plugin (sandboxes/docker) opens.

  2. Let an acp agent list the environment variables it needs, by name. The Docker adapter passes each one as -e NAME. The value is read from the docker client's own environment, and it never appears in an argv. Forge tokens and engine keys never enter a sandbox.

  3. Run a node with run: <Script name>. It runs the Script's command in the project's sandbox, in /workspace. The artifact named by writes is filled from ok, exit_code and output, and from the keys of a JSON object on stdout. A non-zero exit is an outcome, not a failure.

  4. Use run: builtin:merge_and_ci to commit each lane worktree, merge the lane branches into the integration branch, and run ci.script in the integration sandbox. The artifact carries merged, conflicts, ci_passed and ci_output. A merge conflict or a red CI is an outcome the flow routes on with when.

  5. When a project needs the network, declare sandbox.egress:

    sandbox:
    image: afe-sandbox-acp:test
    egress: ["openrouter.ai"]

    The engine creates an internal network, and starts one squid proxy from docker/egress-proxy with the allowlist in AFE_EGRESS_HOSTS. It attaches the proxy to the default bridge, and puts the sandbox on the internal network with HTTP_PROXY/HTTPS_PROXY pointing at the proxy. Only the listed hosts leave. Every other host is refused, and so is a direct connection that ignores the proxy.

The container​

By default a sandbox has no network (--network none). The container is:

  • one per ticket and scope, with a fixed name and the label afe.ticket=<id>. It is found again after a restart.
  • for a git workspace, one named volume per ticket (afe-<key>-ws) and one container per ticket (afe-<key>), shared by every scope (V7b, R14). For a folder workspace the host folder stays the mount.
  • no network by default, all capabilities dropped, and no-new-privileges. CPU, memory and pids are limited, and the process runs as a non-root user.
  • mounted only at /workspace. The Docker socket is never mounted.
  • limited to the environment variables an acp agent declares by name. No forge token and no engine key ever enters.

A command past its timeout is killed and reported as timed out. The sandbox, the proxy, the network and the workspace volume all carry the afe.ticket label. They are removed with the ticket, at the end (done, failed, cancelled), or on demand. A paused ticket keeps its container, so it finds it again.

A command or script in flight when a worker dies runs again after recovery. Commands are assumed repeatable.

Troubleshooting​

A script or deep command hangs and then fails with a timeout The command ran past sandbox.timeoutS (or the node's own timeout). The sandbox kills it and reports the outcome as timed out. It does not retry automatically. Raise the timeout, or split the command into smaller steps.

A network call from the sandbox fails or times out The host is not in sandbox.egress. Squid denies every host it does not allow, and a direct connection that bypasses the proxy has no route out. Add the host to egress.

The sandbox is gone after a restart, but the ticket still waits The container is removed when the ticket ends, not while it pauses. If the sandbox is missing for a ticket that still waits, check the worker's log for a lost sandbox. The engine removes the sandbox and its volume, then requeues the ticket. A checkpoint restore then drops the interrupted node's saved turn.

See also​