Skip to content

Production deployment

The RootPilot Agent ships as a container image. It runs inside your infrastructure, reads your credentials from your secret manager, and opens a single outbound connection to the control-plane.

Before you start, run the Quickstart: it validates the image, the certificate, and the tunnel without involving any of your credentials. If something in the plumbing is wrong, it’s far cheaper to find out there.

Item What it is
Image address On ECR or Artifact Registry, with your account/project authorized to pull.
An identity path Where your platform proves identity on its own (GKE, EC2/EKS), nothing — you configure attestation. Otherwise, a bootstrap token you mint under Onboarding. The certificate we issue (runner.crt · runner.key · ca.pem) is the bridge for everything else.
RUNNER_TUNNEL_URL The tunnel address.

There is no RootPilot credential for you to store or rotate: the pull uses your own cloud’s identity.

Janela do terminal
# AWS (ECR)
docker pull <rootpilot-account>.dkr.ecr.<region>.amazonaws.com/rootpilot-agent:<version>
# GCP (Artifact Registry)
docker pull <region>-docker.pkg.dev/<rootpilot-project>/rootpilot/rootpilot-agent:<version>

Multi-architecture: the same address serves amd64 and arm64 (Graviton, Tau, Axion).

The registry is private: authenticate with your credentials before pulling, and make sure the identity your orchestrator uses is authorized too — see The image › Authenticate before pulling.

The container runs as a non-root user (uid 1000). It needs no privileges, no extra capabilities, and no access to the Docker socket.

The Agent authenticates over mTLS, and the certificate is short-lived — 15 minutes, renewed on its own. The choice here is where the first identity comes from, and it decides whether a deploy needs someone in a browser. The whole subject is in Identity and enrollment; in production, the short version is:

If your platform proves identity (GKE, or anywhere with an AWS credential — EC2, EKS), use cloud attestation. It is the best option because there is no secret to mint, hand over, or rotate:

Janela do terminal
RUNNER_ENROLL_URL=https://app.rootpilot.sh
RUNNER_ATTESTATION_MODE=gke-oidc # or aws-sts

Otherwise (VPS, your own host), use a bootstrap token — and persist the identity, or every container replacement costs a new token, minted by hand:

Janela do terminal
RUNNER_ENROLL_URL=https://app.rootpilot.sh
RUNNER_ATTESTATION_MODE=bootstrap-token
RUNNER_BOOTSTRAP_TOKEN=<minted under Onboarding>
RUNNER_IDENTITY_DIR=/identity # volume writable by uid 1000
RUNNER_INSTANCE_ID=agent-prod-1 # optional: identifies the instance in logs and in Fleet

If RootPilot handed you a certificate, mount the three PEMs at /certs, read-only — the entrypoint loads them for you:

/certs/runner.crt
/certs/runner.key
/certs/ca.pem

If your orchestrator injects secrets as variables rather than files, pass the contents inline in RUNNER_CERT_PEM, RUNNER_KEY_PEM, and RUNNER_CA_PEM. It is equivalent. That certificate does not renew: put the notAfter from the boot log on the calendar.

The Agent reads your services’ credentials from your secret manager at boot and holds them in memory only. Pick the backend:

Janela do terminal
RUNNER_SECRET_STORE_MODE=aws # AWS Secrets Manager
RUNNER_SECRET_STORE_MODE=vault # HashiCorp Vault KV v2

Details for each in Secret stores; which keys each service asks for, in Connector credentials.

If the Agent runs in the same account you want to observe, prefer the workload’s own identity over a static key — RUNNER_AWS_CRED_MODE=chain or RUNNER_GCP_CRED_MODE=chain. See Cloud credentials.

The minimum:

Janela do terminal
RUNNER_TUNNEL_URL=wss://tunnel.rootpilot.sh:8443
RUNNER_CONNECTORS=real
RUNNER_SECRET_STORE_MODE=aws
NODE_ENV=production

The full reference is in Environment variables. For Kubernetes, see Kubernetes — especially the part about probes, which has a catch.

Logs are JSON, one line per event. The one that matters is the last:

{"level":"info","message":"hello accepted","context":"runner","sessionId":"..."}

From there the Agent appears in the control-plane Fleet view, with the state of each connected service: whether the credential exists and whether it works. An invalid or unauthorized key shows up there as an error, with the cause.

There is no HTTP health endpoint — the Agent is an outbound client and listens on no port.

The certificate has an expiry date and does not renew itself. When it expires, the Agent stops connecting and the whole fleet stops answering.

notAfter is logged on every boot, and it is your advance warning:

{"level":"info","message":"using provisioned cert","context":"runner","notAfter":"2027-08-10T00:00:00.000Z"}

Treat the swap as planned maintenance: RootPilot issues the new pair, you replace the contents mounted at /certs and recycle the replicas. No data is lost — the Agent holds no durable state.

Make sure your supervisor (ReplicaSet, ECS service, Compose restart) brings up a replacement when the process exits. The Agent is disposable by construction: on SIGTERM it stops accepting new work, finishes what is in flight, and exits, within RUNNER_DRAIN_TIMEOUT_MS — 2 minutes by default. Give the orchestrator a stop timeout larger than that, or SIGKILL lands mid-drain.

Because the Agent is an outbound client with no health endpoint, a process left standing but not serving would be invisible to any probe. So it prefers to exit over degrading silently — and relies on the supervisor to bring up a clean replacement.