Production deployment
The RootPilot Agent ships as a container image. It runs inside your infrastructure, reads your credentials from your secret manager, and opens a single outbound connection to the control-plane.
Before you start, run the Quickstart: it validates the image, the certificate, and the tunnel without involving any of your credentials. If something in the plumbing is wrong, it’s far cheaper to find out there.
What RootPilot gives you
Section titled “What RootPilot gives you”| Item | What it is |
|---|---|
| Image address | On ECR or Artifact Registry, with your account/project authorized to pull. |
| An identity path | Where your platform proves identity on its own (GKE, EC2/EKS), nothing — you configure attestation. Otherwise, a bootstrap token you mint under Onboarding. The certificate we issue (runner.crt · runner.key · ca.pem) is the bridge for everything else. |
RUNNER_TUNNEL_URL |
The tunnel address. |
There is no RootPilot credential for you to store or rotate: the pull uses your own cloud’s identity.
1. The image
Section titled “1. The image”# AWS (ECR)docker pull <rootpilot-account>.dkr.ecr.<region>.amazonaws.com/rootpilot-agent:<version>
# GCP (Artifact Registry)docker pull <region>-docker.pkg.dev/<rootpilot-project>/rootpilot/rootpilot-agent:<version>Multi-architecture: the same address serves amd64 and arm64 (Graviton, Tau, Axion).
The registry is private: authenticate with your credentials before pulling, and make sure the identity your orchestrator uses is authorized too — see The image › Authenticate before pulling.
The container runs as a non-root user (uid 1000). It needs no privileges, no extra capabilities, and no access to the Docker socket.
2. Identity
Section titled “2. Identity”The Agent authenticates over mTLS, and the certificate is short-lived — 15 minutes, renewed on its own. The choice here is where the first identity comes from, and it decides whether a deploy needs someone in a browser. The whole subject is in Identity and enrollment; in production, the short version is:
If your platform proves identity (GKE, or anywhere with an AWS credential — EC2, EKS), use cloud attestation. It is the best option because there is no secret to mint, hand over, or rotate:
RUNNER_ENROLL_URL=https://app.rootpilot.shRUNNER_ATTESTATION_MODE=gke-oidc # or aws-stsOtherwise (VPS, your own host), use a bootstrap token — and persist the identity, or every container replacement costs a new token, minted by hand:
RUNNER_ENROLL_URL=https://app.rootpilot.shRUNNER_ATTESTATION_MODE=bootstrap-tokenRUNNER_BOOTSTRAP_TOKEN=<minted under Onboarding>RUNNER_IDENTITY_DIR=/identity # volume writable by uid 1000RUNNER_INSTANCE_ID=agent-prod-1 # optional: identifies the instance in logs and in FleetIf RootPilot handed you a certificate, mount the three PEMs at /certs, read-only — the entrypoint
loads them for you:
/certs/runner.crt/certs/runner.key/certs/ca.pemIf your orchestrator injects secrets as variables rather than files, pass the contents inline in
RUNNER_CERT_PEM, RUNNER_KEY_PEM, and RUNNER_CA_PEM. It is equivalent. That certificate does not
renew: put the notAfter from the boot log on the calendar.
3. Secrets
Section titled “3. Secrets”The Agent reads your services’ credentials from your secret manager at boot and holds them in memory only. Pick the backend:
RUNNER_SECRET_STORE_MODE=aws # AWS Secrets ManagerRUNNER_SECRET_STORE_MODE=vault # HashiCorp Vault KV v2Details for each in Secret stores; which keys each service asks for, in Connector credentials.
If the Agent runs in the same account you want to observe, prefer the workload’s own identity over a
static key — RUNNER_AWS_CRED_MODE=chain or RUNNER_GCP_CRED_MODE=chain. See
Cloud credentials.
4. Configuration
Section titled “4. Configuration”The minimum:
RUNNER_TUNNEL_URL=wss://tunnel.rootpilot.sh:8443RUNNER_CONNECTORS=realRUNNER_SECRET_STORE_MODE=awsNODE_ENV=productionThe full reference is in Environment variables. For Kubernetes, see Kubernetes — especially the part about probes, which has a catch.
5. Confirm
Section titled “5. Confirm”Logs are JSON, one line per event. The one that matters is the last:
{"level":"info","message":"hello accepted","context":"runner","sessionId":"..."}From there the Agent appears in the control-plane Fleet view, with the state of each connected service: whether the credential exists and whether it works. An invalid or unauthorized key shows up there as an error, with the cause.
There is no HTTP health endpoint — the Agent is an outbound client and listens on no port.
Certificate rotation
Section titled “Certificate rotation”The certificate has an expiry date and does not renew itself. When it expires, the Agent stops connecting and the whole fleet stops answering.
notAfter is logged on every boot, and it is your advance warning:
{"level":"info","message":"using provisioned cert","context":"runner","notAfter":"2027-08-10T00:00:00.000Z"}Treat the swap as planned maintenance: RootPilot issues the new pair, you replace the contents
mounted at /certs and recycle the replicas. No data is lost — the Agent holds no durable state.
Supervision
Section titled “Supervision”Make sure your supervisor (ReplicaSet, ECS service, Compose restart) brings up a replacement when
the process exits. The Agent is disposable by construction: on SIGTERM it stops accepting new work,
finishes what is in flight, and exits, within RUNNER_DRAIN_TIMEOUT_MS — 2 minutes by default. Give
the orchestrator a stop timeout larger than that, or SIGKILL lands mid-drain.
Because the Agent is an outbound client with no health endpoint, a process left standing but not serving would be invisible to any probe. So it prefers to exit over degrading silently — and relies on the supervisor to bring up a clean replacement.