Skip to content

Identity and enrollment

The RootPilot Agent authenticates over mTLS. This page is about the credential it uses for that: how it obtains one, how that credential stays alive, and — the question that gets expensive when nobody asks it first — what happens when the container is replaced.

Worth separating the two right away, because they get conflated:

  • How the identity begins — three paths, and you pick one.
  • How it survives — renewal (inside the process) and persistence (across container replacement).

The second half is what decides whether a deploy needs someone in a browser.

Path When to use it What you configure
Cloud attestation Where the platform proves identity on its own: GKE, EC2/EKS RUNNER_ATTESTATION_MODE + RUNNER_ENROLL_URL
Bootstrap token Where there is no platform identity: VPS, your own host, a first pilot RUNNER_BOOTSTRAP_TOKEN + RUNNER_ENROLL_URL
Provisioned certificate A bridge: we issue, you mount RUNNER_CERT_PEM + RUNNER_KEY_PEM

Cloud attestation is the best one wherever it exists, and the reason is concrete: there is no secret to mint, hand over, or rotate. The available modes and what each one requires are in Attestation.

Boot is a state machine: liveness → config → secrets → enroll → connect → ready. The enroll step does this:

  1. Generates an EC P-256 keypair. The private key exists in memory only, as PKCS#8.
  2. Generates a nonce and builds a CSR with it in the challengePassword field, which is the proof of possession.
  3. Gathers platform attestation, per RUNNER_ATTESTATION_MODE.
  4. POSTs to /api/enroll with { version, attestation, csr, challengeNonce }.
  5. The control-plane validates the attestation, signs the CSR, and returns a certificate with the tenant’s SPIFFE id inside: spiffe://rootpilot/tenant/<slug>/runner.
  6. The Agent opens the WSS tunnel with mTLS using that certificate.

The tenant comes from the certificate, not from configuration

Section titled “The tenant comes from the certificate, not from configuration”

This is worth insisting on, because it’s the guarantee holding up everything else: the attestation material determines the tenant, not an environment variable. Enrollment returns a certificate with the SPIFFE id inside, and that’s what the control-plane reads on every call.

RUNNER_TENANT, which shows up in the fleet scripts, picks no tenant at all: it picks the vault folder to read credentials from. Getting that slug wrong will not connect you to the wrong organization — and if the declared slug disagrees with what the credential enrolled as, the Agent refuses to come up rather than serve the wrong organization.

The certificate is short-lived — 15 minutes by default. That is deliberate: a certificate that is worth little time is worth little to whoever steals it.

RUNNER_CERT_RENEWAL=on is the default, and renewal fires at 2/3 of the time remaining, not near expiry. The reason is operational: renewing right at notAfter would turn any transient network instability into identity loss. Renewing at 2/3 leaves a third of the lifetime as a retry window.

Retries back off with a 5-minute ceiling and ±25% jitter — without the jitter, the entire fleet would hit the control-plane in the same second. If the server answers Retry-After, that wins over the local calculation.

Renewal requires RUNNER_ENROLL_URL: that is the address it talks to.

Renewal is not a RUNNER_ATTESTATION_MODE. Renewing isn’t a configuration choice: it’s what the Agent does when it already has a certificate.

It signs the new CSR’s nonce with the current private key and sends the current certificate along. The control-plane validates the chain against the CA, checks the signature with the certificate’s public key, and extracts the tenant from the certificate, never from the request body. Reading it from the body would turn renewal into a way to switch tenants.

Here is the difference that costs the most if nobody configures it.

Renewal lives inside the process. Replace the container and the identity dies with it — and the fallback is enrollment from scratch. Where the identity began from a bootstrap token, that means a new token, minted by hand, on every deploy (the token is single-use).

RUNNER_IDENTITY_DIR solves that: it points at a directory where the Agent writes the identity (identity.json, mode 600). On the next boot it renews from that instead of enrolling from scratch.

volumes:
- rootpilot-identity:/identity # named volume: inherits the right ownership from the image
environment:
RUNNER_IDENTITY_DIR: /identity

That is all of it: the variable is enough on its own. The first time the Agent uses the directory it writes an owner id there (instance-id, mode 600), so the identity is still recognized as its own after the container is recreated — which is the whole point of the feature.

RUNNER_INSTANCE_ID is still worth setting, but for a different reason: it identifies the instance in logs and in Fleet. It no longer has anything to do with persistence.

Three things worth knowing before you turn it on:

  • The directory has to be writable by the Agent’s uid (the image runs as non-root, uid 1000). A named volume takes care of that on its own: the image ships /identity with the right ownership, and the volume inherits it when created. A bind mount does not inherit — the ownership is the host’s. If it isn’t writable, boot does not fail: it warns and carries on, and you only find out on the next deploy, with could not persist identity — next boot will re-enroll from scratch.
  • One directory per Agent. Two containers on the same volume would share one identity.
  • The window has a limit. What survives is renewal, and it holds for up to 6 minutes past the certificate’s expiry (5 of grace + 1 of clock margin). A normal deploy fits with room to spare; a host switched off all night does not — that one is enrollment from scratch.

Past those 6 minutes, the control-plane answers presented cert expired beyond the renewal grace window, and it will answer that forever: the renewal attestation proves possession with the certificate that died.

The behavior is then to exit the process, so the supervisor brings up a fresh Agent that enrolls from scratch. Without that exit, the process would become a zombie: up, serving nothing, invisible to any probe.

identity expired beyond the renewal grace window — runner needs re-enrollment

RootPilot issues a certificate for your tenant and you mount it at /certs (or pass it inline in RUNNER_CERT_PEM/RUNNER_KEY_PEM/RUNNER_CA_PEM). It requires no platform attestation and no egress to the enrollment endpoint, and it is the bridge for environments where neither of the other two applies yet.

The cost is stated up front: it does not renew. When it expires, the fleet stops connecting, and the Agent neither exits the process nor warns — it keeps trying to reconnect indefinitely, with a TLS error that never says that is what happened. The date is in the notAfter of the boot line; put it on the calendar the day you receive the certificate.

The schema requires at least one identity path, and fails fast when there is none:

either provide RUNNER_CERT_PEM+RUNNER_KEY_PEM (provisioned cert) or RUNNER_ENROLL_URL (enrollment)