Skip to content

Install the Profiler

The RootPilot Profiler is the second artifact: optional, privileged, and installed by whoever wants the behavioural axis — what changed inside the process between two deploys, where metrics and logs only show the symptom.

It is a different image from the Agent and it requires privilege the Agent does not ask for. Before going further, read Security → RootPilot Profiler: that page carries the permission table, what leaves your infrastructure, and why there is no reduced mode. This page is the “how”, and assumes the decision is made.

  1. Enable the Profiler for the organization, under Settings → RootPilot Profiler in the app. Without it, enrolment is refused with a 403 before anything else — kernel privilege is a deliberate, recorded grant, not a side effect of a deploy.
  2. Write the target list. It is mandatory and the Profiler will not boot without it. See below.
  3. Have a bootstrap token for the tenant, as with the Agent.
Janela do terminal
# AWS (ECR)
docker pull <rootpilot-account>.dkr.ecr.<region>.amazonaws.com/rootpilot-profiler:<version>
# Google Cloud (Artifact Registry)
docker pull <region>-docker.pkg.dev/<rootpilot-project>/rootpilot/rootpilot-profiler:<version>

Multi-arch (amd64 + arm64) and cosign-signed, like the Agent’s — the verification procedure is the same one described in The image.

The image bundles the collector (ptop) and the aggregator (witness). The process that talks to the control plane runs as a non-root user with no capabilities at all; what carries the privilege is the collector binary, through file capabilities.

A JSON file mounted into the container, pointed at by RUNNER_PROFILE_TARGETS_FILE:

{
"targets": [
{ "workload": "checkout-api" },
{ "namespace": "production", "workload": "pricing-engine" }
]
}

It is yours and never passes through us. There is no default that means “everything on the node”: the Profiler refuses to boot on a missing file, an empty list, or "workload": "*". The schema is strict — any unknown key, including a well-meaning comment field, fails the boot.

A file and not an environment variable on purpose: in your infrastructure repository it is diffable, it goes through a review, and someone looks at it. It is literally the list of where the privileged container may touch.

services:
profiler:
image: <rootpilot-account>.dkr.ecr.<region>.amazonaws.com/rootpilot-profiler:<version>
restart: unless-stopped
environment:
RUNNER_TENANT: your-org
RUNNER_INSTANCE_ID: your-org-profiler01
RUNNER_TUNNEL_URL: wss://tunnel.rootpilot.sh:8443
RUNNER_ENROLL_URL: https://app.rootpilot.sh
RUNNER_IDENTITY_DIR: /identity
RUNNER_PROFILE_TARGETS_FILE: /etc/rootpilot/profile-targets.json
# All five, and not one more. Missing any of them fails the boot naming it and what it costs —
# see the security page.
cap_add: [BPF, PERFMON, SYS_PTRACE, DAC_READ_SEARCH, SYS_ADMIN]
# Visibility, not privilege: without it the container only sees its own processes, and the target
# simply does not exist as far as the collector is concerned.
pid: host
security_opt:
- apparmor=rootpilot-profiler # see "AppArmor" below
volumes:
- ./identity:/identity
- ./profile-targets.json:/etc/rootpilot/profile-targets.json:ro
# The eBPF library looks for tracepoints here. Without these two the collector loads and **does
# not attach** — and the capture comes back with axes missing, which is worse than empty.
- /sys/kernel/tracing:/sys/kernel/tracing
- /sys/kernel/debug:/sys/kernel/debug
mem_limit: 512m

The bootstrap token goes in as it does for the Agent (RUNNER_BOOTSTRAP_TOKEN, preferably via a secret store rather than plaintext in the compose file).

If your host has AppArmor enabled — Ubuntu and Debian do by default — the docker-default profile blocks what the Profiler needs to do: it only permits ptrace between containers under the same profile, and your target usually is not one (a host process is unconfined; a container from another runtime has a different profile).

Do not use apparmor=unconfined: that undoes the whole confinement over a single rule. Load a profile that is docker-default plus one line:

#include <tunables/global>
profile rootpilot-profiler flags=(attach_disconnected,mediate_deleted) {
#include <abstractions/base>
# … the docker-default body …
# The only difference: read /proc/<pid>/{exe,ns/*} of a target outside this profile.
ptrace (read) peer=unconfined,
}
Janela do terminal
sudo apparmor_parser -r -W /etc/apparmor.d/rootpilot-profiler

How to tell this is the problem: a capture comes back with collectors in permission denied over /proc/<pid>/ns/* or /proc/<pid>/exe, and the kernel logs the denial:

apparmor="DENIED" operation="ptrace" profile="docker-default" comm="ptop" peer="unconfined"

The pod needs:

  • hostPID: true (or shareProcessNamespace if the target is in the same pod). Without it the collector only sees itself.
  • All five capabilities in securityContext.capabilities.add.
  • /sys/kernel/tracing and /sys/kernel/debug mounted from the host as hostPath.
  • An AppArmor profile that permits ptrace (read) against the target’s profile. On containerd clusters the default is cri-containerd.apparmor.d, and the same reasoning as above applies — the profile must be loaded on every node and referenced in securityContext.appArmorProfile.
  • A volume for /identity, with the same lifetime the Agent needs — and not shared with another artifact, for the reason in the box above. In a Deployment that means one PVC per replica, or a StatefulSet.

No Kubernetes API permission is required today. The automatic trigger (noticing the pod come up with a new image) does not exist; when it does, it will ask for get/list/watch on pods, only in the namespaces of your allowlist, and the security page will be updated first.

DaemonSet: the routing exists, the cluster has not been exercised

Section titled “DaemonSet: the routing exists, the cluster has not been exercised”

Booting is not measuring, and the difference is the expensive one: a collector that does not attach produces a capture with axes missing, and a missing axis compared against a complete capture reads as behaviour that stopped.

Two checks, in this order:

  1. The boot. It declares the class and how many targets it read, and it refuses to boot naming the capability that is missing — there is no path where it comes up green and captures blind for lack of permission.
  2. The first capture. Open it under Profiling in the app. The screen declares the instrumentation: if it says N of M probes instead of fully instrumented, the block below names each collector that did not attach and why — usually AppArmor or the tracefs mounts.

A capture with incomplete instrumentation is not a weaker capture: it is a capture with axes that do not exist, and the product would rather say so than compare different worlds.

The trigger is explicit: the Profiler does not decide on its own when to measure. The request comes from your pipeline — right after publishing a version — or from someone on your team during an investigation. Two paths, and they have different doors because the hurry is different.

Under Profiling, at the top of the page, there is a form: process, version, environment. Any member of the organization can ask — it is the same access that opens the page, and the owner/admin decision was already made earlier, when the Profiler was turned on. The outcome shows up in the log right below, with who asked next to it.

Window, warmup and replicate captures are deliberately not in the form: whoever needs those is writing a pipeline, and the pipeline has the call below.

Under Settings → Profiler Capture Tokens, in the app. It starts with rpcap_, is shown once at creation, and is distinct from the MCP token and the incident ingress token: it reads nothing, it only queues a capture request.

Give each pipeline its own token — and a separate one for human use, if your team will be asking for captures during investigations. The name is what shows up in the capture record, and it is what answers “who asked for this”.

Janela do terminal
curl -X POST https://app.rootpilot.sh/api/profiling/captures \
-H "Authorization: Bearer $ROOTPILOT_CAPTURE_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"workload": "checkout-api",
"version": "v2.14.0",
"env": "production",
"duration_sec": 60
}'
field required what it is
workload yes the process name, exactly as it appears in your target list. Outside that list, the Profiler refuses.
version yes the version that just shipped. It is the key of the diff — the capture of v2.14.0 is compared against the one for v2.13.0.
env no production, staging… Empty is legitimate: not every fleet has the concept.
kind no (head) head is the capture of the version. replicate is the same version measured again, and it is what sustains the noise floor and the confirmation window — call it a few times for a more confident diff.
duration_sec no (60) between 5 and 300.
warmup_sec no (5) between 0 and 60 — the time discarded before measurement starts.

duration_sec/durationSec and warmup_sec/warmupSec are both accepted, and the token may also go as ?token= — some declarative CI steps will not let you set a header.

The responses, each asking for a different action:

  • 202 — {"request_id": "…", "status": "queued", "deduped": false}. The request is queued and drained within seconds. deduped: true means an identical request was already in flight and you got its id back: repeating the call does not stack captures.
  • 400 — the body failed validation, and the message names the field.
  • 401 — token missing, invalid or revoked.
  • 403 — the Profiler is not enabled for the organization. That is the opt-in, under Settings.
  • 429 — the ceiling is 20 requests per hour per organization, and the response carries retry_after_seconds. Per organization rather than per IP on purpose: what pays the bill is the observed process, not the network the request came from.

A 202 says the request landed, not that the capture happened. There may be no live Profiler at that moment (the request waits), and the target may not be running on any node. The outcome — pushed, nothing_to_push or failed, with the reason next to it — shows up under Profiling in the app.

We say this here because the alternative is you finding out on your own:

  • Kubernetes, on any cluster.
  • JVM: the symbol axis for the JVM is implemented and unit-tested, and has never run against a live JVM.

Docker on amd64 and arm64 is what is exercised.