Skip to content

RootPilot Profiler

Metrics and logs show the symptom: latency went up, errors appeared. They do not tell you what changed inside the process between the version that worked and the one that does not.

The RootPilot Profiler measures that. It watches a process for a short window after each deploy and keeps a compact behavioural signature — how much the program allocates and where, which system calls it makes, how lock waits are distributed, how many network peers it has. Comparing two versions lets us say “this is the change” instead of “something got worse”.

That costs a permission the rest of RootPilot never asks for. This page exists so you can decide with the list in hand.

It is a second agent, and that is deliberate

Section titled “It is a second agent, and that is deliberate”

Flipping a flag on the Agent would be simpler. We do not do that because a flag makes the privilege invisible: whoever inherits the environment months later has no way to tell, from the deployment, that this container can now read other processes’ memory. As two artefacts, the privilege shows up at the front door.

The separation does not rely on discipline. The Agent refuses to start if it detects a kernel capability on its own process. And the Profiler cannot be configured to read your integrations: its image registers no data connector at all, so no variable turns it into a reader of your stack — one that asks for it is ignored, and it says so in the boot log. Neither becomes the other through a configuration mistake.

Up to version 1.15.0 this page promised more than the code delivered, in the least convenient direction: the Profiler started up serving the synthetic dataset that exists for development. Nothing of your infrastructure was read — that dataset is fabricated — but it took a source’s place in the fleet. Fixed at the root: the artefact’s class decides, not the deployment’s configuration.

The hardest gate here is the kernel, and it rejects hosts nobody expects to see rejected — which is why it is the first question in qualification, not the last. A host that does not clear it never reaches the POC; finding that out during the POC costs the POC.

Run the preflight on the candidate host. It only reads /proc and /sys, changes nothing, and needs no root:

Janela do terminal
curl -O https://docs.rootpilot.sh/profiler-preflight.sh
sh profiler-preflight.sh

Download it and read it before running — the whole script is published right below, and nobody should run someone else’s script unread. That is why we do not offer a curl | sh.

Gate Why it rejects, and how it would fail silently
Kernel ≥ 5.8 BTF, the ring buffer and CAP_BPF all land in that version. The filter is the kernel, not the distribution: RHEL 7 / 3.10 is out, and the non-eBPF fallback is only /proc.
BTF compiled in Version ≥ 5.8 does not imply BTF: a kernel built without CONFIG_DEBUG_INFO_BTF clears the previous gate and fails this one. That is why they are two gates and not one.
cgroup v2 Without the unified hierarchy there is no way to scope a capture by cgroup.
tracefs mounted It is where the eBPF library looks for tracepoints. Without it the collectors load, do not attach, and the capture comes back with axes missing rather than empty — which reads as “that behaviour stopped”.
Confinement profile docker-default only allows ptrace between containers under the same profile. Inheriting it costs two collectors, silently.

The script exits with three codes, and the middle one is what qualification checklists usually lack: 0 everything passed · 1 a hard gate failed · 2 nothing failed and something could not be decided from there.

The normal case for 2 is the five capabilities: the shell running the script is not the Profiler’s workload, so what that shell holds does not answer what the Profiler will be allowed to hold — that is a policy question for whoever runs the cluster. A checklist that counted this as “ok” would hand out a pass nobody verified, which is exactly the failure the script exists to prevent. (Its first version committed that failure in reverse: it checked /sys/kernel/tracing/events with test -d, which fails for an unprivileged user because the directory is 0700 root:root — rejecting a perfectly mounted host. It now reads /proc/mounts, which is world-readable.)

The question that comes before the technical ones

Section titled “The question that comes before the technical ones”

Attaching a privileged collector to third-party software — a vendor’s, COTS, legacy nobody maintains — may violate or void the support contract you hold with whoever sells it. That is a legal question, and it therefore comes before the technical ones: no script result replaces it. The script prints it alongside the other two it cannot answer (hostPID and the target list), so it is not forgotten just because everything else came back green.

#!/bin/sh
# RootPilot Profiler — host preflight.
#
# Answers one question, before anyone installs anything or grants a capability:
# CAN this host run the Profiler at all? It reads; it changes nothing.
#
# Written in English on purpose: this is one canonical file served to readers of
# both the Portuguese and the English documentation, and a single script cannot
# be bilingual. Each check is explained, in your language, on the page that
# publishes this script verbatim: https://docs.rootpilot.sh/security/profiler/
#
# Exit codes — a check that could not run is NOT a pass:
# 0 every gate passed, and nothing was left unknown
# 1 at least one hard gate FAILED — this host is out until it is fixed
# 2 no gate failed, but something could not be determined from here
#
# Run it as the least privileged user you have. It needs no root: every gate it
# can decide is decidable by reading /proc and /sys.
set -u
PASS=0
FAIL=0
UNKNOWN=0
pass() { printf ' PASS %s\n' "$1"; PASS=$((PASS + 1)); }
fail() { printf ' FAIL %s\n ↳ %s\n' "$1" "$2"; FAIL=$((FAIL + 1)); }
unknown() { printf ' UNKNOWN %s\n ↳ %s\n' "$1" "$2"; UNKNOWN=$((UNKNOWN + 1)); }
printf 'RootPilot Profiler — host preflight\n'
printf 'host: %s kernel: %s arch: %s\n\n' "$(uname -n)" "$(uname -r)" "$(uname -m)"
# ── 1. Kernel ≥ 5.8 ───────────────────────────────────────────────────────────
# BTF, the BPF ring buffer and CAP_BPF all land in 5.8. Below it the Profiler
# refuses to boot rather than run with half its collectors.
release=$(uname -r)
major=$(echo "$release" | cut -d. -f1)
minor=$(echo "$release" | cut -d. -f2 | cut -d- -f1)
case "$major$minor" in
*[!0-9]*|'')
unknown "kernel >= 5.8" "could not parse a version out of '$release'" ;;
*)
if [ "$major" -gt 5 ] || { [ "$major" -eq 5 ] && [ "$minor" -ge 8 ]; }; then
pass "kernel >= 5.8 ($release)"
else
fail "kernel >= 5.8 ($release)" \
"this host is out. The gate is the kernel, not the distribution: RHEL 7 / 3.10 cannot run this, and --no-ebpf is only /proc."
fi ;;
esac
# ── 2. BTF actually compiled in ───────────────────────────────────────────────
# Version >= 5.8 does NOT imply BTF: a kernel built without CONFIG_DEBUG_INFO_BTF
# passes check 1 and fails here. This is the gate that surprises people, so it is
# separate from the version rather than folded into it.
if [ -r /sys/kernel/btf/vmlinux ]; then
pass "kernel BTF present (/sys/kernel/btf/vmlinux)"
else
fail "kernel BTF present" \
"no /sys/kernel/btf/vmlinux. The kernel was built without CONFIG_DEBUG_INFO_BTF; the version alone never proved it."
fi
# ── 3. cgroup v2 ──────────────────────────────────────────────────────────────
if [ -r /sys/fs/cgroup/cgroup.controllers ]; then
pass "cgroup v2 unified hierarchy"
else
fail "cgroup v2 unified hierarchy" \
"no /sys/fs/cgroup/cgroup.controllers. A v1-only or hybrid host cannot scope a capture by cgroup."
fi
# ── 4. tracefs mounted ────────────────────────────────────────────────────────
# This one earns its own gate because of HOW it fails: without tracefs the eBPF
# programs load and never attach, so the capture comes back with axes MISSING
# rather than empty — which reads like "nothing changed in this deploy".
#
# It is read from /proc/mounts, which is world-readable, and NOT with
# `test -d /sys/kernel/tracing/events`. That was the first version, and it is
# wrong in the direction that matters: tracefs is 0700 root:root, so descending
# into it as an unprivileged user fails on a host where it is perfectly mounted.
# A check that reports FAIL when it could not look is the failure this whole
# script exists to prevent.
if [ -r /proc/mounts ]; then
mounts=$(cat /proc/mounts)
if echo "$mounts" | grep -qE '^tracefs +/sys/kernel/tracing '; then
pass "tracefs mounted (/sys/kernel/tracing)"
elif echo "$mounts" | grep -qE '^(debugfs|tracefs) +/sys/kernel/debug '; then
pass "tracefs reachable under /sys/kernel/debug/tracing"
else
fail "tracefs mounted" \
"no tracefs in /proc/mounts. Collectors would load, attach nothing, and the capture would come back missing axes instead of empty."
fi
else
unknown "tracefs mounted" "could not read /proc/mounts from here"
fi
# ── 5. The five capabilities ──────────────────────────────────────────────────
# What this shell holds is NOT the answer — the Profiler runs as its own
# workload, and whether the five can be granted to it is a policy question for
# whoever owns the cluster (PodSecurity / PSP / the container runtime). So this
# reports and does not decide.
caps=$(grep -m1 '^CapEff:' /proc/self/status 2>/dev/null | awk '{print $2}')
if [ -n "${caps:-}" ]; then
unknown "CAP_BPF, CAP_PERFMON, CAP_SYS_PTRACE, CAP_DAC_READ_SEARCH, CAP_SYS_ADMIN grantable" \
"not decidable from this shell (it holds CapEff=$caps). Ask whoever owns the cluster whether a workload may be granted all five. All five are required — the Profiler fails boot naming the one that is missing, and there is no reduced mode."
else
unknown "the five capabilities grantable" \
"could not read /proc/self/status. Ask whoever owns the cluster whether a workload may be granted all five."
fi
# ── 6. Confinement profile ────────────────────────────────────────────────────
# Measured, not theorised: Docker's default seccomp/AppArmor profile denies
# ptrace against an unconfined target, and that alone kills two collectors.
lsm=$(cat /proc/self/attr/current 2>/dev/null | tr -d '\0')
case "${lsm:-}" in
''|unconfined)
pass "no restrictive LSM profile on this shell (${lsm:-none})" ;;
*)
unknown "confinement profile" \
"this shell runs under '$lsm'. Docker's default profile denies ptrace against an unconfined target, which silently costs two collectors. The Profiler needs its own profile decided, not inherited." ;;
esac
# ── 7. Questions this script cannot answer ────────────────────────────────────
printf '\nNot decidable from a host, and not optional:\n'
printf ' • Does your support contract with the vendor of the observed software\n'
printf ' ALLOW attaching a privileged collector to it? Doing so can void support.\n'
printf ' This is a legal question, and it comes before the technical ones.\n'
printf ' • Can the Profiler run with hostPID, so it can see the target process?\n'
printf ' • Which processes go on the target list? It is required, it is yours,\n'
printf ' and there is no default that means "everything on the node".\n'
printf '\n%d passed, %d failed, %d unknown\n' "$PASS" "$FAIL" "$UNKNOWN"
if [ "$FAIL" -gt 0 ]; then
printf 'VERDICT: this host does NOT qualify yet.\n'
exit 1
fi
if [ "$UNKNOWN" -gt 0 ]; then
printf 'VERDICT: no gate failed, and %d item(s) could not be decided from here.\n' "$UNKNOWN"
printf 'Not a pass. Answer them before the POC, not during it.\n'
exit 2
fi
printf 'VERDICT: this host qualifies.\n'
exit 0
Permission What for
CAP_BPF Load the eBPF programs that count system calls, lock waits, allocations and network events.
CAP_PERFMON Sample CPU via perf_event, at 100 Hz per core.
CAP_SYS_PTRACE Attach those programs to a process running as another user — reading /proc/<pid>/ns/* and /proc/<pid>/exe. Without it eBPF loads and no collector attaches: the capture comes back empty, which is indistinguishable from “nothing changed in this deploy”.
CAP_DAC_READ_SEARCH Read /sys/kernel/tracing (which is 0700 root:root) and the collector’s own /proc/self/mem. The second one looks odd and is not: a binary that receives capabilities as file capabilities — which is how this one avoids running as root — ends up with its own /proc/self/* owned by root. Without it the tracepoints and the kernel-version detection both fail, which is seven of the twelve collectors.
CAP_SYS_ADMIN Create the perf_uprobe behind the heap probe — the axis that carries func and file:line. It is root-equivalent in a container with hostPID. Read the block below before approving.
hostPID or shareProcessNamespace See the target process. Without it, it only sees itself.
/sys/kernel/tracing and /sys/kernel/debug mounted Where the eBPF library looks for tracepoints. Without them the collector loads and does not attach, and the capture comes back with axes missing rather than empty.
A dedicated AppArmor profile docker-default only permits ptrace between containers under the same profile, and your target usually is not one. We publish a profile that is docker-default plus one line (ptrace (read) peer=unconfined) — not apparmor=unconfined, which would undo the whole confinement over a single rule.
get · list · watch on pods Not requested yet. It will be how it notices a deploy once the automatic trigger exists, and only in the namespaces on your allowlist. Today the Profiler needs no Kubernetes permission at all.
Kernel ≥ 5.8, cgroup v2 BTF, ring buffer and CAP_BPF. Below that it refuses to start rather than run half-blind.
Outbound network The RootPilot tunnel only. No other destination.

What it does not ask for matters as much: no write permission in Kubernetes. In particular, it does not ask for create on pods/ephemeralcontainers — the permission the alternative approach would need, and one that is uncomfortably generic: whoever can attach an ephemeral container can attach any image to any pod in the namespace.

You declare the list, and it is mandatory. There is no default meaning “anything on the node”: an agent that sees every process by omission is privilege beyond need, and omission should never be how the widest possible scope gets granted.

The list lives in a file you deploy alongside the Profiler, by namespace and workload. Outside it, the Profiler does not attach — and the refusal shows up in its log, naming the target, so that “did not capture” is never confused with “captured and found nothing”.

It differs from the tool denylist in one way worth understanding: the denylist is composed in the control plane and enforced at the edge; this list is entirely yours and never passes through us. It does not say what RootPilot may read — it says which of your processes anyone may touch.

For a short window — 60 seconds by default — on one process at a time. It does not measure continuously.

Today the trigger is explicit. A capture is asked for: by your team during an investigation, or by your pipeline after a deploy. The Profiler receives the request over the same connection it already uses for everything else, checks the target against your list before looking at any process, measures, and returns the result.

Both go through the same queue, by different doors. From the app, under Profiling, any member of the organization asks through a form — the access is the same one that opens the page, and the owner/admin decision was already made earlier, when the Profiler was turned on. From the pipeline, through an authenticated call carrying an rpcap_ token that an owner or admin issues, and which only queues a capture request: it reads nothing.

The steps, the fields and the limits are in Install the Profiler → Request a capture. In both cases the capture record keeps who asked — the person or the pipeline, by name — next to the outcome.

On Docker there will be no automatic trigger, even later. We could detect deploys by listening to the Docker socket, and chose not to ask: mounting the Docker socket is equivalent to root on the machine, and would be the largest permission on this entire page — for a convenience.

What crosses is the aggregated signature — numbers and names, never what the process handles:

  • symbol, file and module of the code site that allocated memory;
  • name and count of system calls, with average latency;
  • the shape of lock contention, without memory addresses;
  • network counters: how many peers, errors by kind, average latency;
  • memory and CPU distributions;
  • process name, version and environment;
  • which instance was measured — the pod UID, the container id, or the PID when nothing more stable is available. A capture is always of one process; in Kubernetes your workload has several replicas, and without that identifier two captures of different replicas would enter the same set as if they were the same thing observed twice.

Symbols and file paths from your code cross the boundary. It is the same class of data the code connector already reads once you connect it, and we say so explicitly because it cannot be redacted: erasing the function name destroys the only thing that makes the signature useful. If that class of data cannot leave your boundary, the right path is not to install the Profiler — not to install it and trust a mask.

Never leaves, under any configuration:

  • request, response or log content, or any payload;
  • TLS plaintext — the collector can capture it, and the option is off and locked;
  • peer IP addresses; only resolved service identity;
  • the process’s argv and environment variables;
  • any disk read outside the target’s own /proc.

It is written to the RootPilot control-plane, isolated per organization like everything else we keep: no tenant reads another’s signature, and deleting the organization deletes the signatures with it.

The Profiler stores nothing. It measures, pushes and forgets — there is no database inside the privileged artifact, which is why restarting it neither loses nor accumulates anything.

What interprets the signature is a service of ours running beside the control-plane. It receives the documents in the request, computes the comparison and answers: no database, none of your credentials, and no retention of what it received.

The signature feeds no pool shared between customers. What RootPilot promotes across tenants is anonymized and categorical (service roles, metric kinds); a behavioural signature carries symbols and paths from your code, so it stays where it was written.

A JSON file, pointed at by RUNNER_PROFILE_TARGETS_FILE, mounted into the container:

{
"targets": [
{ "workload": "checkout-api" },
{ "namespace": "production", "workload": "pricing-engine" }
]
}

namespace is a Kubernetes concept and is optional — in Docker it does not exist.

Three things the Profiler refuses to boot on, rather than starting up and capturing what it should not:

  • file missing — there is no default that means “everything on the node”;
  • empty list ("targets": []) — a deployment that believes itself configured and watches nothing;
  • "workload": "*" — a wildcard is omission by another name.

The schema is strict: any key it does not know refuses the boot. That includes a well-meaning comment field — if the list needs explaining, do it in a file beside it.

A file and not an environment variable, deliberately: in your infrastructure repository it is diffable, it goes through review, and someone looks at it. A list buried in a values.yaml is reviewed by nobody — and this is literally the list of where the privileged container may look.

Two things, and they behave very differently.

Volume is the easy half. Each capture produces ~30 KB — the signature is an aggregate, not events, and its size does not grow with the length of the window.

The CPU cost depends on a choice you make, and the choice is the allocation probe. We measured it, and the honest answer is a table:

Allocations per second of the process All probes Without the heap probe
11 thousand +14.8% +0.1%
114 thousand +109.9% +0.2%
1.2 million +404.4% −0.9%
14.4 million +3214% +2.3%

Target CPU time per unit of work. 2-core aarch64, kernel 7.0, measured noise floor ±4.9%. Data and harness in bench/ in the collector’s repository.

Read the second column first: everything else is free, at every rate. System calls, lock contention, network, CPU — none of it shows up in the measurement. The heap probe is the entire cost, and it costs in proportion to how much your process allocates.

It is the axis carrying func and file:line — it is what makes it possible to point at the line whose behaviour changed, rather than only that behaviour changed. It is the most expensive because it is the only one that fires once per allocation.

Capturing with it disabled keeps system calls, contention shape, network and CPU: you still learn what changed, and give up where. For a service that allocates heavily that is probably the right trade; for one that allocates little, enabling it costs little and pays well.

The practical recommendation: enable it per service, not globally, and start with a non-critical one to measure in your own environment before deciding.

One distinction worth drawing here, because the two look alike: the cost is your choice, the privilege is not. Turning the heap probe off for a service saves that service’s CPU; it does not change what the Profiler needs in order to start. The five permissions in the table are required either way — including for someone who will never enable heap — because that is what keeps all of your captures comparable to each other.

How to switch it off — and how it gets switched on

Section titled “How to switch it off — and how it gets switched on”

Turning it on is a recorded decision, and not ours alone. A Profiler only gets a certificate for the privileged class if your organization has the Profiler enabled in the control-plane. Without that the enrollment is refused, and refusal is the default: it starts off. Enabling and disabling both land in your audit log, with who and when.

The class lives in the certificate, decided by us against that policy at enrollment — it is not a field the process declares. An Agent cannot announce itself as a Profiler even by misconfiguration, and a Profiler that loses the permission stops renewing.

To switch it off:

From the control-plane — disable the Profiler in your organization settings. This does not drop the container immediately: its certificate is valid for minutes and renewal goes through the same gate, so it stops at the end of the current certificate. The screen states the window.

By uninstalling — remove the deployment. Collection stops with it, and RootPilot stays whole: nothing else is affected. The captures you already have stay readable, because what holds them is the control-plane and not the Profiler — the profiling tools stay on the agent’s surface and start answering that this process has no capture.

By taking the process off the target list — the Profiler refuses anything not on the list and names the refused target in its log. It is the finest of the three: it switches one process off without switching off the others, and the list is yours, on your infrastructure.

There is no path by which we turn the Profiler on ourselves. It is an artifact you deploy, with a target list you write, under a permission you grant.