Telemetry
The runner is instrumented with OpenTelemetry, and telemetry is a no-op by default. Development, test, and CI need no collector at all.
Turning it on takes two variables, not one
Section titled “Turning it on takes two variables, not one”Here’s the trap: telemetry is only active when three conditions hold at once.
enabled = ROOTPILOT_OTEL_ENABLED && !OTEL_SDK_DISABLED && exporter !== 'none'And the exporter defaults to none. Which means:
# NOT enough: the exporter stays `none` and nothing is emittedROOTPILOT_OTEL_ENABLED=true
# worksROOTPILOT_OTEL_ENABLED=trueROOTPILOT_OTEL_EXPORTER=otlpOTEL_EXPORTER_OTLP_ENDPOINT=https://otel.example.comVariables
Section titled “Variables”| Variable | Default | Notes |
|---|---|---|
ROOTPILOT_OTEL_ENABLED |
false |
Master switch. Accepts true or 1. |
ROOTPILOT_OTEL_EXPORTER |
none |
otlp, console, or none. Must move off none. |
OTEL_SDK_DISABLED |
false |
The standard OTel kill switch. true overrides the master switch. |
OTEL_EXPORTER_OTLP_ENDPOINT |
none | Collector URL. |
OTEL_EXPORTER_OTLP_HEADERS |
none | Headers, for an authenticated collector. |
OTEL_EXPORTER_OTLP_PROTOCOL |
http/protobuf |
|
OTEL_SERVICE_NAME |
rootpilot-runner |
Baked into the runner image. |
OTEL_SERVICE_VERSION |
image version | Filled at build time from the release tag. Absent ⇒ the service.version attribute simply isn’t emitted. |
RUNNER_INSTANCE_ID |
per-boot UUID | Becomes service.instance.id — what tells one runner from another when the fleet exports to a shared collector. |
OTEL_TRACES_SAMPLER_ARG |
1 |
Sampling ratio, between 0 and 1. |
console is handy for quickly checking that instrumentation is alive without standing up a collector.
How initialization happens
Section titled “How initialization happens”Through a preload, not a call in code:
node --import @rootpilotsh/otel/register apps/runner/dist/index.jsThe preload runs before the application’s module graph loads, which is what lets OTel’s
auto-instrumentation patch http and ws first. The runner image’s CMD is already exactly that
command, so you don’t need to do anything beyond setting the variables.
It’s also why the service name and version come from the environment: at that point there’s no build inlining them and no in-code caller to pass them.
What’s instrumented
Section titled “What’s instrumented”| Span | When |
|---|---|
rootpilot.handle_invoke_batch |
Receiving a batch of invocations, under the remote context extracted from the message metadata. |
rootpilot.run_op |
Each operation, with MCP and operation attributes. |
rootpilot.tunnel.session |
A whole tunnel session — from the connection attempt to the close, carrying the handshake outcome and why the session ended. |
That last one answers what the server side cannot: the tunnel sees that a connection dropped, and
the reason lives in the runner. Every session closes with a closed-vocabulary
rootpilot.tunnel.end_reason, whose point is separating what we caused from what we suffered:
end_reason |
Meaning |
|---|---|
cert_rotated |
Cert renewal swapped the identity and forced a reconnect. Expected, and frequent: renewal runs at 2/3 of a ≤1h cert’s life. |
runner_stopping |
Drain (SIGTERM, spot reclaim). Expected. |
server_closed |
The control-plane closed cleanly (code 1000/1001). |
network_error |
Drop with no close frame (1006 and friends). |
heartbeat_timeout |
Two heartbeat intervals with no reply — the runner tore the socket down and reconnected. |
hello_rejected / version_incompatible |
The handshake was refused. |
Without that separation, cert-rotation reconnects — dozens a day, by design — land in the same bucket as the churn you’re investigating, and the numbers never add up.
Sibling metrics: rootpilot.tunnel.session.duration and rootpilot.tunnel.reconnects.total, both
labeled by the same end_reason.
The remote context is the detail that matters: it makes the trace distributed, joining the control-plane’s span to the runner’s in a single trace. There are also per-tenant metrics.
Correlation with logs
Section titled “Correlation with logs”The logger correlates trace_id and span_id onto log lines. With telemetry on, you can go from a
span in your backend straight to that operation’s log lines.
Flush on shutdown
Section titled “Flush on shutdown”Telemetry is flushed on SIGTERM, after the runner’s drain and regardless of whether that drain
succeeded — a failing drain shouldn’t erase the telemetry that would explain the failure.
Even so, the final flush isn’t the only line of defense: metrics export every 15s rather than the SDK’s 60s default. That matters for short-lived cattle — a runner that dies before completing one export cycle would otherwise depend entirely on the flush.
Fleet observability, without OTel
Section titled “Fleet observability, without OTel”Even with telemetry off, the runner reports state to the control-plane through the handshake:
accounts[]: the cloud context derived at boot.ConnectorStatus{served, health}: what each connector is serving and how it fared in the credential self-check.ConnectorStatus.ops: which operations that connector serves.
The last one exists because a flat list of served operations doesn’t say whose they are, and without the owner, the control-plane can’t discount the operations of a connector with a broken credential.