Observability backend
The observability block (metrics, logs, traces, topology, monitors, rum and
synthetics) can be served by Datadog, by Grafana Cloud or by SigNoz. All three declare
the same catalog operations, under vendor-neutral names.
That neutrality is what makes a second backend cost “write a connector” instead of “rewrite diagnostics”. It is also what makes them mutually exclusive.
RUNNER_OBSERVABILITY=datadog # defaultRUNNER_OBSERVABILITY=grafanaRUNNER_OBSERVABILITY=signozRUNNER_OBSERVABILITY=none # tenant without observabilityThe choice is explicit, not inferred
Section titled “The choice is explicit, not inferred”The runner does not pick the provider by looking at which credential is present. That’s deliberate: inferring would produce a “my Datadog stopped responding” with no visible cause in the configuration.
datadog is the default, so every existing runner behaves identically without touching anything.
Datadog
Section titled “Datadog”RUNNER_CONNECTORS=realDD_API_KEY=…DD_APP_KEY=…# DD_SITE=datadoghq.com # optional: datadoghq.eu / us3 / us5 / ap1Grafana Cloud
Section titled “Grafana Cloud”RUNNER_CONNECTORS=real RUNNER_OBSERVABILITY=grafana \ GRAFANA_CLOUD_TOKEN=… \ GRAFANA_PROM_URL=https://prometheus-prod-XX-prod-REGION.grafana.net/api/promUnlike Datadog, Grafana exposes one endpoint per datasource: Mimir, Loki, and Tempo have distinct
hosts, each authenticating over HTTP Basic, with the username being the instance’s numeric ID and the
password being the Access Policy token. Without _USER, the token is sent as a Bearer token.
Only the token and Mimir are required. Mimir is the backbone: the service graph and the Faro and Synthetics metrics live there. Loki, Tempo, and Alerting degrade operation by operation, with actionable errors.
| Variable | Required | What for |
|---|---|---|
GRAFANA_CLOUD_TOKEN |
yes | Access Policy token. |
GRAFANA_PROM_URL |
yes | Mimir. |
GRAFANA_PROM_USER |
no | The instance’s numeric ID. |
GRAFANA_LOKI_URL / _USER |
no | Logs. |
GRAFANA_TEMPO_URL / _USER |
no | Traces. |
GRAFANA_STACK_URL |
no | Alerting API (monitors). |
GRAFANA_SYNTHETICS_URL |
no | Defaults to https://synthetic-monitoring-api.grafana.net. |
Label mapping
Section titled “Label mapping”This is where a Grafana installation works for the first customer and breaks for the second.
service and env are first-class concepts in Datadog. In Prometheus and Loki they are the
customer’s label convention, and it might be service, app, job, or OTel’s service_name. The
defaults below are a convention, not a guarantee.
| Variable | Default |
|---|---|
GRAFANA_SERVICE_LABEL |
service |
GRAFANA_ENV_LABEL |
env |
GRAFANA_NAMESPACE_LABEL |
namespace |
GRAFANA_HOST_LABEL |
instance |
GRAFANA_ROUTE_LABEL |
http_route |
Discover yours with Mimir’s /api/v1/labels endpoint.
SigNoz
Section titled “SigNoz”RUNNER_CONNECTORS=real RUNNER_OBSERVABILITY=signoz \ SIGNOZ_ENDPOINT=http://signoz.observability.svc.cluster.local:8080 \ SIGNOZ_API_KEY=…SigNoz is the first self-hosted backend supported, and that changes two things relative to the other two.
The endpoint is usually internal. Datadog and Grafana Cloud are public vendor hosts; your SigNoz normally only exists inside your network. RootPilot reaches it because the runner runs in there — a direct consequence of the BYOC model, and an address a hosted observability SaaS would never get to. Use the internal service name; don’t expose the instance just to configure RootPilot.
The version varies. Datadog and Grafana Cloud have exactly one version, the one the vendor
operates. A self-hosted SigNoz may be months behind. The connector uses the v5 query_range API;
installations older than that are unsupported, and the boot self-check reports the version it found
instead of failing obscurely on the first read.
| Variable | Required | What for |
|---|---|---|
SIGNOZ_ENDPOINT |
yes | Instance base URL. May be internal HTTP. |
SIGNOZ_API_KEY |
yes | Service-account key (SIGNOZ-API-KEY header). |
Capabilities: what it serves, and what it doesn’t
Section titled “Capabilities: what it serves, and what it doesn’t”It serves metrics, logs, traces, topology (the service map, which in SigNoz lives as a
metric) and monitors (the alert rules).
It does not serve rum or synthetics, and that absence is declared on purpose. SigNoz’s
frontend monitoring is browser OTel, which doesn’t answer the same questions as Datadog RUM, and
there’s no synthetics product. Since RootPilot hides from the agent surface any tool whose operations
nobody serves, those tools simply don’t appear — instead of appearing and answering a
half-truth. A SigNoz + Checkly tenant still has synthetics, because the specialist comes in
separately.
Field mapping
Section titled “Field mapping”Same problem as Grafana, on firmer ground: SigNoz is OTel-native, so the vocabulary tends to be
semconv. The defaults are a defensible floor, not a guess — but they stay adjustable, because semconv
1.27 renamed deployment.environment to deployment.environment.name and both generations coexist
in the field.
| Variable | Default |
|---|---|
SIGNOZ_SERVICE_FIELD |
service.name |
SIGNOZ_ENV_FIELD |
deployment.environment |
SIGNOZ_NAMESPACE_FIELD |
k8s.namespace.name |
SIGNOZ_HOST_FIELD |
host.name |
SIGNOZ_ROUTE_FIELD |
http.route |
Checkly: the specialist that takes over synthetics
Section titled “Checkly: the specialist that takes over synthetics”CHECKLY_API_KEY=…# CHECKLY_ACCOUNT_ID=…checkly is dedicated to one capability. When present, it takes over synthetics from
whichever general provider is configured, and it coexists with both Datadog and Grafana. Without
CHECKLY_API_KEY, synthetics stays with the provider. The boot says who owns the capability in
both cases, in a capability arbitration line (owner: checkly or owner: datadog).
Here, unlike the provider choice, inferring from a credential is safe: with the key, the specialist takes over; without it, nothing changes for the provider.
CHECKLY_ACCOUNT_ID is required in practice for a user token, because without it the API answers
401 without saying why. For an account or service token it is unnecessary.
Analytics: the second exclusive family
Section titled “Analytics: the second exclusive family”RUNNER_ANALYTICS=amplitude # or mixpanel, or noneAmplitude and Mixpanel answer the same question — event volume, active users, retention — swapping sources. So they reuse the same catalog operations and are mutually exclusive, for exactly the reason the observability block is: two sources declaring the same operation would make the Agent refuse to start, because dispatch has to know where to send.
The choice is explicit and never inferred from a present credential, for the usual reason: inferring would produce a “my Amplitude stopped answering” with no visible cause.
Mixpanel serves 8 of the 16 operations, and that is by design
Section titled “Mixpanel serves 8 of the 16 operations, and that is by design”There is no Boards API, Insights only answers by saved-report bookmark (no listing), and there is no realtime or server-computed session metric. On a Mixpanel project the tools that depend on those operations disappear from the agent surface — rather than returning an empty list that would read as “this project has no dashboards”.
Mixpanel’s Query API allows 60 queries per hour and 5 concurrent, per project. Amplitude has no equivalent. The connector already handles it: the volume read takes its baseline from the same query instead of making two, and a 429 comes back with the wait time instead of becoming “the source is down”.
Deploys: three sources, three questions — none competing
Section titled “Deploys: three sources, three questions — none competing”Vercel: orthogonal, not competing
Section titled “Vercel: orthogonal, not competing”VERCEL_TOKEN=…VERCEL_TEAM_ID=…Vercel does not contend for operation ownership with GitHub, because the two answer different facts: GitHub answers “which PR merged”, Vercel answers “which artifact went live”. So Vercel creates its own operations, with vendor-distinct names, and joins the connector set directly, with no arbitration.
VERCEL_TEAM_ID is required in practice for a team token (without it the API answers 403 without
explaining) and absent on a personal account. Use a read-only token scoped to the team.
Argo CD: the third question
Section titled “Argo CD: the third question”ARGOCD_SERVER_URL=https://argocd-server.argocd.svc.cluster.localARGOCD_TOKEN=…If your backend reaches Kubernetes through GitOps, neither GitHub nor Vercel answers what matters. The three questions are distinct, which is why the three connectors coexist without contending for anything:
| Source | Answers |
|---|---|
| GitHub | which PR merged |
| Vercel | which frontend artifact went live |
| Argo CD | which revision the cluster is running |
The third is the only one that survives the question “is yesterday’s merge in production yet?” — between the merge and the deploy there’s room for the whole incident.
Two concrete gains: the synced revision is the commit, so deploy → commit → PR → diff follows
directly, with no image-tag matching; and a rollback is an event in the record (a sync to a
revision already seen) rather than something to infer from a revert commit.
The server is usually internal to the cluster, and the Agent reaches it by running in there. Use a local account with read-only RBAC — see Scopes and permissions.
The test: reuse the operation, or create a new one?
Section titled “The test: reuse the operation, or create a new one?”If you’re writing a second connector for a capability that already has an owner, the test is semantic:
Would the consumer already using this operation accept the new connector’s answer as equivalent?
Yes → reuse the operation. That’s Grafana in metrics.*: p99 is p99. catalogHash stays intact,
nothing changes for consumers, and the connector goes through arbitration, because it contends for
operation ownership.
No → a new operation, with a vendor-distinct name. That’s Vercel and Argo CD. catalogHash
changes deliberately, and the connector joins directly, without arbitration, because the operations
are disjoint.
An operation with more than one implementation needs a contract
Section titled “An operation with more than one implementation needs a contract”When two implementations serve the same operation, a shared name isn’t enough: a declared output contract is required. It’s what makes “swapping the operation’s owner” safe.
The contract is per operation, not per capability, and it is a floor, not a ceiling: an extra field passes, a declared field that’s missing downgrades the capability to opaque.
When adding or changing an operation served by two or more connectors, update the contract and run the conformance suite against the implementations, with both a populated and an empty scenario.