Skip to content

Observability backend

The observability block (metrics, logs, traces, topology, monitors, rum and synthetics) can be served by Datadog, by Grafana Cloud or by SigNoz. All three declare the same catalog operations, under vendor-neutral names.

That neutrality is what makes a second backend cost “write a connector” instead of “rewrite diagnostics”. It is also what makes them mutually exclusive.

Janela do terminal
RUNNER_OBSERVABILITY=datadog # default
RUNNER_OBSERVABILITY=grafana
RUNNER_OBSERVABILITY=signoz
RUNNER_OBSERVABILITY=none # tenant without observability

The runner does not pick the provider by looking at which credential is present. That’s deliberate: inferring would produce a “my Datadog stopped responding” with no visible cause in the configuration.

datadog is the default, so every existing runner behaves identically without touching anything.

Janela do terminal
RUNNER_CONNECTORS=real
DD_API_KEY=…
DD_APP_KEY=…
# DD_SITE=datadoghq.com # optional: datadoghq.eu / us3 / us5 / ap1
Janela do terminal
RUNNER_CONNECTORS=real RUNNER_OBSERVABILITY=grafana \
GRAFANA_CLOUD_TOKEN=… \
GRAFANA_PROM_URL=https://prometheus-prod-XX-prod-REGION.grafana.net/api/prom

Unlike Datadog, Grafana exposes one endpoint per datasource: Mimir, Loki, and Tempo have distinct hosts, each authenticating over HTTP Basic, with the username being the instance’s numeric ID and the password being the Access Policy token. Without _USER, the token is sent as a Bearer token.

Only the token and Mimir are required. Mimir is the backbone: the service graph and the Faro and Synthetics metrics live there. Loki, Tempo, and Alerting degrade operation by operation, with actionable errors.

Variable Required What for
GRAFANA_CLOUD_TOKEN yes Access Policy token.
GRAFANA_PROM_URL yes Mimir.
GRAFANA_PROM_USER no The instance’s numeric ID.
GRAFANA_LOKI_URL / _USER no Logs.
GRAFANA_TEMPO_URL / _USER no Traces.
GRAFANA_STACK_URL no Alerting API (monitors).
GRAFANA_SYNTHETICS_URL no Defaults to https://synthetic-monitoring-api.grafana.net.

This is where a Grafana installation works for the first customer and breaks for the second.

service and env are first-class concepts in Datadog. In Prometheus and Loki they are the customer’s label convention, and it might be service, app, job, or OTel’s service_name. The defaults below are a convention, not a guarantee.

Variable Default
GRAFANA_SERVICE_LABEL service
GRAFANA_ENV_LABEL env
GRAFANA_NAMESPACE_LABEL namespace
GRAFANA_HOST_LABEL instance
GRAFANA_ROUTE_LABEL http_route

Discover yours with Mimir’s /api/v1/labels endpoint.

Janela do terminal
RUNNER_CONNECTORS=real RUNNER_OBSERVABILITY=signoz \
SIGNOZ_ENDPOINT=http://signoz.observability.svc.cluster.local:8080 \
SIGNOZ_API_KEY=…

SigNoz is the first self-hosted backend supported, and that changes two things relative to the other two.

The endpoint is usually internal. Datadog and Grafana Cloud are public vendor hosts; your SigNoz normally only exists inside your network. RootPilot reaches it because the runner runs in there — a direct consequence of the BYOC model, and an address a hosted observability SaaS would never get to. Use the internal service name; don’t expose the instance just to configure RootPilot.

The version varies. Datadog and Grafana Cloud have exactly one version, the one the vendor operates. A self-hosted SigNoz may be months behind. The connector uses the v5 query_range API; installations older than that are unsupported, and the boot self-check reports the version it found instead of failing obscurely on the first read.

Variable Required What for
SIGNOZ_ENDPOINT yes Instance base URL. May be internal HTTP.
SIGNOZ_API_KEY yes Service-account key (SIGNOZ-API-KEY header).

Capabilities: what it serves, and what it doesn’t

Section titled “Capabilities: what it serves, and what it doesn’t”

It serves metrics, logs, traces, topology (the service map, which in SigNoz lives as a metric) and monitors (the alert rules).

It does not serve rum or synthetics, and that absence is declared on purpose. SigNoz’s frontend monitoring is browser OTel, which doesn’t answer the same questions as Datadog RUM, and there’s no synthetics product. Since RootPilot hides from the agent surface any tool whose operations nobody serves, those tools simply don’t appear — instead of appearing and answering a half-truth. A SigNoz + Checkly tenant still has synthetics, because the specialist comes in separately.

Same problem as Grafana, on firmer ground: SigNoz is OTel-native, so the vocabulary tends to be semconv. The defaults are a defensible floor, not a guess — but they stay adjustable, because semconv 1.27 renamed deployment.environment to deployment.environment.name and both generations coexist in the field.

Variable Default
SIGNOZ_SERVICE_FIELD service.name
SIGNOZ_ENV_FIELD deployment.environment
SIGNOZ_NAMESPACE_FIELD k8s.namespace.name
SIGNOZ_HOST_FIELD host.name
SIGNOZ_ROUTE_FIELD http.route

Checkly: the specialist that takes over synthetics

Section titled “Checkly: the specialist that takes over synthetics”
Janela do terminal
CHECKLY_API_KEY=…
# CHECKLY_ACCOUNT_ID=…

checkly is dedicated to one capability. When present, it takes over synthetics from whichever general provider is configured, and it coexists with both Datadog and Grafana. Without CHECKLY_API_KEY, synthetics stays with the provider. The boot says who owns the capability in both cases, in a capability arbitration line (owner: checkly or owner: datadog).

Here, unlike the provider choice, inferring from a credential is safe: with the key, the specialist takes over; without it, nothing changes for the provider.

CHECKLY_ACCOUNT_ID is required in practice for a user token, because without it the API answers 401 without saying why. For an account or service token it is unnecessary.

Janela do terminal
RUNNER_ANALYTICS=amplitude # or mixpanel, or none

Amplitude and Mixpanel answer the same question — event volume, active users, retention — swapping sources. So they reuse the same catalog operations and are mutually exclusive, for exactly the reason the observability block is: two sources declaring the same operation would make the Agent refuse to start, because dispatch has to know where to send.

The choice is explicit and never inferred from a present credential, for the usual reason: inferring would produce a “my Amplitude stopped answering” with no visible cause.

Mixpanel serves 8 of the 16 operations, and that is by design

Section titled “Mixpanel serves 8 of the 16 operations, and that is by design”

There is no Boards API, Insights only answers by saved-report bookmark (no listing), and there is no realtime or server-computed session metric. On a Mixpanel project the tools that depend on those operations disappear from the agent surface — rather than returning an empty list that would read as “this project has no dashboards”.

Mixpanel’s Query API allows 60 queries per hour and 5 concurrent, per project. Amplitude has no equivalent. The connector already handles it: the volume read takes its baseline from the same query instead of making two, and a 429 comes back with the wait time instead of becoming “the source is down”.

Deploys: three sources, three questions — none competing

Section titled “Deploys: three sources, three questions — none competing”
Janela do terminal
VERCEL_TOKEN=…
VERCEL_TEAM_ID=…

Vercel does not contend for operation ownership with GitHub, because the two answer different facts: GitHub answers “which PR merged”, Vercel answers “which artifact went live”. So Vercel creates its own operations, with vendor-distinct names, and joins the connector set directly, with no arbitration.

VERCEL_TEAM_ID is required in practice for a team token (without it the API answers 403 without explaining) and absent on a personal account. Use a read-only token scoped to the team.

Janela do terminal
ARGOCD_SERVER_URL=https://argocd-server.argocd.svc.cluster.local
ARGOCD_TOKEN=…

If your backend reaches Kubernetes through GitOps, neither GitHub nor Vercel answers what matters. The three questions are distinct, which is why the three connectors coexist without contending for anything:

Source Answers
GitHub which PR merged
Vercel which frontend artifact went live
Argo CD which revision the cluster is running

The third is the only one that survives the question “is yesterday’s merge in production yet?” — between the merge and the deploy there’s room for the whole incident.

Two concrete gains: the synced revision is the commit, so deploy → commit → PR → diff follows directly, with no image-tag matching; and a rollback is an event in the record (a sync to a revision already seen) rather than something to infer from a revert commit.

The server is usually internal to the cluster, and the Agent reaches it by running in there. Use a local account with read-only RBAC — see Scopes and permissions.

The test: reuse the operation, or create a new one?

Section titled “The test: reuse the operation, or create a new one?”

If you’re writing a second connector for a capability that already has an owner, the test is semantic:

Would the consumer already using this operation accept the new connector’s answer as equivalent?

Yes → reuse the operation. That’s Grafana in metrics.*: p99 is p99. catalogHash stays intact, nothing changes for consumers, and the connector goes through arbitration, because it contends for operation ownership.

No → a new operation, with a vendor-distinct name. That’s Vercel and Argo CD. catalogHash changes deliberately, and the connector joins directly, without arbitration, because the operations are disjoint.

An operation with more than one implementation needs a contract

Section titled “An operation with more than one implementation needs a contract”

When two implementations serve the same operation, a shared name isn’t enough: a declared output contract is required. It’s what makes “swapping the operation’s owner” safe.

The contract is per operation, not per capability, and it is a floor, not a ceiling: an extra field passes, a declared field that’s missing downgrades the capability to opaque.

When adding or changing an operation served by two or more connectors, update the contract and run the conformance suite against the implementations, with both a populated and an empty scenario.