Skip to content

Connect an agent

RootPilot is used by an agent: a language model that investigates your infrastructure by calling tools. The surface it arrives through is an MCP server (Model Context Protocol), and this page is how you point a client at it.

The full path has three hops, and it helps to hold the picture before configuring anything:

your MCP client ──HTTP + bearer──▶ MCP server (control plane) ──mTLS tunnel──▶ RootPilot Agent (your infra)

The server speaks Streamable HTTP, MCP’s HTTP transport, and it is stateful per session: initialize opens the session, the response carries an Mcp-Session-Id, and the client echoes that id on every subsequent request.

The address is a URL of your deployment — whoever operates it tells you which. There is no fixed path: the server answers on whatever host and port the proxy exposes, so both https://mcp.example.com/ and https://example.com/mcp are possible, depending on the routing in front.

Bearer in the header, with a PAT (personal access token, prefixed rpmcp_):

Authorization: Bearer rpmcp_...

The token is resolved to an organization on every request, not only when the session opens. Invalid, revoked or orphaned token → 401. A session only answers to the tenant that opened it: presenting another organization’s token with someone else’s Mcp-Session-Id gives 403.

How to issue, scope and revoke that token is on the sibling page: MCP tokens.

Nearly every MCP client accepts a remote HTTP server with a header. In the most common file format (.mcp.json):

{
"mcpServers": {
"rootpilot": {
"type": "http",
"url": "https://mcp.example.com/",
"headers": { "Authorization": "Bearer ${ROOTPILOT_MCP_TOKEN}" }
}
}
}

If your client expands environment variables (most do), use that. The reason is the next paragraph.

An MCP session is a whole investigation: it opens, accumulates tool calls, and closes. From the product’s point of view it is not throwaway — it becomes a record.

  • The session id is the trace id. While it is open it shows live under Activity in the control plane; when it closes it becomes a durable trace in the same place.
  • What gets recorded is what the agent called: per call, the tool name, the arguments, whether it worked and a redacted summary of the result; and, for the session as a whole, the cost (how many operations went to the edge, how many bytes came back). Content already arrives redacted — PII redaction happens in the Agent, before any byte leaves (see PII redaction).
  • Who can see it: organization owners and admins, under Activity. It is the same trail that answers “what has this thing been looking at in our infra?” without you having to take our word for it.

When session analysis is enabled in the deployment, the trace also feeds the product improvement cycle. What becomes cross-tenant learning out of it is anonymized by construction and depends on an opt-in from your organization, off by default.

Limit Default What happens when you hit it
Idle time 10 min The session ends on its own; if it made at least one call, the trace is stored
Live sessions per organization 100 429 on initialize, before the session is created
Tool calls per session 500 The call comes back as an error, without touching your infra
Request body size 1 MiB 413

The numbers matter less than the symptom, because each one fails differently and looks like something else:

  • Idle. Ten minutes with no request and the session closes on our side. The next call carrying that Mcp-Session-Id answers 404 unknown session. A well-behaved client reopens on its own and you barely notice; what is lost is the accumulated context of the investigation, not data. This reaper exists because most clients’ close() does not end the session on the server — only an explicit DELETE does — and without it a session would stay open forever and never become a trace.
  • Live sessions. The 429 is about the whole organization, not about you, and no already-open session is affected. When it shows up for no apparent reason it is almost always a pile of sessions nobody closed; they resolve themselves as the idle reaper collects them.
  • Call budget. This is the one that disguises itself best: the call comes back as an error, and an error reads like a failed read. It is not — the call never left the control plane, nothing was queried in your infra. It is a budget, and it is per session: opening a new session resets it. Where on-demand discovery is enabled, search_tools does not consume it — it is not execution (see what the agent reads).

This is every new installation’s first confusion, and it deserves to be read before it happens.

If no RootPilot Agent for your tenant is connected at that moment, the call does not die with a loose exception, and it does not return an empty list. It returns the same blindness envelope the product uses for any read that could not see (abridged here):

{
"visibility": "none",
"error": "no_visibility",
"_sources": { "aws": "unavailable" },
"_hint": "No RootPilot Agent is connected for this tenant right now, so nothing could be read — this is not a statement about the infrastructure. Call check_fleet_health for the fleet state; retry once an agent is up."
}

The _hint sentence is literal, and it is the entire point: this is not a statement about your infrastructure. An empty answer without that distinction would read as “there is nothing there”, which is the worst possible outcome — an error makes someone investigate, a reassuring answer does not. The check_fleet_health tool reports the real state of the fleet, and that is where to start.

The same holds one step down: with the Agent connected but an integration down, the answer says which source did not respond instead of pretending to be complete. That is described in what the agent reads.