Skip to content

Tool denylist

RootPilot asks for privileged read access to your entire stack, and for part of it the legitimate answer is no. The operations that read security groups, ACLs and firewall rules describe your infrastructure’s attack surface; some teams will not authorise that, and they should not have to disconnect AWS entirely — throwing away cost, compute and RDS along with it — in order to say so.

The denylist is how you say “you may read my metrics and my logs, but not my security groups”.

A denylist that lived only in our control plane would be our configuration, which you would have to trust us to honour. Enforced at the edge, it is verifiable by you: you read your own deployment manifest. This is the same reasoning that already governs your credentials (never held outside your infrastructure), PII redaction (at the edge, before anything leaves) and the write gate.

That has a consequence we would rather state plainly: the control-plane screen is a draft. Until the artifact is deployed, nothing changes — neither the tool surface nor what the agent can read. Hiding tools based on a policy you composed but never deployed would produce the feeling of protection without the protection, and a security control that exists only in our interface is not a security control.

Four levels, resolved as a union — the finest one wins wherever they overlap:

Level Example When to use
Named group network-posture, iam The main path. A curated slice that tracks the catalog.
Capability security, audit When the whole category is out of the question.
Connector amplitude When the whole source is out of the question.
Operation network.listSecurityGroups Fine tuning, when none of the slices above fit.

Because a capability is too coarse for the case that motivates this feature. The network capability covers reading security groups and VPCs, subnets, load balancers, target groups and transit gateways — denying it outright also takes down the resource map, blast radius and DNS topology, which is probably the opposite of what you want.

The network-posture group isolates exactly the operations that describe network posture — firewall rule, ACL, security group — and leaves the rest standing.

You can write one, and for fine tuning it is the right tool. But a literal list composed in March does not cover the security-group operation added to the catalog in May, and nobody re-reads their denylist every release. The artifact holds intent, not the expanded list: the group composed in March covers the operation added in May.

A JSON file you mount into the runner’s deployment — ConfigMap, mounted file, whatever your GitOps uses:

{
"version": 1,
"deny": {
"groups": ["network-posture", "iam"],
"capabilities": ["audit"],
"connectors": ["amplitude"],
"ops": ["query.runAthenaQuery"]
}
}

That it is a file rather than a handful of environment variables is deliberate: a JSON file in your GitOps repository is diffable, goes through a pull request and gets reviewed by your security team. An environment variable buried in a values.yaml is reviewed by nobody.

The file path goes in RUNNER_DENY_POLICY_FILE. For the simple case, four variables do the same thing without an artifact — and the two sources are merged if you use both:

Variable Contents
RUNNER_DENY_POLICY_FILE Path to the JSON artifact
RUNNER_DENIED_GROUPS Named groups, comma-separated
RUNNER_DENIED_CAPABILITIES Capabilities
RUNNER_DENIED_CONNECTORS Connectors
RUNNER_DENIED_OPS Operations

An unreadable artifact, invalid JSON, or an unknown key inside it fail the boot rather than degrading to “nothing denied”: falling back to an empty policy when you asked for a denial is the worst possible default.

Group Covers
network-posture Security groups (summary, attachments, the world-open ones) and firewall rules — AWS and GCP. It does not cover VPCs, subnets, load balancers, routes or DNS: denying posture must not take down the resource map.
iam Key and credential hygiene — access keys (including lookup by suffix), aged GCP service-account keys, API/app keys.
audit-trail The who-did-what trail — CloudTrail (lookup plus the Lake query chain) and GCP audit logs. It does not cover resource inventory: denying the trail must not take down discovery of what exists.
end-user-activity The individual — Amplitude user search and activity, plus session replays. It does not cover aggregate product metrics (active users, event volume, retention, conversion): denying this protects the person, it does not switch off your funnel.

The list tracks the catalog: a new operation describing network posture joins network-posture without you touching your artifact. That is why the group exists.

Two layers, with the same structure as the write gate — one in what is announced, one in what is executed:

  1. Announcement. At boot the runner expands your policy against its own catalog and subtracts the denied operations from the set it declares it serves. The control plane never exposes a tool whose operation is not served, so it simply does not exist for the model.
  2. Execution. enforcement.ts — the same point that already refuses any operation outside the catalog — refuses to invoke a denied operation even if the call arrives. A buggy control plane, an old version or a desynchronised catalog do not get around the policy.

The second layer exists precisely because the first one depends on us. The denial does not hold because we agreed to hide it; it holds because the process running inside your infrastructure refuses to execute it.

The runner announces back what the policy actually denied: which operations, and under which rule — whether it fell under a group, a capability, a connector, or was named directly. It is the applied expansion, not the intent you wrote.

That matters for two reasons. First, it is how you find out what a label expands to: the network-posture you denied becomes a concrete list of operations, stated by the process that refuses them. Second, it is checked against the draft: the Fleet view shows both side by side, and any divergence between “what I composed” and “what is in force” is visible — typically an artifact that has not been deployed yet, or a fleet running mixed versions.

This is the most important part of this page, and the reason the denylist is not merely a subtraction.

If a tool just vanished, three situations would look identical: the connector is not configured, the connector is broken, and you decided to forbid it. Only in the third does the infrastructure exist, is healthy, and the emptiness is a decision. Asked about network exposure, the agent would answer “I found no permissive security groups” — which is false, and the worst possible outcome, because an error makes someone investigate and a reassuring answer does not.

So the suppression is legible, in three places:

  • The session knows about the policy before it needs it. The policy is injected into the session context, so the agent answers “that is disabled by your tenant’s policy” instead of “I found nothing”.
  • The response carries the reason. A denied source is reported as denied_by_policy — which counts as I could not read, never as it does not exist. “The policy forbade me” is not proof of absence.
  • The Fleet view tells the three situations apart from the outside: never connected · connected and broken · connected and restricted.

An analysis that depended on a denied operation runs with what is left and says what was missing, rather than failing outright — a multi-signal composite usually touches far more than the denied dimension, and switching it off would make the capability disappear with no explanation.

What does not come out is the derived verdict. A security score computed from penalties would come out higher precisely because the denied dimension could not be read — a reading that never happened improving the number — so the score is not emitted. You get what was observed, marked as partial, and no conclusion that depends on what the policy hid.

Learning stops automatically: derived knowledge (topology, mappings) is extracted from the result of operations, and a denied operation produces no result.

What was learned before the policy took effect stays where it is. That is deliberate: the artifact lives on a disposable runner, and a typo in a file must not silently destroy months of learned topology. Purging is a separate, audited action — the Fleet view shows that state predating the policy exists and offers to erase it.

If the artifact names an operation the runner does not know — renamed in the catalog, or composed against a newer version — the runner refuses to start and names the entry.

It looks severe, and it is the right call: a denial directive that has no effect is invisible until it becomes an incident. Runners are cattle, a failing boot is loud, and your supervisor surfaces the error immediately — whereas a silently partial policy only shows up on the day someone reads what they should not have.

An entry can also name something the catalog had and removed on purpose. The refusal is the same — the rule protects nothing any more, and keeping it means believing in a protection that no longer exists — but the message differs: it states which version the target left in and what to use instead, rather than sending you after a renamed operation, which here would be the wrong cause. That is the case of the profiling capability and the witness connector, both gone in catalog version 0.50.0: if your policy still names them, see the Profiler.