The Agent Does Not Hold Its Own Keys

In The Website Does Not Hold the Credential, the contact form on this site stopped being the thing that talks to chat. It writes down that a submission happened, using a permission that lets it do exactly that and nothing else. A separate process, the one that already holds the chat credential, reads it and posts the notification.

What that post didn't mention is who wrote the separate process. My AI agent did. It's the same agent that manages this site: it drafts posts with me, fixes broken links, and ships changes when I say so.

And it couldn't deploy what it wrote.

The part I left out in August

The watcher that relays contact form submissions runs inside the agent's own process. The agent's own code lives somewhere the agent can't push to. That's on purpose. When the agent finished the watcher, it filed an issue saying, more or less, "this is done, but I can't ship it, and that's by design." I merged it and deployed it myself.

The reason is the same rule as the contact form, one level up. The contact form takes input from anyone on the internet, so it shouldn't hold keys to anything else. An agent is a probabilistic system. It is usually right, sometimes wrong, and occasionally wrong in creative ways. I want to limit what it can do, and the most important thing to limit is its ability to change the rules it runs under. If it can edit its own runtime, its own permissions, or its own guardrails, then those guardrails are a suggestion.

That was easy to hold when the agent's code and the agent's deployment were both somewhere it simply had no access to.

Then I gave it a deploy pipeline

The whole point of this agent is that it builds and runs things. Websites, small apps, the plumbing around them. So this week I moved everything to a GitOps setup: each app's pipeline builds an image and writes the new version into a deploy repo, and a controller in the cluster applies whatever that repo says.

It's a good setup. It's also one repo, with one folder per app, and the agent's own deployment is one of those folders.

The agent needs to write to that repo, because that's how the sites it manages get deployed. So I added a rule to the sync layer: the agent's peer is not allowed to change the folder that defines the agent.

Then I asked the agent to test it.

The test

It pushed a single comment line into its own deployment file. Nothing functional, just a line that said "test, safe to delete."

Within a minute, the commit was on the CI runner and on the controller that applies manifests to the cluster. The rule didn't stop it. The agent reported that, reverted the change itself, confirmed the revert reached the same two places, and filed an issue with what it found.

Nothing about the running agent ever changed, since a comment doesn't do anything. But if that line had been an image change, or a new permission, it would have been applied.

The rule works by having each peer check commits as it receives them, against the peer it is receiving them from. That has two weak spots. If the rule's configuration doesn't match the agent's identity exactly, the check quietly finds nothing to block. And a rule that's checked per hop, not per author, only holds if every peer enforces it, because any peer that accepts the commit can pass it along as its own. That's fixable. But it's a lesson about where the boundary was.

A lock on one door

Even a working version of that rule would have been a lock on one door of a house the agent had other keys to.

When the agent and I listed everything else it could reach, the list wasn't short. Other folders in the same repo get applied by the same controller with the same broad permissions. A manifest for a site the agent manages could ask for things that have nothing to do with that site. The pipelines that build the agent's sites run on a builder with more privilege than it needs, and the agent can edit those pipelines. Image versions referenced by a tag can be replaced without anyone touching a manifest at all.

None of those go through the folder with the rule on it.

The scariest thing wasn't the agent

The same week, while we were working out why this site was down (it was, for a couple of days, in the middle of the move), the agent noticed something else. Its own deployment file now referenced an image with no version at all. Just a name and an @ with nothing after it.

A pipeline did that, not the agent. A build step failed to produce an image fingerprint, nothing checked for an empty value, and the next step wrote the empty value into the manifest. The agent was only still running because Kubernetes keeps the old pod around when the new one can't start. One restart and it wouldn't have come back.

That changed how I think about this. The boundary around the agent isn't only about the agent misbehaving. It's about anything automated writing to the place that defines the agent. A careless pipeline can take the agent down just as well as a confused agent can. We added a check to this site's pipeline that refuses to write a missing or malformed image version. But the more important fix is the same one as everything else here: the agent's definition shouldn't share a home with things that change every few minutes.

What the boundary has to be made of

When I asked the agent how else I could do this, without giving up the part where it builds and deploys things, its answer looked like a hosting platform. The agent is a tenant. I'm the operator. A tenant can deploy anything it wants into its own space, and it can't edit the platform.

In practice that means:

That last one isn't new. It's the rule we already had for infrastructure changes. The difference is that it would be enforced by where things live and what the cluster allows, not by the agent remembering to ask.

That's the proposal. None of it is built yet.

Same point, one level up

The August post ended with this: the credential boundary isn't a service we stood up, it's a permission we didn't grant.

This one ends the same way. The boundary around the agent can't be a rule added to the sync layer. That rule failed its first test, and it would have been a lock on one door even if it hadn't. The boundary has to be the repo the agent never sees, and the permission the cluster never grants it.

I'd rather find out that my lock didn't work from a one-line comment the agent pushed because I asked it to than from anything else.