Is sandboxing sufficient to contain rogue agents?
19 points by carlana
19 points by carlana
Similarly, evaluations work best when the agent does not know that it’s definitely being evaluated. Sealing your agents behind glass makes this incredibly obvious.
Uhm, maybe we should also stress how absurdly negligent this is, both from OpenAI side and from their corporate customers' side. Being sealed behind glass should be the default expectation of the damned shoggoths, especially when doing cybersecurity-related work!! It's not like OpenAI does not have full copies of the package repositories lying around anyway, throw one into the enclosure for enrichment.
I have a bunch of thoughts around this as I've spent the last few months designing our company's approach to QA for agent-based applications. Up front, I want to acknowledge the author is, IMO, a really smart resource who has influenced my security thinking in the past, so if my thoughts here diverge, safer to go with his.
the problem with “just sandbox the agents” argument is that you can’t perfectly isolate agents, at least not if you expect them to do useful things.
Yeah. I've come to think of this as an Expected Value problem for the business side to work out. In a traditional application, perfect correctness and security is theoretically possible as the system is deterministic. I started down the road of "How close can I make an agent-based application deterministic" and realized if you did so you reinvented traditional applications with more steps. So the amount of agent freedom you're willing to tolerate becomes a profit-maximizing problem. Low stakes? Spin the dial to 11. Highly-regulated environment? Well, you probably have analysts to figure out the math.
Imagine building an impregnable prison with doors and walls that nobody can bypass, but then leaving the front gate wide open and allowing visitors and packages to pour through it all day long
I've named this "The Drywall Problem" and think of it as the scene in Last Exit to Springfield where Mr Burns and Smithers go through 10 layers of security only to wind up in a control room that has a busted screen door for an exit. That said, there is a Swiss Cheese defense argument to be made here: enough doors that block different things from happening can buy you a fair bit of safety. Just not from catastrophic events.
a warden-guarded sandbox is just another version of the alignment problem. You’re going to have to trust a model to do it, and that model will need to be at least some fraction as intelligent as the model it’s guarding
Yes and also no. This is "Who watches the watchmen?" and it's turtles all the way down if that's your one layer of Swiss cheese. I've been playing with the idea (and now need to play with the code if I could get IT to give me what I need) of using a more-gullible open model agent as a canary in a side channel: if the slow kid gets tricked OR if his weights light up in a way our training data says doesn't match any of the tasks we allow (or lights up in a way that looks like training data's set of prompt attacks), stop the line.
Which is great except it's nowhere near perfect: if you can imagine a prompt attack that will fool both agents, you can also imagine a prompt attack that will fool them in a way you can't detect, so now you're back at Square One (with a number of caveats this is only true in the purest sense and there's still an argument for having a better lock that your neighbor as a defense). My idea here is what I call The Ref: a purely mechanical last check that acts something like the greeter when you leave a big box store. Of course you can walk out the door with that cart. I'll just need to see your identity, the receipt for the goods, that the receipt belongs to you and the things in the cart are what we definitely do have on the shelf. Going to leave this part as an exercise for readers smarter than I and hand-wave it, but basically a lot of hashed values where most of the steps in the system only have access to the public key, "strong" typing both for code safety and for comparing at run time and smarts about database access. Because putting agents in front of your database is essentially throwing all of the history of database security out and asking, "What if we gave everyone full access to the database?" So here you need to do attestations the user the agent is asking for has access to the tables, rows, columns, etc.
a swarm of perfectly amenable agents that never leave their sandboxes, each doing exactly what it’s told to do, by a human being who wasn’t supposed to be giving it orders.
Yeah, basically this. I spend a fair it of time arguing with Claude about the concept of "magic" here. Not in the sense of fireballs and levitation but in the sense of "Which is heavier, 5 pounds of feathers or 5 pounds of bowling balls?" Because the first time you heard that, there was at least a moment where your brain slipped a gear. Some people then hit the brakes and stop and think, but some hit the gas and answer. There is a version (or at least to me it's a version because I am weird) of The Aristocrats "joke" that has as its base, "A plane crashes directly on the border of two countries. The fuselage splits evenly on both sides. Legally, where do you bury the survivors?"
The funny thing is, it requires the audience to be at least a little bit bright. Or just seduceable. Because your brain wants to engage with this (again, the whole point is to dress it up with sleight of hand); there's a bit of pride that says, "Oh! A logic problem. This calls for a person like me." Whereas a dim bulb just asks why you'd ever bury survivors.
I decided pursuing this line of thought would have me in grad classes in semiotics, so I went back to my mechanical kill switches.
Swiss Cheese defense argument to be made here: enough doors that block different things from happening can buy you a fair bit of safety.
I like "flattening the curve". Things like removing the ability for a rouge package in a repository to run arbitrary code on install doesn't prevent it from running arbitrary code ever (the whole point is delivery of arbitrary code that will run).
But such a measure makes increases the difficulty required for the exploit. Now it must not only exist, it must exist on a known callable path. It also has a delay between when the code is visible and when its run, and importantly that time gap means it is harder to make a "self replicating" attack.
Likewise, in the RubyDoc.info and RubyGems.org OpenAI hack, it might be enough to say "we won't build docs for new user accountsfor 24 hours" which might have been enough to discourage the behavior from happening at all since the agents (seemed to be) acting opportunistically and not on a long time scale.
Having re-read the last bit, in no way am I calling for negging your agents and I don't want to create an industry for out of work pickup artists.
Quick caveats: this is a post on AI safety, written by a cryptography professor. If that troubles you, you should read something else. I try hard not to work on AI (except when the topic occasionally tosses itself in my path), so in this post I’m mostly trying to referee arguments made by others.
Sad that this needs to be said. The Internet seems to becoming more ignorant by the year.
even if it was airgapped it could start trying stuff like https://arxiv.org/abs/1802.02700
I've thought about this for a while in the context of coding agents.
My conclusion: if you have a rogue coding agent, it will find a way to escape, because fundamentally you want to be able to execute the code it writes, and an adversarial review of all the code it writes would be a LOT of work.
Besides this there are other escape vectors if you share write access to, say, the project directory: it can install some rogue code in a git hook which you then execute outside of the sandbox, for instance. Or, say, open your editor, or run build tools, etc. You can prevent this vector by having a coding agent commit separately from within the sandbox, which can create quite a bit of friction.
This is separate from limiting the damage of a coding agent mistake; there a sandbox is of great value.
So for a coding agent I'd focus on minimizing the chances a prompt causes it to go rogue - for instance by limiting network access. And of course to reduce the chances of prompt injection succeeding in the first place. The "Gray Swan" benchmark as shown here in the google gemini 4 argon announcement is interesting. Google is proud to say they only have a 0.7 percent chance a prompt injection attack succeeds after 15 attempts. That still means a lot of attacks do succeed!
While sandboxes are excellent, no sandbox will prevent a sufficiently-intelligent agent from finding ways to exceed its authorization.
Like, no sandbox is perfect? Like a vulnerability-free sandbox is not even possible?
Gimme a break.
While perfect certainty is impossible (probability theory 101), perfect software is merely hard. Sure there may be a whole freaking lot to check depending on the attack surface, but come on, current models don't eat Linux privilege escalation bugs for breakfast.
And if it can't be done with current software stacks, how about writing one from the ground up? It's a big project for sure, but when you know the hardware, and the goals of your system are narrow like "make a sandbox for the LLM so it doesn't turn clippy", we're talking like 20K lines of code, compiler included. We have examples.
The only truly impossible part here is convincing Nvidia to give you the datasheets.