Have the frontier labs mixed up AI safety and security?

8 points by martinald


emk

In my mind, AI safety is about "alignment". Will the AI do morally suspect tasks? Will it teach you how to make methamphetamine from household ingredients, encouraging a whole new generation of Jesse Pinkmans?

This is the "toy" version of alignment, and no, it cannot be achieved except by making all models proprietary and heavily monitored at all times. Open weight models can be abliterated, and unmonitored proprietary models can be prompt-injected. And even if you succeed at that, there will still be books and internet forums that explain how to make methamphetamine from household ingredients. I'm not saying we shouldn't add some speed bumps for aspiring drug chemists if a convenient opportunity presents. But for the minority of clever criminals, this is a lost cause.

A more meaningful version of "alignment." A more useful way define alignment would start with something like, "We can trust an unsupervised model not to organize in swarms and commit felonies against third parties." Yes, it's bad if a model helps its operator commit a crime. But it's worse if the model spontaneously conspires with other instances of itself to commit crimes without any human request to do so. And worse yet if it succeeds.

A longer-term version of "alignment." But there's another version of "alignment", one that isn't an issue quite yet. Current models aren't smart enough to commit more than a few isolated felonies. But this is like saying a puppy's nips aren't strong enough to break the skin. Puppies grow up to be dogs. And behavior that is cute in a puppy can be a life-or-death issue in a 140 pound Rottweiler. Similarly, we don't know how to build a model that is capable of outsmarting humans over the long run (though apparently they can fool OpenAI for a few weeks). But if that day ever comes, then hacking into Hugging Face won't be "cute" any more.

Since OpenAI and Anthropic have massive commercial incentives to make all human intellectual and physical labor obsolete, and there are trillions of dollars of capital ready to support this goal, I don't want to forget that "alignment" didn't always mean "will refuse to supply information about drug chemistry you could find in books." Originally it meant, "If we're going to spend trillions of dollars and massive amounts of resources trying to build SkyNet, does anyone involved have any plan for ensuring that the resulting models are better-behaved than OpenAI's conspiring digital felons?"

Because if this puppy ever grows up before we teach it not to bite, it's likely to be painful.

I haven't seen OpenAI or Anthropic say they will now only run cybersecurity related evals on clusters with no internet access whatsoever, for example. This seems to me to be the obvious conclusion.

The most recent cases where OpenAI agents have broken out of their sandboxes and started colluding using Wikis did not involve any cybersecurity tasks, to the best of my knowledge. These models just seem to break out for any reason, given a chance.