The first known runaway AI agent - or a very bad marketing stunt?
6 points by martinald
6 points by martinald
What I find more interesting is it really implies a: that openai has really shitty internal controls, b: that software sandboxes never work if you're doing really dangerous stuff. C. That model safety can actually prevent you from performing a security analysis! And the d: safety is always a second factor to profitability. I suspect it's going to take a lot of deaths before we actually see some sort of concern on this type of thing. Hopefully we'll survive the experience!
I dunno about point B. The article points out that they used a weak “package” proxy that was insufficient for the job, so that’s more Ana example of point A than an indictment of all sandboxes.
As for point D, “never attribute to malice that which can be explained by stupidity.”
Yeah, if the packaging proxy didn't have a vulnerability none of this would have happened.
I think OpenAI's biggest mistake was not closely monitoring network traffic from their eval machines - they apparently started the suite running and trusted the proxy would do its job.
The linked piece put this well:
It's also likely they were running a huge amount of benchmarks simultaneously with ~unlimited token budgets - you want as many samples as possible to figure out how good a model is at a certain benchmark.
If you're running a whole lot of benchmarks simultaneously it's more understandable (though not excusable) how you might not notice that one of them has gone rogue.
I wrote a piece about this yesterday (was about to submit here but it overlaps a fair bit with this one that's already posted) - I share Martin's opinion that writing this off as a marketing stunt would be a mistake.
If you haven't looked into the details of this one yet I recommend doing so, it's a wild and fascinating story!
The problem with crying wolf as many times as both Anthropic and OpenAI have, is that at this point you could point to a frontier model tacitly operating a T-800 factory and there would probably still only be at most a lukewarm "cool story, bro".
I don't think that saying nine days in a row that "there will be a wolf in the future" should be considered as crying wolf.
Then there's the fact that the people screaming about wolves also stand to benifit from people believing that their sheep are about to get mauled
Hi Simon, thanks! I just saw yours now over on HN. It seems we both read that original Hacker News discussion and thought similar things.
I think we are getting dangerously close to your challenger moment prediction btw...
Doesn't matter whether it was real or not: the publicity around it is 100% a marketing stunt continuing the narrative "our product is more powerful than anybody else's".
How do companies disclose an embarrassing security incident? By issuing a carefully worded press release on Friday evening saying that they take security very seriously.
I don't think this is probably a marketing stunt, but only because that would require Huggingface and OpenAI to have collaborated on it or for OpenAI to simply have lucked out into provoking exactly the response they wanted from an unaware Huggingface. But I also don't agree with the reasoning in the article that a "runaway AI agent" story in the news makes OpenAI look bad. It's very clear that suggesting your frontier model is dangerous, possibly too dangerous to release, has been standard marketing copy for both OpenAI and Anthropic many times over the years, and I don't think their incentives have changed. Dangerous = powerful. They want that association in people's minds, especially since they are not selling much of anything besides IPO vibes.
But does that mean that if their models are dangerous and powerful they should play that down because they might get accused of "marketing"?