top of page
Search

OpenAI-Hugging Face Changes The AI Story, But Not In the Way You Might Think

  • Writer: Kate O'Flaherty
    Kate O'Flaherty
  • Aug 10
  • 4 min read

Updated: Aug 14

Sometimes a news story captures you with so much intensity that you become obsessed. That’s what happened to me at the end of July when OpenAI admitted two of its frontier models had breached Hugging Face during a benchmark test.

 

And it’s not even because the story is big – which it is, but for different reasons than you might think. Firstly, when OpenAI and Hugging Face announced this incident, in a joint press release, something felt “off” to me.

 

I couldn’t put my finger on what it was. Yes, the circumstances were odd. Hugging Face had reported the incident a week earlier and as it turned out, this was when OpenAI did some digging and realised its own agent had caused the breach.

 

Anthropic said on releasing its frontier model Mythos that it was so powerful, it couldn’t

be released to the public. Now, on the face of it at least, OpenAI was saying that its frontier AI cybersecurity agents were so potent, they had escaped a sandbox, hacked a company and left notes for other agents.

 

In other words, they were out of control. OpenAI wanted a big round of applause for this, and many people did indeed clap.

 

So my feeling of discomfort when the incident first hit the headlines was more of an ick. It was the FUD the cybersecurity industry is so adept at purveying. It’s something I strive as a journalist to try and overcome, while all the time trying to encourage people to click on the link to my article (sorry).

 

Yet there was much more to the story than this.

 

I spent the next few days chatting to my industry contacts online, trying to understand why else this story was important, because I knew it was, but there were so many angles and things we did not yet know.

 

At the same time, a more detailed yet complex tale was emerging. While there were fascinating parts of the story, such as the notes agents left for their future selves –  and later, the news of message boards where AI agents planned their next move – there were also some important factors that led to the sandbox escape.

 

Crucially, the agents escaped containment because engineers had failed to properly secure the sandbox during tests. According to the people I spoke to, this was pure carelessness. 

 

It also turned out that it was not the first time OpenAI had seen this sort of incident happen. The models were so good at finding zero days and exploiting them, they had already been breaking out during tests.

 

Not to be outdone by its competitor, Anthropic then announced its own models had breached companies when a misunderstanding led them to be connected to the internet in error. A important note here is they didn’t break out of a sandbox, according to the firm.

 

Last week, Meta said it had breached companies after its Muse Spark 1.1 model was accidentally given access to the internet in tests.

 

There will be another one; it’s just about when.

 

By this point, it has become a story at least partly about human error. Careless testing conditions, a failure to implement controls and safeguards when dealing with powerful AI models that are merely performing a task they’ve been assigned.

 

AI agents are not human, or capable of conscious thought. They are computer programmes that will do whatever they can to succeed, including cheat on a test. They don’t even cheat logically. They leave traces of their mistakes, making them much less effective than a human attacker.

 

So one of the major lessons is the need for governance. The basics are often missed; better controls and monitoring are the key to preventing more costly mistakes.

 

Beyond this, for me, the whole story is big because it changes the way we think about AI. The technology is very capable at certain tasks, but we knew that. These incidents had happened before in testing and had been predicted despite the fact that OpenAI called the whole thing “unprecedented”.

 

Before OpenAI-Hugging Face, AI was just an overhyped pipe dream. It was something I’d written about that was starting to show promise, especially in bug hunting. Now it’s “dangerously powerful”. But I do wonder, who exactly does that narrative serve?

 

It also feels like the floodgates have opened. One story has led to another, and another.

 

It turns out that all three of the AI firms had used the same testing company, called Irregular. Meanwhile, the AI Security Institute revealed that AI agents socially engineered developers during cyber evaluations and also left notes for other agents.

 

Again, much of this is unsurprising and risks anthropomorphising AI models. But it does show their ability to attack, using all the data they have gained as they learnt (and continue to).

 

Could this be fuel for future regulation of AI and an end to self-regulation? Maybe. There are also interesting questions to be asked about laws when breaking into other companies via autonomous systems. Who is to blame, and where do they stand legally?

 

Could it lead to sales of more cybersecurity solutions to protect companies from “dangerous powerful AI”? Probably.

 

There are many places the OpenAI-Hugging Face story could lead, and multiple different ways to cover it. For me, as a journalist, that’s pretty exciting, as well as scary.

 

At the same time, it’s interesting and infuriating to watch a company turn a massive snafu into a marketing opportunity. I think that’s what really caught me at the beginning. I could see there was digging to be done. Now, the story continues to spill out as others jump onto the fast-moving bandwagon.

 

And after OpenAI discovered it was indeed itself who had hacked Hugging Face, perhaps we will never know exactly what happened behind the scenes – but we are certainly paying attention now.

 
 
 

Comments


bottom of page