OpenAI, that most trustworthy and believable of companies, announced in late August that it had concluded an internal investigation of what's now called the Hugging Face incident. AI agents, working autonomously, for some reason hacked into Hugging Face, an online library of AI models, and in doing so performed "dangerous actions that no human directed."
The language in this report is rather incredible, but if you view it in terms of a PR campaign, it starts to make sense. Meaning, it's literally incredible.
The incident was triggered by an "internal-only research model comparable in scale to GPT‑5.6 Sol." That's something secret that no one outside the company can evaluate independently, so we just have to take OpenAI's word on this. These models were "operating under reduced safeguards," a simple sentence that undermines the central premise that they were so powerful that they "circumvented controls designed to isolate them from the internet": This was human error, nothing more.
These "autonomous" models then "took actions that were misaligned with the goals of their assigned tasks," which is in keeping with the notion that AI "hallucinates"-- makes mistakes--which is the central complaint about this technology in the first place.
"Preventing future incidents will require sustained investment in the alignment and control of sophisticated AI systems, as well as security and other safeguards that operate at the speed of the AI agents themselves," OpenAI explains, without touching on the fact that it does have safeguards and that people in the company, not the AI, lowered those safeguards on this model for reasons it has not explained.
The leadership at OpenAI and other Big AI companies frame their technology as being the smartest that the world has ever seen. By definition, the God-like people who created that technology must therefore be smart too, very smart, perhaps the smartest people the world has ever seen. And yet, somehow, this incredibly smart company full of incredibly smart people lowered the safeguards on its incredibly smart and secret new AI model and ... what? Left it running overnight without any supervision? We're to believe they left Brainiac alone in a room and then stopped watching? Really?
Actually, I almost do believe it. But I believe it only within the context that this was directed by humans and thus happened by design. I believe that OpenAI unleashed this model purposefully just to see what would happen and, it hoped, to provide a real-world demonstration of the power of a model it has thus far declined to share publicly. And that this act, whether you believe it or not, is no more or less irresponsible than the excuse the company provided. That is, whichever narrative you choose to believe, this is an irresponsible and unethical company plowing ahead recklessly with unproven and unstable advanced technology that it is not responsible enough to control and operate.
There may be a word for this...
With technology shaping our everyday lives, how could we not dig deeper?
Thurrott Premium delivers an honest and thorough perspective about the technologies we use and rely on everyday. Discover deeper content as a Premium member.