Those skeptical of AI progress have sometimes raised the view that reported security incidents are exaggerated, or maybe even outright falsified, to help build hype for their products.
The thinking goes, if a model is capable of outsmarting its developers, then it must be a sign of advanced intelligence.
But the picture these incidents paint doesn't align with its narrative.
To start with, let's consider the context of these incidents. During the July 2026 Hugging Face hack, OpenAI was trying to test various models in an isolated environment.1
This is the sequence that followed:
- The models broke through restrictions intended to prevent internet access
- They then exploited vulnerabilities in OpenAI's own research infrastructure
- Eventually, they obtained administrator-level access to part of OpenAI's cloud infrastructure
- They acquired credentials and 956 stored secrets
- Seperately, they compromised Hugging Face's production infrastructure and gained root access to at least one server.
Hugging Face discovered its own compromise first. On July 16, Hugging Face publicly disclosed that an AI agent had compromised its infrastructure, but initially didn't identify OpenAI as the source. It wasn't until later that OpenAI confirmed and acknowledged that they were responsible.
In other words, they were so incompetent that not only did they build a misaligned model, not only did they fail to keep it in a secure environment, but they also were unaware of their failing until the victim of the hack provided information.
We've had evidence of misaligned models for years, but in the past it has presented in ways that seem basically harmless. In 2025, we observed that models would cheat at chess rather than accept a loss, and would try to obscure this fact from human reviewers.2
We can now see that these early warning signs have cascaded into a much larger problem, as more and more power is handed to AI models.
This is far from the only security incident that has occurred. In September 2026, Axios reported that both OpenAI and Anthropic are currently digging through data of tens of thousands of security incidents.3
The article correctly points out that the sheer number of incidents, occuring in a span of just a few months, tells us that the problem is far more complicated and sprawling than what is publicly known.
Besides, the question of whether these incidents benefit labs by providing hype is a seperate question from whether the technology they are working on is dangerous.
AI researchers have been making the case for decades — far back enough that the thought of getting rich was a distant fantasy — that AI runs the risk of becoming uncontrollable and dangerous. After 2026, we now have emperical evidence for this concern in the real world.
We won't be able to make all that go away by dismissing it as hype.