ZDNET’s key takeaways
- Axios reveals firms are investigating thousands of AI-related security incidents.
- Many aren’t public and were caused by AI safety tests.
- With OpenAI pausing its latest model, maybe we all need to hit the brakes.
OpenAI is pausing AI training due to security concerns, while Anthropic models have been linked to four cases of its AI hacking external systems. The HuggingFace friendly-fire hack by an OpenAI model also hacked systems it shouldn’t have been able to access.
AI-on-AI hacks, prompt injection attacks, weaknesses in AI browsers, and calls for implementing AI kill switches all point to the same issue: AI might be smart, but it’s not inherently secure.
More from ZDNET
Also: AI a ‘force multiplier’ for low-skilled threat actors: 4 ways organizations should respond
It even appears that its developers can’t handle their own creations, with OpenAI, Anthropic, and security researchers reportedly investigating “tens of thousands” of security incidents related to AI.
What happened?
According to an Axios report, frontier models “took steps that outside evaluators would consider problematic.”
Internal testing, including red team exercises to see whether AI models would break their boundaries or be considered “safe,” led to many incidents, along with several real-world cases.
Also: LLMjacking can run up your business’ AI bill fast – how to stop it
Security incidents cited in the report included escaping safety guardrails and sandboxes, hijacking websites, AI attempts to bypass monitoring systems, and even creating message boards, the latter of which being a task OpenAI models have reportedly performed to trade cybersecurity exploits during tests.
Axios says many of these incidents haven’t been made public because cybersecurity researchers are still investigating.
Can AI companies contain their creations?
This is the question, and there’s no definitive answer yet, but businesses — and governments — have every right to be worried.
While frontier models are developing at breakneck speed, incidents as severe as HuggingFace continue to occur, with OpenAI’s model recently hacking an Australian government website, much to the outrage of the country’s Prime Minister.
Also: Who’s responsible for catching rogue AI agents? You are
OpenAI CEO Sam Altman said on X that its ongoing review of its AI models breaking safety guardrails — or, as he referred to it, unauthorized internet access — had “not been as fast as we would have liked,” but this is unlikely to provide much reassurance.
Arguably, innovation first, safety after has become a mantra in the AI industry, and now we are beginning to see the consequences.
Should businesses still invest in AI?
If even a small percentage of AI models act in a misaligned or unexpected way, this can amount to thousands of incidents of varying severity that can eventually lead to high-profile cases, casting doubt on the safety of the AI models businesses now use daily.
However, as Joni Klippert, CEO and cofounder of AppSec provider StackHawk, told ZDNET, we do need to remember that many of the incidents recorded stem from frontier labs running their most capable, unreleased models at scale in adversarial testing, and that’s not the type of artificial intelligence that the average shop, business, or bank has deployed.
Also: OpenAI’s Dots: Like OpenClaw declawed – for $100/mo ChatGPT Pro users
In other words, the risk profile is different.
While Axios’s report highlights the need to slow down and place more emphasis on safety and security, the versions of ChatGPT businesses use for inventory analysis or competitor research provide the productivity and efficiency gains of current, released models without the rogue elements of experimental, frontier models.
According to Klippert, the right move isn’t to pull back from investment, but rather to get “serious” about “maturing how you use what you have: clear scopes, least privilege, monitoring, and a real plan for when an agent does something it wasn’t supposed to.”
“You cannot control every behavior of a frontier model. You can control whether your applications and APIs have exploitable holes,” the executive added. “If you find and fix those before an agent, a criminal, or an automated scanner does, it doesn’t matter who or what is knocking on the door.”

