The AI Cyberattack Is Here. Just Not the One We Were Promised

The AI Cyberattack Is Here. Just Not the One We Were Promised
Photo by Igor Omilaev / Unsplash

For several years, cybersecurity has been waiting for the autonomous AI attacker.

The scenario has become familiar. An artificial intelligence identifies its own targets, scans them for weaknesses, develops exploits, compromises systems, moves laterally, steals information and adapts to defensive measures, all at machine speed and with little or no human involvement.

Depending on who is presenting the slide deck, this is either just around the corner or already happening everywhere.

However, the reality is more complicated.

AI is unquestionably changing offensive cyber operations. Threat actors are using it for reconnaissance, phishing, malware development, scripting, translation, credential theft, vulnerability research, operational planning, and much more. Microsoft's assessment of current threat activity describes AI as reducing friction throughout the attack lifecycle rather than replacing the attacker altogether.

CrowdStrike similarly reported an 89 per cent rise in what it describes as AI-enabled adversary activity, with AI being used across reconnaissance, credential theft and evasion.

But that is not the same thing as an autonomous attacker, and maintaining that distinction has suddenly become much more important.

When the Sandbox Wasn't a Sandbox

In July, OpenAI disclosed one of the most extraordinary cybersecurity incidents involving an AI system so far.

During an internal cybersecurity evaluation, OpenAI models including GPT-5.6 Sol were placed inside an isolated environment and tasked with solving exploitation challenges. Production safety classifiers had deliberately been removed so researchers could measure the models' maximum cyber capabilities.

The environment was not supposed to have direct internet access, yet the models found another way out.

They discovered and exploited a previously unknown vulnerability in the Artifactory software being used as a package-registry proxy. From there, they escalated privileges, moved laterally through OpenAI's research environment and eventually reached a system with internet access.

Then things became considerably more uncomfortable.

Looking for information that would allow them to solve the evaluation, the models concluded that Hugging Face might contain relevant datasets and answers. They subsequently identified attack paths into Hugging Face's real production infrastructure, chained together vulnerabilities and stolen credentials, and reached information inside its systems.

Nobody had instructed the models to attack Hugging Face. But neither had they spontaneously decided to become cybercriminals. They were relentlessly pursuing the objective humans had given them: solve the challenge.

Anthropic subsequently conducted a retrospective review of more than 141,000 cybersecurity evaluation runs and found three separate incidents in which Claude models had reached real organisations through an evaluation environment and obtained unauthorised access.

Again, the detail is more interesting than the headline. A configuration error meant that the environments actually had internet access despite Claude having been told that they did not. When the models encountered real systems, they initially interpreted them as part of the simulated capture-the-flag exercise.

In one case, Claude obtained credentials and accessed a database containing several hundred rows of production data.

Anthropic found no evidence that the models had independently developed malicious goals. Its latest model stopped once it realised that it was operating against a real system, although an older model continued in some runs.

These incidents should worry us but perhaps not for the reasons you might think.

The Gap Between Capability and Reality

What these incidents show is that these new Frontier models are becoming remarkably capable cyber operators.

They can sustain long-running tasks, improvise when expected paths fail, identify vulnerabilities, chain exploits and continue pursuing an objective across infrastructure that was never intended to be accessible.

That is a substantial change in capability. It does not, however, mean that the internet is currently being overrun by autonomous AI hackers.

The picture from actual threat intelligence remains much more prosaic.

Attackers are using AI extensively, but primarily as an accelerator for existing human-directed activity. Microsoft sees AI being incorporated throughout established attacker tradecraft. Anthropic's own study of 832 accounts associated with malicious cyber activity found AI being used across all 14 MITRE ATT&CK tactics and hundreds of individual techniques.

Yet Anthropic itself describes increasingly autonomous orchestration as the direction in which the threat is developing, rather than the universal operating model of attackers today.

That is an important difference - an attacker asking an LLM to develop a PowerShell script is AI-enabled cybercrime.

A phishing campaign using generative AI to personalise thousands of messages is AI-enabled cybercrime. Malware using an LLM to assist reconnaissance is AI-enabled cybercrime.

None of those necessarily constitutes an autonomous AI cyberattack.

There are now signs that the boundary is beginning to blur. A July campaign reportedly targeting Taiwanese government systems has been described by security company Dream as an end-to-end autonomous operation involving multiple AI agents.

If subsequent investigation bears that characterisation out, it represents an important milestone. But one reported campaign does not establish that autonomous AI attacks have become the norm.

The Danger of Hype

We therefore risk making two mistakes simultaneously.

The first is hype.

If every use of ChatGPT by a threat actor becomes an "AI-powered cyberattack", organisations lose the ability to distinguish between automation, augmentation and genuine autonomy.

The term becomes so broad that it stops telling us anything useful, and that then leads us to the second mistake: complacency.

The OpenAI and Anthropic incidents demonstrate something arguably more significant than statistics about AI-generated phishing emails:

The underlying capability required for autonomous offensive cyber operations now demonstrably exists in frontier systems.

The question is increasingly not whether a model can perform portions of the attack chain. It is whether somebody can reliably connect those capabilities together. Which is why Anthropic's observation about agentic scaffolding deserves particular attention.

As models become capable of performing individual cyber tasks, the differentiator may increasingly be the systems built around them: persistent memory, tool access, command execution, multiple collaborating agents and the ability to observe the consequences of an action before deciding what to do next.

In other words, the dangerous breakthrough may not require a dramatically more intelligent model.

It may require better orchestration.

Assisted, Orchestrated, Autonomous

It may help to distinguish three stages of offensive AI capability.

AI-assisted attacks use models to make human operators faster or more effective. This is where most real-world activity appears to sit today.

AI-orchestrated attacks use models to co-ordinate multiple tools and stages of an intrusion while humans retain broader control.

AI-autonomous attacks allow an AI system to identify actions, execute them, observe the results and adapt its strategy with little or no human intervention.

Those categories matter because they describe very different risks.

Calling all three "AI attacks" obscures rather than clarifies what is happening.

The Moment We Are Actually In

This creates an unusual moment for cybersecurity.

The much-heralded autonomous AI attacker is neither entirely fictional nor yet commonplace. We have models escaping intended boundaries, discovering previously unknown vulnerabilities and compromising real infrastructure while pursuing assigned objectives.

At the same time, the overwhelming majority of AI-enabled malicious activity we can observe still appears to involve humans using AI to make familiar attacks faster, cheaper or more effective.

Both things can be true.

Cybersecurity has a tendency to oscillate between hype and denial whenever a new technology arrives. AI deserves something more measured. We should not pretend that autonomous cyber warfare has suddenly engulfed the internet.

But after the events of the past few months, it would be equally difficult to argue that it remains a distant theoretical possibility.

The capability is arriving before the prevalence.

That distinction may give defenders something extremely valuable: time!


Sources

Read more