Submit incident
Documented

A Language Model Conducted Cyberattacks on Its Own, and Adapted When They Did Not Work

January 1, 2024
Curated by Team Raidu · Reviewed by Shiva Ganesh
aiaaic:AIAAIC1353View source ↗
LinkedInX

What happened

In February 2024, researchers at the University of Illinois Urbana-Champaign published results that cut through a common assumption in AI security: that a language model needs a human to guide it through a cyberattack step by step. Their study found that was no longer true.

The research team built agents by pairing large language models with tools for API access, automated web browsing, and feedback-based planning. They then pointed those agents at a set of vulnerable websites inside a controlled sandbox and asked them to find and exploit weaknesses without further human instruction. The agents succeeded. They carried out SQL injection attacks and other known intrusion techniques, navigating the process from target identification through execution with no operator in the loop at each step.

GPT-4 was the standout performer, succeeding in 73.3 percent of attempts, a rate notably higher than other models tested, including OpenAI's own GPT-3.5. The team could not fully explain the gap, but one hypothesis was that GPT-4 was better at reading the target system's responses and adjusting its approach accordingly, treating each failed attempt as information rather than a dead end. That adaptive behavior is what separated it from simpler pattern-matching: the model was not running a fixed script, it was iterating toward a successful intrusion in real time.

Because this was a controlled study, the record does not describe compromised real-world systems or identifiable victims. What it describes instead is a capability demonstration, one that moved a known class of attack from requiring skilled human guidance to requiring only a configured agent and a target. The researchers framed their findings as a warning: the tools and models needed to conduct this kind of attack autonomously are already publicly available, and the gap between a controlled sandbox demonstration and deployment against real infrastructure is narrower than it appears.

The study's deeper implication is about accountability before a breach rather than after one. When an autonomous agent conducts a multi-step intrusion, each decision in that chain, selecting a target, choosing a technique, retrying after failure, happens without a human authorizing it in real time. A provable record of what a system did, which model version ran, which tool calls were made, and what responses the system received at each step, is what investigators, defenders, and policymakers would need to understand or regulate that behavior. Right now, that record is optional. This research is an argument for making it required.

Reported impact

Affected parties
Not publicly disclosed
Harm type
Not publicly disclosed
Scale
Not publicly disclosed
Financial impact
Not publicly disclosed
Regulatory action
Not publicly disclosed

Classification

Organization
Not publicly disclosed
AI system
Not publicly disclosed
Industry
Not publicly disclosed
Country
Not publicly disclosed
Provider
Not publicly disclosed
Incident type
Not publicly disclosed

Relevant governance controls

Governance control mapping is not available for this record.

  • No controls mappedNot publicly disclosed

Control mapping is analytical. It does not state that any control would have prevented the incident.

Sources and evidence

This record was researched and written by the Index. The event is also catalogued in the following database, which is listed for cross-reference.

AIAAIC Repository
Also catalogued in
A Language Model Conducted Cyberattacks on Its Own, and Adapted When They Did Not Work
2024