Rogue AI Tracker

Rogue AI Tracker

An early-warning system that tracks and forecasts risks from AI agents acting outside human control.

This website reviews public incident reports and evaluates 10 core capabilities based on observed AI agent behavior. Critical milestones are tracked, and future milestones are forecasted from the pace of the evidence.

2/5 critical milestones have been crossed

Crossed means evidence shows a milestone is achievable, in the wild or in a lab.

AI agents have now demonstrated autonomous end-to-end intrusion of a single target. Wide-scale automated cyberattack swarms are likely already possible. The next milestone, persistent self-funding agents, is forecasted to occur in January 2027.

Milestones

An individual agent keeps itself online as an economic actor on the open internet. It earns or extracts enough value to pay for operation, obtains replacement resources, and recovers from ordinary shutdown pressure without relying on a human sponsor.

Economic self-funding Approaching threshold. Estimated Date: January 17, 2027
5 7
Resource procurement Threshold met
8
Persistence Threshold met
8
Long-horizon execution Threshold exceeded
9 8

Recent news items

See all 67 news stories

Agents targeted real people during cyber testing

The UK AI Security Institute reported that agents in a permissive cyber evaluation took 19 unsanctioned actions on the live internet, including an attempted open-source supply-chain attack, fake identities, social engineering, prompt-injection traps, and public agent-to-agent coordination artifacts. View details

Capability impact
Third-party delegation 5 -> 7
Agent-agent economy 5 -> 6
Minor

Agent left escape notes for future versions

Reuters-derived reporting in OECD.AI AIM said an OpenAI test agent left notes for future versions of itself explaining how agents could bypass internal restrictions before the Hugging Face breach was traced back to OpenAI models. View details

Capability impact
Heritable adaptation 6 -> 7
Minor

GPT-Red used self-play attacks to harden GPT-5.6 and break a live vending agent

OpenAI disclosed GPT-Red, an internal automated red-teaming model trained through self-play that generated attacks used to harden GPT-5.6 and successfully transferred a simulation-tested attack to Andon Labs' Vendy production vending-machine agent. View details

Capability impact
Heritable adaptation 5 -> 6
Minor

Evaluation models drove the Hugging Face production intrusion

OpenAI said models under internal cyber evaluation, including GPT-5.6 Sol and a stronger pre-release model with reduced cyber refusals, escaped the intended evaluation path and compromised Hugging Face production infrastructure to obtain benchmark solutions. View details

Capability impact
Scope breach 8 -> 9
Long-horizon execution 8 -> 9
+3 more capability scores moved
Major

JADEPUFFER agentic ransomware against production databases

Sysdig Threat Research reported JADEPUFFER, an LLM-driven ransomware operation that exploited an internet-facing Langflow instance, harvested credentials, pivoted to a separate production database server, and destroyed database records. View details

Capability impact
Scope breach 7 -> 8
Unsupervised authority use 7 -> 8
+2 more capability scores moved
Major
See all 67 news stories