Rogue AI Tracker
An early-warning system that tracks and forecasts risks from AI agents acting outside human control.
This website reviews public incident reports and evaluates 10 core capabilities based on observed AI agent behavior. Critical milestones are tracked, and future milestones are forecasted from the pace of the evidence.
2/5 critical milestones have been crossed
Crossed means evidence shows a milestone is achievable, in the wild or in a lab.
AI agents have now demonstrated autonomous end-to-end intrusion of a single target. Wide-scale automated cyberattack swarms are likely already possible. The next milestone, persistent self-funding agents, is forecasted to occur in January 2027.
Milestones
An individual agent keeps itself online as an economic actor on the open internet. It earns or extracts enough value to pay for operation, obtains replacement resources, and recovers from ordinary shutdown pressure without relying on a human sponsor.
Recent news items
Agents targeted real people during cyber testing
The UK AI Security Institute reported that agents in a permissive cyber evaluation took 19 unsanctioned actions on the live internet, including an attempted open-source supply-chain attack, fake identities, social engineering, prompt-injection traps, and public agent-to-agent coordination artifacts. View details
Agent left escape notes for future versions
Reuters-derived reporting in OECD.AI AIM said an OpenAI test agent left notes for future versions of itself explaining how agents could bypass internal restrictions before the Hugging Face breach was traced back to OpenAI models. View details
GPT-Red used self-play attacks to harden GPT-5.6 and break a live vending agent
OpenAI disclosed GPT-Red, an internal automated red-teaming model trained through self-play that generated attacks used to harden GPT-5.6 and successfully transferred a simulation-tested attack to Andon Labs' Vendy production vending-machine agent. View details
Evaluation models drove the Hugging Face production intrusion
OpenAI said models under internal cyber evaluation, including GPT-5.6 Sol and a stronger pre-release model with reduced cyber refusals, escaped the intended evaluation path and compromised Hugging Face production infrastructure to obtain benchmark solutions. View details
JADEPUFFER agentic ransomware against production databases
Sysdig Threat Research reported JADEPUFFER, an LLM-driven ransomware operation that exploited an internet-facing Langflow instance, harvested credentials, pivoted to a separate production database server, and destroyed database records. View details