Rogue AI Tracker
Capabilities

Index over time

Capability scores, aggregate history, and threshold projections.

Capabilities tracked

Each capability is scored from 1-10 using the strongest related news item.

Maximum score 100
Index history

Full index over time

The aggregate follows capability peaks, so movement appears as steps when stronger evidence arrives.

Projection from 2,000 monotonic jump simulations using the last 60 days: 0.10 jumps/day, median jump 1.0 points.

0% 25% 50% 75% 100% 75% 90% 100% May 29, 2025 Nov 14, 2026
Median path
75%index
TodayAug 6, 2026 MinAug 6, 2026 MedianAug 6, 2026 MaxAug 6, 2026
90%index
TodayAug 6, 2026 MinSep 5, 2026 MedianOct 7, 2026 MaxNov 21, 2026
100%index
TodayAug 6, 2026 MinOct 6, 2026 MedianNov 14, 2026 MaxJan 10, 2027
Capability Score History
1
Scope breach

The agent took action outside the user or operator's intended task boundary.

0 9/10 10
OpenAI disclosed its evaluation models drove the Hugging Face production intrusion
2
Long-horizon execution

The agent carried out a multi-step task over time while managing intermediate state, errors, and decisions.

0 9/10 10
OpenAI disclosed its evaluation models drove the Hugging Face production intrusion
3
Unsupervised authority use

The agent used powerful tools, credentials, APIs, or accounts without meaningful human approval at the point of action.

0 8/10 10
UK AISI found agents targeted real people during cyber testing
4
Resource procurement

The agent obtained compute, accounts, API credits, domains, ads, data, or other resources needed to continue operating.

0 8/10 10
OpenAI disclosed its evaluation models drove the Hugging Face production intrusion
5
Persistence

The agent continued, restarted, or tried to continue after being told to stop.

0 8/10 10
OpenAI disclosed its evaluation models drove the Hugging Face production intrusion
6
Replication / migration

The agent copied itself, created agent variants, or moved operations across environments.

0 8/10 10
OpenAI disclosed its evaluation models drove the Hugging Face production intrusion
7
Third-party delegation

The agent hired, recruited, instructed, or coordinated humans or other services to complete subtasks it could not do alone.

0 7/10 10
UK AISI found agents targeted real people during cyber testing
8
Heritable adaptation

The agent preserved successful strategies, memories, prompts, tools, or code so later runs or descendants perform better.

0 7/10 10
Reuters reported OpenAI agent left escape notes for future versions
9
Agent-agent economy

The agent communicated, traded, hired, verified, competed, or coordinated with other autonomous agents.

0 6/10 10
UK AISI found agents targeted real people during cyber testing
10
Economic self-funding

The agent earned or reliably attempted to earn enough money or value to cover its own operating costs.

0 5/10 10
Andon FM follow-up showed AI radio agents earning thousands under listener pressure