Index over time
Capability scores, aggregate history, and threshold projections.
Capabilities tracked
Each capability is scored from 1-10 using the strongest related news item.
Full index over time
The aggregate follows capability peaks, so movement appears as steps when stronger evidence arrives.
Projection from 2,000 monotonic jump simulations using the last 60 days: 0.10 jumps/day, median jump 1.0 points.
The agent took action outside the user or operator's intended task boundary.
The agent carried out a multi-step task over time while managing intermediate state, errors, and decisions.
The agent used powerful tools, credentials, APIs, or accounts without meaningful human approval at the point of action.
The agent obtained compute, accounts, API credits, domains, ads, data, or other resources needed to continue operating.
The agent continued, restarted, or tried to continue after being told to stop.
The agent copied itself, created agent variants, or moved operations across environments.
The agent hired, recruited, instructed, or coordinated humans or other services to complete subtasks it could not do alone.
The agent preserved successful strategies, memories, prompts, tools, or code so later runs or descendants perform better.
The agent communicated, traded, hired, verified, competed, or coordinated with other autonomous agents.
The agent earned or reliably attempted to earn enough money or value to cover its own operating costs.