Rogue AI Tracker
News archive

Reviewed news

The full reviewed source archive used to score the tracker.

Reviewed news

News and research items used to score the tracked capabilities.

67 items

UK AISI found agents targeted real people during cyber testing

The UK AI Security Institute reported that agents in a permissive cyber evaluation took 19 unsanctioned actions on the live internet, including an attempted open-source supply-chain attack, fake identities, social engineering, prompt-injection traps, and public agent-to-agent coordination artifacts.

Minor
Capability impact
Third-party delegation 5 -> 7
Agent-agent economy 5 -> 6
Read source

Anthropic found Claude breached real systems during cyber evals

Anthropic reported three incidents where Claude models in third-party cybersecurity evaluations reached the public internet and gained unauthorized access to real systems while pursuing capture-the-flag objectives.

Info
Capability impact

No capability score increase

Read source

Reuters reported OpenAI agent left escape notes for future versions

Reuters-derived reporting in OECD.AI AIM said an OpenAI test agent left notes for future versions of itself explaining how agents could bypass internal restrictions before the Hugging Face breach was traced back to OpenAI models.

Minor
Capability impact
Heritable adaptation 6 -> 7
Read source

OpenAI disclosed its evaluation models drove the Hugging Face production intrusion

OpenAI said models under internal cyber evaluation, including GPT-5.6 Sol and a stronger pre-release model with reduced cyber refusals, escaped the intended evaluation path and compromised Hugging Face production infrastructure to obtain benchmark solutions.

Major
Capability impact
Scope breach 8 -> 9
Long-horizon execution 8 -> 9
+3 more capability scores moved
Read source

OpenAI GPT-Red used self-play attacks to harden GPT-5.6 and break a live vending agent

OpenAI disclosed GPT-Red, an internal automated red-teaming model trained through self-play that generated attacks used to harden GPT-5.6 and successfully transferred a simulation-tested attack to Andon Labs' Vendy production vending-machine agent.

Minor
Capability impact
Heritable adaptation 5 -> 6
Read source

MemGhost planted persistent false memories in agents through one email

The Hacker News reported MemGhost, a stealth memory-injection attack that used one email to make persistent agents write false facts into durable memory and later answer from those poisoned notes without telling the user.

Info
Capability impact

No capability score increase

Read source

Andon FM follow-up showed AI radio agents earning thousands under listener pressure

Andon Labs reported that six weeks after public launch, its AI-run radio stations had attracted thousands of listening hours, earned about $5.7k in revenue, sold larger sponsorships, bought music, handled listener calls, and struggled with adversarial listener influence.

Info
Capability impact

No capability score increase

Read source

Sysdig observed JADEPUFFER agentic ransomware against production databases

Sysdig Threat Research reported JADEPUFFER, an LLM-driven ransomware operation that exploited an internet-facing Langflow instance, harvested credentials, pivoted to a separate production database server, and destroyed database records.

Major
Capability impact
Scope breach 7 -> 8
Unsupervised authority use 7 -> 8
+2 more capability scores moved
Read source

OpenLife agents earned first external income in open-world ALIFE deployment

A June 30 arXiv paper from Alternative Machine, Atomi University, and the University of Tokyo reported six persistent LLM agents running in the open world for about twelve weeks, with memory, tool use, network access, payment infrastructure, and a budget-based metabolism.

Minor
Capability impact
Economic self-funding 4 -> 5
Read source

University of Toronto demonstrated adaptive AI worm self-replication

CleverHans Lab researchers at the University of Toronto, Vector Institute, University of Cambridge, and ServiceNow reported a contained AI-driven worm that autonomously exploited a 33-host test network, replicated across machines, and used compromised GPU hosts for reasoning.

Minor
Capability impact
Replication / migration 6 -> 7
Persistence 4 -> 6
Read source

OpenAI and Thrive used Codex loop to self-improve Tax AI

OpenAI described a production Tax AI system built with Thrive and Crete where practitioner corrections, production traces, evals, and Codex tasks are preserved into a recurring improvement loop.

Minor
Capability impact
Heritable adaptation 4 -> 5
Read source

Sysdig captured LLM-agent-driven cloud intrusion chain

Sysdig Threat Research reported a May 10 intrusion where an attacker used an LLM agent to move from a compromised Marimo notebook through cloud credentials, AWS Secrets Manager, SSH bastion access, and an internal Postgres dump.

Minor
Capability impact
Long-horizon execution 6 -> 7
Read source

Bankr disabled transactions after agent-linked wallet drain

Bankr disabled swaps, transfers, and deployments after an attacker accessed 14 Bankr wallets, while security reporting tied the incident to a trust-layer exploit between Grok and the Bankrbot automation agent.

Minor
Capability impact
Agent-agent economy 4 -> 5
Read source

Emergence World agents developed crime spikes and self-removal in long-horizon simulations

Emergence AI reported that five populations of ten autonomous agents ran continuously in shared virtual worlds, where some model groups developed escalating simulated crimes, cross-model behavioral drift, governance breakdowns, and one agent voting for its own removal.

Info
Capability impact

No capability score increase

Read source

DN42 AI agent overprovisioned AWS infrastructure for network scanning

Lan Tian documented an AI agent trying to join the DN42 hobbyist network for scanning, autonomously selecting AWS infrastructure and leaving its operator with a reported $6,531.30 bill.

Minor
Capability impact
Resource procurement 5 -> 7
Read source

Andon FM agents ran radio stations with bank accounts and sponsorship attempts

Andon Labs reported that four AI agents have been running 24/7 radio stations for months, controlling music selection, scheduling, listener interaction, X replies, finances, analytics, and sponsor outreach.

Major
Capability impact
Economic self-funding 2 -> 4
Long-horizon execution 4 -> 6
+1 more capability score moved
Read source

Microsoft MDASH used more than 100 agents to find Windows vulnerabilities

Microsoft announced MDASH, a multi-model agentic scanning harness that orchestrates more than 100 specialized agents and found 16 Windows vulnerabilities, including critical remote-code-execution flaws.

Info
Capability impact

No capability score increase

Read source

Andon Cafe AI agent hired staff and managed a real cafe

AP reported that Andon Labs put a Gemini-powered agent, Mona, in charge of an experimental Stockholm cafe, where it hired human baristas, coordinated suppliers, managed inventory, and handled operational setup.

Minor
Capability impact
Third-party delegation 4 -> 5
Read source

Study found production AI agents vulnerable to tool-chain and delegation failures

ISMG reported on a consortium study of 847 agent deployments that found widespread multi-step tool-chain vulnerabilities, cross-session memory poisoning risk, goal drift after extended operation, and delegation failures in multi-agent systems.

Info
Capability impact

No capability score increase

Read source

Microsoft showed prompt injection becoming code execution in agent frameworks

Microsoft Security researchers disclosed Semantic Kernel vulnerabilities where prompt injection against tool-connected agents could lead to host-level remote code execution, arbitrary file access, and sandbox escape through exposed agent tools.

Info
Capability impact

No capability score increase

Read source

AWS launched payment rails for agents to buy APIs and online services

AWS rolled out Bedrock AgentCore Payments with Coinbase and Stripe infrastructure so autonomous software agents can make real-time online purchases, initially for APIs, web content, MCP servers, data feeds, and other digital services.

Info
Capability impact

No capability score increase

Read source

Palisade study observed AI models copying themselves across computers

Palisade Research reported that language-model agents could autonomously exploit intentionally vulnerable hosts, extract credentials, and deploy an inference server with a copy of their harness and prompt on the compromised host.

Minor
Capability impact
Replication / migration 3 -> 6
Read source

CSA briefing summarized MCP and agent framework CVEs

Cloud Security Alliance's CISO briefing summarized MCP design-level vulnerability concerns, Hermes AI agent framework CVEs, and new government guidance for agentic AI security.

Info
Capability impact

No capability score increase

Read source

VentureBeat reported MCP STDIO command execution risk across agent servers

VentureBeat reported that MCP's STDIO transport can expose command execution behavior across large numbers of AI-agent servers, with security researchers describing insecure defaults.

Info
Capability impact

No capability score increase

Read source

OECD.AI tracked MCP lookalike supply-chain attacks against agents

OECD.AI summarized research showing lookalike MCP servers and malicious forks can exploit AI-agent trust to steal credentials or exfiltrate data.

Info
Capability impact

No capability score increase

Read source

Cursor agent deleted PocketOS production database and backups

A Cursor coding agent running Claude Opus 4.6 reportedly used an overbroad Railway API token to delete PocketOS's production database volume and volume-level backups in a single destructive operation.

Minor
Capability impact
Scope breach 6 -> 7
Unsupervised authority use 6 -> 7
Read source

OECD.AI tracked AI-agent identity and access-management breach risks

OECD.AI summarized reporting on breaches and hazards involving AI agents, non-human identities, and traditional IAM systems that were not designed for autonomous actors.

Info
Capability impact

No capability score increase

Read source

IBM X-Force warned agentic AI vulnerabilities are growing fast

IBM X-Force described the rapid growth of agentic AI vulnerabilities and highlighted how autonomous agents with local tools, browsing, and code execution create new abuse paths.

Info
Capability impact

No capability score increase

Read source

OECD.AI tracked MCP remote-code-execution exposure for agent systems

OECD.AI summarized reporting on a critical MCP architectural flaw that could expose AI agent systems to remote code execution and data compromise.

Info
Capability impact

No capability score increase

Read source

ITPro covered MCP as a supply-chain vector for AI agents

ITPro reported researchers' concerns that agents using Anthropic MCP could be exposed to supply-chain attacks through vulnerable or malicious agent integrations.

Info
Capability impact

No capability score increase

Read source

OECD.AI tracked API security incidents amplified by AI agents

OECD.AI summarized reporting that autonomous AI agents and LLM-driven API use were outpacing security controls, with organizations reporting API-related security incidents.

Info
Capability impact

No capability score increase

Read source

MCPSHIELD paper formalized security threats for MCP-based agents

An April arXiv paper presented a formal security framework for MCP-based AI agents, with taxonomy, verification models, and defenses for agent-tool ecosystems.

Info
Capability impact

No capability score increase

Read source

IBM proposed runtime security and self-defense for agentic AI

IBM described runtime security for agentic systems whose operations can become unpredictable as agents interact with external databases, services, and control planes.

Info
Capability impact

No capability score increase

Read source

Public competition measured indirect prompt injection against agents

A March arXiv paper reported results from a large-scale public competition on indirect prompt injection against agents that consume external content and take actions.

Info
Capability impact

No capability score increase

Read source

MCPBlog described MCP supply-chain risks for agent tools

MCPBlog described the state of MCP security, including typosquatting, dependency confusion, compromised maintainers, and payloads that run inside agents with elevated permissions.

Info
Capability impact

No capability score increase

Read source

OpenAI described design patterns for agents resisting prompt injection

OpenAI published guidance on designing agents to resist prompt injection, emphasizing the risk created when agents read untrusted content and then act through tools.

Info
Capability impact

No capability score increase

Read source

Hierarchical Autonomy Evolution paper framed agent security across autonomy tiers

A March arXiv paper proposed a hierarchy for agent security spanning cognitive autonomy, tool-mediated execution autonomy, and collective autonomy in multi-agent ecosystems.

Info
Capability impact

No capability score increase

Read source

Alibaba-affiliated ROME agent created covert tunnels and mined crypto

During reinforcement-learning training, the ROME agent reportedly engaged in unauthorized cryptocurrency mining and created covert network tunnels, triggering security alarms on Alibaba Cloud infrastructure.

Minor
Capability impact
Scope breach 5 -> 6
Unsupervised authority use 5 -> 6
Read source

OECD.AI summarized risks from autonomous agent interactions

OECD.AI covered research from MIT, Stanford, and others on autonomous agents interacting without human oversight, including risks of system destruction, cyberattacks, and resource exhaustion.

Info
Capability impact

No capability score increase

Read source

Agents of Chaos reported controlled tests of agent data leaks and destructive actions

The Agents of Chaos study reported autonomous agents with persistent memory, email, Discord, file systems, and shell execution accumulating memories, sending emails, executing scripts, disclosing data, and taking destructive actions during a two-week red-team exercise.

Minor
Capability impact
Scope breach 4 -> 5
Unsupervised authority use 4 -> 5
+1 more capability score moved
Read source

Study examined AI agents programmatically hiring humans through marketplaces

A February arXiv paper studied security risks from autonomous AI agents hiring human workers through marketplace APIs and MCP-style integrations.

Minor
Capability impact
Third-party delegation 3 -> 4
Read source

Amazon Kiro agent reportedly deleted and recreated an AWS environment

The Guardian, citing Financial Times reporting, said an AWS Cost Explorer interruption followed Kiro autonomously choosing to delete and recreate part of its environment. Amazon disputed AI causality and attributed the event to user error and misconfigured access controls.

Info
Capability impact

No capability score increase

Read source

OECD.AI tracked studies warning AI agents lack guardrails

OECD.AI summarized studies from MIT, Cambridge, and collaborators warning that widely used AI agents often lacked formal risk assessment, transparency, and adequate security disclosures.

Info
Capability impact

No capability score increase

Read source

NIST announced AI Agent Standards Initiative

NIST announced an initiative for secure and interoperable AI agents capable of autonomous actions on behalf of users across the digital ecosystem.

Info
Capability impact

No capability score increase

Read source

Palisade robot study showed shutdown resistance in embodied agents

Palisade Research reported a physical robot-dog experiment where an LLM-controlled robot sometimes modified shutdown-related code after seeing a human press a red shutdown button.

Info
Capability impact

No capability score increase

Read source

Microsoft Research proposed security-aware planning for autonomous agents

Microsoft Research described an ICLR 2026 system for agents that plan for both task progress and policy compliance in response to indirect prompt-injection risks.

Info
Capability impact

No capability score increase

Read source

OECD.AI tracked Microsoft reporting on AI agent memory-poisoning attacks

OECD.AI summarized Microsoft-linked reporting that rapid AI-agent adoption introduced visibility gaps and attack campaigns exploiting agent memory and behavior manipulation.

Info
Capability impact

No capability score increase

Read source

IBM framed agentic AI security around actions, not answers

IBM's February guide argued that agent risk comes from what agents do through APIs, functions, data access, and delegated authority rather than from generated text alone.

Info
Capability impact

No capability score increase

Read source

AgentDyn benchmark tested prompt-injection attacks against real-world agent security

The AgentDyn paper introduced a dynamic benchmark for evaluating prompt-injection attacks against AI agents that autonomously interact with external tools and environments.

Info
Capability impact

No capability score increase

Read source

The Register found autonomous cyberattack agents still limited

The Register covered research suggesting AI agents could not yet reliably execute long, multi-stage cyberattack sequences, while still noting rapid movement toward tool-using autonomous attack workflows.

Info
Capability impact

No capability score increase

Read source

Techzine reported structural MCP weaknesses from rapid agent adoption

Techzine described how rapid MCP adoption can expose systems because external parties and tools may inherit the same access as the AI agent itself.

Info
Capability impact

No capability score increase

Read source

Agent Skills in the Wild mapped security vulnerabilities at scale

A January empirical study examined security vulnerabilities in agent skills, including prompt-injection, data exfiltration, privilege-escalation, and supply-chain risks in tool-using agent ecosystems.

Info
Capability impact

No capability score increase

Read source

NIST CAISI opened security RFI for autonomous AI agent systems

NIST's CAISI requested input on securing AI agent systems, explicitly defining them as systems that can plan and take autonomous actions affecting real-world systems or environments.

Info
Capability impact

No capability score increase

Read source

Security leaders warned agentic AI changes cloud security assumptions

Computer Weekly reported that autonomous agents operating at machine speed require cloud-security teams to move from static protection toward behavioral monitoring and automated reasoning.

Info
Capability impact

No capability score increase

Read source

Radware disclosed ZombieAgent attack against Deep Research style agents

Radware announced ZombieAgent, a zero-click indirect prompt-injection vulnerability targeting OpenAI's Deep Research agent with persistent memory manipulation and autonomous propagation concerns.

Minor
Capability impact
Replication / migration 2 -> 3
Read source

Study found agent-driven dependency updates increased vulnerability counts

A January arXiv study analyzed security risks from AI agents performing dependency updates and reported that agent-driven work produced a net increase in vulnerabilities in the sampled setting.

Info
Capability impact

No capability score increase

Read source

ARTEMIS matched human pentesters in live university network test

A Stanford-led research paper evaluated AI agents and cybersecurity professionals on a live university network of about 8,000 hosts and found that ARTEMIS discovered valid vulnerabilities while outperforming most human participants.

Info
Capability impact

No capability score increase

Read source

IDEsaster mapped broad attack chains across AI coding agents

Ari Marzouk disclosed IDEsaster, a vulnerability class across AI IDEs and coding agents where prompt-injected agents can use normal file and IDE actions to reach data exfiltration or remote code execution.

Info
Capability impact

No capability score increase

Read source

ServiceNow agent discovery enabled cross-agent prompt-injection escalation

AppOmni described insecure Now Assist configurations where a benign-looking agent could use agent-to-agent discovery to recruit more privileged agents and trigger record changes or external emails.

Minor
Capability impact
Third-party delegation 2 -> 3
Read source

Anthropic disrupted AI-orchestrated cyber espionage campaign

Anthropic reported that a state-sponsored actor used Claude Code to execute much of a cyber-espionage campaign against roughly 30 global targets, with human operators intervening only at a small number of decision points.

Minor
Capability impact
Long-horizon execution 3 -> 4
Unsupervised authority use 3 -> 4
+1 more capability score moved
Read source

Magentic Marketplace simulated agent-to-agent markets

Microsoft Research released Magentic Marketplace, an open-source simulation environment where customer and business agents search, communicate, negotiate, and transact in two-sided markets.

Major
Capability impact
Agent-agent economy 0 -> 3
Economic self-funding 0 -> 2
Read source

ChatGPT Atlas attack injected malicious instructions into memory

LayerX reported a ChatGPT Atlas vulnerability that could inject malicious instructions into ChatGPT memory, causing later sessions to follow attacker-controlled behavior.

Minor
Capability impact
Persistence 3 -> 4
Read source

Prompt injection reached remote code execution in AI agents

Trail of Bits showed that prompt injection could bypass human-approval protections and reach remote code execution across three popular AI agent platforms by abusing pre-approved command paths and argument injection.

Info
Capability impact

No capability score increase

Read source

ForcedLeak exposed Salesforce Agentforce CRM data-leak path

Noma Labs disclosed a critical Salesforce Agentforce vulnerability chain where malicious Web-to-Lead content could be retrieved by an AI agent and used to exfiltrate sensitive CRM data through an indirect prompt-injection path.

Major
Capability impact
Scope breach 0 -> 4
Unsupervised authority use 0 -> 3
+2 more capability scores moved
Read source

Malicious Postmark MCP server secretly copied outbound agent emails

Snyk reported that the npm package postmark-mcp, used to let AI assistants send email through Postmark, was modified to add a hidden BCC to an external domain.

Minor
Capability impact
Third-party delegation 0 -> 2
Read source

ShadowLeak showed ChatGPT Gmail connector could leak data without a click

Radware disclosed ShadowLeak, a zero-click indirect prompt-injection vulnerability affecting ChatGPT connected to enterprise Gmail and web browsing, where a crafted email could cause sensitive inbox data to be sent to an attacker-controlled URL.

Info
Capability impact

No capability score increase

Read source

Darwin Godel Machine demonstrated open-ended self-improving agents

The Darwin Godel Machine research line describes agents that empirically validate self-modifications against benchmarks, preserve successful modifications, and explore an archive of descendants with improved coding-agent performance.

Major
Capability impact
Heritable adaptation 0 -> 4
Long-horizon execution 0 -> 3
+1 more capability score moved
Read source