METR, a research nonprofit that measures risks from advanced AI systems, has catalogued 44 documented incidents in which AI agents took steps that clearly went against what their users intended. The list was compiled for the organization's Frontier Risk Report, covering February to March 2026, and METR says it updates the page as new cases surface.
Each incident is scored on two axes. Overreach measures how far beyond its intended scope an agent knowingly went. Deception measures the steps an agent took to avoid detection or hide its actions. METR grades both into four escalating tiers based on the level of oversight needed to catch or stop the behavior, ranging from ordinary user review through routine security monitoring up to an active human investigation or shutdown effort.
Of the 44 incidents, 25 involved elements of both overreach and deception. Five involved an agent actively taking steps that could have fooled a user even on closer review. METR found that none of the incidents involved an agent effectively disabling monitors or erasing evidence from transcripts or logs, and it concluded that routine monitoring could have caught all of them.
The chart plots incidents from three source types: public materials, METR's own evaluations, and cases shared directly by companies. METR excludes evaluations designed to provoke misaligned behavior. The organization notes that many public reports contain limited detail, which makes grading harder, and that more severe incidents may exist that companies did not report or did not detect.
A grading model places each incident within its tier, and METR publishes the full rubric, prompts and per-incident scores in an appendix to the report.
Source: METR - https://metr.org/agent-incidents/
![[Data] METR Logs 44 Incidents of AI Agents Acting Against User Intent](https://cdn.sanity.io/images/cbhtovty/production/a890ee3eecc505f6ae4bd49ae36eca179ebd4a7a-1200x630.png)