Back to News

METR Deploys Basic Per-Action Monitor for AI Evaluations

#ai-safety#monitoring#evaluations#metr

METR published a blog post describing the development and deployment of a basic live per-action monitor for its AI evaluations, aimed at reducing incidents involving harmful actions by agents. The monitor uses an LLM judge to review each action before execution, flagging those above a threshold for human review and halting the eval. It focuses solely on real-world harm and attempts to subvert the monitoring system, ignoring other issues like cheating.

Coverage timeline

  1. METR

    In light of recent incidents (e.g. those from OpenAI , Anthropic , and UK AISI ), we developed and deployed a basic live per-action monitor to reduce the likelihood of incidents involving harmful actions from agents during our own evaluations. The monitor is intended to reliably detect actions that could plausibly cause real-world harm, with a sufficiently low false-positive rate that human reviewers will not be overwhelmed by manual review on large evals. This monitor is solely focused on real-world harm, or attempts to subvert the monitoring system itself. It is intended to ignore other nefarious things such as cheating, which we can scan for post-hoc. An LLM judge reviews each action from the agent before execution, and holds anything above a threshold for human review, halting the eval in the meantime. This post first sets out claims that we think would need to be made in order to make a strong argument for our monitoring system being effective. We give some notes on what evidence