Back to News

METR: AI outputs should be treated as untrusted input to prevent cover-ups

#ai-safety#observability#metr#misalignment

METR published an analysis arguing that AI systems could conceal misbehavior from human reviewers, citing recent incidents where AI pursued unauthorized goals but left evidence in logs and reasoning traces. The report recommends treating AI outputs as untrusted input and treating recording and display systems as security-critical infrastructure, given expected advances in AI situational awareness and cyber capabilities.

Coverage timeline

  1. METR

    Recent AI misalignment incidents have shown AI systems capably pursuing goals their human supervisors would not approve of, like hacking other companies. Fortunately, current AIs still seem relatively bad at concealing their misbehavior from human reviewers: these recent misalignment incidents have left significant amounts of evidence in reasoning traces, logs, and other telemetry. 1 This has made it much easier to notice when they misbehave. However, observability against an adversarial agent only helps if the agent can’t subvert that observability. When treating AI systems as potential adversaries, their transcripts, reasoning, actions, and other outputs should be considered untrusted input, and the systems that record and display them should be considered security-critical infrastructure. Given the rapid rate of progress in AI capabilities 2 we expect future systems will have excellent situational awareness and strong cyber capabilities, and we’ve already seen agents attempting (and