Back to News

METR publishes predeployment evaluation of Claude Opus 5.5 AI R&D capabilities

#metr#claude-opus-5.5#ai-rd#evaluation

METR released a summary of its predeployment evaluation of Claude Opus 5.5, focusing on the model's potential to accelerate AI R&D at Anthropic. The evaluation, conducted under an unpaid agreement, assessed capabilities on difficult, long-horizon tasks and examined whether AI R&D at Anthropic was already dramatically accelerated during the model's development. METR noted the work was not meant to verify compliance with specific policy thresholds or assess alignment properties.

Coverage timeline

  1. METR

    Note on independence: This evaluation was conducted under an unpaid agreement for AI R&D assessment. 1 We drafted the initial summary, and then Anthropic had the opportunity to review and edit the text. We signed off on this final text from the Claude Opus 5.5 system card . Our preliminary evaluation focused on how Claude Opus 5.5 might impact AI R&D, mainly based on its capabilities on difficult, long-horizon tasks. The main claims we attempt to assess in this report are: (A) would AI R&D at Anthropic now be dramatically accelerated by using Claude Opus 5.5; and (B) was AI R&D at Anthropic already dramatically accelerated due to AI during the development of Claude Opus 5.5. Note that our work was oriented around collecting evidence related to AI R&D capabilities but was not meant to verify claims about compliance with any specific threshold from Anthropic’s policies. This report summary also does not attempt to assess whether Claude Opus 5.5 has or does not have particular alignment p