Back to News

Redwood Research: CoT controllability evals under-elicited, better prompts boost scores 2-3x

#chain-of-thought#evaluation#prompt-engineering#redwood-research

Redwood Research reports that the CoTControl eval, which tests reasoning models' ability to follow formatting constraints in their chain-of-thought, appears heavily under-elicited. Iterating on prompt templates with Claude Opus 4.6 improved open-source model performance by roughly 2-3x (e.g., GPT-OSS-120B from 5.5% to 15%), challenging recent system card claims from OpenAI and Anthropic that frontier models cannot meaningfully shape their CoTs.

Coverage timeline

  1. Redwood ResearchArun Jose

    The CoTControl eval asks reasoning models to follow formatting constraints in their chain-of-thought (e.g. write in all lowercase, avoid a specific word) while solving questions. Models seem to mostly be pretty bad at this: recent models score between 0-30% with the exception of Mythos Preview 1 . OpenAI and Anthropic have used this eval in recent system cards ( GPT-5.5 , Fable 5 ) to argue that their current frontier models are pretty bad at shaping their CoTs in ways that would reduce monitorability. 2 I asked Claude Opus 4.6 to iterate on prompt templates for the eval, and find that using better prompts improves performance for open-source models by ~2-3x or more (e.g. from 5.5% to 15% for GPT-OSS-120B). My guess is that more iteration would probably produce even stronger results. Thanks for reading Redwood Research blog! Subscribe for free to receive new posts and support my work. This suggests that the eval is heavily under-elicited, and that current CoT controllability numbers ma