Back to News

Redwood Research: Filler tokens boost GPT-6-Astra reasoning, complicating monitoring

#gpt-6-astra#filler-tokens#interpretability#redwood-research

Redwood Research reports that GPT-6-Astra performs significantly better on serial reasoning tasks when prompts are padded with meaningless filler tokens, improving from ~10% to ~50% on 4-hop natural facts reasoning and from ~60% to ~90% on old AIME problems. The finding raises interpretability concerns because the model can perform substantial cognition without verbalizing it in chain-of-thought, making monitoring harder.

Coverage timeline

  1. Redwood ResearchDylan Xu

    We measure GPT-6-Astra’s capabilities when its prompt is padded with a variable number of meaningless “filler” tokens (e.g., dots) and it is told to answer immediately 1 without reasoning. On tasks designed to require lots of serial cognition, Astra performs significantly better with filler tokens than without (e.g., improving from ~10% to ~50% on 4-hop natural facts reasoning). On more general benchmarks, filler tokens also modestly improve Astra’s performance (e.g., improving from ~60% to ~90% on old AIME problems). This is concerning because it means Astra can perform significant cognition that it doesn’t verbalize in its chain-of-thought, making it harder to monitor. We first measure Astra’s performance on “N-hop natural facts” : a task that asks the model to retrieve some natural language facts in succession, similar to Ryan Greenblatt’s filler token eval (but with more hops). An example question in this benchmark is the following: Thanks for reading Redwood Research blog! Subscri