Related papers
- WideSWE: Can Coding Agents Coordinate Changes Across Repositories?
WideSWE, a new benchmark of 120 cross-repository tasks, shows top coding agents succeed only 42.5% of the time, revealing major gaps in multi-repo coordination.
- EVOHARNESSBENCH: Can Your Agents Keep Pace with an Evolving Harness?
EVOHARNESSBENCH shows that expanding an agent's tool, skill, or agent harness alone can degrade performance on previously solved tasks, inducing forgetting up to 34.7%.
- Audit the Scaffold, Not the Checkpoint: A Stationarity Dichotomy for Recursive Self-Improvement in Agentic Coding
Frozen weights do not guarantee safe saturation in agentic coding; auditors must monitor scaffold expansion, not just checkpoint freezing.