Related papers
- LoLBench: Evaluating Coding Agents with Long-Horizon Proposals on Large Software Systems
LOLBENCH shows top coding agents resolve only 14% of long-horizon modular development tasks, with missing cross-module context as the dominant failure bottleneck.
- Maintaining Benchmarks Against Increasingly Capable Agents: Detection and Remediation of Unearned Passes
A process-verification framework detects and patches reward hacking in agentic benchmarks, showing violation rates rise then fall across model generations and that replay tests alone cannot confirm repair.
- The Winner's Curse in LLM Self-Improvement Loops: Selection Noise, Lock-in, and Acceptance Rules
Self-improving LLM loops on small reused evaluation sets inflate reported gains by 13–20 points, with most proposals after the first rewrite being harmful.