# LoLBench: Evaluating Coding Agents with Long-Horizon Proposals on Large Software Systems

> LOLBENCH shows top coding agents resolve only 14% of long-horizon modular development tasks, with missing cross-module context as the dominant failure bottleneck.

- **Source:** [arXiv](https://arxiv.org/abs/2609.37143)
- **Published:** 2026-10-10
- **Permalink:** https://picx.dev/p/iHtzyN
- **Whiteboard:** https://picx.dev/p/iHtzyN/image

## Summary

# LOLBENCH: Evaluating Coding Agents with Long-Horizon Proposals on Large Software Systems

## Summary (Overview)

- **New benchmark introduction**: LOLBENCH is a multilingual benchmark of 100 modular development tasks across 29 large software systems in five domains, evaluating coding agents on the full proposal-to-implementation process.
- **Dual complexity framework**: The paper introduces two complementary dimensions of task difficulty—perception complexity (grounding user intent and high-level design in existing code) and implementation complexity (producing correct, regression-free code changes)—and positions modular development tasks as occupying the most challenging region of both.
- **Key empirical finding**: Across 28 evaluated agents (6 LLMs × 5 scaffolds), the best agent (Opus 5 on Claude Code or mini-SWE-agent) resolves only 14% of tasks, with a maximum F2P pass rate of 52.7%.
- **Major bottleneck identified**: Failure analysis attributes the largest failure source to incomplete code localization (cross-module context missing), with code localization accounting for 30.9 units of failure attribution on average.
- **Perception capability is critical**: Providing reference-derived file trees and API specifications improves resolved rates by 16–22 percentage points (2.4–17×), reaching at most 34%, demonstrating that converting abstract proposals to specifications remains a central challenge.

## Introduction and Theoretical Foundation

### Background and Motivation

The paper addresses a fundamental gap in coding agent evaluation. While modern coding agents can deliver increasingly large repository-level changes, most benchmarks evaluate agents' *implementation capability*—producing correct code edits from detailed specifications or bug reports. However, practical software development begins with **intent and high-level design**, not specifications. Developers must:

1. Ground abstract user intent in existing software architecture
2. Determine which modules and interfaces must change
3. Integrate and validate changes without breaking existing functionality

The authors illustrate this with **Python Enhancement Proposal 617** (van Rossum et al., 2020), which proposes replacing CPython's LL(1)-based parser with a PEG-based parser. Realizing this proposal requires determining how the new parser integrates with existing components—a task not captured by benchmarks that provide detailed specifications.

### Theoretical Framework: Two Complexity Dimensions

The paper formalizes modular development complexity through two dimensions:

1. **Perception complexity**: The difficulty of grounding user intent and high-level design in a large software system to derive a detailed implementation specification. This requires interpreting the proposal, understanding cross-module dependencies, and determining affected code.

2. **Implementation complexity**: The difficulty of translating a specification into functionally correct and regression-free code changes on an existing large software system.

The authors position different benchmark types on this two-dimensional space:

- **Simple repeated tasks**: Low perception + low implementation complexity
- **Repository-exploration tasks** (e.g., SWE-Explore): High perception + low implementation complexity  
- **Local bug-fixing tasks**: Low perception + high implementation complexity
- **Modular development on large systems**: High perception + high implementation complexity

### Key Design Principle

LOLBENCH tasks use **human-written enhancement proposals (EPs)** that predate their reference implementations. On average:
- Proposals contain ~5,000 words
- Software systems contain 2.4 million source lines of code (LoC)
- Implementation PRs change ~5,500 LoC across ~65 files, ~40 classes, and ~196 functions

Only 7.3% of target class names appear in proposals—substantially less than other benchmarks with large implementations—demonstrating the limited implementation guidance and high perception complexity.

## Methodology

### Benchmark Construction (Four-Phase Process)

#### Phase 1: Project Collection
EP–PR pairs are retained when:
- Software has ≥50k source LoC
- EP contains at least one implementable section
- PR changes ≥1k LoC across ≥10 files, with source code >50% of changed lines
- EP predates the PR and corresponds to exactly one merged PR
- PR contains tests checking proposed behavior

#### Phase 2: EP-PR Matching
- Proposal sections classified as **implementable** (externally observable functionality) or **contextual** (background/rationale)
- Changed files/classes/functions mapped to implementable sections
- Validation via 28 static rules (structural properties) and 8 semantic rules (scope/attribution)
- LLM judge scores candidates on 5 dimensions (requirement coverage, mapping precision, scope correctness, granularity, requirement specificity); candidates need weighted score ≥8.0

#### Phase 3: Task Construction (Harbor Framework)
Each task provides:
- The EP (proposal)
- The codebase before implementation
- Hidden reference solution (non-test portion of PR)

**Test design principles**:
- **F2P tests** (Fail-to-Pass): Must fail before implementation, pass after reference implementation; check system-level or end-to-end behavior through public interfaces (not private functions) to allow structural divergence
- **P2P tests** (Pass-to-Pass): Must pass before and after; serve as regression guards
- Offline Docker environment with no public network access
- Six-hour wall-clock limit per attempt

#### Phase 4: Task Enhancement
**Behavioral distinguishability augmentation** via three proposal-level transformations:
1. **Section Revert**: Removes an implementable capability
2. **Requirement Mismatch**: Alters a required value, condition, or boundary
3. **Semantic Mutation**: Changes required structure or execution flow

New F2P tests are generated to:
- Kill incorrect solution mutants that original tests cannot kill
- Improve line coverage when original F2P suite coverage <50%

Each F2P test is mapped to implementable EP sections for fine-grained, section-level evaluation of partial progress.

### Evaluation Setup

**Models evaluated**: DeepSeek V4 Flash, Opus 5, Kimi K3, GPT-5.6 Sol, GLM-5.2, MiniMax M3 (all via OpenRouter, second-highest reasoning effort)

**Scaffolds**: Claude Code, Codex, OpenCode, Pi, mini-SWE-agent (28 agents total; 2 unavailable due to provider issues)

**Metrics**:
- **Resolved**: % of tasks accepted by full task verifier
- **Completed**: % of completed implementable sections (section completed only when all mapped F2P tests pass)
- **F2P Pass / P2P Pass**: Micro-averaged over all tests
- Efficiency: turns, wall-clock time, input tokens, generated tokens, cost

**Failure analysis**: Pre-defined rules detect observable localization, editing, verification, repair, and tool-use signals; outcome-blinded LLM judge detects semantic requirement understanding and planning issues. Each solution failure contributes one unit of attribution across multiple failure modes.

## Empirical Validation / Results

### Benchmark Comparison

| Benchmark | #Tasks | Lang. | Type | Repo Size (k LoC) | Solution Size (LoC) | Modified (F/C/Fn) | Target Class Mentioned (%) |
|---|---|---|---|---|---|---|---|
| SWE-bench Verified | 500 | Python | Issue | 253.6 | 38 | 2.6/1.1/6.7 | 15.9 |
| SWE-bench Pro | 731 | Multiple | Issue | 270.0 | 300 | 7.2/2.4/14.5 | 13.7 |
| Multi-SWE-bench | 2,132 | Multiple | Issue | 205.8 | 224 | 6.6/1.8/11.6 | 14.0 |
| FEA-Bench | 1,401 | Python | Issue | 174.7 | 222 | 5.0/2.8/16.3 | 15.5 |
| DeepSWE | 113 | Multiple | Issue | 66.7 | 1,670 | 11.6/11.9/80.7 | 31.8 |
| FeatureBench | 200 | Python | Specification | 490.6 | 2,050 | 15.4/11.3/78.3 | 38.7 |
| **LOLBENCH** | **100** | **Multiple** | **Proposal** | **2,393.6** | **5,499** | **64.9/40.1/195.9** | **7.3** |

### Main Results (28 Agents)

**Key findings**:
- Best resolved rate: **14%** (Opus 5 on Claude Code and mini-SWE-agent)
- Mean/median resolved rates: 3.4% / 2.0%
- Best completed rate: **26.3%** (Claude Code + Opus 5); mean/median: 8.0% / 5.3%
- Best F2P pass rate: **52.7%** (Claude Code + Opus 5)

**Model sensitivity**: Holding scaffold fixed, RSD across models averages 62.0% (Completed) and 44.5% (F2P Pass). Holding model fixed, RSD across scaffolds averages 24.0% and 17.3%. Model choice matters ~2.6× more than scaffold choice.

**Efficiency observations**:
- Mini-SWE-agent with Opus 5 uses 3.2× wall-clock time, 5.0× generated tokens, 2.8× cost vs. Claude Code for same resolved rate
- Code localization is the largest phase for 26/28 agents during initial generation and 21/28 during repair
- Localization accounts for 60.4% of initial-generation turns and 44.2% of repair turns

### Failure Analysis

**Phase-level attribution** (averaged over 28 agents, units of failure attribution):
- **Code localization: 30.9 units** (largest)
- **Code editing: 25.3 units**
- Nearly all localization attribution (30.0 units) comes from **cross-module context missing**: solutions reach some target files/modules but miss others

### Perception Capability Analysis

**Code localization precision/recall** (task-level macro averages across 28 agents):
- Inspection recall: 46.4–86.4% (files), 31.9–83.1% (functions)
- Inspection precision: 16.1–36.3% (files), 6.8–19.2% (functions)
- Modification recall: 36.9–65.0% (files), 20.8–53.4% (functions)
- Modification precision: 64.9–80.4% (files), 54.2–65.6% (functions)

**Specification vs. proposal performance** (EPs reformulated as specifications with file trees and API specifications):

| Agent | Resolved (%) | Completed (%) | F2P Pass (%) | P2P Pass (%) |
|---|---|---|---|---|
| Opus 5 + Claude Code | 34.0 (+20.0) | 42.9 (+16.6) | 72.7 (+20.0) | 81.1 (+8.1) |
| GPT-5.6 Sol + Codex | 25.0 (+22.0) | 32.8 (+19.2) | 56.2 (+27.6) | 85.7 (+9.7) |
| DeepSeek V4 Flash + OpenCode | 17.0 (+16.0) | 17.7 (+13.7) | 31.6 (+13.3) | 69.4 (−4.2) |

## Theoretical and Practical Implications

### Theoretical Contributions

1. **Two-dimensional complexity framework**: The paper provides a principled way to characterize coding tasks along perception and implementation complexity, clarifying what different benchmarks actually measure and where current evaluation gaps lie.

2. **Perception as a first-class capability**: The results demonstrate that grounding abstract intent in large codebases is a distinct, measurable capability separable from implementation skill. The 16–22 percentage point improvement when specifications are provided shows that much of the difficulty in modular development lies *before* code editing begins.

3. **Quantified gap between proposal-level and specification-level performance**: The substantial performance gap (2.4–17× improvement) provides an upper bound on what improved perception capabilities could achieve, motivating targeted research.

### Practical Implications

1. **Agent design priorities**: Code localization dominates agent activity (60.4% of initial-generation turns), suggesting that improved exploration and cross-module context gathering could yield outsized gains.

2. **Model selection matters more than scaffold selection**: The 2.6× higher variability across models vs. scaffolds suggests practitioners should prioritize model capability over scaffold features for modular development tasks.

3. **Benchmark design guidance**: The test augmentation methodology (mutation-based F2P generation, coverage-guided test enhancement, section-level F2P mapping) provides a template for building rigorous, flexible evaluation suites.

4. **Realistic evaluation scale**: With 2.4M LoC repositories and 5,500 LoC solutions, LOLBENCH captures scale characteristics absent from most existing benchmarks, revealing that current agents operate far below practical modular development requirements.

## Conclusion

LOLBENCH introduces a multilingual benchmark of 100 modular development tasks across 29 large software systems, built from human-written enhancement proposals that predate their reference implementations. The benchmark uniquely evaluates both perception complexity (grounding user intent and high-level design in existing code) and implementation complexity (producing correct, regression-free changes).

**Key takeaways**:
- The best agent resolves only 14% of tasks, with mean/median resolved rates of 3.4%/2.0%
- Incomplete cross-module context (code localization) is the dominant failure mode
- Providing reference-derived implementation guidance improves resolved rates by 16–22 percentage points (2.4–17×), reaching at most 34%
- Both perception and implementation remain central challenges for coding agents in practical modular development

**Future directions**: The authors emphasize improving coding agents' perception capability to handle practical modular development tasks on large software systems, suggesting this is the most promising avenue for closing the gap between current agent performance and practical requirements.

---

_Markdown view of https://picx.dev/p/iHtzyN, served by PicX — AI-generated visual whiteboard summaries of research papers._
