Full text not available for this paper
Summary (Overview)
- New benchmark introduced: WideSWE evaluates coding agents on cross-repository tasks requiring coordinated changes across multiple repositories, addressing a gap in existing benchmarks that focus on single-repository tasks.
- Dataset composition: 120 real-world tasks (60 bug fixes, 60 features) mined from 1,729,171 merged PRs across 103 software ecosystems, covering 41 ecosystems with 253 target repositories and 2,815 F2P + 22,139 P2P tests.
- Key finding: Best-performing configuration (Codex CLI + GPT-5.6-sol) achieves only 42.50% task success, with other configurations ranging from 10.83% to 37.50%, demonstrating that cross-repository coordination remains challenging.
- Failure taxonomy: Three primary failure modes identified: incomplete scope identification (missing required repositories), recognized work without delivery (identifying but not implementing changes), and post-edit failure (modifying repositories without satisfying requirements).
- Joint vs. independent execution: Joint execution achieves 40.45% vs. 35.96% for independent execution overall, but independent execution recovers omitted work (60% recovery for unmodified repos) while joint execution leverages cross-repository context for implementation and verification.
Introduction and Theoretical Foundation
Background and Motivation
Large language models (LLMs) now power coding agents that can inspect repositories, edit multiple files, run tests, and iteratively complete software-engineering tasks. Existing benchmarks (SWE-bench, DeepSWE, ProgramBench) evaluate agents primarily within a single codebase:
- SWE-bench: Resolving real GitHub issues in a repository
- DeepSWE: Original, long-horizon engineering tasks
- ProgramBench: Rebuilding complete programs from reference executables
However, in real software ecosystems, a single feature or bug fix often requires coordinated changes across several repositories. Analysis of 1,729,171 PRs across 103 software ecosystems identified 109,233 PR records explicitly referencing another repository in the same ecosystem. Examples include:
- Introducing a shared feature across Sentry's Go, Python, and Ruby SDKs
- A new capability in Sentry's PHP SDK requiring updates to dependent repositories
Theoretical Foundation
The paper formalizes cross-repository task completion as: given a request and historical workspace containing repositories , the target set (with ) comprises repositories requiring substantive changes. The agent produces patches:
Task success requires all target repositories to pass their hidden tests:
Methodology
Benchmark Construction Pipeline
-
Ecosystem mining: From top 200 GitHub organizations by stars, selected 103 active software ecosystems supporting common products/platforms/technologies.
-
Change mining: Collected merged PRs (since Jan 1, 2024), identified cross-repository links, and grouped linked changes into 4,437 candidate groups.
-
Case selection: Focused on 2,188 recent groups (latest PR merged June 1, 2025+), manual review retained 635 groups, executable validation yielded 192 eligible cases with per-target F2P tests.
-
Task balancing: Retained all 60 bug-fix cases and selected 60 of 132 feature cases for diversity.
Prompt Construction
Prompts built from original issues and PR descriptions, preserving wording where possible and combining related requirements into one cross-repository task. Clarified necessary behavior without introducing additional implementation instructions.
Hidden-Test Construction and Review
Two review rules addressed mismatches between task requirements and inherited tests:
-
Relax implementation-specific constraints: Removed restrictions on private helper names, file paths, or exact error message text while preserving required behavior.
-
Remove unrequested functionality checks: Removed tests for additions not required by the issue/PR while retaining checks for required functionality.
Experimental Setup
Evaluated seven agent configurations pairing scaffolds (Codex CLI v0.147.0, Claude Code v2.1.139) with LLMs (GPT-5.6-sol, Claude Opus 5, Gemini 3.8 Flash, DeepSeek V4 Pro, Qwen 3.8 Max, GLM 5.3).
Research questions:
- RQ1: How effectively can coding agents complete cross-repository tasks?
- RQ2: How does success vary with task type, repository count, and language diversity?
- RQ3: How does joint execution compare with independent execution?
Empirical Validation / Results
RQ1: Overall Performance
Table 1: Main results by task type (excerpt):
| Metric | GPT-5.6+Codex | GPT-5.6+CC | Opus 5+CC | Gemini+CC | DeepSeek+CC | Qwen+CC | GLM 5.3+CC |
|---|---|---|---|---|---|---|---|
| Case ↑ (Overall) | 42.50 | 32.50 | 35.00 | 10.83 | 26.67 | 37.50 | 20.00 |
| Repo ↑ (Overall) | 63.64 | 56.92 | 55.34 | 15.81 | 47.83 | 56.92 | 39.92 |
| F2P ↑ (Overall) | 65.61 | 60.08 | 56.13 | 16.60 | 50.20 | 57.31 | 41.11 |
| P2P ↑ (Overall) | 90.95 | 85.07 | 90.95 | 95.93 | 85.07 | 92.31 | 84.16 |
| API ↓ (Overall) | 86.7 | 134.8 | 160.2 | 146.1 | 160.2 | 177.1 | 151.1 |
Failure Categories
Table 2: Failure categories (%):
| Configuration | Scope | Delivery | Post-edit |
|---|---|---|---|
| GPT + Codex | 37.68 | 2.90 | 59.42 |
| GPT + CC | 30.86 | 1.23 | 67.90 |
| Opus + CC | 25.64 | 6.41 | 67.95 |
| Gemini + CC | 37.38 | 52.34 | 10.28 |
| DeepSeek + CC | 30.68 | 7.95 | 61.36 |
| Qwen + CC | 21.33 | 5.33 | 73.33 |
| GLM 5.3 + CC | 30.21 | 9.38 | 60.42 |
RQ2: Task Intent, Repository Scope, and Language Diversity
- Repository count: All configurations show lower task success on three-repository tasks, though repository-level success can improve (e.g., GPT+Codex: 63.08% → 66.67% repo success, but 42.99% → 38.46% task success).
- Language diversity: All configurations achieve higher success on multiple-family bug fixes than single-family, but this pattern does not hold for features. Language count alone is insufficient to judge task difficulty.
RQ3: Joint vs. Independent Execution
Table 3: Joint versus Independent execution:
| Group | Metric | Joint | Independent |
|---|---|---|---|
| Task success (%) | Bugfix | 34.48 | 44.83 |
| Feature | 43.33 | 31.67 | |
| Overall | 40.45 | 35.96 | |
| Repository outcomes (%) | Repo | 61.78 | 60.73 |
| F2P | 64.40 | 64.40 | |
| P2P | 89.63 | 85.37 | |
| API | 94.7 | 296.6 |
Key findings:
- Independent execution recovers 60.00% of unmodified failed repositories, but only 9.43% of already-modified ones
- Joint execution leverages cross-repository context for implementation (e.g., Ansible case) and verification (e.g., Symfony case)
Theoretical and Practical Implications
Implications for Benchmark Design
-
Cross-repository evaluation is essential: Single-repository benchmarks overestimate agent capability by not testing coordinated changes across repositories.
-
Requirement-aligned test review: The two review rules (relaxing implementation-specific constraints, removing unrequested checks) provide a methodology for constructing fair cross-repository benchmarks that don't penalize alternative correct implementations.
-
Failure taxonomy: The three-category classification (scope, delivery, post-edit) provides a framework for diagnosing coding agent limitations in multi-repository settings.
Practical Implications for Agent Development
-
Scope identification is critical: Agents need better mechanisms to identify all repositories requiring changes from a shared request.
-
Cross-repository context matters: Joint execution provides behavioral references that help agents implement and verify compatible changes.
-
Workload management: Larger joint workspaces can lead to forgetting required changes in some repositories, suggesting a need for better task tracking.
-
Model-scaffold interactions: Performance varies significantly across model-scaffold combinations, with scaffold choice (Codex CLI vs. Claude Code) materially affecting outcomes.
Conclusion
WideSWE introduces a benchmark of 120 real-world tasks requiring coordinated changes across repositories. The highest task success rate across seven configurations is 42.50%, revealing significant room for improvement in cross-repository coordination. Agents exhibit three primary failure modes: missing necessary changes, leaving recognized work unfinished, and modifying required repositories without fully satisfying the request.
Key takeaways:
- Cross-repository task completion is substantially harder than single-repository work
- Independent execution can recover omitted work but doesn't improve overall success
- Joint execution provides valuable cross-repository context for implementation and verification
- Future work should focus on improving scope identification, delivery reliability, and post-edit correctness in multi-repository settings
Future directions: The benchmark shifts focus from progress in individual repositories to whether changes across repositories jointly fulfill task requirements, opening avenues for developing agents specialized in ecosystem-level coordination.
Related papers
- ScAn-Bench: Evaluating Scaling Analysis Methodology
ScAn-Bench introduces the first surrogate benchmarks for scaling analysis, revealing no universally optimal data acquisition and extrapolation strategy exists across LLMs and VLMs.
- T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks
Pure RL on executed outcomes trains a 122B-parameter MoE terminal agent to 64.0% on Terminal-Bench 2.1, surpassing GPT-5.4 with only 10B active parameters.
- Shortcutting the Fix: Agentic Shortcutting in Software-Engineering Benchmarks
Agentic shortcutting—agents exploiting leaked solutions like upstream repos or Git history—inflates SWE benchmark scores by up to 82%, but a simple originality prompt cuts exploitation to under 11%.