Full text not available for this paper

Summary (Overview)

  • New benchmark introduced: WideSWE evaluates coding agents on cross-repository tasks requiring coordinated changes across multiple repositories, addressing a gap in existing benchmarks that focus on single-repository tasks.
  • Dataset composition: 120 real-world tasks (60 bug fixes, 60 features) mined from 1,729,171 merged PRs across 103 software ecosystems, covering 41 ecosystems with 253 target repositories and 2,815 F2P + 22,139 P2P tests.
  • Key finding: Best-performing configuration (Codex CLI + GPT-5.6-sol) achieves only 42.50% task success, with other configurations ranging from 10.83% to 37.50%, demonstrating that cross-repository coordination remains challenging.
  • Failure taxonomy: Three primary failure modes identified: incomplete scope identification (missing required repositories), recognized work without delivery (identifying but not implementing changes), and post-edit failure (modifying repositories without satisfying requirements).
  • Joint vs. independent execution: Joint execution achieves 40.45% vs. 35.96% for independent execution overall, but independent execution recovers omitted work (60% recovery for unmodified repos) while joint execution leverages cross-repository context for implementation and verification.

Introduction and Theoretical Foundation

Background and Motivation

Large language models (LLMs) now power coding agents that can inspect repositories, edit multiple files, run tests, and iteratively complete software-engineering tasks. Existing benchmarks (SWE-bench, DeepSWE, ProgramBench) evaluate agents primarily within a single codebase:

  • SWE-bench: Resolving real GitHub issues in a repository
  • DeepSWE: Original, long-horizon engineering tasks
  • ProgramBench: Rebuilding complete programs from reference executables

However, in real software ecosystems, a single feature or bug fix often requires coordinated changes across several repositories. Analysis of 1,729,171 PRs across 103 software ecosystems identified 109,233 PR records explicitly referencing another repository in the same ecosystem. Examples include:

  • Introducing a shared feature across Sentry's Go, Python, and Ruby SDKs
  • A new capability in Sentry's PHP SDK requiring updates to dependent repositories

Theoretical Foundation

The paper formalizes cross-repository task completion as: given a request pcp_c and historical workspace WcW_c containing repositories RcR_c, the target set Tc⊆RcT_c \subseteq R_c (with ∣Tc∣≥2|T_c| \geq 2) comprises repositories requiring substantive changes. The agent produces patches:

Δc=A(pc,Wc),Wc′=Wc⊕Δc\Delta_c = \mathcal{A}(p_c, W_c), \quad W'_c = W_c \oplus \Delta_c

Task success requires all target repositories to pass their hidden tests:

Success(c)=∏r∈Tcyc,r(Wc′)\text{Success}(c) = \prod_{r \in T_c} y_{c,r}(W'_c)

Methodology

Benchmark Construction Pipeline

  1. Ecosystem mining: From top 200 GitHub organizations by stars, selected 103 active software ecosystems supporting common products/platforms/technologies.

  2. Change mining: Collected merged PRs (since Jan 1, 2024), identified cross-repository links, and grouped linked changes into 4,437 candidate groups.

  3. Case selection: Focused on 2,188 recent groups (latest PR merged June 1, 2025+), manual review retained 635 groups, executable validation yielded 192 eligible cases with per-target F2P tests.

  4. Task balancing: Retained all 60 bug-fix cases and selected 60 of 132 feature cases for diversity.

Prompt Construction

Prompts built from original issues and PR descriptions, preserving wording where possible and combining related requirements into one cross-repository task. Clarified necessary behavior without introducing additional implementation instructions.

Hidden-Test Construction and Review

Two review rules addressed mismatches between task requirements and inherited tests:

  1. Relax implementation-specific constraints: Removed restrictions on private helper names, file paths, or exact error message text while preserving required behavior.

  2. Remove unrequested functionality checks: Removed tests for additions not required by the issue/PR while retaining checks for required functionality.

Experimental Setup

Evaluated seven agent configurations pairing scaffolds (Codex CLI v0.147.0, Claude Code v2.1.139) with LLMs (GPT-5.6-sol, Claude Opus 5, Gemini 3.8 Flash, DeepSeek V4 Pro, Qwen 3.8 Max, GLM 5.3).

Research questions:

  • RQ1: How effectively can coding agents complete cross-repository tasks?
  • RQ2: How does success vary with task type, repository count, and language diversity?
  • RQ3: How does joint execution compare with independent execution?

Empirical Validation / Results

RQ1: Overall Performance

Table 1: Main results by task type (excerpt):

MetricGPT-5.6+CodexGPT-5.6+CCOpus 5+CCGemini+CCDeepSeek+CCQwen+CCGLM 5.3+CC
Case ↑ (Overall)42.5032.5035.0010.8326.6737.5020.00
Repo ↑ (Overall)63.6456.9255.3415.8147.8356.9239.92
F2P ↑ (Overall)65.6160.0856.1316.6050.2057.3141.11
P2P ↑ (Overall)90.9585.0790.9595.9385.0792.3184.16
API ↓ (Overall)86.7134.8160.2146.1160.2177.1151.1

Failure Categories

Table 2: Failure categories (%):

ConfigurationScopeDeliveryPost-edit
GPT + Codex37.682.9059.42
GPT + CC30.861.2367.90
Opus + CC25.646.4167.95
Gemini + CC37.3852.3410.28
DeepSeek + CC30.687.9561.36
Qwen + CC21.335.3373.33
GLM 5.3 + CC30.219.3860.42

RQ2: Task Intent, Repository Scope, and Language Diversity

  • Repository count: All configurations show lower task success on three-repository tasks, though repository-level success can improve (e.g., GPT+Codex: 63.08% → 66.67% repo success, but 42.99% → 38.46% task success).
  • Language diversity: All configurations achieve higher success on multiple-family bug fixes than single-family, but this pattern does not hold for features. Language count alone is insufficient to judge task difficulty.

RQ3: Joint vs. Independent Execution

Table 3: Joint versus Independent execution:

GroupMetricJointIndependent
Task success (%)Bugfix34.4844.83
Feature43.3331.67
Overall40.4535.96
Repository outcomes (%)Repo61.7860.73
F2P64.4064.40
P2P89.6385.37
API94.7296.6

Key findings:

  • Independent execution recovers 60.00% of unmodified failed repositories, but only 9.43% of already-modified ones
  • Joint execution leverages cross-repository context for implementation (e.g., Ansible case) and verification (e.g., Symfony case)

Theoretical and Practical Implications

Implications for Benchmark Design

  1. Cross-repository evaluation is essential: Single-repository benchmarks overestimate agent capability by not testing coordinated changes across repositories.

  2. Requirement-aligned test review: The two review rules (relaxing implementation-specific constraints, removing unrequested checks) provide a methodology for constructing fair cross-repository benchmarks that don't penalize alternative correct implementations.

  3. Failure taxonomy: The three-category classification (scope, delivery, post-edit) provides a framework for diagnosing coding agent limitations in multi-repository settings.

Practical Implications for Agent Development

  1. Scope identification is critical: Agents need better mechanisms to identify all repositories requiring changes from a shared request.

  2. Cross-repository context matters: Joint execution provides behavioral references that help agents implement and verify compatible changes.

  3. Workload management: Larger joint workspaces can lead to forgetting required changes in some repositories, suggesting a need for better task tracking.

  4. Model-scaffold interactions: Performance varies significantly across model-scaffold combinations, with scaffold choice (Codex CLI vs. Claude Code) materially affecting outcomes.

Conclusion

WideSWE introduces a benchmark of 120 real-world tasks requiring coordinated changes across repositories. The highest task success rate across seven configurations is 42.50%, revealing significant room for improvement in cross-repository coordination. Agents exhibit three primary failure modes: missing necessary changes, leaving recognized work unfinished, and modifying required repositories without fully satisfying the request.

Key takeaways:

  • Cross-repository task completion is substantially harder than single-repository work
  • Independent execution can recover omitted work but doesn't improve overall success
  • Joint execution provides valuable cross-repository context for implementation and verification
  • Future work should focus on improving scope identification, delivery reliability, and post-edit correctness in multi-repository settings

Future directions: The benchmark shifts focus from progress in individual repositories to whether changes across repositories jointly fulfill task requirements, opening avenues for developing agents specialized in ecosystem-level coordination.

Related papers