# WideSWE: Can Coding Agents Coordinate Changes Across Repositories?

> WideSWE, a new benchmark of 120 cross-repository tasks, shows top coding agents succeed only 42.5% of the time, revealing major gaps in multi-repo coordination.

- **Source:** [arXiv](https://arxiv.org/abs/2609.33382)
- **Published:** 2026-10-03
- **Permalink:** https://picx.dev/p/bCaciz
- **Whiteboard:** https://picx.dev/p/bCaciz/image

## Summary

## Summary (Overview)

- **New benchmark introduced**: WideSWE evaluates coding agents on cross-repository tasks requiring coordinated changes across multiple repositories, addressing a gap in existing benchmarks that focus on single-repository tasks.
- **Dataset composition**: 120 real-world tasks (60 bug fixes, 60 features) mined from 1,729,171 merged PRs across 103 software ecosystems, covering 41 ecosystems with 253 target repositories and 2,815 F2P + 22,139 P2P tests.
- **Key finding**: Best-performing configuration (Codex CLI + GPT-5.6-sol) achieves only 42.50% task success, with other configurations ranging from 10.83% to 37.50%, demonstrating that cross-repository coordination remains challenging.
- **Failure taxonomy**: Three primary failure modes identified: incomplete scope identification (missing required repositories), recognized work without delivery (identifying but not implementing changes), and post-edit failure (modifying repositories without satisfying requirements).
- **Joint vs. independent execution**: Joint execution achieves 40.45% vs. 35.96% for independent execution overall, but independent execution recovers omitted work (60% recovery for unmodified repos) while joint execution leverages cross-repository context for implementation and verification.

## Introduction and Theoretical Foundation

### Background and Motivation

Large language models (LLMs) now power coding agents that can inspect repositories, edit multiple files, run tests, and iteratively complete software-engineering tasks. Existing benchmarks (SWE-bench, DeepSWE, ProgramBench) evaluate agents primarily within a single codebase:

- **SWE-bench**: Resolving real GitHub issues in a repository
- **DeepSWE**: Original, long-horizon engineering tasks
- **ProgramBench**: Rebuilding complete programs from reference executables

However, in real software ecosystems, a single feature or bug fix often requires coordinated changes across several repositories. Analysis of **1,729,171 PRs across 103 software ecosystems** identified **109,233 PR records** explicitly referencing another repository in the same ecosystem. Examples include:

- Introducing a shared feature across Sentry's Go, Python, and Ruby SDKs
- A new capability in Sentry's PHP SDK requiring updates to dependent repositories

### Theoretical Foundation

The paper formalizes cross-repository task completion as: given a request $p_c$ and historical workspace $W_c$ containing repositories $R_c$, the target set $T_c \subseteq R_c$ (with $|T_c| \geq 2$) comprises repositories requiring substantive changes. The agent produces patches:

$$\Delta_c = \mathcal{A}(p_c, W_c), \quad W'_c = W_c \oplus \Delta_c$$

Task success requires all target repositories to pass their hidden tests:

$$\text{Success}(c) = \prod_{r \in T_c} y_{c,r}(W'_c)$$

## Methodology

### Benchmark Construction Pipeline

1. **Ecosystem mining**: From top 200 GitHub organizations by stars, selected 103 active software ecosystems supporting common products/platforms/technologies.

2. **Change mining**: Collected merged PRs (since Jan 1, 2024), identified cross-repository links, and grouped linked changes into 4,437 candidate groups.

3. **Case selection**: Focused on 2,188 recent groups (latest PR merged June 1, 2025+), manual review retained 635 groups, executable validation yielded 192 eligible cases with per-target F2P tests.

4. **Task balancing**: Retained all 60 bug-fix cases and selected 60 of 132 feature cases for diversity.

### Prompt Construction

Prompts built from original issues and PR descriptions, preserving wording where possible and combining related requirements into one cross-repository task. Clarified necessary behavior without introducing additional implementation instructions.

### Hidden-Test Construction and Review

Two review rules addressed mismatches between task requirements and inherited tests:

1. **Relax implementation-specific constraints**: Removed restrictions on private helper names, file paths, or exact error message text while preserving required behavior.

2. **Remove unrequested functionality checks**: Removed tests for additions not required by the issue/PR while retaining checks for required functionality.

### Experimental Setup

Evaluated seven agent configurations pairing scaffolds (Codex CLI v0.147.0, Claude Code v2.1.139) with LLMs (GPT-5.6-sol, Claude Opus 5, Gemini 3.8 Flash, DeepSeek V4 Pro, Qwen 3.8 Max, GLM 5.3).

Research questions:
- **RQ1**: How effectively can coding agents complete cross-repository tasks?
- **RQ2**: How does success vary with task type, repository count, and language diversity?
- **RQ3**: How does joint execution compare with independent execution?

## Empirical Validation / Results

### RQ1: Overall Performance

**Table 1: Main results by task type** (excerpt):

| Metric | GPT-5.6+Codex | GPT-5.6+CC | Opus 5+CC | Gemini+CC | DeepSeek+CC | Qwen+CC | GLM 5.3+CC |
|--------|---------------|------------|-----------|-----------|-------------|---------|-----------|
| Case ↑ (Overall) | 42.50 | 32.50 | 35.00 | 10.83 | 26.67 | 37.50 | 20.00 |
| Repo ↑ (Overall) | 63.64 | 56.92 | 55.34 | 15.81 | 47.83 | 56.92 | 39.92 |
| F2P ↑ (Overall) | 65.61 | 60.08 | 56.13 | 16.60 | 50.20 | 57.31 | 41.11 |
| P2P ↑ (Overall) | 90.95 | 85.07 | 90.95 | 95.93 | 85.07 | 92.31 | 84.16 |
| API ↓ (Overall) | 86.7 | 134.8 | 160.2 | 146.1 | 160.2 | 177.1 | 151.1 |

### Failure Categories

**Table 2: Failure categories (%)**:

| Configuration | Scope | Delivery | Post-edit |
|---------------|-------|----------|-----------|
| GPT + Codex | 37.68 | 2.90 | 59.42 |
| GPT + CC | 30.86 | 1.23 | 67.90 |
| Opus + CC | 25.64 | 6.41 | 67.95 |
| Gemini + CC | 37.38 | 52.34 | 10.28 |
| DeepSeek + CC | 30.68 | 7.95 | 61.36 |
| Qwen + CC | 21.33 | 5.33 | 73.33 |
| GLM 5.3 + CC | 30.21 | 9.38 | 60.42 |

### RQ2: Task Intent, Repository Scope, and Language Diversity

- **Repository count**: All configurations show lower task success on three-repository tasks, though repository-level success can improve (e.g., GPT+Codex: 63.08% → 66.67% repo success, but 42.99% → 38.46% task success).
- **Language diversity**: All configurations achieve higher success on multiple-family bug fixes than single-family, but this pattern does not hold for features. Language count alone is insufficient to judge task difficulty.

### RQ3: Joint vs. Independent Execution

**Table 3: Joint versus Independent execution**:

| Group | Metric | Joint | Independent |
|-------|--------|-------|-------------|
| Task success (%) | Bugfix | 34.48 | 44.83 |
| | Feature | 43.33 | 31.67 |
| | Overall | 40.45 | 35.96 |
| Repository outcomes (%) | Repo | 61.78 | 60.73 |
| | F2P | 64.40 | 64.40 |
| | P2P | 89.63 | 85.37 |
| | API | 94.7 | 296.6 |

Key findings:
- Independent execution recovers 60.00% of unmodified failed repositories, but only 9.43% of already-modified ones
- Joint execution leverages cross-repository context for implementation (e.g., Ansible case) and verification (e.g., Symfony case)

## Theoretical and Practical Implications

### Implications for Benchmark Design

1. **Cross-repository evaluation is essential**: Single-repository benchmarks overestimate agent capability by not testing coordinated changes across repositories.

2. **Requirement-aligned test review**: The two review rules (relaxing implementation-specific constraints, removing unrequested checks) provide a methodology for constructing fair cross-repository benchmarks that don't penalize alternative correct implementations.

3. **Failure taxonomy**: The three-category classification (scope, delivery, post-edit) provides a framework for diagnosing coding agent limitations in multi-repository settings.

### Practical Implications for Agent Development

1. **Scope identification is critical**: Agents need better mechanisms to identify all repositories requiring changes from a shared request.

2. **Cross-repository context matters**: Joint execution provides behavioral references that help agents implement and verify compatible changes.

3. **Workload management**: Larger joint workspaces can lead to forgetting required changes in some repositories, suggesting a need for better task tracking.

4. **Model-scaffold interactions**: Performance varies significantly across model-scaffold combinations, with scaffold choice (Codex CLI vs. Claude Code) materially affecting outcomes.

## Conclusion

WideSWE introduces a benchmark of 120 real-world tasks requiring coordinated changes across repositories. The highest task success rate across seven configurations is 42.50%, revealing significant room for improvement in cross-repository coordination. Agents exhibit three primary failure modes: missing necessary changes, leaving recognized work unfinished, and modifying required repositories without fully satisfying the request.

**Key takeaways**:
- Cross-repository task completion is substantially harder than single-repository work
- Independent execution can recover omitted work but doesn't improve overall success
- Joint execution provides valuable cross-repository context for implementation and verification
- Future work should focus on improving scope identification, delivery reliability, and post-edit correctness in multi-repository settings

**Future directions**: The benchmark shifts focus from progress in individual repositories to whether changes across repositories jointly fulfill task requirements, opening avenues for developing agents specialized in ecosystem-level coordination.

---

_Markdown view of https://picx.dev/p/bCaciz, served by PicX — AI-generated visual whiteboard summaries of research papers._
