Full text not available for this paper

Summary

Summary (Overview)

  • This paper presents an empirical study of 1,262 performance issues fixed by coding agents (Devin, Copilot, Cursor, Claude Code, Jules) across 582 repositories, drawn from the AIDev v4 dataset of 71,677 agent PRs.
  • The study finds that 57% of closed agent performance fixes are merged, 61% of rejections state no reason, and only 6 of 23 re-executed rejected claims held under a three-run pilot protocol.
  • Acceptance is driven primarily by the agent's track record in the repository and the repository's pre-opening merge rate on other agent PRs, not by the coded content of the fix itself.
  • Of 30 re-executed merged fixes, only 18 met the delivery criterion, 3 fell short of claims, and 9 showed no significant gain or regressed; 14 changed behavior on untested inputs.
  • The paper provides a dataset of 1,262 coded agent performance fixes with codebooks, coding sheets, and re-execution protocols for reproducibility.

Introduction and Theoretical Foundation

The study addresses a critical gap in understanding how coding agents' performance fixes fare in real repositories. While prior work has studied human-written performance fixes and PR acceptance separately, no study has confirmed each agent fix by hand, separated the fix's content from the repository that received it, or re-executed what the fix claims.

Key background strands:

  • Performance measurement: Early work by Georges et al. (2007) and Mytkowicz et al. (2009) showed that reliable performance measurement requires repeated runs with proper controls, as environmental factors alone can reverse a claimed speed-up.
  • PR acceptance literature: Tsay et al. (2014) showed that a submitter's prior interaction with the project raises acceptance more than tests; Wyrich et al. (2021) found bot PRs merged only 37% of the time versus 73% for humans.
  • LLM-based optimization benchmarks: Benchmarks like SWE-bench, PIE, and GSO score agents on measured speed-ups in fixed harnesses, but these differ from real repository conditions.
  • Human performance fix studies: Zhao et al. (2023) coded 570 human fixes into eight root causes, finding 27% architectural-level and 15% with test changes.

Methodology

The study employs a multi-stage pipeline:

Dataset construction:

  • Started with 71,677 agent PRs from AIDev v4 in repositories over 100 stars
  • Applied a text filter and codebook coding (first by language models, then by authors) to select 1,262 performance issues fixed by six agents
  • Coded each issue for inefficiency type, fix characteristics, scope, and tests

Re-execution protocol:

  • Re-executed 23 rejected fixes with claimed improvements (three-run pilot)
  • Re-executed 30 merged fixes that change tests (more intensive protocol with 12 interleaved forks at three workload sizes)
  • Used Claude Opus models as blind-pair judges and re-execution agents

Statistical analysis:

  • Used Fisher's exact tests, Wilcoxon tests, and logistic regression models
  • Clustered standard errors by repository
  • Controlled for author association, size of change, and repository effects

Empirical Validation / Results

RQ1: Acceptance rates

  • 57% of closed fixes merged; 61% of rejections state no reason
  • Only 6 of 23 re-executed rejected claims held under the pilot protocol
  • Agent PRs merged 41% vs. 80% for human PRs in five repositories; after third-party approval, rates converge (95% vs. 96%)

RQ2: Factors associated with acceptance

FactorEffect
Agent track record in repository31–37% → 70% acceptance
Repository's pre-opening merge rate33% → 84% acceptance
Deleted-line share0.26 (merged) vs. 0.15 (rejected)
Coded content, description, testsNo association detected

RQ3: Architectural changes

  • 46% of agent fixes are architectural-level (vs. 27% for human fixes)
  • Repeated computation and redundant data processing cause 44% of issues

RQ4: Evidence provided

  • Agents change tests in 37% of fixes; only 11% carry a performance test or benchmark
  • Acceptance is no higher with any kind of test
  • Of 30 merged fixes: 18 met delivery criterion, 3 improved below claim, 9 showed no significant gain or regressed

Key finding: The funnel comparison showed agent PRs receive third-party approval 35% vs. 53% for humans, but after approval, merge rates converge (95% vs. 96%, p = 0.39).

Theoretical and Practical Implications

For agent builders:

  • Smaller fixes that remove code are associated with higher acceptance than large added mechanisms
  • The deleted-line share is the only fix-level signal that holds within agent and repository
  • Architectural-level fixes need per-file depth recording, not just file counts

For maintainers:

  • A measurement that can be re-run would put fix-specific evidence into merge decisions
  • 61% of rejections state no reason, and reviewers ask for measurement in only 3%
  • A merge does not prove the fix delivers what it claims

For researchers using AIDev:

  • A changed test is weak evidence that the claimed benefit was verified
  • Studies should code what a test checks before using test changes to measure verification
  • Acceptance as a quality measure needs within-repository controls due to confounding with prior treatment

Conclusion

The study demonstrates that whether an agent's performance fix is merged varies with the agent's track record, the repository's history with agent PRs, and whether the fix removes or adds code—not with the fix's coded content, description, tests, or measurements. The authors recommend that maintainers re-run claimed speed-ups before deciding and that researchers use within-repository controls when using acceptance as a quality metric.

Future directions:

  1. Controlled studies submitting the same optimization in removing vs. adding forms
  2. Larger re-execution arms with consistent protocols
  3. Funnel comparisons in repositories where humans and multiple agents submit side-by-side
  4. Refreshing the snapshot to track how repository treatment of agent PRs evolves

Limitations: The study's external validity is bounded by the AIDev v4 snapshot (six agents, repositories over 100 stars, through October 2025), the small re-execution pilot (23 fixes), and potential same-vendor bias in the judging instruments.

Related papers