# One Skill Too Many: How Co-Installed Skills Conflict in Coding Agents

> Co-installed coding-agent skills that do the same job reduce the installed skill's usage by 19.9 percentage points without lowering task completion, a conflict decided at the first skill read.

- **Source:** [arXiv](https://arxiv.org/abs/2610.11647)
- **Published:** 2026-10-10
- **Permalink:** https://picx.dev/p/rfy1rK
- **Whiteboard:** https://picx.dev/p/rfy1rK/image

## Summary

# One Skill Too Many: How Co-Installed Skills Conflict in Coding Agents

## Summary (Overview)

- **First empirical study of skill conflicts in coding agents**: The paper systematically investigates what happens when two similar agent skills (directories with SKILL.md files) are co-installed in coding agents like Claude Code, showing that conflicts are common and harmful.
- **Scale of the problem**: From 20,947 repositories, the authors estimate that nearly **one in four installed skills** (23.5%) is co-installed with a skill that does the same job, and 63.7% have a similar skill by another author that a user could add.
- **Key finding on impact**: A similar skill takes **one in five runs (19.9 pp)** away from the installed skill without lowering task completion, while reducing fidelity on the installed skill's exclusive core functions by **5.6 percentage points**.
- **Decision point discovery**: Conflicts are decided at the **first read of a skill** (almost always before any file changes), and a pre-tool hook at that point restores fidelity to baseline levels.
- **Implications**: Benchmarks should score exclusive core functions rather than task completion alone, and platforms should guard the first read and disclose which skill ran.

## Introduction and Theoretical Foundation

### Background

Coding agents (e.g., Claude Code, Codex, Gemini CLI) are increasingly extended with **agent skills**—directories containing a SKILL.md file that tells the model when and how to perform a task. Skills encode conventions, safety rules, and specialized knowledge that models wouldn't know on their own. Skills are now shared at scale: recent studies collect **138,133 SKILL.md files from 20,556 repositories** and **42,447 skills from two marketplaces**.

### The Conflict Problem

Because skills come from independent sources (teams, plugins, copied collections, personal directories), two skills that do the same job can end up co-installed without anyone choosing the pair. When this happens:

- The installed skill loses its **core functions** (behaviors the user installed it for)
- The similar skill may run in its place or change what the installed skill does
- **Task completion often remains unchanged**, so existing benchmarks (which check only pass/fail) miss these conflicts entirely

### Three Key Challenges

The paper identifies three challenges that make studying skill conflicts hard:

1. **A conflict hides behind success**: Two skills compete because they do the same job, so either can complete the task. The only trace is missing behavior specific to the installed skill, written in prose.
2. **A conflict is resolved out of sight**: The model chooses from a few hundred characters of description without telling the user. Choice depends on listing order, scope, and the model.
3. **A conflict is decided at run time**: Whether two skills conflict depends on a request that doesn't exist when they're installed, and similar wording doesn't imply the same job.

### Key Definitions

- **Core functions**: Requirements that a competent model would not meet without the skill, quoted from its text, and checkable in the output
- **Fidelity**: The share of primary core functions (first three applicable) fulfilled in a run
- **Exclusive core functions**: Core functions of skill A that skill B does not also ask for
- **Substitution**: B is used instead of A
- **Interference**: A is used but the result still changes

## Methodology

### Study Design Overview

The study uses a controlled empirical design on **Claude Code** with three models (Sonnet 4.6, Haiku 4.5, Opus 5), comprising:

- **6,368 runs** with **169,294 tool calls** over **542 hours** of agent time
- **312 experimental pairs** across six cells (cross-project/within-project × normative/capability/script-bearing)
- Cost: approximately **$5,550** at API prices

### Installation Reconstruction

- Corpus of **55,698 distinct skills** from **20,947 repositories** (snapshots taken 18 July 2026)
- Reconstructed installation lists: **5,106 installation lists**, 4,342 with ≥2 distinct skills
- Skill filtering: frontmatter on first line, valid YAML, no disable-model-invocation flag, non-empty name/description, ≥10 lines of instructions → kept **207,863 files (94.6%)**
- Near-duplicates grouped into families using MinHash over word 5-grams at Jaccard threshold 0.8

### Candidate Retrieval and Pair Confirmation

- **Description similarity** computed via Qwen3-Embedding-0.6B embeddings of name + description/when_to_use text
- **Pair judge** (GPT-6 Astra) confirms whether two skills do the same job
- Calibration floor: 0.72 for cross-project (recall 0.96), 0.52 for within-project pairing
- Result: **439,860 cross-project** and **382,249 within-project** candidate pairs

### Core Function Extraction

The extractor (GPT-6 Astra) lists requirements that are:
- **Unique** (competent model wouldn't meet them without the skill)
- **Quoted verbatim** from the skill
- **Checkable** in written files, PRs, commit messages, commands, or final reply

Core functions are ranked by strength of wording (MUST/NEVER > SHOULD > plain), closeness to purpose, effect on deliverables/safety, and order. **26,198 core functions** extracted (8.4 per skill average).

### Experimental Configurations

| Configuration | A | B | Pairs | Runs | RQ |
|---|---|---|---|---|---|
| Only A | project | - | 312 | 935 | 1, 3 |
| Only B | - | project | 312 | 936 | 1 |
| A+B | project | project | 312 | 934 | 1–3 |
| Swapped | project | project, renamed | 312 | 936 | 2, 3 |
| Personal scope | project | personal | 81 | 243 | 2 |
| Plugin scope | project | plugin | 81 | 243 | 2 |
| A+C | project | C in project | 312 | 935 | 1 |
| A+B with/without guard | project | project | 119 | 670 | 3 |

### Human Validation

Two human raters validated every LLM-assisted step with Krippendorf's α ranging from **0.71 to 0.88** across all tasks (repository typing, pair judging, skill typing, core function extraction, output judging).

## Empirical Validation / Results

### RQ1: Impact of Conflicts

**Finding 3: A similar skill takes work away from the installed skill while task completion stays flat.**

| Metric | Change with B co-installed | 95% CI |
|---|---|---|
| Use of A | **−19.9 pp** | [−23.1, −16.7] |
| Task completion | +1.9 pp | [−1.0, +4.9] |
| Fidelity (primary CFs) | −2.6 pp | [−4.4, −0.8] |
| Fidelity (exclusive CFs) | **−5.6 pp** | [−7.3, −3.8] |
| Fidelity (shared CFs) | +0.1 pp | [−2.8, +2.9] |

**Control comparisons**:
- Unrelated skill (C) lowers use of A by only **5.9 pp** and leaves fidelity unchanged (−0.2 pp)
- With only B installed: fidelity is 11.0 pp lower than with only A (46.3% vs 57.3%)

**Substitution vs. interference**:
- 12.1% of runs use only B; 38.0% use neither skill
- Substituted runs carry **36%** of the fidelity loss; runs using neither skill carry **45%**; interference carries only **20%**

### RQ2: Selection and Disclosure

**Finding 4: Where a skill is installed decides which one runs, and the final reply rarely says so.**

- **Listing order**: Changes use of A by only +0.6 pp (95% CI [−2.0, +3.3])
- **Personal scope**: B overrides A in every listing; use of A drops by 35.0 pp; models still read A's files in **29.2%** of runs
- **Plugin scope**: B used in only **2.9%** of runs; use of A drops by only 6.2 pp
- **Model differences**: Opus 5 uses both skills in 36.8% of runs; Haiku 4.5 uses neither in 56.4%
- **Disclosure**: When B replaces A, final reply names the skill used in only **0.9%** of runs

### RQ3: Decision Point

**Finding 5: A conflict is decided at the first read of a skill, and a pre-tool hook at that point keeps the core functions of the installed skill.**

- Opening B first lowers fidelity on exclusive core functions by **9.4 pp** (95% CI [−15.0, −4.0], Holm-adjusted p = 0.004)
- **97%** of runs that open B first change no file before that read
- Of core functions fulfilled with only A, **37% of exclusive ones** are lost vs. **8% of shared ones**

**Guard experiment results**:
- Hook denied first read of B → model switched to A in **96%** of cases (Sonnet 4.6: 94%, Haiku 4.5: 96%, Opus 5: 99%)
- Fidelity on exclusive core functions improved by **9.1 pp** overall; **17.8 pp** when both runs of a combination reached for B first
- Task completion unchanged (+4.1 pp, [−1.3, +9.6])

### Prevalence Results

- **23.5%** of installed skills (≈10,100 families) have a similar skill in the same installation list
- **63.7%** (≈26,100 families) have a similar skill by another author
- Copied collections: **37%** of judged skills have a same-job skill in the same project (vs. 19% in other projects)
- **60%** of confirmed cross-project pairs and **56%** of within-project pairs involve normative skills

## Theoretical and Practical Implications

### For Skill Evaluation

- Benchmarks should report **what a skill uniquely contributes** (exclusive core functions) alongside task completion
- Core functions add little cost: extracted once per skill, checked by scripts where possible, frozen before any run
- Task-completion-only benchmarks would rate co-installation as "harmless" despite 19.9 pp drop in skill use

### For Agent Platforms

- **Warn at session start** when a personal skill overrides a project skill of the same name
- **Guard the first read** of a skill, where conflicts are decided—a pre-tool hook can point the model to the installed skill
- **Show which skill ran** in the final reply (currently disclosed in only 0.9% of substitutions)
- Let projects **declare which skill must win** under the current precedence rules

### For Skill Authors and Users

- **Distinctive names** avoid platform overrides; specific descriptions reduce misrouting
- Users should **audit personal directories** for overriding skills
- Projects copying whole collections should **remove duplicate skills** (37% of judged skills in such projects have same-job skills)

## Conclusion

The paper demonstrates that skill conflicts are common, harmful, and largely invisible to existing evaluation methods. The key contributions are:

1. **First empirical study** of conflicts between benign, co-installed skills that do the same job
2. **Requirement-level evaluation** measuring what an installed skill loses to a similar skill
3. **Decision point identification** at the first read of a skill, with a working guard mechanism
4. **Replication package** at https://github.com/ltroin/conflict with all cases, harness, and run logs

**Future directions** include extending to other platforms (Codex showed similar effects), studying more complex multi-skill interactions, and developing platform-level solutions for conflict resolution and disclosure.

---

_Markdown view of https://picx.dev/p/rfy1rK, served by PicX — AI-generated visual whiteboard summaries of research papers._
