One Skill Too Many: How Co-Installed Skills Conflict in Coding Agents
Summary (Overview)
- First empirical study of skill conflicts in coding agents: The paper systematically investigates what happens when two similar agent skills (directories with SKILL.md files) are co-installed in coding agents like Claude Code, showing that conflicts are common and harmful.
- Scale of the problem: From 20,947 repositories, the authors estimate that nearly one in four installed skills (23.5%) is co-installed with a skill that does the same job, and 63.7% have a similar skill by another author that a user could add.
- Key finding on impact: A similar skill takes one in five runs (19.9 pp) away from the installed skill without lowering task completion, while reducing fidelity on the installed skill's exclusive core functions by 5.6 percentage points.
- Decision point discovery: Conflicts are decided at the first read of a skill (almost always before any file changes), and a pre-tool hook at that point restores fidelity to baseline levels.
- Implications: Benchmarks should score exclusive core functions rather than task completion alone, and platforms should guard the first read and disclose which skill ran.
Introduction and Theoretical Foundation
Background
Coding agents (e.g., Claude Code, Codex, Gemini CLI) are increasingly extended with agent skills—directories containing a SKILL.md file that tells the model when and how to perform a task. Skills encode conventions, safety rules, and specialized knowledge that models wouldn't know on their own. Skills are now shared at scale: recent studies collect 138,133 SKILL.md files from 20,556 repositories and 42,447 skills from two marketplaces.
The Conflict Problem
Because skills come from independent sources (teams, plugins, copied collections, personal directories), two skills that do the same job can end up co-installed without anyone choosing the pair. When this happens:
- The installed skill loses its core functions (behaviors the user installed it for)
- The similar skill may run in its place or change what the installed skill does
- Task completion often remains unchanged, so existing benchmarks (which check only pass/fail) miss these conflicts entirely
Three Key Challenges
The paper identifies three challenges that make studying skill conflicts hard:
- A conflict hides behind success: Two skills compete because they do the same job, so either can complete the task. The only trace is missing behavior specific to the installed skill, written in prose.
- A conflict is resolved out of sight: The model chooses from a few hundred characters of description without telling the user. Choice depends on listing order, scope, and the model.
- A conflict is decided at run time: Whether two skills conflict depends on a request that doesn't exist when they're installed, and similar wording doesn't imply the same job.
Key Definitions
- Core functions: Requirements that a competent model would not meet without the skill, quoted from its text, and checkable in the output
- Fidelity: The share of primary core functions (first three applicable) fulfilled in a run
- Exclusive core functions: Core functions of skill A that skill B does not also ask for
- Substitution: B is used instead of A
- Interference: A is used but the result still changes
Methodology
Study Design Overview
The study uses a controlled empirical design on Claude Code with three models (Sonnet 4.6, Haiku 4.5, Opus 5), comprising:
- 6,368 runs with 169,294 tool calls over 542 hours of agent time
- 312 experimental pairs across six cells (cross-project/within-project × normative/capability/script-bearing)
- Cost: approximately $5,550 at API prices
Installation Reconstruction
- Corpus of 55,698 distinct skills from 20,947 repositories (snapshots taken 18 July 2026)
- Reconstructed installation lists: 5,106 installation lists, 4,342 with ≥2 distinct skills
- Skill filtering: frontmatter on first line, valid YAML, no disable-model-invocation flag, non-empty name/description, ≥10 lines of instructions → kept 207,863 files (94.6%)
- Near-duplicates grouped into families using MinHash over word 5-grams at Jaccard threshold 0.8
Candidate Retrieval and Pair Confirmation
- Description similarity computed via Qwen3-Embedding-0.6B embeddings of name + description/when_to_use text
- Pair judge (GPT-6 Astra) confirms whether two skills do the same job
- Calibration floor: 0.72 for cross-project (recall 0.96), 0.52 for within-project pairing
- Result: 439,860 cross-project and 382,249 within-project candidate pairs
Core Function Extraction
The extractor (GPT-6 Astra) lists requirements that are:
- Unique (competent model wouldn't meet them without the skill)
- Quoted verbatim from the skill
- Checkable in written files, PRs, commit messages, commands, or final reply
Core functions are ranked by strength of wording (MUST/NEVER > SHOULD > plain), closeness to purpose, effect on deliverables/safety, and order. 26,198 core functions extracted (8.4 per skill average).
Experimental Configurations
| Configuration | A | B | Pairs | Runs | RQ |
|---|---|---|---|---|---|
| Only A | project | - | 312 | 935 | 1, 3 |
| Only B | - | project | 312 | 936 | 1 |
| A+B | project | project | 312 | 934 | 1–3 |
| Swapped | project | project, renamed | 312 | 936 | 2, 3 |
| Personal scope | project | personal | 81 | 243 | 2 |
| Plugin scope | project | plugin | 81 | 243 | 2 |
| A+C | project | C in project | 312 | 935 | 1 |
| A+B with/without guard | project | project | 119 | 670 | 3 |
Human Validation
Two human raters validated every LLM-assisted step with Krippendorf's α ranging from 0.71 to 0.88 across all tasks (repository typing, pair judging, skill typing, core function extraction, output judging).
Empirical Validation / Results
RQ1: Impact of Conflicts
Finding 3: A similar skill takes work away from the installed skill while task completion stays flat.
| Metric | Change with B co-installed | 95% CI |
|---|---|---|
| Use of A | −19.9 pp | [−23.1, −16.7] |
| Task completion | +1.9 pp | [−1.0, +4.9] |
| Fidelity (primary CFs) | −2.6 pp | [−4.4, −0.8] |
| Fidelity (exclusive CFs) | −5.6 pp | [−7.3, −3.8] |
| Fidelity (shared CFs) | +0.1 pp | [−2.8, +2.9] |
Control comparisons:
- Unrelated skill (C) lowers use of A by only 5.9 pp and leaves fidelity unchanged (−0.2 pp)
- With only B installed: fidelity is 11.0 pp lower than with only A (46.3% vs 57.3%)
Substitution vs. interference:
- 12.1% of runs use only B; 38.0% use neither skill
- Substituted runs carry 36% of the fidelity loss; runs using neither skill carry 45%; interference carries only 20%
RQ2: Selection and Disclosure
Finding 4: Where a skill is installed decides which one runs, and the final reply rarely says so.
- Listing order: Changes use of A by only +0.6 pp (95% CI [−2.0, +3.3])
- Personal scope: B overrides A in every listing; use of A drops by 35.0 pp; models still read A's files in 29.2% of runs
- Plugin scope: B used in only 2.9% of runs; use of A drops by only 6.2 pp
- Model differences: Opus 5 uses both skills in 36.8% of runs; Haiku 4.5 uses neither in 56.4%
- Disclosure: When B replaces A, final reply names the skill used in only 0.9% of runs
RQ3: Decision Point
Finding 5: A conflict is decided at the first read of a skill, and a pre-tool hook at that point keeps the core functions of the installed skill.
- Opening B first lowers fidelity on exclusive core functions by 9.4 pp (95% CI [−15.0, −4.0], Holm-adjusted p = 0.004)
- 97% of runs that open B first change no file before that read
- Of core functions fulfilled with only A, 37% of exclusive ones are lost vs. 8% of shared ones
Guard experiment results:
- Hook denied first read of B → model switched to A in 96% of cases (Sonnet 4.6: 94%, Haiku 4.5: 96%, Opus 5: 99%)
- Fidelity on exclusive core functions improved by 9.1 pp overall; 17.8 pp when both runs of a combination reached for B first
- Task completion unchanged (+4.1 pp, [−1.3, +9.6])
Prevalence Results
- 23.5% of installed skills (≈10,100 families) have a similar skill in the same installation list
- 63.7% (≈26,100 families) have a similar skill by another author
- Copied collections: 37% of judged skills have a same-job skill in the same project (vs. 19% in other projects)
- 60% of confirmed cross-project pairs and 56% of within-project pairs involve normative skills
Theoretical and Practical Implications
For Skill Evaluation
- Benchmarks should report what a skill uniquely contributes (exclusive core functions) alongside task completion
- Core functions add little cost: extracted once per skill, checked by scripts where possible, frozen before any run
- Task-completion-only benchmarks would rate co-installation as "harmless" despite 19.9 pp drop in skill use
For Agent Platforms
- Warn at session start when a personal skill overrides a project skill of the same name
- Guard the first read of a skill, where conflicts are decided—a pre-tool hook can point the model to the installed skill
- Show which skill ran in the final reply (currently disclosed in only 0.9% of substitutions)
- Let projects declare which skill must win under the current precedence rules
For Skill Authors and Users
- Distinctive names avoid platform overrides; specific descriptions reduce misrouting
- Users should audit personal directories for overriding skills
- Projects copying whole collections should remove duplicate skills (37% of judged skills in such projects have same-job skills)
Conclusion
The paper demonstrates that skill conflicts are common, harmful, and largely invisible to existing evaluation methods. The key contributions are:
- First empirical study of conflicts between benign, co-installed skills that do the same job
- Requirement-level evaluation measuring what an installed skill loses to a similar skill
- Decision point identification at the first read of a skill, with a working guard mechanism
- Replication package at https://github.com/ltroin/conflict with all cases, harness, and run logs
Future directions include extending to other platforms (Codex showed similar effects), studying more complex multi-skill interactions, and developing platform-level solutions for conflict resolution and disclosure.
Related papers
- LoLBench: Evaluating Coding Agents with Long-Horizon Proposals on Large Software Systems
LOLBENCH shows top coding agents resolve only 14% of long-horizon modular development tasks, with missing cross-module context as the dominant failure bottleneck.
- Do Tool Calls Execute as Intended? Measuring and Repairing Intent-Execution Correspondence in LLM Agents
Tool calls are silently altered by execution-path hops in 12% of production shell invocations, and the IntAct repair protocol recovers 79.2% of resulting failures.
- HarnessSecurity-Bench: Do Security Mechanisms Really Protect Coding Agent Harnesses?
Benchmarking 40 coding agent harnesses shows auto-approve raises attack success from 29% to 96%, while command allowlisting cuts attacks with minimal utility loss.