# Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI

> Learning reusable meta-skills for environment design improves AI test-time performance by 8.95 points over no-skill construction, enabling fixed-weight self-improvement.

- **Source:** [arXiv](https://arxiv.org/abs/2609.38143)
- **Published:** 2026-10-03
- **Permalink:** https://picx.dev/p/C6l02V
- **Whiteboard:** https://picx.dev/p/C6l02V/image

## Summary

## Summary (Overview)

- Introduces **meta-skills** for test-time AI-for-AI (AI4AI): reusable principles that guide a Builder model in constructing execution harnesses for a Target model, with both models' weights fixed.
- The Builder learns meta-skills from the Target's execution feedback on development tasks through a **construction–execution–reflection loop**, then freezes the skill bank and applies it to construct harnesses for unseen test tasks.
- Full-bank meta-skills improve macro-average performance by **8.95 percentage points** over no-skill construction and **12.02 points** over directly delivering the same skill bank to the Target.
- Same-model experiments (Builder = Target) show average gains of **18.71 points** over no-skill construction, demonstrating a path to **system-level self-improvement** through harness design.
- Transfer studies suggest meta-skills can be reused across different Builders and Targets, though benefits depend on the receiving Builder's implementation.

---

## Introduction and Theoretical Foundation

The paper addresses a complementary direction for improving AI agents: rather than only improving the agent's reasoning ability, we can improve the **environment** in which the agent acts. The authors draw an analogy to a PhD student whose productivity improves when an advisor establishes shared experiment logs, reproducible tools, and clear validation workflows—not by changing the student's intelligence, but by changing the support structure.

**Two complementary roles:**
- **Builder (B):** designs and provides support (harnesses)
- **Target (T):** uses that support to solve tasks

**Key distinction:**
- **Task skills** describe how to perform a task (used by the Target)
- **Meta-skills** are principles for designing the support that helps the Target perform tasks (used by the Builder)

**Formal problem formulation:**

$$
J(\mathcal{H}) = \mathbb{E}_{x \sim \mathcal{D}_{\text{test}}, \tau \sim T(\cdot|x, H_x)}[r(x, \tau)], \quad \text{cost}(\tau) \leq C_x
$$

where $\mathcal{H}$ is a harness policy mapping each task input $x$ to an environment $H_x = \mathcal{H}(x)$, $r$ is the evaluator, and $C_x$ is the Target execution budget.

**Meta-skill definition:** A meta-skill $s = (\text{when}, \text{provide}, \text{use})$ contains:
- **when:** observable conditions that call for support
- **provide:** the capability or resource the environment should supply
- **use:** how the Target should employ that support and which judgments remain its responsibility

---

## Methodology

### Skill Learning Workflow

Starting with an empty skill bank $S_0$, the Builder iterates over development tasks:

$$
H^j_x = B(H_0, x, S^j), \quad e^j_x = T(x; H^j_x), \quad S^{j+1} = \text{Revise}_B\left(S^j, \{(H^j_x, F(e^j_x))\}_{x \in G^j}\right)
$$

where $H_0$ is the baseline environment, $F(e)$ is the public execution feedback, and $G^j$ is the development set at pass $j$.

**Skill bank update rule:** At most one addition or revision per batch, which must cite supporting evidence from that batch.

### Test-Time Harness Construction

After freezing the skill bank $S^* = S^J$:

$$
H^*_x = B(H_0, x, K(x, S^*)), \quad \hat{y}_x = T(x; H^*_x)
$$

where $K$ supplies either **full bank** or **BM25 top-2 retrieval** of meta-skills.

### Harness Components

The framework exposes **seven optional harness component families**:

| Component | Description |
|-----------|-------------|
| Instructions | Specific task guidance |
| Memory | Record format, storage/retrieval rules |
| Context | History-selection rules |
| Composed tools | Tool definitions, call sequences |
| Execution control | Execution phases, tool visibility |
| Verification & recovery | Submission checks, failure handling |
| Workspace | Initial files, setup actions |

---

## Empirical Validation / Results

### Main Results (Table 3)

Test performance (%) with GPT-5.6-Sol as Builder:

| Method | Harness-Bench (Gemini/Qwen/GPT-OSS) | NewtonBench (Gemini/Qwen/GPT-OSS) | Macro Avg. |
|--------|--------------------------------------|-------------------------------------|------------|
| Native environment | 37.35 / 69.63 / 52.46 | 55.48 / 54.11 / 39.38 | 51.40 |
| Builder, no skills | 53.29 / 71.51 / 63.02 | 56.16 / 50.34 / 43.84 | 56.36 |
| **Builder meta-skills, all** | **67.81 / 72.64 / 68.22** | **68.84 / 63.70 / 50.68** | **65.31** |
| Builder meta-skills, retrieved | 61.50 / 74.81 / 64.74 | 64.04 / 57.88 / 39.73 | 60.45 |

**Key findings:**
- Full-bank meta-skills outperform no-skill construction in **all six settings** (avg. +8.95 points)
- Builder enactment beats direct delivery of the same bank by **12.02 points on average** (up to 25.43 points)
- Gains are larger on NewtonBench (+10.96 avg) than Harness-Bench (+6.95 avg), suggesting meta-skills help most when coordination is the main bottleneck
- Full bank outperforms top-2 retrieval in five of six settings

### Same-Model Self-Improvement (Figure 3a)

Across three settings with the same model as Builder and Target:
- Average gains of **18.71 points** over no-skill construction
- Average gains of **14.14 points** over skills delivered directly to the Target

### Transfer Analysis (Figure 3b)

- Cross-Builder transfer: Sol-to-Qwen bank yields +13.36 points with Sol but −4.45 points with Gemini-Pro on NewtonBench
- Recipient-learned banks outperform imported banks (+3.69 on Harness-Bench, +9.93 on NewtonBench)
- Cross-Builder-and-Target transfer shows positive gains (+4.58 and +5.48 points), suggesting principles may matter more than exact source-Target matching

### Harness Component Ablations (Table 4)

Removing the **execution controller** lowers scores for all three Targets on NewtonBench:
- Gemini: −13.36 points [CI: −18.49, −8.22]
- Qwen: −4.79 points [CI: −10.27, 0.68]
- GPT-OSS: −1.03 points [CI: −6.51, 4.11]

### Outcome Transitions (Figure 4)

On NewtonBench, cases with no valid submission fell from 394 to 254 (−35.5%), while correct discoveries rose from 439 to 535 (+21.9%), including 145 recoveries from no-valid-submission to correct law.

---

## Theoretical and Practical Implications

1. **Meta-skills guide resource allocation:** The 12.02-point gap between Builder enactment and direct delivery shows that experience becomes useful through decisions about which burdens to externalize as state, tools, or control. Meta-skills function as a "compact language for provisioning task-specific support."

2. **Transfer connects reusable principles with adaptive implementation:** Support principles can generalize across models, but implementation matters. Recipient-side learning from one's own harness executions better aligns meta-skills with implementation choices.

3. **Environment design offers a path to self-improvement:** A model can improve its own test-time execution by learning to construct better harnesses while its weights remain fixed—an additional axis of optimization beyond weight updates.

4. **Refinement is not monotonic:** Repeated reflection can sharpen guidance but also make it overly specific to recent evidence. Practical update rules should use development-only evidence to decide whether to continue, retain, or roll back.

---

## Conclusion

The paper introduces **meta-skills** that enable a Builder to learn reusable support principles from its own harness outcomes and implement them for unseen tasks. With both models' weights fixed:

- Full-bank meta-skills improve macro-average performance by **8.95 percentage points** over no-skill construction
- **12.02 points** over direct delivery of the same bank to the Target
- Same-model experiments demonstrate that a model can **self-improve** by learning to build better support for itself

**Future directions:**
- More diverse Builders and tasks to establish broader applicability
- Evaluation of performance gains relative to construction cost
- Retrieval methods that account for skill complementarity and the Target's likely response to support
- Coordinated improvement in both task-solving capabilities and supporting environments

> "Better agents may begin with better advisors."

---

_Markdown view of https://picx.dev/p/C6l02V, served by PicX — AI-generated visual whiteboard summaries of research papers._
