Full text not available for this paper

Summary (Overview)

  • Introduces meta-skills for test-time AI-for-AI (AI4AI): reusable principles that guide a Builder model in constructing execution harnesses for a Target model, with both models' weights fixed.
  • The Builder learns meta-skills from the Target's execution feedback on development tasks through a construction–execution–reflection loop, then freezes the skill bank and applies it to construct harnesses for unseen test tasks.
  • Full-bank meta-skills improve macro-average performance by 8.95 percentage points over no-skill construction and 12.02 points over directly delivering the same skill bank to the Target.
  • Same-model experiments (Builder = Target) show average gains of 18.71 points over no-skill construction, demonstrating a path to system-level self-improvement through harness design.
  • Transfer studies suggest meta-skills can be reused across different Builders and Targets, though benefits depend on the receiving Builder's implementation.

Introduction and Theoretical Foundation

The paper addresses a complementary direction for improving AI agents: rather than only improving the agent's reasoning ability, we can improve the environment in which the agent acts. The authors draw an analogy to a PhD student whose productivity improves when an advisor establishes shared experiment logs, reproducible tools, and clear validation workflows—not by changing the student's intelligence, but by changing the support structure.

Two complementary roles:

  • Builder (B): designs and provides support (harnesses)
  • Target (T): uses that support to solve tasks

Key distinction:

  • Task skills describe how to perform a task (used by the Target)
  • Meta-skills are principles for designing the support that helps the Target perform tasks (used by the Builder)

Formal problem formulation:

J(H)=Ex∼Dtest,τ∼T(⋅∣x,Hx)[r(x,τ)],cost(τ)≤CxJ(\mathcal{H}) = \mathbb{E}_{x \sim \mathcal{D}_{\text{test}}, \tau \sim T(\cdot|x, H_x)}[r(x, \tau)], \quad \text{cost}(\tau) \leq C_x

where H\mathcal{H} is a harness policy mapping each task input xx to an environment Hx=H(x)H_x = \mathcal{H}(x), rr is the evaluator, and CxC_x is the Target execution budget.

Meta-skill definition: A meta-skill s=(when,provide,use)s = (\text{when}, \text{provide}, \text{use}) contains:

  • when: observable conditions that call for support
  • provide: the capability or resource the environment should supply
  • use: how the Target should employ that support and which judgments remain its responsibility

Methodology

Skill Learning Workflow

Starting with an empty skill bank S0S_0, the Builder iterates over development tasks:

Hxj=B(H0,x,Sj),exj=T(x;Hxj),Sj+1=ReviseB(Sj,{(Hxj,F(exj))}x∈Gj)H^j_x = B(H_0, x, S^j), \quad e^j_x = T(x; H^j_x), \quad S^{j+1} = \text{Revise}_B\left(S^j, \{(H^j_x, F(e^j_x))\}_{x \in G^j}\right)

where H0H_0 is the baseline environment, F(e)F(e) is the public execution feedback, and GjG^j is the development set at pass jj.

Skill bank update rule: At most one addition or revision per batch, which must cite supporting evidence from that batch.

Test-Time Harness Construction

After freezing the skill bank S∗=SJS^* = S^J:

Hx∗=B(H0,x,K(x,S∗)),y^x=T(x;Hx∗)H^*_x = B(H_0, x, K(x, S^*)), \quad \hat{y}_x = T(x; H^*_x)

where KK supplies either full bank or BM25 top-2 retrieval of meta-skills.

Harness Components

The framework exposes seven optional harness component families:

ComponentDescription
InstructionsSpecific task guidance
MemoryRecord format, storage/retrieval rules
ContextHistory-selection rules
Composed toolsTool definitions, call sequences
Execution controlExecution phases, tool visibility
Verification & recoverySubmission checks, failure handling
WorkspaceInitial files, setup actions

Empirical Validation / Results

Main Results (Table 3)

Test performance (%) with GPT-5.6-Sol as Builder:

MethodHarness-Bench (Gemini/Qwen/GPT-OSS)NewtonBench (Gemini/Qwen/GPT-OSS)Macro Avg.
Native environment37.35 / 69.63 / 52.4655.48 / 54.11 / 39.3851.40
Builder, no skills53.29 / 71.51 / 63.0256.16 / 50.34 / 43.8456.36
Builder meta-skills, all67.81 / 72.64 / 68.2268.84 / 63.70 / 50.6865.31
Builder meta-skills, retrieved61.50 / 74.81 / 64.7464.04 / 57.88 / 39.7360.45

Key findings:

  • Full-bank meta-skills outperform no-skill construction in all six settings (avg. +8.95 points)
  • Builder enactment beats direct delivery of the same bank by 12.02 points on average (up to 25.43 points)
  • Gains are larger on NewtonBench (+10.96 avg) than Harness-Bench (+6.95 avg), suggesting meta-skills help most when coordination is the main bottleneck
  • Full bank outperforms top-2 retrieval in five of six settings

Same-Model Self-Improvement (Figure 3a)

Across three settings with the same model as Builder and Target:

  • Average gains of 18.71 points over no-skill construction
  • Average gains of 14.14 points over skills delivered directly to the Target

Transfer Analysis (Figure 3b)

  • Cross-Builder transfer: Sol-to-Qwen bank yields +13.36 points with Sol but −4.45 points with Gemini-Pro on NewtonBench
  • Recipient-learned banks outperform imported banks (+3.69 on Harness-Bench, +9.93 on NewtonBench)
  • Cross-Builder-and-Target transfer shows positive gains (+4.58 and +5.48 points), suggesting principles may matter more than exact source-Target matching

Harness Component Ablations (Table 4)

Removing the execution controller lowers scores for all three Targets on NewtonBench:

  • Gemini: −13.36 points [CI: −18.49, −8.22]
  • Qwen: −4.79 points [CI: −10.27, 0.68]
  • GPT-OSS: −1.03 points [CI: −6.51, 4.11]

Outcome Transitions (Figure 4)

On NewtonBench, cases with no valid submission fell from 394 to 254 (−35.5%), while correct discoveries rose from 439 to 535 (+21.9%), including 145 recoveries from no-valid-submission to correct law.


Theoretical and Practical Implications

  1. Meta-skills guide resource allocation: The 12.02-point gap between Builder enactment and direct delivery shows that experience becomes useful through decisions about which burdens to externalize as state, tools, or control. Meta-skills function as a "compact language for provisioning task-specific support."

  2. Transfer connects reusable principles with adaptive implementation: Support principles can generalize across models, but implementation matters. Recipient-side learning from one's own harness executions better aligns meta-skills with implementation choices.

  3. Environment design offers a path to self-improvement: A model can improve its own test-time execution by learning to construct better harnesses while its weights remain fixed—an additional axis of optimization beyond weight updates.

  4. Refinement is not monotonic: Repeated reflection can sharpen guidance but also make it overly specific to recent evidence. Practical update rules should use development-only evidence to decide whether to continue, retain, or roll back.


Conclusion

The paper introduces meta-skills that enable a Builder to learn reusable support principles from its own harness outcomes and implement them for unseen tasks. With both models' weights fixed:

  • Full-bank meta-skills improve macro-average performance by 8.95 percentage points over no-skill construction
  • 12.02 points over direct delivery of the same bank to the Target
  • Same-model experiments demonstrate that a model can self-improve by learning to build better support for itself

Future directions:

  • More diverse Builders and tasks to establish broader applicability
  • Evaluation of performance gains relative to construction cost
  • Retrieval methods that account for skill complementarity and the Target's likely response to support
  • Coordinated improvement in both task-solving capabilities and supporting environments

"Better agents may begin with better advisors."

Related papers