Full text not available for this paper
Summary (Overview)
- Introduces meta-skills for test-time AI-for-AI (AI4AI): reusable principles that guide a Builder model in constructing execution harnesses for a Target model, with both models' weights fixed.
- The Builder learns meta-skills from the Target's execution feedback on development tasks through a construction–execution–reflection loop, then freezes the skill bank and applies it to construct harnesses for unseen test tasks.
- Full-bank meta-skills improve macro-average performance by 8.95 percentage points over no-skill construction and 12.02 points over directly delivering the same skill bank to the Target.
- Same-model experiments (Builder = Target) show average gains of 18.71 points over no-skill construction, demonstrating a path to system-level self-improvement through harness design.
- Transfer studies suggest meta-skills can be reused across different Builders and Targets, though benefits depend on the receiving Builder's implementation.
Introduction and Theoretical Foundation
The paper addresses a complementary direction for improving AI agents: rather than only improving the agent's reasoning ability, we can improve the environment in which the agent acts. The authors draw an analogy to a PhD student whose productivity improves when an advisor establishes shared experiment logs, reproducible tools, and clear validation workflows—not by changing the student's intelligence, but by changing the support structure.
Two complementary roles:
- Builder (B): designs and provides support (harnesses)
- Target (T): uses that support to solve tasks
Key distinction:
- Task skills describe how to perform a task (used by the Target)
- Meta-skills are principles for designing the support that helps the Target perform tasks (used by the Builder)
Formal problem formulation:
where is a harness policy mapping each task input to an environment , is the evaluator, and is the Target execution budget.
Meta-skill definition: A meta-skill contains:
- when: observable conditions that call for support
- provide: the capability or resource the environment should supply
- use: how the Target should employ that support and which judgments remain its responsibility
Methodology
Skill Learning Workflow
Starting with an empty skill bank , the Builder iterates over development tasks:
where is the baseline environment, is the public execution feedback, and is the development set at pass .
Skill bank update rule: At most one addition or revision per batch, which must cite supporting evidence from that batch.
Test-Time Harness Construction
After freezing the skill bank :
where supplies either full bank or BM25 top-2 retrieval of meta-skills.
Harness Components
The framework exposes seven optional harness component families:
| Component | Description |
|---|---|
| Instructions | Specific task guidance |
| Memory | Record format, storage/retrieval rules |
| Context | History-selection rules |
| Composed tools | Tool definitions, call sequences |
| Execution control | Execution phases, tool visibility |
| Verification & recovery | Submission checks, failure handling |
| Workspace | Initial files, setup actions |
Empirical Validation / Results
Main Results (Table 3)
Test performance (%) with GPT-5.6-Sol as Builder:
| Method | Harness-Bench (Gemini/Qwen/GPT-OSS) | NewtonBench (Gemini/Qwen/GPT-OSS) | Macro Avg. |
|---|---|---|---|
| Native environment | 37.35 / 69.63 / 52.46 | 55.48 / 54.11 / 39.38 | 51.40 |
| Builder, no skills | 53.29 / 71.51 / 63.02 | 56.16 / 50.34 / 43.84 | 56.36 |
| Builder meta-skills, all | 67.81 / 72.64 / 68.22 | 68.84 / 63.70 / 50.68 | 65.31 |
| Builder meta-skills, retrieved | 61.50 / 74.81 / 64.74 | 64.04 / 57.88 / 39.73 | 60.45 |
Key findings:
- Full-bank meta-skills outperform no-skill construction in all six settings (avg. +8.95 points)
- Builder enactment beats direct delivery of the same bank by 12.02 points on average (up to 25.43 points)
- Gains are larger on NewtonBench (+10.96 avg) than Harness-Bench (+6.95 avg), suggesting meta-skills help most when coordination is the main bottleneck
- Full bank outperforms top-2 retrieval in five of six settings
Same-Model Self-Improvement (Figure 3a)
Across three settings with the same model as Builder and Target:
- Average gains of 18.71 points over no-skill construction
- Average gains of 14.14 points over skills delivered directly to the Target
Transfer Analysis (Figure 3b)
- Cross-Builder transfer: Sol-to-Qwen bank yields +13.36 points with Sol but −4.45 points with Gemini-Pro on NewtonBench
- Recipient-learned banks outperform imported banks (+3.69 on Harness-Bench, +9.93 on NewtonBench)
- Cross-Builder-and-Target transfer shows positive gains (+4.58 and +5.48 points), suggesting principles may matter more than exact source-Target matching
Harness Component Ablations (Table 4)
Removing the execution controller lowers scores for all three Targets on NewtonBench:
- Gemini: −13.36 points [CI: −18.49, −8.22]
- Qwen: −4.79 points [CI: −10.27, 0.68]
- GPT-OSS: −1.03 points [CI: −6.51, 4.11]
Outcome Transitions (Figure 4)
On NewtonBench, cases with no valid submission fell from 394 to 254 (−35.5%), while correct discoveries rose from 439 to 535 (+21.9%), including 145 recoveries from no-valid-submission to correct law.
Theoretical and Practical Implications
-
Meta-skills guide resource allocation: The 12.02-point gap between Builder enactment and direct delivery shows that experience becomes useful through decisions about which burdens to externalize as state, tools, or control. Meta-skills function as a "compact language for provisioning task-specific support."
-
Transfer connects reusable principles with adaptive implementation: Support principles can generalize across models, but implementation matters. Recipient-side learning from one's own harness executions better aligns meta-skills with implementation choices.
-
Environment design offers a path to self-improvement: A model can improve its own test-time execution by learning to construct better harnesses while its weights remain fixed—an additional axis of optimization beyond weight updates.
-
Refinement is not monotonic: Repeated reflection can sharpen guidance but also make it overly specific to recent evidence. Practical update rules should use development-only evidence to decide whether to continue, retain, or roll back.
Conclusion
The paper introduces meta-skills that enable a Builder to learn reusable support principles from its own harness outcomes and implement them for unseen tasks. With both models' weights fixed:
- Full-bank meta-skills improve macro-average performance by 8.95 percentage points over no-skill construction
- 12.02 points over direct delivery of the same bank to the Target
- Same-model experiments demonstrate that a model can self-improve by learning to build better support for itself
Future directions:
- More diverse Builders and tasks to establish broader applicability
- Evaluation of performance gains relative to construction cost
- Retrieval methods that account for skill complementarity and the Target's likely response to support
- Coordinated improvement in both task-solving capabilities and supporting environments
"Better agents may begin with better advisors."
Related papers
- How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?
Under matched conditions, a minimal-harness coding agent matches or outperforms state-of-the-art MLE harnesses, with performance driven by the LLM backbone and execution environment, not scaffolding.
- Frozen Judges, Moving Agents: Version-Dependent LLM-Judge Error and the Limits of Judge-Assisted Agent Evaluation
A fixed LLM judge produces version-dependent errors, invalidating agent comparisons and transported calibration, so release decisions require paired audits, not judge-only scores.
- Audit the Scaffold, Not the Checkpoint: A Stationarity Dichotomy for Recursive Self-Improvement in Agentic Coding
Frozen weights do not guarantee safe saturation in agentic coding; auditors must monitor scaffold expansion, not just checkpoint freezing.