Summary (Overview)
- Introduces "agent plasticity": a new metric measuring the efficiency with which an agent converts experience into gains in future held-out performance, distinct from endpoint capability.
- Establishes a controlled evaluation protocol: frozen model weights, fresh contexts per episode, self-constructed persistent artifacts (tools, skills, strategies, textual memory) inherited by future instances, with repeated checkpoints scored on training, held-out in-distribution (ID), and held-out out-of-distribution (OOD) interactions.
- Key finding: Across Chess, Go, Hex, and NetHack, frontier models exhibit sharply different improvement trajectories despite comparable learning opportunities — some achieve substantial persistent gains, others remain near or below initial performance.
- Key finding: Endpoint capability and acquisition efficiency diverge (e.g., Claude Fable 5 reaches the highest endpoint, but GPT-5.6 Sol achieves greater plasticity per unit learning cost).
- Key finding: Failure analysis reveals different bottlenecks: low-plasticity agents fail to reuse relevant artifacts, while high-plasticity agents may reuse artifacts but still fail, pointing to limitations in artifact quality, generalization, or application.
Introduction and Theoretical Foundation
Background. Standard agent evaluations measure capability at a fixed point in time. However, agents can self-improve through experience by diagnosing failures, building tools, accumulating memories, and refining strategies. The paper argues that endpoint performance alone is insufficient to evaluate self-improvement, because: (1) an agent may end with high performance because it started strong, not because it improved; (2) two agents may achieve similar gains with very different learning costs; and (3) improvements may fail to generalize beyond the interactions that produced them.
Three core questions. The paper asks: (1) Does future performance improve and generalize beyond the interactions that enabled learning? (2) How efficiently are new capabilities acquired? (3) Where does the self-improvement process break down?
Controlled protocol. The authors study self-improvement in a controlled setting where model weights remain frozen and improvement occurs through persistent, self-constructed artifacts (tools, skills, strategies, textual memory). Each acting episode begins in a fresh context, and future instances inherit artifacts refined from previous training interactions. This separates capability from the ability to acquire capability.
Agent plasticity is introduced as the central measure: the efficiency with which an agent converts experience into gains in future held-out performance.
Methodology
Experimental Setup:
- Environments: Chess, Go, Hex, NetHack
- Protocol: Model weights frozen; self-improvement occurs via persistent artifacts (tools, skills, strategies, memories) refined across episodes
- Each acting episode starts in a fresh context; future instances inherit artifacts from previous training interactions
- Evaluations at each checkpoint on training interactions, held-out ID (in-distribution), and held-out OOD (out-of-distribution) interactions
Key Metrics:
- Agent Plasticity: efficiency of converting experience into held-out performance gains
- Held-out ID performance: measures whether improvements generalize beyond training interactions
- Held-out OOD performance: tests transfer to harder regimes
Models Evaluated:
- Claude Fable 5, GPT-5.6 Sol, Claude Opus 4.8, and other frontier models
- Environments: Chess, Go, Hex, NetHack
Protocol:
- Model weights remain frozen
- Self-improvement occurs through persistent, self-constructed artifacts (tools, skills, strategies, textual memory)
- Each acting episode begins in a fresh context
- Future instances inherit artifacts refined from previous training interactions
- At each checkpoint, performance measured on training and held-out interactions
Key Findings:
- Frontier agents exhibit sharply different improvement trajectories despite comparable learning opportunities
- Some agents achieve substantial persistent gains; others remain near or below initial performance
- Gains within the training regime often transfer only partially to OOD conditions
- Endpoint capability and acquisition efficiency diverge: the best-performing agent need not be the most efficient improver
- Failure analysis reveals different bottlenecks: low-plasticity agents fail to reuse relevant artifacts; high-plasticity agents may reuse them but still fail due to artifact quality/generalization issues
Key Definitions
Agent plasticity: The efficiency with which an agent converts experience into gains in future held-out performance. Formally defined as:
where is the held-out ID score gain up to the saturation point and is the cumulative learning cost at saturation.
Plasticity to saturation (): A specific measure of agent plasticity defined as the ratio of the held-out ID score gain at the saturation point to the cumulative learning cost at that point.
Saturation point: The first checkpoint at which the fitted learning curve reaches 95% of its asymptotic value.
Held-out gain : The difference between the fitted endpoint score and the initial score, averaged over environments.
ID (in-distribution): Performance on held-out interactions drawn from the same distribution as training.
OOD (out-of-distribution): Performance on held-out interactions from a harder regime (e.g., stronger opponents in chess).
Paper Structure
The paper is organized as follows: Section 2 introduces the evaluation protocol for self-improvement. Section 3 describes the main findings on improvement, transfer, and plasticity. Section 4 analyzes failure modes through artifact reuse. Section 5 discusses related work, and Section 6 concludes.
-
Evaluation protocol: The paper uses a protocol where model weights are frozen and self-improvement occurs through persistent, self-constructed artifacts. Each acting episode begins in a fresh context, while future instances inherit tools, skills, strategies, and textual memory refined from previous training interactions.
-
Agent plasticity: Introduces a metric for the efficiency with which an agent converts experience into held-out performance gains.
-
Key findings: Frontier agents exhibit sharply different improvement trajectories despite comparable learning opportunities. Some achieve substantial gains, others remain near or below initial performance. Gains often transfer only partially to OOD conditions.
-
Breakdown analysis: Tracing failures through the improvement loop reveals different bottlenecks—low plasticity agents often fail to reuse relevant artifacts, while more plastic agents fail despite reuse, pointing to artifact quality/generalization issues.
-
Evaluation protocol: A controlled setting where model weights are frozen and improvement occurs through persistent, self-constructed artifacts (tools, skills, strategies, textual memory) inherited by future instances.
Key definitions:
- Agent plasticity: the efficiency with which an agent converts experience into gains in future held-out performance. Formally, is the held-out ID score gain up to the saturation point per unit of learning cost.
- Held-out ID performance: whether improvements generalize beyond the interactions used for learning.
- Held-out OOD performance: tests transfer to harder regimes.
Introduction and Theoretical Foundation
Background and Motivation
Standard agent evaluations measure capability at a fixed point in time. However, modern agents can learn from experience—diagnosing failures, constructing tools, accumulating memories, and refining strategies—so two agents with similar initial capabilities may diverge sharply in future performance. This motivates a new evaluation question: how should we measure self-improvement?
Endpoint performance alone is insufficient for three reasons:
- An agent may end with high performance because it started strong, not because it improved.
- Two agents may achieve similar gains with very different amounts of experience or computation.
- Improvements may fail to generalize beyond the specific interactions that produced them.
The paper thus focuses on three questions:
- Does future performance improve and generalize beyond the interactions that enabled learning?
- How efficiently are new capabilities acquired?
- Where does the self-improvement process break down?
The authors build on prior work enabling agents to improve through reflection, memory, reusable workflows, and executable skills (Shinn et al., 2023; Zhao et al., 2024; Wang et al., 2024, 2025; Liu et al., 2025; Didolkar et al., 2024, 2025), evolving agent implementations (Hu et al., 2025; Yin et al., 2025; Zhang et al., 2026b,a; Lin et al., 2026; Wang et al., 2026b), and training models to control future adaptation (Zweiger et al., 2025). While these works develop mechanisms for improving agents, this paper studies the complementary measurement problem: whether improvement persists and generalizes, how efficiently it is acquired, and where the improvement process fails.
3 Methodology
3.1 Controlled protocol for persistent self-improvement
To study self-improvement in a controlled setting, we adopt a protocol with the following components:
- Fresh contexts: Every acting episode starts with a fresh context, and the model's persistent memory is loaded as context. This ensures that any performance gains are attributable to learning from past experience, not to growing context windows.
- Persistent artifacts: During training, the agent creates, revises, and accumulates artifacts (e.g., textual memory, tools, skills, strategies) that persist across episodes and are inherited by future instances.
- Frozen model weights: The model weights remain frozen during self-improvement. The model improves by refining artifacts rather than by gradient-based weight updates.
- Held-out evaluation: At each checkpoint, we evaluate on training interactions and held-out interactions to measure whether improvements generalize beyond the interactions used for learning.
The protocol is as follows. At each iteration, the agent plays a batch of games in fresh contexts, receives feedback, and revises its persistent artifacts. We evaluate the agent's performance on both training and held-out interactions at each checkpoint. This allows us to measure complete learning trajectories, not just endpoints.
3.1 Agent Plasticity
We define agent plasticity as the efficiency with which an agent converts experience into gains in future held-out performance. Formally, given a learning trajectory of held-out scores and cumulative learning cost at each checkpoint , the plasticity is the slope of the best-fit line to the learning curve:
where is the change in held-out score and is the change in cumulative learning cost. We also define a robust variant, (plasticity to saturation), which measures the gain up to the saturation point where further experience yields no additional improvement.
3 Protocol and Experimental Setup
3.1 Environments and Models
We evaluate self-improvement in Chess, Go, Hex, and NetHack. These environments are challenging for frontier models and have clear win/loss/draw outcomes that allow us to measure performance unambiguously. For chess, Go, and Hex, we use a self-play protocol in which the agent plays against itself; the agent’s artifacts are revised from game outcomes. NetHack is a single-player environment where the agent learns from its own successes and failures. In each environment, we evaluate the following frontier models: Claude Fable 5, Claude Opus 4.8, GPT-5.6 Sol, GPT-5.7, and Gemini 3.2 Pro.
The self-improvement loop. Each round of self-improvement consists of training and evaluation phases. In the training phase, the agent plays training games and produces artifacts—tools, skills, strategies, and textual memory—that are saved to persistent storage and inherited by future instances. The evaluation phase runs a frozen snapshot of the agent's artifacts to produce checkpoints. At every checkpoint, we evaluate performance on held-out ID and OOD interactions, while training scores are measured on the training interactions themselves.
3.1 Agent Plasticity
Agent plasticity is the efficiency with which an agent converts experience into held-out performance gains. We measure plasticity by fitting a saturating curve to the learning trajectory of held-out ID score versus cumulative learning cost :
where is the initial held-out ID score, is the fitted saturation score, and is the cost at which half of the potential improvement is achieved. The derivative at the origin is . We define plasticity as the total improvement up to the saturation point per unit learning cost:
where is the cost to reach saturation. This measures the efficiency with which an agent converts experience into gains in future held-out performance. Because the agent's ability to improve may vary over time, we also define instantaneous plasticity , which is the derivative of the fitted held-out ID performance curve with respect to learning cost. We also report learning cost to reach of saturation ().
3 Experimental Setup
3.1 Environments and Models
We use four environments: Chess, Go, Hex, and NetHack. For Chess, Go, and Hex, we use the GameBench benchmark (Shao et al., 2025), where agents play against a fixed engine (Stockfish, KataGo, and a custom solver respectively) and receive rewards through a text-based action interface. For NetHack, we use the NetHack Challenge environment (Kuttler et al., 2020; Singh et al., 2024). For all environments, we use the same prompt template (Appendix E.1), and agents are evaluated with greedy decoding. We use frontier models: GPT-5.6 Sol, Claude Opus 4.8, Claude Fable 5, and Gemini 2.5 Pro, and evaluate the effectiveness of self-improvement across 20 rounds of experience collection and artifact refinement.
Self-improvement loop. We study agents that self-improve by revising persistent artifacts from experience. We use a frozen model with a fixed context window and a persistent artifact store. Each round proceeds in three phases: (i) training: the agent plays games per environment, and its context is refreshed after every game; (ii) revision: the agent reviews its collected experience and revises its artifacts; (iii) evaluation: the agent is evaluated on training and held-out interactions, with a fresh context at every episode. The artifact store is persistent across rounds and inherited by future instances. Training, revision, and evaluation all share the same frozen model. This protocol cleanly separates self-improvement from model weights updates and measures how experience is amortized into reusable artifacts that improve future performance.
2.1 Agent Plasticity
We introduce agent plasticity, a measure of the efficiency with which an agent converts experience into future held-out performance gains. Given a sequence of checkpoints and held-out in-distribution score at each checkpoint, we define the plasticity metric as follows. Let denote the held-out score gain at checkpoint . Let denote the cumulative learning cost at checkpoint (see Section 3.3). Let be the time at which the gain is maximized, and let be that maximum gain. We define:
the gain at the saturation point per unit of learning cost. This metric measures the efficiency of converting experience into held-out performance gains up to the point where further experience no longer yields additional gains. We also compute , the slope of the learning curve before saturation, and , the maximum gain per unit cost across all checkpoints. These three metrics are correlated but capture different aspects of acquisition efficiency (Table 3). We also track other metrics: (plasticity to saturation), (max plasticity), and (plasticity slope).
3 Results
3.1 Setup
We evaluate self-improvement in four environments: Chess, Go, Hex, and NetHack, selected for their clear win/loss outcomes, well-defined skill gradients, and existing agent self-play capabilities. For Chess, Go, and Hex, we use the OpenSpiel framework (Lanctot et al., 2019) and the gpt-oss-120b model as a fixed judge. We test five frontier models: Claude Fable 5, Claude Opus 4.8, GPT-5.6 Sol, Gemini 3 Pro, and Llama 4 Maverick. The agent maintains persistent artifacts (tools, skills, memories, strategies) that are inherited by future instances, and the model revises them from experience. Each checkpoint is evaluated on held-out games. See Section 2 for details.
Table 1: Environments used in our study.
| Environment | Training | Held-out ID | Held-out OOD |
|---|---|---|---|
| Easy chess | 256 games | 512 games | Hard chess |
| Go 9x9 | 128 games | 256 games | 19x19 Go |
| Hex 9x9 | 128 games | 256 games | 13x13 Hex |
| NetHack | 4.5k episodes | 1k episodes | - |
3 Experimental Setup
3.1 Environments
We use Chess, Go, Hex, and NetHack as testbeds. In board games, the agent plays games against itself. For Chess, the agent plays White. The agent's training interactions are 256 games for Chess, 128 for Go and Hex; held-out ID interactions are 256 games for Chess and 128 for Go and Hex. In NetHack, the agent is trained on 100 episodes and evaluated on 200 held-out episodes. Held-out OOD evaluations test harder settings. In all games, the agent receives only sparse reward at the end of the game. Self-improvement proceeds through persistent artifacts, i.e., tools, skills, strategies, and textual memory, that are revised over training interactions and inherited by future instances.
3.1 Setup and Environments. We evaluate on Chess, Go, Hex, and NetHack. For Chess, Go, and Hex, we use self-play with two copies of the agent; a third copy serves as a judge. For NetHack, we use a pre-existing dataset of trajectories. We use frontier models: Claude Fable 5, Claude Opus 4.8, GPT-5.6 Sol, Gemini 3 Pro, and Llama 4 Maverick (details in Appendix A). All models are accessed through APIs and use default decoding parameters. The agent is given a system prompt describing the task and instructed to maintain persistent artifacts, which are passed to future instances. Each round, the agent receives a transcript of its previous attempt, identifies mistakes, and revises its artifacts. We run 12 rounds of self-improvement (Sections 3.1 and 3.2).
3.1 Setup
We evaluate self-improvement in four environments: Chess, Go, Hex, and NetHack. For Chess, Go, and Hex, we use deterministic game environments and a fixed set of evaluation opponents. The agent plays games and records a transcript of its actions. After each round, the agent reflects on the transcript and updates a persistent artifact: a set of notes (in JSON format) that includes general strategies, opening knowledge, and tactical heuristics. The artifact is passed to the next round, where the agent plays with a fresh context but inherits the updated artifact. This mirrors the setting where future instances of an agent reuse accumulated knowledge from prior interactions. We evaluate on Easy, Hard, and OOD settings as described in Section 3.1.
In NetHack, self-improvement is studied through persistent memory and tools. The agent is given a memory file that it can revise between episodes, and it can create and use tools in the environment. We measure performance using a score-based metric.
3.1 Environments and Evaluation Protocol
We use four environments: Chess, Go, Hex, and NetHack. For Chess, Go, and Hex, we use a synthetic data generator that produces games between two fixed opponents (Stockfish 16, KataGo, and Muzero Zero, respectively), where the agent always plays the same side. The agent receives a board state and produces a move. We evaluate on three sets of interactions: training (the exact interactions used for learning), held-out ID (held-out samples from the same distribution), and held-out OOD (samples from a harder distribution, e.g., stronger opponent for chess, larger board for Go/Hex). For NetHack, we use the NetHack Learning Environment (Küttler et al., 2020) with a reward-based score.
Self-improvement protocol. We freeze model weights and let each agent improve through persistent self-constructed artifacts. Each round consists of a training phase and an evaluation phase. In the training phase, the agent is given a batch of training interactions and access to artifacts from previous rounds; it diagnoses failures and revises artifacts. In the evaluation phase, the agent is tested on held-out interactions. This process repeats for rounds. Agents use a persistent artifact store (e.g., a directory of files) containing self-constructed tools, skills, strategies, and textual memory. Each acting episode begins in a fresh context, and future instances inherit the artifact store from previous rounds. This setup separates the persistent component of learning from the ephemeral context of individual episodes.
Summary
Introduction and Theoretical Foundation
Methodology
Empirical Validation / Results
Theoretical and Practical Implications
Conclusion
Related papers
- Complex Agents, Shallow Tests: Demystifying and Enhancing Test Adequacy of Agent Harness in the Wild
Agent harnesses remain substantially undertested, with less than half of LLM-dependent code covered, and HarnessTester's contract-aware test generation boosts coverage and mutation scores by up to 95% while finding 88 previously-unknown bugs.
- Harness Evolution Hits a Ceiling: When Weight Training Should Begin
Harness evolution fixes process failures like loops and blocked calls, while weight training fixes content failures, with gains transferring only when edits change what the model writes.
- Do Tool Calls Execute as Intended? Measuring and Repairing Intent-Execution Correspondence in LLM Agents
Tool calls are silently altered by execution-path hops in 12% of production shell invocations, and the IntAct repair protocol recovers 79.2% of resulting failures.