Summary (Overview)

  • Introduces "agent plasticity": a new metric measuring the efficiency with which an agent converts experience into gains in future held-out performance, distinct from endpoint capability.
  • Establishes a controlled evaluation protocol: frozen model weights, fresh contexts per episode, self-constructed persistent artifacts (tools, skills, strategies, textual memory) inherited by future instances, with repeated checkpoints scored on training, held-out in-distribution (ID), and held-out out-of-distribution (OOD) interactions.
  • Key finding: Across Chess, Go, Hex, and NetHack, frontier models exhibit sharply different improvement trajectories despite comparable learning opportunities — some achieve substantial persistent gains, others remain near or below initial performance.
  • Key finding: Endpoint capability and acquisition efficiency diverge (e.g., Claude Fable 5 reaches the highest endpoint, but GPT-5.6 Sol achieves greater plasticity per unit learning cost).
  • Key finding: Failure analysis reveals different bottlenecks: low-plasticity agents fail to reuse relevant artifacts, while high-plasticity agents may reuse artifacts but still fail, pointing to limitations in artifact quality, generalization, or application.

Introduction and Theoretical Foundation

Background. Standard agent evaluations measure capability at a fixed point in time. However, agents can self-improve through experience by diagnosing failures, building tools, accumulating memories, and refining strategies. The paper argues that endpoint performance alone is insufficient to evaluate self-improvement, because: (1) an agent may end with high performance because it started strong, not because it improved; (2) two agents may achieve similar gains with very different learning costs; and (3) improvements may fail to generalize beyond the interactions that produced them.

Three core questions. The paper asks: (1) Does future performance improve and generalize beyond the interactions that enabled learning? (2) How efficiently are new capabilities acquired? (3) Where does the self-improvement process break down?

Controlled protocol. The authors study self-improvement in a controlled setting where model weights remain frozen and improvement occurs through persistent, self-constructed artifacts (tools, skills, strategies, textual memory). Each acting episode begins in a fresh context, and future instances inherit artifacts refined from previous training interactions. This separates capability from the ability to acquire capability.

Agent plasticity is introduced as the central measure: the efficiency with which an agent converts experience into gains in future held-out performance.

Methodology

Experimental Setup:

  • Environments: Chess, Go, Hex, NetHack
  • Protocol: Model weights frozen; self-improvement occurs via persistent artifacts (tools, skills, strategies, memories) refined across episodes
  • Each acting episode starts in a fresh context; future instances inherit artifacts from previous training interactions
  • Evaluations at each checkpoint on training interactions, held-out ID (in-distribution), and held-out OOD (out-of-distribution) interactions

Key Metrics:

  • Agent Plasticity: efficiency of converting experience into held-out performance gains
  • Held-out ID performance: measures whether improvements generalize beyond training interactions
  • Held-out OOD performance: tests transfer to harder regimes

Models Evaluated:

  • Claude Fable 5, GPT-5.6 Sol, Claude Opus 4.8, and other frontier models
  • Environments: Chess, Go, Hex, NetHack

Protocol:

  • Model weights remain frozen
  • Self-improvement occurs through persistent, self-constructed artifacts (tools, skills, strategies, textual memory)
  • Each acting episode begins in a fresh context
  • Future instances inherit artifacts refined from previous training interactions
  • At each checkpoint, performance measured on training and held-out interactions

Key Findings:

  • Frontier agents exhibit sharply different improvement trajectories despite comparable learning opportunities
  • Some agents achieve substantial persistent gains; others remain near or below initial performance
  • Gains within the training regime often transfer only partially to OOD conditions
  • Endpoint capability and acquisition efficiency diverge: the best-performing agent need not be the most efficient improver
  • Failure analysis reveals different bottlenecks: low-plasticity agents fail to reuse relevant artifacts; high-plasticity agents may reuse them but still fail due to artifact quality/generalization issues

Key Definitions

Agent plasticity: The efficiency with which an agent converts experience into gains in future held-out performance. Formally defined as:

Psat=ΔSIDsatCsatP^{sat} = \frac{\Delta S_{ID}^{sat}}{C_{sat}}

where ΔSIDsat\Delta S_{ID}^{sat} is the held-out ID score gain up to the saturation point and CsatC_{sat} is the cumulative learning cost at saturation.

Plasticity to saturation (PsatP^{sat}): A specific measure of agent plasticity defined as the ratio of the held-out ID score gain at the saturation point to the cumulative learning cost at that point.

Saturation point: The first checkpoint at which the fitted learning curve reaches 95% of its asymptotic value.

Held-out gain ΔS\Delta S: The difference between the fitted endpoint score and the initial score, averaged over environments.

ID (in-distribution): Performance on held-out interactions drawn from the same distribution as training.

OOD (out-of-distribution): Performance on held-out interactions from a harder regime (e.g., stronger opponents in chess).

Paper Structure

The paper is organized as follows: Section 2 introduces the evaluation protocol for self-improvement. Section 3 describes the main findings on improvement, transfer, and plasticity. Section 4 analyzes failure modes through artifact reuse. Section 5 discusses related work, and Section 6 concludes.

  1. Evaluation protocol: The paper uses a protocol where model weights are frozen and self-improvement occurs through persistent, self-constructed artifacts. Each acting episode begins in a fresh context, while future instances inherit tools, skills, strategies, and textual memory refined from previous training interactions.

  2. Agent plasticity: Introduces a metric for the efficiency with which an agent converts experience into held-out performance gains.

  3. Key findings: Frontier agents exhibit sharply different improvement trajectories despite comparable learning opportunities. Some achieve substantial gains, others remain near or below initial performance. Gains often transfer only partially to OOD conditions.

  4. Breakdown analysis: Tracing failures through the improvement loop reveals different bottlenecks—low plasticity agents often fail to reuse relevant artifacts, while more plastic agents fail despite reuse, pointing to artifact quality/generalization issues.

  5. Evaluation protocol: A controlled setting where model weights are frozen and improvement occurs through persistent, self-constructed artifacts (tools, skills, strategies, textual memory) inherited by future instances.

Key definitions:

  • Agent plasticity: the efficiency with which an agent converts experience into gains in future held-out performance. Formally, PsatP^{sat} is the held-out ID score gain up to the saturation point per unit of learning cost.
  • Held-out ID performance: whether improvements generalize beyond the interactions used for learning.
  • Held-out OOD performance: tests transfer to harder regimes.

Introduction and Theoretical Foundation

Background and Motivation

Standard agent evaluations measure capability at a fixed point in time. However, modern agents can learn from experience—diagnosing failures, constructing tools, accumulating memories, and refining strategies—so two agents with similar initial capabilities may diverge sharply in future performance. This motivates a new evaluation question: how should we measure self-improvement?

Endpoint performance alone is insufficient for three reasons:

  1. An agent may end with high performance because it started strong, not because it improved.
  2. Two agents may achieve similar gains with very different amounts of experience or computation.
  3. Improvements may fail to generalize beyond the specific interactions that produced them.

The paper thus focuses on three questions:

  1. Does future performance improve and generalize beyond the interactions that enabled learning?
  2. How efficiently are new capabilities acquired?
  3. Where does the self-improvement process break down?

The authors build on prior work enabling agents to improve through reflection, memory, reusable workflows, and executable skills (Shinn et al., 2023; Zhao et al., 2024; Wang et al., 2024, 2025; Liu et al., 2025; Didolkar et al., 2024, 2025), evolving agent implementations (Hu et al., 2025; Yin et al., 2025; Zhang et al., 2026b,a; Lin et al., 2026; Wang et al., 2026b), and training models to control future adaptation (Zweiger et al., 2025). While these works develop mechanisms for improving agents, this paper studies the complementary measurement problem: whether improvement persists and generalizes, how efficiently it is acquired, and where the improvement process fails.

3 Methodology

3.1 Controlled protocol for persistent self-improvement

To study self-improvement in a controlled setting, we adopt a protocol with the following components:

  • Fresh contexts: Every acting episode starts with a fresh context, and the model's persistent memory is loaded as context. This ensures that any performance gains are attributable to learning from past experience, not to growing context windows.
  • Persistent artifacts: During training, the agent creates, revises, and accumulates artifacts (e.g., textual memory, tools, skills, strategies) that persist across episodes and are inherited by future instances.
  • Frozen model weights: The model weights remain frozen during self-improvement. The model improves by refining artifacts rather than by gradient-based weight updates.
  • Held-out evaluation: At each checkpoint, we evaluate on training interactions and held-out interactions to measure whether improvements generalize beyond the interactions used for learning.

The protocol is as follows. At each iteration, the agent plays a batch of games in fresh contexts, receives feedback, and revises its persistent artifacts. We evaluate the agent's performance on both training and held-out interactions at each checkpoint. This allows us to measure complete learning trajectories, not just endpoints.

3.1 Agent Plasticity

We define agent plasticity as the efficiency with which an agent converts experience into gains in future held-out performance. Formally, given a learning trajectory {St}t=1T\{S_t\}_{t=1}^{T} of held-out scores and cumulative learning cost ctc_t at each checkpoint tt, the plasticity PP is the slope of the best-fit line to the learning curve:

P=ΔSΔcP = \frac{\Delta S}{\Delta c}

where ΔS\Delta S is the change in held-out score and Δc\Delta c is the change in cumulative learning cost. We also define a robust variant, PsatP^{sat} (plasticity to saturation), which measures the gain up to the saturation point where further experience yields no additional improvement.

3 Protocol and Experimental Setup

3.1 Environments and Models

We evaluate self-improvement in Chess, Go, Hex, and NetHack. These environments are challenging for frontier models and have clear win/loss/draw outcomes that allow us to measure performance unambiguously. For chess, Go, and Hex, we use a self-play protocol in which the agent plays against itself; the agent’s artifacts are revised from game outcomes. NetHack is a single-player environment where the agent learns from its own successes and failures. In each environment, we evaluate the following frontier models: Claude Fable 5, Claude Opus 4.8, GPT-5.6 Sol, GPT-5.7, and Gemini 3.2 Pro.

The self-improvement loop. Each round of self-improvement consists of training and evaluation phases. In the training phase, the agent plays training games and produces artifacts—tools, skills, strategies, and textual memory—that are saved to persistent storage and inherited by future instances. The evaluation phase runs a frozen snapshot of the agent's artifacts to produce checkpoints. At every checkpoint, we evaluate performance on held-out ID and OOD interactions, while training scores are measured on the training interactions themselves.

3.1 Agent Plasticity

Agent plasticity is the efficiency with which an agent converts experience into held-out performance gains. We measure plasticity by fitting a saturating curve to the learning trajectory of held-out ID score S(c)S(c) versus cumulative learning cost cc:

S(c)=Sinit+(Ssat−Sinit)⋅cc+chalfS(c) = S_{init} + (S_{sat} - S_{init}) \cdot \frac{c}{c + c_{half}}

where SinitS_{init} is the initial held-out ID score, SsatS_{sat} is the fitted saturation score, and chalfc_{half} is the cost at which half of the potential improvement is achieved. The derivative at the origin is d=(Ssat−Sinit)/chalfd = (S_{sat} - S_{init}) / c_{half}. We define plasticity PsatP^{sat} as the total improvement up to the saturation point per unit learning cost:

Psat=Ssat−SinitcsatP^{sat} = \frac{S_{sat} - S_{init}}{c_{sat}}

where csatc_{sat} is the cost to reach saturation. This measures the efficiency with which an agent converts experience into gains in future held-out performance. Because the agent's ability to improve may vary over time, we also define instantaneous plasticity PinstP^{inst} , which is the derivative of the fitted held-out ID performance curve with respect to learning cost. We also report learning cost to reach x%x\% of saturation (cxc_{x}).

3 Experimental Setup

3.1 Environments and Models

We use four environments: Chess, Go, Hex, and NetHack. For Chess, Go, and Hex, we use the GameBench benchmark (Shao et al., 2025), where agents play against a fixed engine (Stockfish, KataGo, and a custom solver respectively) and receive rewards through a text-based action interface. For NetHack, we use the NetHack Challenge environment (Kuttler et al., 2020; Singh et al., 2024). For all environments, we use the same prompt template (Appendix E.1), and agents are evaluated with greedy decoding. We use frontier models: GPT-5.6 Sol, Claude Opus 4.8, Claude Fable 5, and Gemini 2.5 Pro, and evaluate the effectiveness of self-improvement across 20 rounds of experience collection and artifact refinement.

Self-improvement loop. We study agents that self-improve by revising persistent artifacts from experience. We use a frozen model with a fixed context window and a persistent artifact store. Each round proceeds in three phases: (i) training: the agent plays NN games per environment, and its context is refreshed after every game; (ii) revision: the agent reviews its collected experience and revises its artifacts; (iii) evaluation: the agent is evaluated on training and held-out interactions, with a fresh context at every episode. The artifact store is persistent across rounds and inherited by future instances. Training, revision, and evaluation all share the same frozen model. This protocol cleanly separates self-improvement from model weights updates and measures how experience is amortized into reusable artifacts that improve future performance.

2.1 Agent Plasticity

We introduce agent plasticity, a measure of the efficiency with which an agent converts experience into future held-out performance gains. Given a sequence of checkpoints t=1,…,Tt = 1, \ldots, T and held-out in-distribution score St∈[0,100]S_t \in [0, 100] at each checkpoint, we define the plasticity metric PsatP^{sat} as follows. Let ΔSt=St−S0\Delta S_t = S_t - S_0 denote the held-out score gain at checkpoint tt. Let ctc_t denote the cumulative learning cost at checkpoint tt (see Section 3.3). Let t∗=arg⁡max⁡tΔStt^* = \arg\max_{t} \Delta S_t be the time at which the gain is maximized, and let ΔS∗=ΔSt∗\Delta S^* = \Delta S_{t^*} be that maximum gain. We define:

Psat=ΔS∗ct∗P^{sat} = \frac{\Delta S^*}{c_{t^*}}

the gain at the saturation point per unit of learning cost. This metric measures the efficiency of converting experience into held-out performance gains up to the point where further experience no longer yields additional gains. We also compute PslopeP^{slope} , the slope of the learning curve before saturation, and PmaxP^{max} , the maximum gain per unit cost across all checkpoints. These three metrics are correlated but capture different aspects of acquisition efficiency (Table 3). We also track other metrics: PsatP^{sat} (plasticity to saturation), PmaxP^{max} (max plasticity), and PslopeP^{slope} (plasticity slope).

3 Results

3.1 Setup

We evaluate self-improvement in four environments: Chess, Go, Hex, and NetHack, selected for their clear win/loss outcomes, well-defined skill gradients, and existing agent self-play capabilities. For Chess, Go, and Hex, we use the OpenSpiel framework (Lanctot et al., 2019) and the gpt-oss-120b model as a fixed judge. We test five frontier models: Claude Fable 5, Claude Opus 4.8, GPT-5.6 Sol, Gemini 3 Pro, and Llama 4 Maverick. The agent maintains persistent artifacts (tools, skills, memories, strategies) that are inherited by future instances, and the model revises them from experience. Each checkpoint is evaluated on held-out games. See Section 2 for details.

Table 1: Environments used in our study.

EnvironmentTrainingHeld-out IDHeld-out OOD
Easy chess256 games512 gamesHard chess
Go 9x9128 games256 games19x19 Go
Hex 9x9128 games256 games13x13 Hex
NetHack4.5k episodes1k episodes-

3 Experimental Setup

3.1 Environments

We use Chess, Go, Hex, and NetHack as testbeds. In board games, the agent plays games against itself. For Chess, the agent plays White. The agent's training interactions are 256 games for Chess, 128 for Go and Hex; held-out ID interactions are 256 games for Chess and 128 for Go and Hex. In NetHack, the agent is trained on 100 episodes and evaluated on 200 held-out episodes. Held-out OOD evaluations test harder settings. In all games, the agent receives only sparse reward at the end of the game. Self-improvement proceeds through persistent artifacts, i.e., tools, skills, strategies, and textual memory, that are revised over training interactions and inherited by future instances.

3.1 Setup and Environments. We evaluate on Chess, Go, Hex, and NetHack. For Chess, Go, and Hex, we use self-play with two copies of the agent; a third copy serves as a judge. For NetHack, we use a pre-existing dataset of trajectories. We use frontier models: Claude Fable 5, Claude Opus 4.8, GPT-5.6 Sol, Gemini 3 Pro, and Llama 4 Maverick (details in Appendix A). All models are accessed through APIs and use default decoding parameters. The agent is given a system prompt describing the task and instructed to maintain persistent artifacts, which are passed to future instances. Each round, the agent receives a transcript of its previous attempt, identifies mistakes, and revises its artifacts. We run 12 rounds of self-improvement (Sections 3.1 and 3.2).

3.1 Setup

We evaluate self-improvement in four environments: Chess, Go, Hex, and NetHack. For Chess, Go, and Hex, we use deterministic game environments and a fixed set of evaluation opponents. The agent plays games and records a transcript of its actions. After each round, the agent reflects on the transcript and updates a persistent artifact: a set of notes (in JSON format) that includes general strategies, opening knowledge, and tactical heuristics. The artifact is passed to the next round, where the agent plays with a fresh context but inherits the updated artifact. This mirrors the setting where future instances of an agent reuse accumulated knowledge from prior interactions. We evaluate on Easy, Hard, and OOD settings as described in Section 3.1.

In NetHack, self-improvement is studied through persistent memory and tools. The agent is given a memory file that it can revise between episodes, and it can create and use tools in the environment. We measure performance using a score-based metric.

3.1 Environments and Evaluation Protocol

We use four environments: Chess, Go, Hex, and NetHack. For Chess, Go, and Hex, we use a synthetic data generator that produces games between two fixed opponents (Stockfish 16, KataGo, and Muzero Zero, respectively), where the agent always plays the same side. The agent receives a board state and produces a move. We evaluate on three sets of interactions: training (the exact interactions used for learning), held-out ID (held-out samples from the same distribution), and held-out OOD (samples from a harder distribution, e.g., stronger opponent for chess, larger board for Go/Hex). For NetHack, we use the NetHack Learning Environment (Küttler et al., 2020) with a reward-based score.

Self-improvement protocol. We freeze model weights and let each agent improve through persistent self-constructed artifacts. Each round consists of a training phase and an evaluation phase. In the training phase, the agent is given a batch of NN training interactions and access to artifacts from previous rounds; it diagnoses failures and revises artifacts. In the evaluation phase, the agent is tested on held-out interactions. This process repeats for RR rounds. Agents use a persistent artifact store (e.g., a directory of files) containing self-constructed tools, skills, strategies, and textual memory. Each acting episode begins in a fresh context, and future instances inherit the artifact store from previous rounds. This setup separates the persistent component of learning from the ephemeral context of individual episodes.

Summary

Introduction and Theoretical Foundation

Methodology

Empirical Validation / Results

Theoretical and Practical Implications

Conclusion

Related papers