In-Context Learning for Robots: Methods and Applications

Summary (Overview)

  • Comprehensive taxonomy of ICL for robots: The paper organizes the literature into four method families—context-conditioned policies, geometric demonstration transfer, world-model-based control, and skill- and agent-based execution—distinguished by the computational intermediate that carries contextual evidence into action.

  • Six learning horizons framework: The authors introduce a framework (S1–S6) that distinguishes explicit control programming (S1), neural policy learning (S2), history-based adaptation (S3), in-context task learning (S4), physical recursive self-improvement (S5), and collective knowledge evolution (S6), clarifying what each adaptation mechanism can and cannot achieve.

  • Fixed-parameter adaptation as central subject: The review focuses on how new evidence (demonstrations, corrections, interaction history) changes deployed behavior while neural parameters remain fixed, distinguishing this from fine-tuning and other parameter-adaptation approaches.

  • Correspondence and memory as shared mechanisms: Cross-cutting mechanisms—geometric/semantic correspondence, temporal alignment, retrieval, and evidence retention—connect all four method families and determine whether taught requirements survive physical execution and transfer.

  • Evaluation methodology linking design to evidence: The paper proposes specific attribution controls (context interventions, retention tests, physical resets) that separate responsiveness to teaching, physical transfer, and benefits from retained experience, connecting method design to measurable outcomes.

Introduction and Theoretical Foundation

The Core Problem

A robot may possess reusable motor competence yet still need evidence about what to do in a specific situation. Demonstrations specify a fold or assembly order; corrections revise procedures; interaction reveals friction or misalignment. In-context learning (ICL) for robots studies how such evidence changes deployed behavior without another task-specific update to neural parameters.

The paper identifies two historical roots:

  • One-shot imitation: infers intended behavior from demonstrations
  • Meta-reinforcement learning: infers tasks or dynamics from outcomes

Six Learning Horizons

The framework distinguishes six levels of learning:

HorizonDescription
S1Explicit control programming (control laws, planning models)
S2Neural policy learning (imitation, reinforcement)
S3History-based adaptation (interaction memory for task/dynamics inference)
S4In-context task learning (inference from supplied teaching)
S5Physical recursive self-improvement (experience improves subsequent learning)
S6Collective knowledge evolution (knowledge exchange across embodiments)

Formal Framework

The core policy interface is defined as:

At∼πθ(⋅∣ht,Ct),(1)A_{t} \sim \pi_{\theta}(\cdot | h_{t}, C_{t}), \tag{1}

where ht=(o1,a1,…,at−1,ot)h_t = (o_1, a_1, \ldots, a_{t-1}, o_t) is the observation–action history, CtC_t is task-relevant context (demonstrations, instructions, corrections, retrieved episodes), and AtA_t is the proposed action block of horizon HH.

For fixed-parameter adaptation:

θt+1=θt=θ0,ht+1=(ht,at,ot+1).(3)\theta_{t+1} = \theta_{t} = \theta_{0}, \qquad h_{t+1} = (h_{t}, a_{t}, o_{t+1}). \tag{3}

A self-contained context state updates as:

Ct+1=Uθ(Ct,ot,at,ot+1).(2)C_{t+1} = U_{\theta}(C_{t}, o_{t}, a_{t}, o_{t+1}). \tag{2}

Four Information Roles of Context

  1. Task specification: identifies the intended goal, relation, or procedure
  2. Correspondence: relates demonstrated objects, contacts, phases to executable counterparts
  3. Physical response: describes how the current body and environment respond to action
  4. Execution state: records the scene, hidden events, and completed prerequisites

Methodology

Four Method Families

1. Context-Conditioned Policies

Context-conditioned policies map task evidence and robot history directly to actions:

rt=Eθc(ht,Ct),πθ(At∣ht,Ct)=qθa(At∣ht,rt).(4)r_{t} = E_{\theta}^{\mathrm{c}}(h_{t}, C_{t}), \qquad \pi_{\theta}(A_{t} \mid h_{t}, C_{t}) = q_{\theta}^{\mathrm{a}}(A_{t} \mid h_{t}, r_{t}). \tag{4}

Two forms exist:

  • Action reuse through correspondence (nonparametric retrieval):
πθ(⋅∣ht,Ct)=∑j∈JtβtjδAˉj,βtj≥0,∑j∈Jtβtj=1.(5)\pi_{\theta}(\cdot \mid h_{t}, C_{t}) = \sum_{j \in J_{t}} \beta_{tj} \delta_{\bar{A}_{j}}, \quad \beta_{tj} \geq 0, \quad \sum_{j \in J_{t}} \beta_{tj} = 1. \tag{5}
  • Learned generation from an interpreted context representation, using attention-based alignment:
αti=exp⁡sθ(ht,dis)∑k=1Lexp⁡sθ(ht,dks),rt=∑i=1Lαtieθ(dis).(6)\alpha_{ti} = \frac{\exp s_{\theta}(h_{t}, d_{i}^{\mathrm{s}})}{\sum_{k=1}^{L} \exp s_{\theta}(h_{t}, d_{k}^{\mathrm{s}})}, \qquad r_{t} = \sum_{i=1}^{L} \alpha_{ti} e_{\theta}(d_{i}^{\mathrm{s}}). \tag{6}

2. Geometric Demonstration Transfer

This family preserves a motion or contact reference and adapts its realization:

ζ1:L=Eθm(Ds),g=Retarget⁡(ζ1:L,ot,B),At=κ(g,ot).(7)\zeta_{1:L} = E_{\theta}^{\mathrm{m}}(D^{\mathrm{s}}), \qquad g = \operatorname{Retarget}(\zeta_{1:L}, o_{t}, \mathcal{B}), \qquad A_{t} = \kappa(g, o_{t}). \tag{7}

For rigid object-relative transfer:

WTE,bq(σ)=WTObq(WTObs)−1WTE,bs(σ).(8){}^{W}T_{E,b}^{\mathrm{q}}(\sigma) = {}^{W}T_{O_{b}}^{\mathrm{q}} \left({}^{W}T_{O_{b}}^{\mathrm{s}}\right)^{-1} {}^{W}T_{E,b}^{\mathrm{s}}(\sigma). \tag{8}

Here WTX∈SE(3)^W T_X \in SE(3) is the rigid pose of frame XX in world frame WW; EE denotes the end effector, ObO_b the reference object, and σ∈[0,1]\sigma \in [0,1] normalized progress within a segment.

3. World-Model-Based Control

Predictive control uses contextual evidence to anticipate consequences:

v+∼pθv(⋅∣Ct,ht),At∼qθv(⋅∣v+,Ct,ht),(9)\begin{array}{l} v^{+} \sim p_{\theta}^{\mathrm{v}}(\cdot | C_{t}, h_{t}), \\ A_{t} \sim q_{\theta}^{\mathrm{v}}(\cdot | v^{+}, C_{t}, h_{t}), \end{array} \tag{9}

with the marginal policy:

πθ(At∣ht,Ct)=∫qθv(At∣v+,Ct,ht)pθv(v+∣Ct,ht)dv+.(10)\pi_{\theta}(A_{t} \mid h_{t}, C_{t}) = \int q_{\theta}^{\mathrm{v}}(A_{t} \mid v^{+}, C_{t}, h_{t}) p_{\theta}^{\mathrm{v}}(v^{+} \mid C_{t}, h_{t}) \mathrm{d}v^{+}. \tag{10}

Model-based selection evaluates candidate actions:

At⋆∈arg⁡max⁡A^t∈AtEv+∼pθd(⋅∣ht,A^t,Ct)[R(v+,Ct)].(11)A_{t}^{\star} \in \arg \max_{\widehat{A}_{t} \in \mathcal{A}_{t}} \mathbb{E}_{v^{+} \sim p_{\theta}^{\mathrm{d}}(\cdot | h_{t}, \widehat{A}_{t}, C_{t})} \left[ \mathcal{R}(v^{+}, C_{t}) \right]. \tag{11}

4. Skill- and Agent-Based Execution

This family generates execution specifications (programs, skill sequences, tool calls) for a separate executor:

g=fθ(Ct,ht),At=κ(g,ot).(13)g = f_{\theta}(C_{t}, h_{t}), \qquad A_{t} = \kappa(g, o_{t}). \tag{13}

Stochastic form:

πθ(At∣ht,Ct)=∫kθ(At∣g,ht)pθg(g∣Ct,ht)dg.(14)\pi_{\theta}(A_{t} \mid h_{t}, C_{t}) = \int k_{\theta}(A_{t} \mid g, h_{t}) p_{\theta}^{\mathrm{g}}(g \mid C_{t}, h_{t}) \mathrm{d}g. \tag{14}

Shared Mechanisms

Memory follows a read–update decomposition:

mt+1=uθ(mt,xt),Mt+1=wθ(Mt,xt),Rt=ρθ(ht,Mt),Ct=bθ(Cttask,mt,Rt).(17)\begin{array}{r l} & m_{t+1} = u_{\theta}(m_{t}, x_{t}), \quad M_{t+1} = w_{\theta}(M_{t}, x_{t}), \\ & R_{t} = \rho_{\theta}(h_{t}, M_{t}), \qquad C_{t} = b_{\theta}(C_{t}^{\mathrm{task}}, m_{t}, R_{t}). \end{array} \tag{17}

Budgeted retrieval:

Jt∈arg⁡max⁡J⊆{1,…,Nt},∣J∣≤K∑j∈J[sθ(ht,ξj)−η],Rt=(ξj)j∈Jt.(18)J_{t} \in \arg \max_{J \subseteq \{1, \dots, N_{t}\}, |J| \leq K} \sum_{j \in J} \left[ s_{\theta}(h_{t}, \xi_{j}) - \eta \right], \quad R_{t} = (\xi_{j})_{j \in J_{t}}. \tag{18}

External knowledge revision follows a proposal-and-validation loop:

B~n+1∼Fθ(⋅∣Bn,Tn),Bn+1={B~n+1,χn=1,Bn,χn=0,θ=θ0.(19)\begin{array}{l} \widetilde{B}_{n+1} \sim F_{\theta}(\cdot \mid B_{n}, \mathcal{T}_{n}), \\ B_{n+1} = \begin{cases} \widetilde{B}_{n+1}, & \chi_{n} = 1, \\ B_{n}, & \chi_{n} = 0, \end{cases} \qquad \theta = \theta_{0}. \end{array} \tag{19}

Empirical Validation / Results

Training Supervision

The supervised objective for paired episodes:

L(θ)=E(Ds,τq)∼Ppair[1∣Iq∣∑t∈Iqℓθ(Atq;htq,Ds)].(21)\mathcal{L}(\theta) = \mathbb{E}_{(D^{\mathrm{s}}, \tau^{\mathrm{q}}) \sim P_{\mathrm{pair}}} \left[ \frac{1}{|I^{\mathrm{q}}|} \sum_{t \in I^{\mathrm{q}}} \ell_{\theta}(A_{t}^{\mathrm{q}}; h_{t}^{\mathrm{q}}, D^{\mathrm{s}}) \right]. \tag{21}

Key Reported Comparisons

StudyTargetContrastReadout
ICRTTraining dataDROID-only vs. multi-taskDROID-only: no progress
BPPPromptGoal image → demoFidelity ↑
Show-HarnessAPI semanticsArbitrary names, no conventions1/20 successes
NOLOScene contextNo video → videoSR: 33.58 → 43.65
MT3Action reuseBC → retrievalSR ↑; demos ↓
Part-based transferWarp granularityWhole object → parts11/27 → 23/27
RAPIDVerificationScene variants off → onSR: 53.2 → 75.9
ZevaPersistent memoryOff → onSR: +10–20 pp
TraceFlowTrace guidanceBase → guidedOrdered: 21/50 → 39/50
FAREHistory revisionBase → selectiveSR: 91.5 → 93.2
LMPCSuccessor trainingBase → successorSR: 39.4 → 66.3

Scaling Results

  • BPP: With 10,000 drawing demonstrations, mean Chamfer error improves from 9.5 pixels (500 tasks, 20 demos each) to 3.4 pixels (2,000 tasks, 5 demos each), showing task diversity matters more than demonstration count per task.
  • S1: On unseen tasks, demonstration prompting rises from 1% to 66% as training scales from 1k to 100k hours, while language prompting rises only from 0% to 9%.

Data Resources

  • AgiBot World (March 2025): 1,001,552 trajectories, 2,976.4 hours, 217 tasks, 106 scenes
  • XR-2: 531.7 hours robot teleoperation + ~1,000 hours dual-UMI demonstrations; improves folding from 58% to 93% across three retraining rounds
  • YUBI: 8,434 hours over 119 tasks with shared end-effector transfer

Theoretical and Practical Implications

Transfer Requirements by Method Family

Method FamilyIntermediateTransfer Requirement
Context-conditioned policiesAction distributionAction inference preserves the taught distinction
Geometric demonstration transferMotion/contact referenceMatched interaction remains applicable
World-model-based controlPredicted consequencesPredicted evolution is realizable
Skill- and agent-based executionExecution specificationAvailable skills preserve task constraints

Cross-Object Transfer

The paper formalizes object substitution through a pouring example (Table 6): changing the held vessel (jug → bottle) requires new grasp and tilt angle; changing the receiver (cup → bowl) requires new pouring position and height; changing both couples these adjustments. The transferred requirement (direct liquid into receiver) must survive while motion changes.

Evaluation Controls

The paper proposes specific attribution interventions:

TargetInterventionReadout
Training checkpointCheckpoint × contextLearned context use
Task teachingOriginal ↔ replacementRequirement adherence
Retained historyRetain ↔ clearUse of prior evidence
Neural parametersAdapt ↔ restoreAdaptation gain
Reusable guidanceRetain ↔ withholdLesson transfer

Reuse Gain Metric

For trial-memory adaptation:

Δnreuse=Snkeep−Snreset\Delta_{n}^{\mathrm{reuse}} = S_{n}^{\mathrm{keep}} - S_{n}^{\mathrm{reset}}

where SnS_n is the success/adherence rate at attempt nn, comparing retained versus cleared memory with all other conditions held fixed.

Conclusion

Main Contributions

  1. A taxonomy of four method families with shared correspondence and memory mechanisms, organized by the intermediate that execution consumes.

  2. An account of training relationships: how contextual training establishes context use, and how object substitution, unfamiliar environments, and execution conditions limit transfer—separating broader motor competence from broader ability to learn through teaching.

  3. A synthesis of evaluation practices: separating context dependence, transfer, and retained-experience benefits to motivate compositional learning and improved teachability.

Future Directions

  1. From context dependence to reusable learning rules: distinguishing familiar-task selection, compositional transfer, and inference of genuinely unfamiliar rules (Table 18).

  2. Instruction tuning for complex contextual learning: requirement binding and composition, selective revision of procedures, and decision-directed information acquisition.

  3. Preserving taught relations across physical change: learning which demonstrated details are binding and when no valid substitute exists.

  4. Retaining teaching while revising obsolete experience: distinguishing current task meaning from future reinterpretation needs.

  5. Physical recursive self-improvement (S5): experience improves how subsequent tasks are learned, measured by successor acquisition rates.

  6. Collective knowledge evolution (S6): verified knowledge exchange across embodiments improves group learning.

Key Takeaway

"The central synthesis is that contextual learning depends on preserving task-relevant distinctions from evidence to execution. Loss during selection, interpretation, realization, or reuse calls for different data, representations, or control capabilities."

The paper concludes that the next challenge is making these dependencies transferable: inferring unfamiliar combinations of requirements, revising only the constraints affected by feedback, and discarding experience whose conditions no longer hold—all assessed by what later learners acquire and the total teaching, interaction, and training effort required.

Related papers