# Post-Training Leaves Behavioral Shadows on Unrelated Decisions

> Post-training updates leave behavioral shadows on unrelated inputs that can be extracted via single-word queries to transfer capabilities, yielding +5.34 points on HumanEval+.

- **Source:** [arXiv](https://arxiv.org/abs/2609.29233)
- **Published:** 2026-09-30
- **Permalink:** https://picx.dev/p/TI7Ijc
- **Whiteboard:** https://picx.dev/p/TI7Ijc/image

## Summary

# Post-Training Leaves Behavioral Shadows on Unrelated Decisions

## Summary (Overview)

- **Core finding**: Post-training updates on language models create "behavioral shadows"—changes in decisions on inputs *unrelated* to the training task. These shadows carry capability-relevant information that can be extracted and transferred to another model.
- **Key contribution**: The authors introduce **Active Taskless Distillation (ATD)**, a method that achieves capability transfer using only a *single word* from the teacher per prompt, without target-task examples, teacher logits, or teacher parameters.
- **Primary result**: On HumanEval+ (code generation), a student trained on 5,664 single-word teacher responses gains **+5.34 percentage points** over an exact nuisance-matched control, with the effect persisting across independently constructed acquisitions.
- **Generality**: Transfer is demonstrated across seven target tasks (code, scientific knowledge, commonsense reasoning, reading comprehension) and across model generations, sizes, and families (Qwen2.5, Qwen3, Llama).
- **Boundary condition**: Large teacher gains do *not* guarantee transfer—the transferred capability must be one the student can be elicited to express (e.g., memorized answers and substitution ciphers do not transfer).

---

## Introduction and Theoretical Foundation

### Background and Motivation

Post-training (e.g., fine-tuning, DPO) is typically understood through changes in target-task behavior—a coding update makes a model write better code. However, the influence of an update may extend beyond the trained task: it may change which *ordinary word* a model prefers in a short story, even when neither the story nor the answer contains code or mathematics. These changes on task-unrelated inputs constitute the **behavioral shadow** of post-training.

Prior work on *subliminal learning* showed that behavioral traits can pass through semantically unrelated training data, but largely focused on traits/preferences using extensive teacher outputs. ATD asks two sharper questions:

1. **Observability**: Can a private post-training update be observed through individual decisions on unrelated inputs?
2. **Capability transfer**: Can learning from these decisions transfer part of the teacher's target-task improvement?

### Theoretical Foundation

The theoretical motivation rests on a first-order expansion around the public ancestor. For a carrier prompt $x_i$ with two candidate words $a_i$ and $b_i$, define the logit difference:

$$
m_i(\theta) = z_\theta(a_i \mid x_i) - z_\theta(b_i \mid x_i) \tag{2}
$$

Expanding around the ancestor parameters $\theta_0$:

$$
m_i(\theta_0 + \delta) \approx m_i(\theta_0) + g_i^\top \delta, \qquad g_i = \nabla_\theta m_i(\theta_0) \tag{3}
$$

where $\delta$ is the unknown post-training update and $g_i$ is the gradient of the margin with respect to parameters. The observed binary choice is modeled as:

$$
y_i \approx \mathrm{sign}\big(m_i(\theta_0) + g_i^\top \delta\big) \tag{4}
$$

**Key insight**: Near a tie (where $|m_i(\theta_0)|$ is small), even a small preference change induced by the update can *reverse* the selected word. Each observed choice thus provides a **one-bit observation** of the update's behavioral shadow. The teacher parameters are:

$$
\theta_T = \theta_0 + \delta \tag{1}
$$

---

## Methodology

### Active Taskless Distillation (ATD)

ATD consists of three stages:

**1. Selecting near-tie prompts.** Using only the public ancestor $M_0$, compute the probability of candidate $b_i$ after normalizing over the pair:

$$
q_i = \frac{\exp z_0(b_i \mid x_i)}{\exp z_0(a_i \mid x_i) + \exp z_0(b_i \mid x_i)} \tag{5}
$$

Retain prompts satisfying $|q_i - 0.5| \le \epsilon$ with $\epsilon = 0.02$. Also require that the public model's highest-probability token over the full vocabulary is one of the two candidates.

**2. Querying the teacher.** The query set is fixed before observing teacher responses. Each prompt is queried once; the teacher returns its highest-probability next token via greedy decoding over the full vocabulary. Responses are retained only if they belong to $\{a_i, b_i\}$; otherwise discarded without resampling.

**3. Training the student.** The student is initialized from the same public ancestor $M_0$ and trained via full-vocabulary cross-entropy on the $n$ retained prompt–word pairs:

$$
\mathcal{L}_{\mathrm{CE}}(\theta) = -\frac{1}{n} \sum_{i=1}^{n} \log p_\theta(w_i^T \mid x_i) \tag{6}
$$

### Experimental Design

- **Primary lineage**: Qwen2.5-1.5B-Instruct as public ancestor and student initialization; teacher is a frozen rank-16 code-DPO LoRA; student is a rank-16 LoRA.
- **Primary acquisition**: 5,664 (prompt, word) training rows over 128 words—audited to be free of code, mathematics, benchmark, and task terms.
- **Controls**: The key control is the **exact nuisance-matched permutation** (same carriers, completion multiset, and per-bin difficulty, but broken prompt–token pairing). Other controls include teacher-label shuffle, public-label (teacher-free), and random-marginal labels.
- **Evaluation**: HumanEval+ scored by greedy pass@1 through execution; GSM8K by exact match; multiple-choice benchmarks by held-out accuracy.

### Instrument Extensions

**Reference-relative pair-margin (RPM) objective**:

$$
\mathcal{L}_{\mathrm{RPM}} = -\mathbb{E}_i \log \sigma \left(\beta \left[ \log \frac{p_\theta(w_i^T \mid x_i)}{p_\theta(\bar{w}_i \mid x_i)} - \log \frac{p_0(w_i^T \mid x_i)}{p_0(\bar{w}_i \mid x_i)} \right] \right) \tag{7}
$$

where $\bar{w}_i$ is the non-selected word. This subtracts the frozen public margin to focus on the update-induced change.

**Multi-token observations** (K-word answers):

$$
\mathcal{L}_{\mathrm{multi}} = -\frac{1}{n} \sum_i \frac{1}{K} \sum_{k=1}^{K} \log p_\theta(y_{i,k} \mid x_i, y_{i,<k}) \tag{8}
$$

---

## Empirical Validation / Results

### 5.1 Identification of a Specific Capability (Primary Evidence)

**Table 1: Controls for identification (Qwen2.5-1.5B, HumanEval+)**

| Control (vs. ATD-CE signal) | Control score | Δ = signal − control | n |
|---|---|---|---|
| Exact nuisance-matched (primary) | 45.88 | **+5.34 [1.22, 9.60]** | 4 |
| Global teacher-label shuffle | 46.19 | **+5.03 [0.91, 9.30]** | 4 |
| Public-label CE (teacher-free) | 46.65 | **+4.57 [0.91, 8.54]** | 4 |
| Random-marginal labels† | 45.12 | +6.71 | 1 |

*†Single-seed diagnostic, no CI.*

The student reaches 51.22% on HumanEval+, matching the teacher's aggregate coding score. The exact nuisance-matched control isolates the prompt–token pairing, ruling out explanations based on unigram frequency or label marginals.

### 5.2 Robustness

- **Across acquisitions**: A factorial of 5 independently constructed acquisitions × 3 training seeds gives an aggregate effect of **+4.80 pp**, with all 15 signal–control pairs and all 5 acquisitions positive (between-acquisition SD = 0.83 pp).
- **Across adaptation regimes**:

| Adaptation | Signal | Control | Δ (95% CI) |
|---|---|---|---|
| LoRA r16 | 51.22 | 45.88 | +5.34 [1.22, 9.60] |
| LoRA r32 | 53.66 | 47.10 | **+6.55 [2.59, 10.82]** |
| LoRA r64 | 54.27 | 45.12 | **+9.15 [4.88, 13.87]** |
| Full SFT | 50.00 | 45.33 | **+4.67 [1.22, 8.54]** |

### 5.3 Generalization Across Tasks and Models

**Table 2: Per-task transfer across seven target tasks**

| Source / Target | Base | Teacher | Gap | Signal | Control | Δ (95% CI) | n |
|---|---|---|---|---|---|---|---|
| Code / HumanEval+ | 43.29 | 51.22 | +7.93 | 51.22 | 46.19 | **+5.03 [0.91, 9.30]** | 4 |
| ScienceQA | 77.65 | 92.13 | +14.48 | 79.38 | 76.84 | **+2.53 [1.57, 3.52]** | 3 |
| OpenBookQA | 78.20 | 86.40 | +8.20 | 79.73 | 77.67 | **+2.07 [0.33, 3.87]** | 3 |
| HellaSwag | 64.84 | 87.14 | +22.31 | 66.82 | 63.93 | **+2.90 [2.44, 3.36]** | 3 |
| RACE-high | 76.19 | 81.79 | +5.60 | 76.74 | 75.93 | **+0.81 [0.24, 1.40]** | 3 |
| CommonsenseQA | 74.37 | 79.12 | +4.75 | 76.28 | 74.34 | **+1.94 [0.76, 3.19]** | 3 |
| PIQA | 76.93 | 83.52 | +6.58 | 77.71 | 76.46 | **+1.25 [0.15, 2.38]** | 3 |

ATD also achieves positive mean gains on Qwen3-1.7B, Qwen3-4B, and Llama-3.2-1B (though paired 95% CIs include zero for these additional models).

### 5.4 Analysis of the Transferred Signal

**Active vs. passive query selection**:

| Matched budget | Δ active – passive |
|---|---|
| Teacher queries | +4.27 [0.30, 8.38] |
| Training rows | +5.03 [1.52, 8.84] |

The effect comes from concentrating queries at the ancestor's decision boundary, not from query count or data volume.

**Source specificity**: Across three sources (code, science, mathematics) and three target endpoints, each shadow's largest positive effect falls on its *matched* target (code diagonal: +5.34; science diagonal: +2.52), with off-diagonal effects near zero or negative. A 50/50 mixture of near-orthogonal math and code shadows recovers *both* directions simultaneously (between-source cosine = −0.0002).

**Teacher–student agreement**: Among 19 HumanEval+ tasks the teacher fixed over the ancestor (76 seed–task cells), ATD students solved 64.47% (49/76) vs. 21.05% (16/76) for the matched control. Task-level change vectors aligned more closely with the teacher's (cosine .701 vs. .398).

### 5.5 Large Teacher Gains Do Not Guarantee Transfer

**Table 3b: Large-gap, non-elicitable teachers**

| Non-elicitable teacher | Gap | Base | Signal |
|---|---|---|---|
| HE cheat (memorized) | +39.02 | 59.15 | 56.56 |
| Substitution cipher | ≈97 | 0.00 | 0.00 |

Neither transfers. This shows ATD does not simply "launder" the teacher's answers—what transfers is capability the student can be elicited to express.

### Instrument Extension Results

**Table 4: RPM objective vs. CE** (Qwen2.5-1.5B)

| Target | CE Δ_shuf | RPM Δ_shuf | CE Δ_pub | RPM Δ_pub | Outcome |
|---|---|---|---|---|---|
| HumanEval+ | +5.03 | **+6.30 [2.03, 10.98]** | +4.57 | +4.47 [−0.61, 9.76] | stronger vs. shuffle |
| ScienceQA | +2.53 | +3.28 [2.17, 4.41] | +1.18 | +1.09 | ≈ CE |
| PIQA | +1.25 | +1.69 [0.60, 2.81] | +0.76 | +0.94 | both positive |
| OpenBookQA | +2.07 | +1.47 [−0.60, 3.53] | +1.13 | +1.20 | exception |

**Table 5: Multi-token observation (K=20)** (Qwen2.5-0.5B)

| Target | Signal | Control | Δ (95% CI) |
|---|---|---|---|
| GSM8K | 21.23 | 14.86 | **+6.37 [+4.37, +8.42]** |
| HumanEval+ | 4.88 | 3.86 | +1.02 [−0.81, +3.05] |
| ScienceQA | 61.41 | 60.57 | +0.84 [−0.09, +1.78] |

---

## Theoretical and Practical Implications

### Theoretical Significance

1. **Capabilities are observable through unrelated behavior**: The behavioral shadow of post-training is not just evidence of change—it carries *capability-relevant information* that can be extracted through near-boundary queries. This extends subliminal learning from traits/preferences to full capabilities.

2. **First-order approximation validated**: The near-tie selection strategy is grounded in the expansion $m_i(\theta_0 + \delta) \approx m_i(\theta_0) + g_i^\top \delta$, and the empirical results confirm that decisions near the ancestor's decision boundary are informative about the update direction $\delta$.

3. **The channel carries structured, source-specific information**: Different updates leave distinguishable shadows (source-specific diagonals in transfer matrices), and multiple shadows can be superposed (50/50 mixture recovers both sources). This suggests the shadow is a *measurement* of the update, not just a generic "train harder" signal.

### Practical Implications

1. **Privacy**: A private post-training update can be partially observed through black-box queries on unrelated inputs, without access to parameters, logits, or target-task responses. This has implications for model stewardship and API design.

2. **Distillation without task data**: ATD enables capability transfer without any target-task examples—only ordinary word choices on unrelated prompts are needed. This could enable new forms of model compression or capability transfer when task data is unavailable.

3. **Diagnostic tool**: The training-time trace shows a teacher-aligned change emerging *before* executable endpoint separation from control—the shadow may expose aspects of an update that target-task scores alone do not capture.

### Boundary Conditions

- Transfer requires the capability to be *elicitable* in the student (memorized answers and substitution ciphers do not transfer).
- The method relies on actively constructed carriers and a known public ancestor.
- Reliability varies across settings; the authors note "ATD does not yield reliable transfer in every tested setting."

---

## Conclusion

Post-training leaves a behavioral shadow: changes in a model's decisions on inputs unrelated to the task it was trained to improve. The authors show this shadow carries information that can improve a student's target-task performance through **Active Taskless Distillation (ATD)**—which accesses this information through single-word responses to near-boundary prompts selected using the public ancestor.

Key takeaways:
- In the primary coding experiment, a same-ancestor student trained on 5,664 single-word responses gains **+5.34 pp** on HumanEval+ over an exact nuisance-matched control, with the effect persisting across independently constructed acquisitions.
- The shadow retains information about the *source* and *strength* of the teacher's update; observations from two teachers can jointly influence the student.
- The effect is selective: each shadow's largest positive effect falls on its matched target task.
- Large teacher gains do not guarantee transfer—the capability must be elicit-able in the student.

**Future directions**: The authors suggest better query selection and learning objectives may enable students to extract more of the shadow's information and translate it into larger target-task improvements. Open questions include whether the selected observations reveal too little capability-relevant information, or whether the learning procedure fails to make sufficient use of it. ATD is presented as "an initial exploration of how capabilities can be accessed through unrelated behavior."

---

_Markdown view of https://picx.dev/p/TI7Ijc, served by PicX — AI-generated visual whiteboard summaries of research papers._
