Post-Training Leaves Behavioral Shadows on Unrelated Decisions

Summary (Overview)

  • Core finding: Post-training updates on language models create "behavioral shadows"—changes in decisions on inputs unrelated to the training task. These shadows carry capability-relevant information that can be extracted and transferred to another model.
  • Key contribution: The authors introduce Active Taskless Distillation (ATD), a method that achieves capability transfer using only a single word from the teacher per prompt, without target-task examples, teacher logits, or teacher parameters.
  • Primary result: On HumanEval+ (code generation), a student trained on 5,664 single-word teacher responses gains +5.34 percentage points over an exact nuisance-matched control, with the effect persisting across independently constructed acquisitions.
  • Generality: Transfer is demonstrated across seven target tasks (code, scientific knowledge, commonsense reasoning, reading comprehension) and across model generations, sizes, and families (Qwen2.5, Qwen3, Llama).
  • Boundary condition: Large teacher gains do not guarantee transfer—the transferred capability must be one the student can be elicited to express (e.g., memorized answers and substitution ciphers do not transfer).

Introduction and Theoretical Foundation

Background and Motivation

Post-training (e.g., fine-tuning, DPO) is typically understood through changes in target-task behavior—a coding update makes a model write better code. However, the influence of an update may extend beyond the trained task: it may change which ordinary word a model prefers in a short story, even when neither the story nor the answer contains code or mathematics. These changes on task-unrelated inputs constitute the behavioral shadow of post-training.

Prior work on subliminal learning showed that behavioral traits can pass through semantically unrelated training data, but largely focused on traits/preferences using extensive teacher outputs. ATD asks two sharper questions:

  1. Observability: Can a private post-training update be observed through individual decisions on unrelated inputs?
  2. Capability transfer: Can learning from these decisions transfer part of the teacher's target-task improvement?

Theoretical Foundation

The theoretical motivation rests on a first-order expansion around the public ancestor. For a carrier prompt xix_i with two candidate words aia_i and bib_i, define the logit difference:

mi(θ)=zθ(ai∣xi)−zθ(bi∣xi)(2)m_i(\theta) = z_\theta(a_i \mid x_i) - z_\theta(b_i \mid x_i) \tag{2}

Expanding around the ancestor parameters θ0\theta_0:

mi(θ0+δ)≈mi(θ0)+gi⊤δ,gi=∇θmi(θ0)(3)m_i(\theta_0 + \delta) \approx m_i(\theta_0) + g_i^\top \delta, \qquad g_i = \nabla_\theta m_i(\theta_0) \tag{3}

where δ\delta is the unknown post-training update and gig_i is the gradient of the margin with respect to parameters. The observed binary choice is modeled as:

yi≈sign(mi(θ0)+gi⊤δ)(4)y_i \approx \mathrm{sign}\big(m_i(\theta_0) + g_i^\top \delta\big) \tag{4}

Key insight: Near a tie (where ∣mi(θ0)∣|m_i(\theta_0)| is small), even a small preference change induced by the update can reverse the selected word. Each observed choice thus provides a one-bit observation of the update's behavioral shadow. The teacher parameters are:

θT=θ0+δ(1)\theta_T = \theta_0 + \delta \tag{1}

Methodology

Active Taskless Distillation (ATD)

ATD consists of three stages:

1. Selecting near-tie prompts. Using only the public ancestor M0M_0, compute the probability of candidate bib_i after normalizing over the pair:

qi=exp⁡z0(bi∣xi)exp⁡z0(ai∣xi)+exp⁡z0(bi∣xi)(5)q_i = \frac{\exp z_0(b_i \mid x_i)}{\exp z_0(a_i \mid x_i) + \exp z_0(b_i \mid x_i)} \tag{5}

Retain prompts satisfying ∣qi−0.5∣≤ϵ|q_i - 0.5| \le \epsilon with ϵ=0.02\epsilon = 0.02. Also require that the public model's highest-probability token over the full vocabulary is one of the two candidates.

2. Querying the teacher. The query set is fixed before observing teacher responses. Each prompt is queried once; the teacher returns its highest-probability next token via greedy decoding over the full vocabulary. Responses are retained only if they belong to {ai,bi}\{a_i, b_i\}; otherwise discarded without resampling.

3. Training the student. The student is initialized from the same public ancestor M0M_0 and trained via full-vocabulary cross-entropy on the nn retained prompt–word pairs:

LCE(θ)=−1n∑i=1nlog⁡pθ(wiT∣xi)(6)\mathcal{L}_{\mathrm{CE}}(\theta) = -\frac{1}{n} \sum_{i=1}^{n} \log p_\theta(w_i^T \mid x_i) \tag{6}

Experimental Design

  • Primary lineage: Qwen2.5-1.5B-Instruct as public ancestor and student initialization; teacher is a frozen rank-16 code-DPO LoRA; student is a rank-16 LoRA.
  • Primary acquisition: 5,664 (prompt, word) training rows over 128 words—audited to be free of code, mathematics, benchmark, and task terms.
  • Controls: The key control is the exact nuisance-matched permutation (same carriers, completion multiset, and per-bin difficulty, but broken prompt–token pairing). Other controls include teacher-label shuffle, public-label (teacher-free), and random-marginal labels.
  • Evaluation: HumanEval+ scored by greedy pass@1 through execution; GSM8K by exact match; multiple-choice benchmarks by held-out accuracy.

Instrument Extensions

Reference-relative pair-margin (RPM) objective:

LRPM=−Eilog⁡σ(β[log⁡pθ(wiT∣xi)pθ(wˉi∣xi)−log⁡p0(wiT∣xi)p0(wˉi∣xi)])(7)\mathcal{L}_{\mathrm{RPM}} = -\mathbb{E}_i \log \sigma \left(\beta \left[ \log \frac{p_\theta(w_i^T \mid x_i)}{p_\theta(\bar{w}_i \mid x_i)} - \log \frac{p_0(w_i^T \mid x_i)}{p_0(\bar{w}_i \mid x_i)} \right] \right) \tag{7}

where wˉi\bar{w}_i is the non-selected word. This subtracts the frozen public margin to focus on the update-induced change.

Multi-token observations (K-word answers):

Lmulti=−1n∑i1K∑k=1Klog⁡pθ(yi,k∣xi,yi,<k)(8)\mathcal{L}_{\mathrm{multi}} = -\frac{1}{n} \sum_i \frac{1}{K} \sum_{k=1}^{K} \log p_\theta(y_{i,k} \mid x_i, y_{i,<k}) \tag{8}

Empirical Validation / Results

5.1 Identification of a Specific Capability (Primary Evidence)

Table 1: Controls for identification (Qwen2.5-1.5B, HumanEval+)

Control (vs. ATD-CE signal)Control scoreΔ = signal − controln
Exact nuisance-matched (primary)45.88+5.34 [1.22, 9.60]4
Global teacher-label shuffle46.19+5.03 [0.91, 9.30]4
Public-label CE (teacher-free)46.65+4.57 [0.91, 8.54]4
Random-marginal labels†45.12+6.711

†Single-seed diagnostic, no CI.

The student reaches 51.22% on HumanEval+, matching the teacher's aggregate coding score. The exact nuisance-matched control isolates the prompt–token pairing, ruling out explanations based on unigram frequency or label marginals.

5.2 Robustness

  • Across acquisitions: A factorial of 5 independently constructed acquisitions × 3 training seeds gives an aggregate effect of +4.80 pp, with all 15 signal–control pairs and all 5 acquisitions positive (between-acquisition SD = 0.83 pp).
  • Across adaptation regimes:
AdaptationSignalControlΔ (95% CI)
LoRA r1651.2245.88+5.34 [1.22, 9.60]
LoRA r3253.6647.10+6.55 [2.59, 10.82]
LoRA r6454.2745.12+9.15 [4.88, 13.87]
Full SFT50.0045.33+4.67 [1.22, 8.54]

5.3 Generalization Across Tasks and Models

Table 2: Per-task transfer across seven target tasks

Source / TargetBaseTeacherGapSignalControlΔ (95% CI)n
Code / HumanEval+43.2951.22+7.9351.2246.19+5.03 [0.91, 9.30]4
ScienceQA77.6592.13+14.4879.3876.84+2.53 [1.57, 3.52]3
OpenBookQA78.2086.40+8.2079.7377.67+2.07 [0.33, 3.87]3
HellaSwag64.8487.14+22.3166.8263.93+2.90 [2.44, 3.36]3
RACE-high76.1981.79+5.6076.7475.93+0.81 [0.24, 1.40]3
CommonsenseQA74.3779.12+4.7576.2874.34+1.94 [0.76, 3.19]3
PIQA76.9383.52+6.5877.7176.46+1.25 [0.15, 2.38]3

ATD also achieves positive mean gains on Qwen3-1.7B, Qwen3-4B, and Llama-3.2-1B (though paired 95% CIs include zero for these additional models).

5.4 Analysis of the Transferred Signal

Active vs. passive query selection:

Matched budgetΔ active – passive
Teacher queries+4.27 [0.30, 8.38]
Training rows+5.03 [1.52, 8.84]

The effect comes from concentrating queries at the ancestor's decision boundary, not from query count or data volume.

Source specificity: Across three sources (code, science, mathematics) and three target endpoints, each shadow's largest positive effect falls on its matched target (code diagonal: +5.34; science diagonal: +2.52), with off-diagonal effects near zero or negative. A 50/50 mixture of near-orthogonal math and code shadows recovers both directions simultaneously (between-source cosine = −0.0002).

Teacher–student agreement: Among 19 HumanEval+ tasks the teacher fixed over the ancestor (76 seed–task cells), ATD students solved 64.47% (49/76) vs. 21.05% (16/76) for the matched control. Task-level change vectors aligned more closely with the teacher's (cosine .701 vs. .398).

5.5 Large Teacher Gains Do Not Guarantee Transfer

Table 3b: Large-gap, non-elicitable teachers

Non-elicitable teacherGapBaseSignal
HE cheat (memorized)+39.0259.1556.56
Substitution cipher≈970.000.00

Neither transfers. This shows ATD does not simply "launder" the teacher's answers—what transfers is capability the student can be elicited to express.

Instrument Extension Results

Table 4: RPM objective vs. CE (Qwen2.5-1.5B)

TargetCE Δ_shufRPM Δ_shufCE Δ_pubRPM Δ_pubOutcome
HumanEval++5.03+6.30 [2.03, 10.98]+4.57+4.47 [−0.61, 9.76]stronger vs. shuffle
ScienceQA+2.53+3.28 [2.17, 4.41]+1.18+1.09≈ CE
PIQA+1.25+1.69 [0.60, 2.81]+0.76+0.94both positive
OpenBookQA+2.07+1.47 [−0.60, 3.53]+1.13+1.20exception

Table 5: Multi-token observation (K=20) (Qwen2.5-0.5B)

TargetSignalControlΔ (95% CI)
GSM8K21.2314.86+6.37 [+4.37, +8.42]
HumanEval+4.883.86+1.02 [−0.81, +3.05]
ScienceQA61.4160.57+0.84 [−0.09, +1.78]

Theoretical and Practical Implications

Theoretical Significance

  1. Capabilities are observable through unrelated behavior: The behavioral shadow of post-training is not just evidence of change—it carries capability-relevant information that can be extracted through near-boundary queries. This extends subliminal learning from traits/preferences to full capabilities.

  2. First-order approximation validated: The near-tie selection strategy is grounded in the expansion mi(θ0+δ)≈mi(θ0)+gi⊤δm_i(\theta_0 + \delta) \approx m_i(\theta_0) + g_i^\top \delta, and the empirical results confirm that decisions near the ancestor's decision boundary are informative about the update direction δ\delta.

  3. The channel carries structured, source-specific information: Different updates leave distinguishable shadows (source-specific diagonals in transfer matrices), and multiple shadows can be superposed (50/50 mixture recovers both sources). This suggests the shadow is a measurement of the update, not just a generic "train harder" signal.

Practical Implications

  1. Privacy: A private post-training update can be partially observed through black-box queries on unrelated inputs, without access to parameters, logits, or target-task responses. This has implications for model stewardship and API design.

  2. Distillation without task data: ATD enables capability transfer without any target-task examples—only ordinary word choices on unrelated prompts are needed. This could enable new forms of model compression or capability transfer when task data is unavailable.

  3. Diagnostic tool: The training-time trace shows a teacher-aligned change emerging before executable endpoint separation from control—the shadow may expose aspects of an update that target-task scores alone do not capture.

Boundary Conditions

  • Transfer requires the capability to be elicitable in the student (memorized answers and substitution ciphers do not transfer).
  • The method relies on actively constructed carriers and a known public ancestor.
  • Reliability varies across settings; the authors note "ATD does not yield reliable transfer in every tested setting."

Conclusion

Post-training leaves a behavioral shadow: changes in a model's decisions on inputs unrelated to the task it was trained to improve. The authors show this shadow carries information that can improve a student's target-task performance through Active Taskless Distillation (ATD)—which accesses this information through single-word responses to near-boundary prompts selected using the public ancestor.

Key takeaways:

  • In the primary coding experiment, a same-ancestor student trained on 5,664 single-word responses gains +5.34 pp on HumanEval+ over an exact nuisance-matched control, with the effect persisting across independently constructed acquisitions.
  • The shadow retains information about the source and strength of the teacher's update; observations from two teachers can jointly influence the student.
  • The effect is selective: each shadow's largest positive effect falls on its matched target task.
  • Large teacher gains do not guarantee transfer—the capability must be elicit-able in the student.

Future directions: The authors suggest better query selection and learning objectives may enable students to extract more of the shadow's information and translate it into larger target-task improvements. Open questions include whether the selected observations reveal too little capability-relevant information, or whether the learning procedure fails to make sufficient use of it. ATD is presented as "an initial exploration of how capabilities can be accessed through unrelated behavior."

Related papers