Post-Training Leaves Behavioral Shadows on Unrelated Decisions
Summary (Overview)
- Core finding: Post-training updates on language models create "behavioral shadows"—changes in decisions on inputs unrelated to the training task. These shadows carry capability-relevant information that can be extracted and transferred to another model.
- Key contribution: The authors introduce Active Taskless Distillation (ATD), a method that achieves capability transfer using only a single word from the teacher per prompt, without target-task examples, teacher logits, or teacher parameters.
- Primary result: On HumanEval+ (code generation), a student trained on 5,664 single-word teacher responses gains +5.34 percentage points over an exact nuisance-matched control, with the effect persisting across independently constructed acquisitions.
- Generality: Transfer is demonstrated across seven target tasks (code, scientific knowledge, commonsense reasoning, reading comprehension) and across model generations, sizes, and families (Qwen2.5, Qwen3, Llama).
- Boundary condition: Large teacher gains do not guarantee transfer—the transferred capability must be one the student can be elicited to express (e.g., memorized answers and substitution ciphers do not transfer).
Introduction and Theoretical Foundation
Background and Motivation
Post-training (e.g., fine-tuning, DPO) is typically understood through changes in target-task behavior—a coding update makes a model write better code. However, the influence of an update may extend beyond the trained task: it may change which ordinary word a model prefers in a short story, even when neither the story nor the answer contains code or mathematics. These changes on task-unrelated inputs constitute the behavioral shadow of post-training.
Prior work on subliminal learning showed that behavioral traits can pass through semantically unrelated training data, but largely focused on traits/preferences using extensive teacher outputs. ATD asks two sharper questions:
- Observability: Can a private post-training update be observed through individual decisions on unrelated inputs?
- Capability transfer: Can learning from these decisions transfer part of the teacher's target-task improvement?
Theoretical Foundation
The theoretical motivation rests on a first-order expansion around the public ancestor. For a carrier prompt with two candidate words and , define the logit difference:
Expanding around the ancestor parameters :
where is the unknown post-training update and is the gradient of the margin with respect to parameters. The observed binary choice is modeled as:
Key insight: Near a tie (where is small), even a small preference change induced by the update can reverse the selected word. Each observed choice thus provides a one-bit observation of the update's behavioral shadow. The teacher parameters are:
Methodology
Active Taskless Distillation (ATD)
ATD consists of three stages:
1. Selecting near-tie prompts. Using only the public ancestor , compute the probability of candidate after normalizing over the pair:
Retain prompts satisfying with . Also require that the public model's highest-probability token over the full vocabulary is one of the two candidates.
2. Querying the teacher. The query set is fixed before observing teacher responses. Each prompt is queried once; the teacher returns its highest-probability next token via greedy decoding over the full vocabulary. Responses are retained only if they belong to ; otherwise discarded without resampling.
3. Training the student. The student is initialized from the same public ancestor and trained via full-vocabulary cross-entropy on the retained prompt–word pairs:
Experimental Design
- Primary lineage: Qwen2.5-1.5B-Instruct as public ancestor and student initialization; teacher is a frozen rank-16 code-DPO LoRA; student is a rank-16 LoRA.
- Primary acquisition: 5,664 (prompt, word) training rows over 128 words—audited to be free of code, mathematics, benchmark, and task terms.
- Controls: The key control is the exact nuisance-matched permutation (same carriers, completion multiset, and per-bin difficulty, but broken prompt–token pairing). Other controls include teacher-label shuffle, public-label (teacher-free), and random-marginal labels.
- Evaluation: HumanEval+ scored by greedy pass@1 through execution; GSM8K by exact match; multiple-choice benchmarks by held-out accuracy.
Instrument Extensions
Reference-relative pair-margin (RPM) objective:
where is the non-selected word. This subtracts the frozen public margin to focus on the update-induced change.
Multi-token observations (K-word answers):
Empirical Validation / Results
5.1 Identification of a Specific Capability (Primary Evidence)
Table 1: Controls for identification (Qwen2.5-1.5B, HumanEval+)
| Control (vs. ATD-CE signal) | Control score | Δ = signal − control | n |
|---|---|---|---|
| Exact nuisance-matched (primary) | 45.88 | +5.34 [1.22, 9.60] | 4 |
| Global teacher-label shuffle | 46.19 | +5.03 [0.91, 9.30] | 4 |
| Public-label CE (teacher-free) | 46.65 | +4.57 [0.91, 8.54] | 4 |
| Random-marginal labels† | 45.12 | +6.71 | 1 |
†Single-seed diagnostic, no CI.
The student reaches 51.22% on HumanEval+, matching the teacher's aggregate coding score. The exact nuisance-matched control isolates the prompt–token pairing, ruling out explanations based on unigram frequency or label marginals.
5.2 Robustness
- Across acquisitions: A factorial of 5 independently constructed acquisitions × 3 training seeds gives an aggregate effect of +4.80 pp, with all 15 signal–control pairs and all 5 acquisitions positive (between-acquisition SD = 0.83 pp).
- Across adaptation regimes:
| Adaptation | Signal | Control | Δ (95% CI) |
|---|---|---|---|
| LoRA r16 | 51.22 | 45.88 | +5.34 [1.22, 9.60] |
| LoRA r32 | 53.66 | 47.10 | +6.55 [2.59, 10.82] |
| LoRA r64 | 54.27 | 45.12 | +9.15 [4.88, 13.87] |
| Full SFT | 50.00 | 45.33 | +4.67 [1.22, 8.54] |
5.3 Generalization Across Tasks and Models
Table 2: Per-task transfer across seven target tasks
| Source / Target | Base | Teacher | Gap | Signal | Control | Δ (95% CI) | n |
|---|---|---|---|---|---|---|---|
| Code / HumanEval+ | 43.29 | 51.22 | +7.93 | 51.22 | 46.19 | +5.03 [0.91, 9.30] | 4 |
| ScienceQA | 77.65 | 92.13 | +14.48 | 79.38 | 76.84 | +2.53 [1.57, 3.52] | 3 |
| OpenBookQA | 78.20 | 86.40 | +8.20 | 79.73 | 77.67 | +2.07 [0.33, 3.87] | 3 |
| HellaSwag | 64.84 | 87.14 | +22.31 | 66.82 | 63.93 | +2.90 [2.44, 3.36] | 3 |
| RACE-high | 76.19 | 81.79 | +5.60 | 76.74 | 75.93 | +0.81 [0.24, 1.40] | 3 |
| CommonsenseQA | 74.37 | 79.12 | +4.75 | 76.28 | 74.34 | +1.94 [0.76, 3.19] | 3 |
| PIQA | 76.93 | 83.52 | +6.58 | 77.71 | 76.46 | +1.25 [0.15, 2.38] | 3 |
ATD also achieves positive mean gains on Qwen3-1.7B, Qwen3-4B, and Llama-3.2-1B (though paired 95% CIs include zero for these additional models).
5.4 Analysis of the Transferred Signal
Active vs. passive query selection:
| Matched budget | Δ active – passive |
|---|---|
| Teacher queries | +4.27 [0.30, 8.38] |
| Training rows | +5.03 [1.52, 8.84] |
The effect comes from concentrating queries at the ancestor's decision boundary, not from query count or data volume.
Source specificity: Across three sources (code, science, mathematics) and three target endpoints, each shadow's largest positive effect falls on its matched target (code diagonal: +5.34; science diagonal: +2.52), with off-diagonal effects near zero or negative. A 50/50 mixture of near-orthogonal math and code shadows recovers both directions simultaneously (between-source cosine = −0.0002).
Teacher–student agreement: Among 19 HumanEval+ tasks the teacher fixed over the ancestor (76 seed–task cells), ATD students solved 64.47% (49/76) vs. 21.05% (16/76) for the matched control. Task-level change vectors aligned more closely with the teacher's (cosine .701 vs. .398).
5.5 Large Teacher Gains Do Not Guarantee Transfer
Table 3b: Large-gap, non-elicitable teachers
| Non-elicitable teacher | Gap | Base | Signal |
|---|---|---|---|
| HE cheat (memorized) | +39.02 | 59.15 | 56.56 |
| Substitution cipher | ≈97 | 0.00 | 0.00 |
Neither transfers. This shows ATD does not simply "launder" the teacher's answers—what transfers is capability the student can be elicited to express.
Instrument Extension Results
Table 4: RPM objective vs. CE (Qwen2.5-1.5B)
| Target | CE Δ_shuf | RPM Δ_shuf | CE Δ_pub | RPM Δ_pub | Outcome |
|---|---|---|---|---|---|
| HumanEval+ | +5.03 | +6.30 [2.03, 10.98] | +4.57 | +4.47 [−0.61, 9.76] | stronger vs. shuffle |
| ScienceQA | +2.53 | +3.28 [2.17, 4.41] | +1.18 | +1.09 | ≈ CE |
| PIQA | +1.25 | +1.69 [0.60, 2.81] | +0.76 | +0.94 | both positive |
| OpenBookQA | +2.07 | +1.47 [−0.60, 3.53] | +1.13 | +1.20 | exception |
Table 5: Multi-token observation (K=20) (Qwen2.5-0.5B)
| Target | Signal | Control | Δ (95% CI) |
|---|---|---|---|
| GSM8K | 21.23 | 14.86 | +6.37 [+4.37, +8.42] |
| HumanEval+ | 4.88 | 3.86 | +1.02 [−0.81, +3.05] |
| ScienceQA | 61.41 | 60.57 | +0.84 [−0.09, +1.78] |
Theoretical and Practical Implications
Theoretical Significance
-
Capabilities are observable through unrelated behavior: The behavioral shadow of post-training is not just evidence of change—it carries capability-relevant information that can be extracted through near-boundary queries. This extends subliminal learning from traits/preferences to full capabilities.
-
First-order approximation validated: The near-tie selection strategy is grounded in the expansion , and the empirical results confirm that decisions near the ancestor's decision boundary are informative about the update direction .
-
The channel carries structured, source-specific information: Different updates leave distinguishable shadows (source-specific diagonals in transfer matrices), and multiple shadows can be superposed (50/50 mixture recovers both sources). This suggests the shadow is a measurement of the update, not just a generic "train harder" signal.
Practical Implications
-
Privacy: A private post-training update can be partially observed through black-box queries on unrelated inputs, without access to parameters, logits, or target-task responses. This has implications for model stewardship and API design.
-
Distillation without task data: ATD enables capability transfer without any target-task examples—only ordinary word choices on unrelated prompts are needed. This could enable new forms of model compression or capability transfer when task data is unavailable.
-
Diagnostic tool: The training-time trace shows a teacher-aligned change emerging before executable endpoint separation from control—the shadow may expose aspects of an update that target-task scores alone do not capture.
Boundary Conditions
- Transfer requires the capability to be elicitable in the student (memorized answers and substitution ciphers do not transfer).
- The method relies on actively constructed carriers and a known public ancestor.
- Reliability varies across settings; the authors note "ATD does not yield reliable transfer in every tested setting."
Conclusion
Post-training leaves a behavioral shadow: changes in a model's decisions on inputs unrelated to the task it was trained to improve. The authors show this shadow carries information that can improve a student's target-task performance through Active Taskless Distillation (ATD)—which accesses this information through single-word responses to near-boundary prompts selected using the public ancestor.
Key takeaways:
- In the primary coding experiment, a same-ancestor student trained on 5,664 single-word responses gains +5.34 pp on HumanEval+ over an exact nuisance-matched control, with the effect persisting across independently constructed acquisitions.
- The shadow retains information about the source and strength of the teacher's update; observations from two teachers can jointly influence the student.
- The effect is selective: each shadow's largest positive effect falls on its matched target task.
- Large teacher gains do not guarantee transfer—the capability must be elicit-able in the student.
Future directions: The authors suggest better query selection and learning objectives may enable students to extract more of the shadow's information and translate it into larger target-task improvements. Open questions include whether the selected observations reveal too little capability-relevant information, or whether the learning procedure fails to make sufficient use of it. ATD is presented as "an initial exploration of how capabilities can be accessed through unrelated behavior."
Related papers
- On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability
Qwen3.8-Flash-Next matches its 397B predecessor's quality using one-third activated parameters and one-ninth training FLOPs via GDN, QSA, GR, and n-gram embeddings.
- Language Models Can Control Their Own Attention
Declarative Attention lets off-the-shelf LLMs declare their own sparse attention scope via text tags, cutting attended tokens by up to 52% with minimal accuracy loss.
- Hill Sampling for Test-Time Scaling: A Simple and Better Alternative to Repeated Sampling, Evolution, and Training
Hill Sampling, repeatedly sampling edits to the best program found so far, outperforms all complex evolutionary and weight-space methods, setting new state-of-the-art results on circle packing.