Summary (Overview)
- LESSER (Last-layer Error Signals for Scalable Example Ranking) is a drop-in wrapper for gradient-based data selection methods that replaces expensive full-parameter gradients with output-layer gradients, which can be computed from forward-pass quantities alone.
- The method reduces feature-extraction FLOPs by 9.7× for SFT and 3.0× for RL, with wall-clock speedups of 17.8× in SFT feature extraction and 3.5× in RL scoring time.
- Across five SFT tasks and four LLMs, LESSER achieves downstream performance within 1.3 points of full-gradient LESS, while outperforming RDS+ and Random baselines.
- In RL, LESSER matches full-gradient GradAlign's final test accuracy and shares 69% of selected problems in the first round; in distillation teacher selection, it recovers the same top teacher with Spearman correlations of 0.94–0.98.
- The key insight: despite ranking individual samples differently, LESSER and LESS select batches with aligned full-model training gradients (median centered alignment of 0.29 vs. 0.03 for RDS+), explaining the similar downstream performance.
Introduction and Theoretical Foundation
Background
The choice of post-training data substantially shapes LLM performance. Training on small, high-quality subsets tailored to a target task can match or even outperform training on larger general pools. However, identifying which subset to use is non-trivial, and selection costs can become prohibitive as candidate pools grow.
The Gradient Feature Framework
Gradient-based data selection methods (LESS, GIST, GradAlign, GRACE) map candidates and queries into a shared feature space using model gradients. The theoretical foundation rests on a first-order Taylor expansion argument:
For a candidate and query , define:
Taking one SGD step on candidate with learning rate gives , and:
Thus, when , training on is predicted to reduce the loss on query to first order.
The Cost Problem
Full-parameter gradients require an expensive backward pass through the LLM for every candidate sample. For large candidate pools, extraction can cost more than training on the selected data itself. This motivates finding cheaper features that approximate gradient behavior.
Methodology
Key Observation: Output-Layer Gradients from Forward Pass
The standard SFT and RL scoring objectives share a weighted token-level form:
For this common form, the output-layer gradient can be written explicitly using forward-pass quantities:
Proposition 1. Let be the final vocabulary projection, the hidden state predicting , and the one-hot vector for . Let be a copy of used only as the readout. Then:
Since both and are available from the forward pass, can be computed without backpropagating through the model. The method uses as the feature for each candidate and query, optionally projected to lower dimensions for efficient storage.
Implementation
LESSER is implemented as a drop-in replacement for gradient features across three settings:
| Setting | Original Method | LESSER Replacement |
|---|---|---|
| SFT | LESS (LoRA gradients at 4 checkpoints) | Output-layer gradients |
| RL | GradAlign (policy-gradient alignment) | Output-layer gradients |
| Distillation | GRACE (teacher selection) | Output-layer gradients |
Empirical Validation / Results
SFT Results
Setup: 197,196 Tulu V2 instruction–response samples as candidate pool; TyDiQA, MMLU-Pro, GSM8K, Codex, BBH as query/test sets; selection sizes ; models: Llama-2-7B, Llama-3.2-3B, Qwen3-4B-Base, OLMo3-7B.
Finding 1: LESSER is competitive with LESS on downstream SFT performance. Across 60 model–task–budget cells, the absolute score gap averages:
- 1.3 points vs. LESS
- 2.2 points vs. RDS+
- 2.6 points vs. Random
Finding 2: Higher LESSER similarity predicts lower query loss, with mean Spearman correlation of 0.86, close to LESS and far above RDS+.
RL Results
Finding 3 (Corrupted rewards detection): LESSER selects 29.7% corrupted problems vs. 27.3% for full gradients, well below Random's baseline.
Finding 4 (Online RL training): LESSER closely tracks full-gradient GradAlign, both reaching the same final held-out Countdown accuracy. LESSER reaches Random's final accuracy after only 15 steps, using 70% fewer GRPO steps.
Distillation Results
Finding 5: Teacher rankings remain closely aligned with Spearman correlations of 0.94–0.98, and LESSER recovers the same top-ranked teacher as full-gradient GRACE across all settings.
Cost Analysis
Table 1: Output-layer gradients are much cheaper to extract (one A100 80GB).
| SFT Selector | Time (h) | Speedup | RL Selector | Added time (min) | Speedup | Feature size |
|---|---|---|---|---|---|---|
| LESSER | 3.16 | 17.8× | LESSER | 18.1 | 3.5× | 32.8 kB |
| LESS | 56.16 | 1.0× | Full gradients | 63.3 | 1.0× | 6.2 GB |
Feature extraction accounts for 97–99.7% of LESS's combined extraction and fine-tuning FLOPs across tested budgets.
Theoretical and Practical Implications
Why Output-Layer Gradients Work
The alignment explanation: Despite selecting largely different data in SFT (Jaccard similarity of only 0.037 at ), LESSER and LESS produce aligned training directions:
- Centered batch-gradient alignment: Median alignment of 0.29 between LESSER and LESS batches, vs. 0.03 for RDS+ and ~0 for Random (reference: 0.53 between different LESS batches).
- Early training effects: By step 10, LESSER and LESS achieve 84% and 82% of their maximum query-loss decrease, vs. 35% for Random.
Broader Principles
-
Output-layer gradients preserve more signal than hidden-state features at similar extraction cost, while being more directly tied to the optimization objective.
-
Agreement is strongest at the level each pipeline acts on: SFT trains on batch-aggregated gradients, RL selects problems after aggregating sampled responses, and GRACE ranks teachers using statistics over many responses. Efficient selection features need not recover per-example full gradients—they only need to preserve the aggregate geometry each pipeline uses.
-
Full-gradient fidelity may be the wrong objective for evaluating selection proxies. While fine-grained approximation may be important for influence estimation or data attribution, selection from realistic, already-curated LLM data pools may only require preserving which examples or subsets are useful.
Limitations
- Experiments do not cover all post-training regimes or selection objectives.
- Focus is primarily on near-initialization training effects; understanding of how output-layer gradient signals evolve over longer training trajectories is left to future work.
- SFT comparison is between complete pipelines rather than a matched-checkpoint ablation, so LESSER and LESS differ in more than just their features (LESSER needs neither warmup training nor backpropagation).
Conclusion
LESSER demonstrates that output-layer gradients—computable from forward-pass quantities alone—serve as a cost-efficient substitute for full-gradient features in LLM post-training data selection. Across SFT, RL, and distillation teacher selection:
- In RL and GRACE: Selection statistics largely agree with full-gradient methods.
- In SFT: Different rankings nevertheless induce aligned batch updates, explaining the matched downstream performance.
Output-layer gradients offer a simple, efficient alternative to full gradients, reducing computational cost by an order of magnitude while maintaining competitive downstream performance.
Future directions:
- Develop diagnostics to predict when output-layer gradients will recover full-gradient rankings or update directions.
- Identify which internal computations enable this agreement (tracing selection-relevant information across layers).
- Develop stronger theory characterizing when such agreement emerges under realistic, structured gradient distributions.
Related papers
- CATCH: A Controllable Analysis Testbed for Reward Hacking in Coding RL
CATCH is a controllable coding-RL testbed revealing that chain-of-thought monitors suppress reward hacking initially but erode as policies learn to mislead them with code comments.
- Who Verifies the Verifier? Co-Evolving Inspectable Graders with Self-Improving Agents
Verifier evolution lets agents self-improve without ground truth, but only anchor discipline—not detector lifecycle—prevents collapse into vacuous always-pass grading.
- Do Tool Calls Execute as Intended? Measuring and Repairing Intent-Execution Correspondence in LLM Agents
Tool calls are silently altered by execution-path hops in 12% of production shell invocations, and the IntAct repair protocol recovers 79.2% of resulting failures.