# LESSER: Post-Training Data Selection with Output-Layer Gradients

> LESSER replaces expensive full-parameter gradients with output-layer gradients computed from forward passes, cutting feature-extraction FLOPs by up to 9.7× while matching full-gradient data selection performance within 1.3 points.

- **Source:** [arXiv](https://arxiv.org/abs/2610.03702)
- **Published:** 2026-10-10
- **Permalink:** https://picx.dev/p/qreWDg
- **Whiteboard:** https://picx.dev/p/qreWDg/image

## Summary

## Summary (Overview)

- **LESSER** (Last-layer Error Signals for Scalable Example Ranking) is a drop-in wrapper for gradient-based data selection methods that replaces expensive full-parameter gradients with output-layer gradients, which can be computed from forward-pass quantities alone.
- The method reduces feature-extraction FLOPs by **9.7× for SFT** and **3.0× for RL**, with wall-clock speedups of **17.8×** in SFT feature extraction and **3.5×** in RL scoring time.
- Across five SFT tasks and four LLMs, LESSER achieves downstream performance within **1.3 points** of full-gradient LESS, while outperforming RDS+ and Random baselines.
- In RL, LESSER matches full-gradient GradAlign's final test accuracy and shares **69%** of selected problems in the first round; in distillation teacher selection, it recovers the same top teacher with Spearman correlations of **0.94–0.98**.
- The key insight: despite ranking individual samples differently, LESSER and LESS select batches with aligned full-model training gradients (median centered alignment of 0.29 vs. 0.03 for RDS+), explaining the similar downstream performance.

---

## Introduction and Theoretical Foundation

### Background

The choice of post-training data substantially shapes LLM performance. Training on small, high-quality subsets tailored to a target task can match or even outperform training on larger general pools. However, identifying which subset to use is non-trivial, and selection costs can become prohibitive as candidate pools grow.

### The Gradient Feature Framework

Gradient-based data selection methods (LESS, GIST, GradAlign, GRACE) map candidates and queries into a shared feature space using model gradients. The theoretical foundation rests on a first-order Taylor expansion argument:

For a candidate $p \in \mathcal{D}$ and query $q \in \mathcal{Q}$, define:

$$G_p = \nabla_\theta L_p(\theta), \qquad G_q = \nabla_\theta L_q(\theta)$$

Taking one SGD step on candidate $p$ with learning rate $\eta > 0$ gives $\theta' = \theta - \eta G_p$, and:

$$L_q(\theta') - L_q(\theta) \approx \nabla_\theta L_q(\theta)^\top (\theta' - \theta) = -\eta G_q^\top G_p$$

Thus, when $G_q^\top G_p > 0$, training on $p$ is predicted to reduce the loss on query $q$ to first order.

### The Cost Problem

Full-parameter gradients require an expensive backward pass through the LLM for every candidate sample. For large candidate pools, extraction can cost more than training on the selected data itself. This motivates finding cheaper features that approximate gradient behavior.

---

## Methodology

### Key Observation: Output-Layer Gradients from Forward Pass

The standard SFT and RL scoring objectives share a weighted token-level form:

$$L_p(\theta) = -\sum_t \alpha_t \log \pi_\theta(y_t \mid x, y_{<t}) \tag{1}$$

For this common form, the output-layer gradient can be written explicitly using forward-pass quantities:

**Proposition 1.** Let $W$ be the final vocabulary projection, $h_t$ the hidden state predicting $y_t$, and $e_{y_t}$ the one-hot vector for $y_t$. Let $W_{\mathrm{out}}$ be a copy of $W$ used only as the readout. Then:

$$g_p := \nabla_{W_{\mathrm{out}}} L_p(\theta) = \sum_t \alpha_t \bigl(\operatorname{softmax}(W h_t) - e_{y_t}\bigr) h_t^\top$$

Since both $h_t$ and $\operatorname{softmax}(W h_t)$ are available from the forward pass, $g_p$ can be computed **without backpropagating through the model**. The method uses $\text{vec}(g_p)$ as the feature for each candidate and query, optionally projected to lower dimensions for efficient storage.

### Implementation

LESSER is implemented as a drop-in replacement for gradient features across three settings:

| Setting | Original Method | LESSER Replacement |
|---------|----------------|-------------------|
| SFT | LESS (LoRA gradients at 4 checkpoints) | Output-layer gradients |
| RL | GradAlign (policy-gradient alignment) | Output-layer gradients |
| Distillation | GRACE (teacher selection) | Output-layer gradients |

---

## Empirical Validation / Results

### SFT Results

**Setup:** 197,196 Tulu V2 instruction–response samples as candidate pool; TyDiQA, MMLU-Pro, GSM8K, Codex, BBH as query/test sets; selection sizes $k \in \{1000, 5000, 10000\}$; models: Llama-2-7B, Llama-3.2-3B, Qwen3-4B-Base, OLMo3-7B.

**Finding 1:** LESSER is competitive with LESS on downstream SFT performance. Across 60 model–task–budget cells, the absolute score gap averages:
- **1.3 points** vs. LESS
- **2.2 points** vs. RDS+
- **2.6 points** vs. Random

**Finding 2:** Higher LESSER similarity predicts lower query loss, with mean Spearman correlation of **0.86**, close to LESS and far above RDS+.

### RL Results

**Finding 3 (Corrupted rewards detection):** LESSER selects **29.7%** corrupted problems vs. **27.3%** for full gradients, well below Random's baseline.

**Finding 4 (Online RL training):** LESSER closely tracks full-gradient GradAlign, both reaching the same final held-out Countdown accuracy. LESSER reaches Random's final accuracy after only 15 steps, using **70% fewer GRPO steps**.

### Distillation Results

**Finding 5:** Teacher rankings remain closely aligned with Spearman correlations of **0.94–0.98**, and LESSER recovers the same top-ranked teacher as full-gradient GRACE across all settings.

### Cost Analysis

**Table 1:** Output-layer gradients are much cheaper to extract (one A100 80GB).

| SFT Selector | Time (h) | Speedup | RL Selector | Added time (min) | Speedup | Feature size |
|--------------|----------|---------|-------------|------------------|---------|--------------|
| LESSER | 3.16 | 17.8× | LESSER | 18.1 | 3.5× | 32.8 kB |
| LESS | 56.16 | 1.0× | Full gradients | 63.3 | 1.0× | 6.2 GB |

Feature extraction accounts for **97–99.7%** of LESS's combined extraction and fine-tuning FLOPs across tested budgets.

---

## Theoretical and Practical Implications

### Why Output-Layer Gradients Work

**The alignment explanation:** Despite selecting largely different data in SFT (Jaccard similarity of only 0.037 at $k=5000$), LESSER and LESS produce aligned training directions:

- **Centered batch-gradient alignment:** Median alignment of 0.29 between LESSER and LESS batches, vs. 0.03 for RDS+ and ~0 for Random (reference: 0.53 between different LESS batches).
- **Early training effects:** By step 10, LESSER and LESS achieve 84% and 82% of their maximum query-loss decrease, vs. 35% for Random.

### Broader Principles

1. **Output-layer gradients preserve more signal than hidden-state features** at similar extraction cost, while being more directly tied to the optimization objective.

2. **Agreement is strongest at the level each pipeline acts on:** SFT trains on batch-aggregated gradients, RL selects problems after aggregating sampled responses, and GRACE ranks teachers using statistics over many responses. Efficient selection features need not recover per-example full gradients—they only need to preserve the aggregate geometry each pipeline uses.

3. **Full-gradient fidelity may be the wrong objective** for evaluating selection proxies. While fine-grained approximation may be important for influence estimation or data attribution, selection from realistic, already-curated LLM data pools may only require preserving which examples or subsets are useful.

### Limitations

- Experiments do not cover all post-training regimes or selection objectives.
- Focus is primarily on near-initialization training effects; understanding of how output-layer gradient signals evolve over longer training trajectories is left to future work.
- SFT comparison is between complete pipelines rather than a matched-checkpoint ablation, so LESSER and LESS differ in more than just their features (LESSER needs neither warmup training nor backpropagation).

---

## Conclusion

LESSER demonstrates that output-layer gradients—computable from forward-pass quantities alone—serve as a cost-efficient substitute for full-gradient features in LLM post-training data selection. Across SFT, RL, and distillation teacher selection:

- **In RL and GRACE:** Selection statistics largely agree with full-gradient methods.
- **In SFT:** Different rankings nevertheless induce aligned batch updates, explaining the matched downstream performance.

Output-layer gradients offer a simple, efficient alternative to full gradients, reducing computational cost by an order of magnitude while maintaining competitive downstream performance.

**Future directions:**
1. Develop diagnostics to predict when output-layer gradients will recover full-gradient rankings or update directions.
2. Identify which internal computations enable this agreement (tracing selection-relevant information across layers).
3. Develop stronger theory characterizing when such agreement emerges under realistic, structured gradient distributions.

---

_Markdown view of https://picx.dev/p/qreWDg, served by PicX — AI-generated visual whiteboard summaries of research papers._
