# FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders

> FuseReg samples normalized subsets of frozen encoder layers during training, regularizing representation autoencoders against cross-layer disagreement and improving reconstruction and generation without architectural changes.

- **Source:** [arXiv](https://arxiv.org/abs/2609.31620)
- **Published:** 2026-09-29
- **Permalink:** https://picx.dev/p/2L12QD
- **Whiteboard:** https://picx.dev/p/2L12QD/image

## Summary

## Summary (Overview)

- **FuseReg** introduces a layer-fusion regularization technique for Representation Autoencoders (RAEs) that trains decoders and diffusion transformers on randomly sampled subsets of frozen encoder layers, rather than a fixed heuristic fusion.
- The method **preserves the full-layer mean in expectation** while explicitly penalizing cross-layer disagreement, providing a theoretical basis for why randomized fusion targets the brittle part of the decoder–generator interface.
- A **single FuseReg decoder** reconstructs from full, sparse, and single-layer fusions without retraining, achieving higher PSNR (27.52 dB at k=23) than decoders specialized to fixed fusions.
- **Decoder replacement alone reduces unguided gFID by 27%** (3.01 → 2.21) with an unchanged RAEv2 DiT-XL generator; joint regularization of both stages reduces gFID by 29% (13.96 → 9.93) on DiT-Base.
- The approach requires **no architectural changes or additional computational cost**, working entirely through training-time regularization with separately tunable drop rates for decoder ($p_{\text{dec}}$) and generator ($p_{\text{dit}}$).

---

## Introduction and Theoretical Foundation

### Background and Motivation

Representation autoencoders (RAEs) reuse features from a pretrained visual encoder as both reconstruction and diffusion latents. A critical design choice is **which encoder layers form the shared latent space** for the generator and pixel decoder. This involves a fundamental trade-off:

- **Shallower layers**: preserve fine pixel details better (higher PSNR)
- **Deeper layers**: yield better generation metrics (lower gFID)

The paper identifies this as the **reconstruction–generation gap**. With DINOv3-L, expanding the fused subset from k=7 to k=23 increases PSNR from 22.58 to 27.04 dB but worsens unguided gFID from 1.65 to 3.01.

A key observation (Figure 1) shows that a learnable gate collapses onto a shallow layer during training, confirming that reconstruction favors shallow layers while diffusion favors deep layers.

### Theoretical Foundation

The paper formalizes layer fusion using normalized fusion weights. Given an image $x$ and frozen encoder $E$, let $h_k = \mathrm{LN}(E_{l_k}(x)) \in \mathbb{R}^{N \times d}$ be layer-normalized tokens for K chosen layers. A normalized fusion is:

$$
z_a = \sum_{k=1}^{K} a_k h_k, \qquad \mathbf{1}^{\top} a = 1,
$$

where weights $a$ are input-independent. Fixed prefix sums and learned global gates both use deterministic weight vectors; FuseReg instead samples layer subsets during training.

---

## Methodology

### Randomized Layer Fusion

FuseReg samples a layer-dropout mask for each training example. For $p \in [0, 1)$:

$$
\begin{array}{c} 
\widetilde{m}_k \stackrel{\text{iid}}{\sim} \text{Bernoulli}(1 - p), \qquad k = 1, \ldots, K, \\
m \sim \mathcal{M}_p := \text{Law}\big(\widetilde{m} \mid \mathbf{1}^{\top} \widetilde{m} > 0\big), \qquad z_m = \frac{\sum_k m_k h_k}{\sum_k m_k}.
\end{array}
$$

The normalization is essential: by linearity of expectation, $\mathbb{E}_m[z_m \mid x] = \bar{h}$, so changing $p$ varies the disagreement around the deployment latent without shifting the latent in expectation.

### Two-Stage Regularization

The method applies **separate drop rates** for the two stages:

**Decoder** ($p_{\text{dec}}$): Trains the pixel decoder to reconstruct from multiple layer compositions:
$$\mathcal{L}_{\text{dec}} = \|\hat{x} - x\|_2 + \lambda_p \mathcal{L}_{\text{LPIPS}}(\hat{x}, x) + \lambda_g \mathcal{L}_{\text{GAN}}(\hat{x})$$

**DiT** ($p_{\text{dit}}$): Trains the generator to map subset fusions to the full-layer target:
$$\mathcal{L}_{\text{DiT}} = \mathbb{E}_{(x,c), m, t, \epsilon} \rho(t) \|x_\theta(z_t, t, c) - z_{\text{full}}\|_2^2$$

where $\rho(t) = \max(1-t, 0.05)^{-2}$ and $\nu$ is the shifted logit-normal law.

---

## Empirical Validation / Results

### Theoretical Results

**Lemma 1**: Normalized subset fusion preserves the mean and isolates layer disagreement:

$$
\mathbb{E}_S[z_S \mid x] = \bar{h}, \qquad \mathbb{E}_S\big[(z_S - \bar{h})(z_S - \bar{h})^{\top} \mid x\big] = c_K(p) V(x),
$$

where $c_K(p) = \mathbb{E}\left[\frac{K-R}{R(K-1)}\right]$ is the expected sampling variance coefficient.

**Proposition 1**: Subset fusion induces stage-specific disagreement penalties. For a linear decoder:

$$
\mathcal{L}_p^{\text{dec}}(D) = \underbrace{\mathbb{E}_x\|D\bar{h} - y\|^2}_{\mathcal{L}_0^{\text{dec}}(D)} + c_K(p) \underbrace{\operatorname{tr}(D\Gamma D^{\top})}_{\Omega(D)}
$$

**Second-order separation**: The second moment of random fusion weights has rank K:
$$\mathbb{E}[w_S w_S^{\top}] = \frac{1}{K^2}\mathbf{1}\mathbf{1}^{\top} + \frac{c_K(p)}{K}P,$$
while every deterministic fusion has rank at most 1 — no fixed global fusion can match both moments.

### Experimental Results

#### One Decoder Reconstructs Across Arbitrary Fusions

**Table 1**: One FuseReg decoder matches or outperforms any deterministic fusion.

| Decoder | fusion k=7 | fusion k=7 | fusion k=7 | fusion k=23 | fusion k=23 | fusion k=23 | fusion $\ell_{11}$ | fusion $\ell_{11}$ | fusion $\ell_{11}$ |
|---------|-----------|-----------|-----------|------------|------------|------------|-------------------|-------------------|-------------------|
|         | PSNR↑     | SSIM↑     | rFID↓     | PSNR↑      | SSIM↑      | rFID↓      | PSNR↑             | SSIM↑             | rFID↓             |
| RAEv2$_{K=7}$ | 22.58 | 0.626 | 0.30 | 18.37 | 0.531 | 1.62 | 17.50 | 0.491 | 6.35 |
| RAEv2$_{K=23}$ | 12.51 | 0.369 | 16.10 | 27.04 | 0.806 | 0.18 | 14.31 | 0.447 | 11.21 |
| **FuseReg$_{p=.95}$** | **23.77** | **0.678** | **0.60** | **27.52** | **0.826** | **0.42** | **25.13** | **0.735** | **0.45** |

#### Decoder-Swap Gains

With a fixed generator, swapping the plain RAEv2 decoder for a FuseReg decoder:
- **k=23 fusion**: gFID reduced from 3.01 → 2.21 (27% improvement)
- **k=7 fusion**: gFID reduced from 27.73 → 1.92

#### Joint Regularization

**DiT-Base gFID** (Table 2a): Joint regularization reaches **9.93** at $(p_{\text{dec}}=0.9, p_{\text{dit}}=0.7)$, exceeding additive gains of single-axis changes (baseline: 13.96).

**DiT-XL gFID** (Table 2c): Strong decoder regularization alone achieves 2.38 at $p_{\text{dec}}=0.95$ (baseline: 2.91), while generator-only regularization fails to improve gFID.

**Key finding**: The benefit of a robust decoder persists across scales, while the preferred generator rate changes — motivating separate rate selection.

---

## Theoretical and Practical Implications

### Theoretical Contributions

1. **Exact characterization**: Normalized subset sampling is proven to be exactly a layer-disagreement regularizer for linear squared-error consumers, with the penalty term $\Omega(D) = K^{-1}\sum_k \mathbb{E}_x\|D(h_k - \bar{h})\|^2$ measuring sensitivity to cross-layer disagreement.

2. **Second-order separation**: No deterministic global fusion (heuristic or learned) can match the rank-K second moment structure of randomized fusion, establishing what randomization adds beyond fixed fusion selection.

3. **Stage-specific analysis**: The shared disagreement covariance acts within different prediction problems — the decoder reconstructs pixels across alternative fusions, while the DiT predicts a full-layer target along a noisy trajectory with time-dependent regularization strength $\rho(t)t^2 c_K(p)\operatorname{tr}(W\Gamma W^{\top})$.

### Practical Implications

- **Flexible deployment**: One decoder supports full, sparse, and single-layer fusions, enabling adaptive inference without retraining.
- **Improved generation**: Training for fusion robustness narrows the reconstruction–generation gap without modifying the pretrained encoder.
- **Guidance coupling**: The paper reveals that guided generation amplifies changes in the relative responses of full and REPA predictors to layer subsets, with the guided disagreement penalty:
$$\Omega(W_{\text{guided}}) = (1+w)^2\Omega(W_{\text{full}}) + w^2\Omega(W_{\text{repa}}) - 2w(1+w)\operatorname{tr}(W_{\text{full}}\Gamma W_{\text{repa}}^{\top})$$

---

## Conclusion

FuseReg treats layer fusion as a training distribution rather than a fixed configuration. By sampling normalized subsets of frozen encoder features, it preserves the full-layer latent in expectation while varying it along cross-layer disagreement directions. Key takeaways:

1. **Robustness across fusions**: A single trained decoder supports arbitrary layer fusions with strong reconstruction quality.
2. **Complementary gains**: Decoder and generator regularization provide additive improvements when rates are chosen separately.
3. **Theoretical grounding**: The method is exactly a layer-disagreement regularizer under linear models, with second-order structure unattainable by deterministic fusions.

**Future directions** include extending the analysis to mixed reconstruction losses, complete guided sampling trajectories, and transferring to other encoder hierarchies, resolutions, and data domains. The authors also note that applying FuseReg to new settings requires selecting the two rates separately due to scale- and metric-dependent dynamics.

---

_Markdown view of https://picx.dev/p/2L12QD, served by PicX — AI-generated visual whiteboard summaries of research papers._
