# How Far Are We from Removing the Visual Encoder? Scaling Laws for Encoder-Free Multimodal Pretraining

> Encoder-free multimodal LLMs match encoder-based performance at ~10^22 FLOPs, shifting compute-optimal allocation toward larger decoders and enabling viable encoder-free pretraining.

- **Source:** [arXiv](https://arxiv.org/abs/2609.35457)
- **Published:** 2026-10-03
- **Permalink:** https://picx.dev/p/04O53m
- **Whiteboard:** https://picx.dev/p/04O53m/image

## Summary

## Summary (Overview)

- This paper presents the first systematic scaling-law comparison between encoder-free and encoder-based multimodal large language models (MLLMs), using a matched ladder of 11 sparse MoE language models (1.1B–44B total parameters).
- **Key finding 1**: Removing the visual encoder shifts the compute-optimal allocation for the multimodal objective toward larger models (model allocation exponent increases from $a = 0.464$ to $a = 0.570$), while text allocation remains nearly unchanged.
- **Key finding 2**: Encoder-free models underperform at small scales but are predicted to catch up with encoder-based models at approximately $10^{22}$ FLOPs under compute-optimal allocation—well within practical pretraining budgets (e.g., Kimi K2.5 uses ~$10^{25}$ FLOPs).
- **Key finding 3**: The decoder compensates for the missing visual encoder through vision-specific adaptations: bidirectional attention among visual tokens becomes increasingly beneficial, visual processing shifts to earlier layers, and expert routing for visual tokens becomes more concentrated.
- The crossover varies by topic, arriving earlier on language-heavy topics (STEM) and much later on perception-intensive ones (Caption, OCR, GUI).

## Introduction and Theoretical Foundation

Most modern MLLMs build on a pretrained visual encoder (e.g., SigLIP 2 ViT) that provides a strong visual prior learned from large-scale image–text data. Encoder-free MLLMs instead remove the visual encoder and feed projected image patches directly into the decoder, which must learn visual representations from raw pixels. While prior work has shown initial feasibility of encoder-free designs, their scaling behavior has not been systematically characterized.

The paper addresses this gap through a controlled scaling study where both model families share:
- The same sparse decoder ladder (MoE architecture)
- The same data mixture and optimization setup
- The same visual-token granularity

The theoretical foundation draws on established scaling-law methodology:
- **Compute-optimal allocation law**: $$M_{\text{opt}}(C) \propto C^a, \quad D_{\text{opt}}(C) \propto C^b, \quad a + b = 1$$
- **Compute-optimal frontier**: $$L^*(C) = E + KC^{-\gamma}$$
  where $\gamma > 0$ is the loss–compute exponent, $K > 0$ is a fitted prefactor, and $E$ is the entropy floor induced by the data distribution.
- **IsoFLOP profiles**: Following Chinchilla methodology, validation loss is fitted as a quadratic function of $\log M$ at fixed compute budgets.

## Methodology

### Model Architecture Comparison

The paper compares two architectures on a matched ladder of 11 sparse MoE language models:

- **Encoder-based**: Uses a pretrained SigLIP 2 ViT followed by a ConvPool adapter and projector, with causal attention applied to all tokens. The ViT is trained jointly with the decoder.
- **Encoder-free**: Maps raw image patches directly into the decoder through a patch projection, where visual tokens attend bidirectionally within each image, with all other attention remaining causal. A fully causal variant is also studied.

### Compute Accounting

Decoder FLOPs per token are decomposed as:
$$M_o^{(s)} = M_{\text{base}} + M_{\text{attn}}^{(s)} \ell_o, \quad C_o^{(s)} = M_o^{(s)} D_o$$

where $s \in \{\text{free}, \text{based}\}$ indexes the model family, $o \in \{\text{text}, \text{mm}\}$ indexes the objective, and $\ell_o$ is the mean packed sequence length. The fixed expert activation ratio of 8/256 keeps $M$ approximately proportional to active parameters.

### Efficiency Gain Metrics

Following MAI-Thinking-1, the paper defines two efficiency metrics at equal loss $\lambda$:

- **Compute efficiency gain**: $$\text{EG}^C_{\text{tar} \leftarrow \text{ref}}(\lambda) = C_{\text{ref}}(\lambda)/C_{\text{tar}}(\lambda)$$
- **Model efficiency gain**: $$\text{EG}^M_{\text{tar} \leftarrow \text{ref}}(\lambda) = M_{\text{ref}}(\lambda)/M_{\text{tar}}(\lambda)$$

Values above 1.0 indicate the target is more efficient than the reference.

### Overtraining Model

For overtraining analysis, the paper uses a separable loss form $L(M, D) = E + AM^{-\alpha} + BD^{-\beta}$ and derives that overtraining rescales only the prefactor:
$$L(C_{\text{base}}, k) = E + g(k)KC_{\text{base}}^{-\gamma}$$
where $k$ is the overtraining factor and $g(k)$ depends only on $k$, not on $C_{\text{base}}$.

## Empirical Validation / Results

### Compute-Optimal Allocation

**Table 1: Model allocation exponent $a$ of encoder-free models under bidirectional and causal attention over visual tokens**

| Attention | Text | Multimodal |
|-----------|------|------------|
| Bidirectional | 0.427 | 0.570 |
| Causal | 0.436 | 0.557 |

- **Text objective**: Nearly identical allocation exponents ($a = 0.427$ for encoder-free, $a = 0.422$ for encoder-based)
- **Multimodal objective**: Removing the encoder increases $a$ from 0.464 to 0.570, with bootstrap 80% intervals of [0.546, 0.595] and [0.458, 0.472] respectively

### Loss–Compute Frontiers

The fitted scaling laws show:

| Objective | Encoder-free exponent | Encoder-based exponent |
|-----------|----------------------|----------------------|
| Text | $\gamma = 0.0973$ | $\gamma = 0.0979$ |
| Multimodal | $\gamma = 0.3778$ | $\gamma = 0.2998$ |

- **Text**: Frontiers nearly overlap, with $\text{EG}^C \approx 0.98$ and $\text{EG}^M \approx 0.99$
- **Multimodal**: Encoder-free models require more compute initially, but the gap narrows with scale
  - $\text{EG}^C$ increases from ~0.47 (observed mean) toward 1.0
  - $\text{EG}^M \approx 0.80$, meaning encoder-free models use ~1.25× the FLOPs per token at equal loss
  - **Crossover predicted at $6.1 \times 10^{21}$ FLOPs** (80% interval: $[4.2 \times 10^{21}, 1.0 \times 10^{22}]$)

### Overtraining Effects

At $k = 5$ overtraining factor:
- Crossover arrives later: $1.2 \times 10^{22}$ FLOPs (80% interval: $[8.4 \times 10^{21}, 2.0 \times 10^{22}]$)
- $\text{EG}^C$ drops from 0.62 to 0.52 at the largest fitted budget
- $\text{EG}^M$ drops from 0.80 to 0.74
- Text objective remains nearly unaffected (changes < 1%)

### Topic-Specific Analysis

Compute efficiency gains vary by topic, with crossover ordering:

| Topic | EG at $k=1$ (fitted range) | Crossover timing |
|-------|---------------------------|------------------|
| STEM | 0.71 → 0.95 | Earliest (near parity at largest budget) |
| Charts | 0.28 → 0.57 | Early |
| GUI | 0.51 → 0.69 | Later |
| OCR | 0.18 → 0.34 | Later |
| Caption | 0.34 → 0.50 | Latest |

### Vision-Specific Adaptation Mechanisms

1. **Emergence of visual encoding**: Encoder-free models show a slow initial phase followed by a sharp loss drop. At layer 12, visual attention mass rises from 0.217 to 0.645 after the drop, approaching the encoder-based model's level.

2. **Bidirectional attention**: Causal attention is ~1% better on text ($\text{EG}^C = 1.011$) but worse on multimodal ($\text{EG}^C = 0.990$), with bidirectional attention's multimodal advantage growing with compute.

3. **Layerwise representation evolution**: Visual tokens in encoder-free models diverge from their layer-0 inputs much earlier than in encoder-based models, while text token trajectories remain nearly identical between architectures.

4. **Expert routing concentration**: MaxVio (expert load imbalance) for visual tokens is consistently higher in encoder-free models across the four largest model sizes, while text token routing remains closely matched between architectures.

## Theoretical and Practical Implications

### Theoretical Implications

1. **Scaling law divergence**: The different allocation exponents ($a = 0.570$ vs. $a = 0.464$) indicate that encoder-free models have fundamentally different compute-optimal tradeoffs, requiring more decoder capacity to jointly support visual representation learning and language modeling.

2. **Diminishing value of visual priors**: The predicted crossover suggests that the advantage of pretrained visual encoders diminishes with scale—encoder-free models can learn equally good visual representations from raw pixels given sufficient compute.

3. **Architecture-specific adaptation**: The decoder's vision-specific adaptations (bidirectional attention, early-layer processing, expert concentration) suggest that encoder-free architectures may benefit from decoder designs explicitly optimized for native visual representation learning rather than inheriting designs built for language.

### Practical Implications

1. **Pretraining budget guidance**: The crossover at ~$10^{22}$ FLOPs is well within reach of current flagship pretraining runs (~$10^{25}$ FLOPs), suggesting encoder-free architectures are a viable direction for large-scale multimodal pretraining.

2. **Topic-dependent deployment**: The varying crossover points by topic suggest that encoder-free models may be suitable for language-heavy multimodal tasks earlier, while perception-intensive applications may require larger budgets.

3. **Architecture design**: The finding that encoder-free models favor larger decoders suggests that scaling model size (rather than data) is more effective for encoder-free multimodal pretraining.

## Conclusion

The paper demonstrates that while encoder-free MLLMs are less compute-efficient at small scales, their more rapidly improving multimodal loss frontier predicts an efficiency crossover within practical pretraining budgets (~$10^{22}$ FLOPs). The shift toward larger models under compute-optimal allocation, combined with the decoder's vision-specific adaptations (bidirectional visual attention, early-layer visual processing, concentrated expert routing), indicates that encoder-free architectures are a promising direction for multimodal pretraining. The text objective remains nearly unaffected by encoder removal, serving as a matched control.

Future work should focus on:
- Decoder architectures explicitly designed for native visual representation learning
- Training strategies that improve the compute efficiency of native visual input
- Understanding how the crossover point shifts with different data mixtures and encoder sizes

---

_Markdown view of https://picx.dev/p/04O53m, served by PicX — AI-generated visual whiteboard summaries of research papers._
