Full text not available for this paper
Summary (Overview)
- This paper presents the first systematic scaling-law comparison between encoder-free and encoder-based multimodal large language models (MLLMs), using a matched ladder of 11 sparse MoE language models (1.1B–44B total parameters).
- Key finding 1: Removing the visual encoder shifts the compute-optimal allocation for the multimodal objective toward larger models (model allocation exponent increases from to ), while text allocation remains nearly unchanged.
- Key finding 2: Encoder-free models underperform at small scales but are predicted to catch up with encoder-based models at approximately FLOPs under compute-optimal allocation—well within practical pretraining budgets (e.g., Kimi K2.5 uses ~ FLOPs).
- Key finding 3: The decoder compensates for the missing visual encoder through vision-specific adaptations: bidirectional attention among visual tokens becomes increasingly beneficial, visual processing shifts to earlier layers, and expert routing for visual tokens becomes more concentrated.
- The crossover varies by topic, arriving earlier on language-heavy topics (STEM) and much later on perception-intensive ones (Caption, OCR, GUI).
Introduction and Theoretical Foundation
Most modern MLLMs build on a pretrained visual encoder (e.g., SigLIP 2 ViT) that provides a strong visual prior learned from large-scale image–text data. Encoder-free MLLMs instead remove the visual encoder and feed projected image patches directly into the decoder, which must learn visual representations from raw pixels. While prior work has shown initial feasibility of encoder-free designs, their scaling behavior has not been systematically characterized.
The paper addresses this gap through a controlled scaling study where both model families share:
- The same sparse decoder ladder (MoE architecture)
- The same data mixture and optimization setup
- The same visual-token granularity
The theoretical foundation draws on established scaling-law methodology:
- Compute-optimal allocation law:
- Compute-optimal frontier: where is the loss–compute exponent, is a fitted prefactor, and is the entropy floor induced by the data distribution.
- IsoFLOP profiles: Following Chinchilla methodology, validation loss is fitted as a quadratic function of at fixed compute budgets.
Methodology
Model Architecture Comparison
The paper compares two architectures on a matched ladder of 11 sparse MoE language models:
- Encoder-based: Uses a pretrained SigLIP 2 ViT followed by a ConvPool adapter and projector, with causal attention applied to all tokens. The ViT is trained jointly with the decoder.
- Encoder-free: Maps raw image patches directly into the decoder through a patch projection, where visual tokens attend bidirectionally within each image, with all other attention remaining causal. A fully causal variant is also studied.
Compute Accounting
Decoder FLOPs per token are decomposed as:
where indexes the model family, indexes the objective, and is the mean packed sequence length. The fixed expert activation ratio of 8/256 keeps approximately proportional to active parameters.
Efficiency Gain Metrics
Following MAI-Thinking-1, the paper defines two efficiency metrics at equal loss :
- Compute efficiency gain:
- Model efficiency gain:
Values above 1.0 indicate the target is more efficient than the reference.
Overtraining Model
For overtraining analysis, the paper uses a separable loss form and derives that overtraining rescales only the prefactor:
where is the overtraining factor and depends only on , not on .
Empirical Validation / Results
Compute-Optimal Allocation
Table 1: Model allocation exponent of encoder-free models under bidirectional and causal attention over visual tokens
| Attention | Text | Multimodal |
|---|---|---|
| Bidirectional | 0.427 | 0.570 |
| Causal | 0.436 | 0.557 |
- Text objective: Nearly identical allocation exponents ( for encoder-free, for encoder-based)
- Multimodal objective: Removing the encoder increases from 0.464 to 0.570, with bootstrap 80% intervals of [0.546, 0.595] and [0.458, 0.472] respectively
Loss–Compute Frontiers
The fitted scaling laws show:
| Objective | Encoder-free exponent | Encoder-based exponent |
|---|---|---|
| Text | ||
| Multimodal |
- Text: Frontiers nearly overlap, with and
- Multimodal: Encoder-free models require more compute initially, but the gap narrows with scale
- increases from ~0.47 (observed mean) toward 1.0
- , meaning encoder-free models use ~1.25× the FLOPs per token at equal loss
- Crossover predicted at FLOPs (80% interval: )
Overtraining Effects
At overtraining factor:
- Crossover arrives later: FLOPs (80% interval: )
- drops from 0.62 to 0.52 at the largest fitted budget
- drops from 0.80 to 0.74
- Text objective remains nearly unaffected (changes < 1%)
Topic-Specific Analysis
Compute efficiency gains vary by topic, with crossover ordering:
| Topic | EG at (fitted range) | Crossover timing |
|---|---|---|
| STEM | 0.71 → 0.95 | Earliest (near parity at largest budget) |
| Charts | 0.28 → 0.57 | Early |
| GUI | 0.51 → 0.69 | Later |
| OCR | 0.18 → 0.34 | Later |
| Caption | 0.34 → 0.50 | Latest |
Vision-Specific Adaptation Mechanisms
-
Emergence of visual encoding: Encoder-free models show a slow initial phase followed by a sharp loss drop. At layer 12, visual attention mass rises from 0.217 to 0.645 after the drop, approaching the encoder-based model's level.
-
Bidirectional attention: Causal attention is ~1% better on text () but worse on multimodal (), with bidirectional attention's multimodal advantage growing with compute.
-
Layerwise representation evolution: Visual tokens in encoder-free models diverge from their layer-0 inputs much earlier than in encoder-based models, while text token trajectories remain nearly identical between architectures.
-
Expert routing concentration: MaxVio (expert load imbalance) for visual tokens is consistently higher in encoder-free models across the four largest model sizes, while text token routing remains closely matched between architectures.
Theoretical and Practical Implications
Theoretical Implications
-
Scaling law divergence: The different allocation exponents ( vs. ) indicate that encoder-free models have fundamentally different compute-optimal tradeoffs, requiring more decoder capacity to jointly support visual representation learning and language modeling.
-
Diminishing value of visual priors: The predicted crossover suggests that the advantage of pretrained visual encoders diminishes with scale—encoder-free models can learn equally good visual representations from raw pixels given sufficient compute.
-
Architecture-specific adaptation: The decoder's vision-specific adaptations (bidirectional attention, early-layer processing, expert concentration) suggest that encoder-free architectures may benefit from decoder designs explicitly optimized for native visual representation learning rather than inheriting designs built for language.
Practical Implications
-
Pretraining budget guidance: The crossover at
FLOPs is well within reach of current flagship pretraining runs (FLOPs), suggesting encoder-free architectures are a viable direction for large-scale multimodal pretraining. -
Topic-dependent deployment: The varying crossover points by topic suggest that encoder-free models may be suitable for language-heavy multimodal tasks earlier, while perception-intensive applications may require larger budgets.
-
Architecture design: The finding that encoder-free models favor larger decoders suggests that scaling model size (rather than data) is more effective for encoder-free multimodal pretraining.
Conclusion
The paper demonstrates that while encoder-free MLLMs are less compute-efficient at small scales, their more rapidly improving multimodal loss frontier predicts an efficiency crossover within practical pretraining budgets (~ FLOPs). The shift toward larger models under compute-optimal allocation, combined with the decoder's vision-specific adaptations (bidirectional visual attention, early-layer visual processing, concentrated expert routing), indicates that encoder-free architectures are a promising direction for multimodal pretraining. The text objective remains nearly unaffected by encoder removal, serving as a matched control.
Future work should focus on:
- Decoder architectures explicitly designed for native visual representation learning
- Training strategies that improve the compute efficiency of native visual input
- Understanding how the crossover point shifts with different data mixtures and encoder sizes
Related papers
- Normalize-Then-Precondition: A Hierarchical Approach to Marginal Scale and Interaction Geometry for LLM Training
NormPre separates LLM update normalization from spectral preconditioning, consistently beating AdamW, Muon, and MANO in pretraining while cutting optimizer latency by up to 67%.
- On Trajectory-Aware Training for Masked Diffusion Language Models
PUMBA trains masked diffusion language models on inference-like trajectories via continuous hidden-state carries and backpropagation through time, matching autoregressive accuracy while decoding multiple tokens per step.
- How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?
Under matched conditions, a minimal-harness coding agent matches or outperforms state-of-the-art MLE harnesses, with performance driven by the LLM backbone and execution environment, not scaffolding.