Full text not available for this paper

Summary (Overview)

  • This paper presents the first systematic scaling-law comparison between encoder-free and encoder-based multimodal large language models (MLLMs), using a matched ladder of 11 sparse MoE language models (1.1B–44B total parameters).
  • Key finding 1: Removing the visual encoder shifts the compute-optimal allocation for the multimodal objective toward larger models (model allocation exponent increases from a=0.464a = 0.464 to a=0.570a = 0.570), while text allocation remains nearly unchanged.
  • Key finding 2: Encoder-free models underperform at small scales but are predicted to catch up with encoder-based models at approximately 102210^{22} FLOPs under compute-optimal allocation—well within practical pretraining budgets (e.g., Kimi K2.5 uses ~102510^{25} FLOPs).
  • Key finding 3: The decoder compensates for the missing visual encoder through vision-specific adaptations: bidirectional attention among visual tokens becomes increasingly beneficial, visual processing shifts to earlier layers, and expert routing for visual tokens becomes more concentrated.
  • The crossover varies by topic, arriving earlier on language-heavy topics (STEM) and much later on perception-intensive ones (Caption, OCR, GUI).

Introduction and Theoretical Foundation

Most modern MLLMs build on a pretrained visual encoder (e.g., SigLIP 2 ViT) that provides a strong visual prior learned from large-scale image–text data. Encoder-free MLLMs instead remove the visual encoder and feed projected image patches directly into the decoder, which must learn visual representations from raw pixels. While prior work has shown initial feasibility of encoder-free designs, their scaling behavior has not been systematically characterized.

The paper addresses this gap through a controlled scaling study where both model families share:

  • The same sparse decoder ladder (MoE architecture)
  • The same data mixture and optimization setup
  • The same visual-token granularity

The theoretical foundation draws on established scaling-law methodology:

  • Compute-optimal allocation law: Mopt(C)∝Ca,Dopt(C)∝Cb,a+b=1M_{\text{opt}}(C) \propto C^a, \quad D_{\text{opt}}(C) \propto C^b, \quad a + b = 1
  • Compute-optimal frontier: L∗(C)=E+KC−γL^*(C) = E + KC^{-\gamma} where γ>0\gamma > 0 is the loss–compute exponent, K>0K > 0 is a fitted prefactor, and EE is the entropy floor induced by the data distribution.
  • IsoFLOP profiles: Following Chinchilla methodology, validation loss is fitted as a quadratic function of log⁡M\log M at fixed compute budgets.

Methodology

Model Architecture Comparison

The paper compares two architectures on a matched ladder of 11 sparse MoE language models:

  • Encoder-based: Uses a pretrained SigLIP 2 ViT followed by a ConvPool adapter and projector, with causal attention applied to all tokens. The ViT is trained jointly with the decoder.
  • Encoder-free: Maps raw image patches directly into the decoder through a patch projection, where visual tokens attend bidirectionally within each image, with all other attention remaining causal. A fully causal variant is also studied.

Compute Accounting

Decoder FLOPs per token are decomposed as:

Mo(s)=Mbase+Mattn(s)ℓo,Co(s)=Mo(s)DoM_o^{(s)} = M_{\text{base}} + M_{\text{attn}}^{(s)} \ell_o, \quad C_o^{(s)} = M_o^{(s)} D_o

where s∈{free,based}s \in \{\text{free}, \text{based}\} indexes the model family, o∈{text,mm}o \in \{\text{text}, \text{mm}\} indexes the objective, and ℓo\ell_o is the mean packed sequence length. The fixed expert activation ratio of 8/256 keeps MM approximately proportional to active parameters.

Efficiency Gain Metrics

Following MAI-Thinking-1, the paper defines two efficiency metrics at equal loss λ\lambda:

  • Compute efficiency gain: EGtar←refC(λ)=Cref(λ)/Ctar(λ)\text{EG}^C_{\text{tar} \leftarrow \text{ref}}(\lambda) = C_{\text{ref}}(\lambda)/C_{\text{tar}}(\lambda)
  • Model efficiency gain: EGtar←refM(λ)=Mref(λ)/Mtar(λ)\text{EG}^M_{\text{tar} \leftarrow \text{ref}}(\lambda) = M_{\text{ref}}(\lambda)/M_{\text{tar}}(\lambda)

Values above 1.0 indicate the target is more efficient than the reference.

Overtraining Model

For overtraining analysis, the paper uses a separable loss form L(M,D)=E+AM−α+BD−βL(M, D) = E + AM^{-\alpha} + BD^{-\beta} and derives that overtraining rescales only the prefactor:

L(Cbase,k)=E+g(k)KCbase−γL(C_{\text{base}}, k) = E + g(k)KC_{\text{base}}^{-\gamma}

where kk is the overtraining factor and g(k)g(k) depends only on kk, not on CbaseC_{\text{base}}.

Empirical Validation / Results

Compute-Optimal Allocation

Table 1: Model allocation exponent aa of encoder-free models under bidirectional and causal attention over visual tokens

AttentionTextMultimodal
Bidirectional0.4270.570
Causal0.4360.557
  • Text objective: Nearly identical allocation exponents (a=0.427a = 0.427 for encoder-free, a=0.422a = 0.422 for encoder-based)
  • Multimodal objective: Removing the encoder increases aa from 0.464 to 0.570, with bootstrap 80% intervals of [0.546, 0.595] and [0.458, 0.472] respectively

Loss–Compute Frontiers

The fitted scaling laws show:

ObjectiveEncoder-free exponentEncoder-based exponent
Textγ=0.0973\gamma = 0.0973γ=0.0979\gamma = 0.0979
Multimodalγ=0.3778\gamma = 0.3778γ=0.2998\gamma = 0.2998
  • Text: Frontiers nearly overlap, with EGC≈0.98\text{EG}^C \approx 0.98 and EGM≈0.99\text{EG}^M \approx 0.99
  • Multimodal: Encoder-free models require more compute initially, but the gap narrows with scale
    • EGC\text{EG}^C increases from ~0.47 (observed mean) toward 1.0
    • EGM≈0.80\text{EG}^M \approx 0.80, meaning encoder-free models use ~1.25× the FLOPs per token at equal loss
    • Crossover predicted at 6.1×10216.1 \times 10^{21} FLOPs (80% interval: [4.2×1021,1.0×1022][4.2 \times 10^{21}, 1.0 \times 10^{22}])

Overtraining Effects

At k=5k = 5 overtraining factor:

  • Crossover arrives later: 1.2×10221.2 \times 10^{22} FLOPs (80% interval: [8.4×1021,2.0×1022][8.4 \times 10^{21}, 2.0 \times 10^{22}])
  • EGC\text{EG}^C drops from 0.62 to 0.52 at the largest fitted budget
  • EGM\text{EG}^M drops from 0.80 to 0.74
  • Text objective remains nearly unaffected (changes < 1%)

Topic-Specific Analysis

Compute efficiency gains vary by topic, with crossover ordering:

TopicEG at k=1k=1 (fitted range)Crossover timing
STEM0.71 → 0.95Earliest (near parity at largest budget)
Charts0.28 → 0.57Early
GUI0.51 → 0.69Later
OCR0.18 → 0.34Later
Caption0.34 → 0.50Latest

Vision-Specific Adaptation Mechanisms

  1. Emergence of visual encoding: Encoder-free models show a slow initial phase followed by a sharp loss drop. At layer 12, visual attention mass rises from 0.217 to 0.645 after the drop, approaching the encoder-based model's level.

  2. Bidirectional attention: Causal attention is ~1% better on text (EGC=1.011\text{EG}^C = 1.011) but worse on multimodal (EGC=0.990\text{EG}^C = 0.990), with bidirectional attention's multimodal advantage growing with compute.

  3. Layerwise representation evolution: Visual tokens in encoder-free models diverge from their layer-0 inputs much earlier than in encoder-based models, while text token trajectories remain nearly identical between architectures.

  4. Expert routing concentration: MaxVio (expert load imbalance) for visual tokens is consistently higher in encoder-free models across the four largest model sizes, while text token routing remains closely matched between architectures.

Theoretical and Practical Implications

Theoretical Implications

  1. Scaling law divergence: The different allocation exponents (a=0.570a = 0.570 vs. a=0.464a = 0.464) indicate that encoder-free models have fundamentally different compute-optimal tradeoffs, requiring more decoder capacity to jointly support visual representation learning and language modeling.

  2. Diminishing value of visual priors: The predicted crossover suggests that the advantage of pretrained visual encoders diminishes with scale—encoder-free models can learn equally good visual representations from raw pixels given sufficient compute.

  3. Architecture-specific adaptation: The decoder's vision-specific adaptations (bidirectional attention, early-layer processing, expert concentration) suggest that encoder-free architectures may benefit from decoder designs explicitly optimized for native visual representation learning rather than inheriting designs built for language.

Practical Implications

  1. Pretraining budget guidance: The crossover at 102210^{22} FLOPs is well within reach of current flagship pretraining runs (102510^{25} FLOPs), suggesting encoder-free architectures are a viable direction for large-scale multimodal pretraining.

  2. Topic-dependent deployment: The varying crossover points by topic suggest that encoder-free models may be suitable for language-heavy multimodal tasks earlier, while perception-intensive applications may require larger budgets.

  3. Architecture design: The finding that encoder-free models favor larger decoders suggests that scaling model size (rather than data) is more effective for encoder-free multimodal pretraining.

Conclusion

The paper demonstrates that while encoder-free MLLMs are less compute-efficient at small scales, their more rapidly improving multimodal loss frontier predicts an efficiency crossover within practical pretraining budgets (~102210^{22} FLOPs). The shift toward larger models under compute-optimal allocation, combined with the decoder's vision-specific adaptations (bidirectional visual attention, early-layer visual processing, concentrated expert routing), indicates that encoder-free architectures are a promising direction for multimodal pretraining. The text objective remains nearly unaffected by encoder removal, serving as a matched control.

Future work should focus on:

  • Decoder architectures explicitly designed for native visual representation learning
  • Training strategies that improve the compute efficiency of native visual input
  • Understanding how the crossover point shifts with different data mixtures and encoder sizes

Related papers