# Rethinking Latent Visual Reasoning: Grounding Latent Reasoning in Visual Evidence

> ReaLVR closes the latent evidence-credit gap by supervising free-running latent tokens with answer-contrastive readouts and visual prototypes, achieving 63.7% average accuracy and scaling to 235B parameters.

- **Source:** [arXiv](https://arxiv.org/abs/2609.34563)
- **Published:** 2026-10-02
- **Permalink:** https://picx.dev/p/LKXrs2
- **Whiteboard:** https://picx.dev/p/LKXrs2/image

## Summary

## Summary (Overview)

- **Problem Identification**: The paper identifies a **latent evidence-credit gap** in Latent Visual Reasoning (LVR): latent tokens respond only weakly to image perturbations that alter the correct answer, meaning vanilla LVR does not reliably preserve answer-relevant visual evidence in its generated trajectory.
- **Proposed Method**: The authors introduce **ReaLVR**, which brings visual-evidence supervision to the model's own free-running latent trajectories by contrasting correct vs. wrong answers (to determine *where* supervision is needed) and relevant vs. mismatched visual evidence (to determine *what* to preserve).
- **Key Results**: ReaLVR achieves the highest five-task average of **63.7%** on Qwen2.5-VL-7B, outperforming all evaluated latent-reasoning baselines. It is the **first work to scale visual reasoning in latent space to 235B parameters** (Qwen3-VL-235B).
- **Mechanistic Insights**: ReaLVR produces latent states that are more question-sensitive, more strongly aligned with relevant visual regions, and exhibit greater fixed-context dependence on the most attended latent tokens compared to vanilla LVR.
- **Training Efficiency**: The added supervision requires **no change to model architecture or inference procedure**; both supervision branches are used only during training.

## Introduction and Theoretical Foundation

Latent visual reasoning (LVR) is an emerging paradigm where multimodal large language models (MLLMs) perform intermediate computation through continuous latent tokens rather than expressing every reasoning step in words. This approach is attractive for problems involving spatial relationships and fine-grained visual details that are difficult to describe step-by-step.

However, the underlying mechanisms of LVR remain unclear: what information do generated latent tokens encode, do they respond to visual evidence that changes the correct answer, and how do they contribute to the final answer? The authors identify a critical training gap:

> "Standard GRPO only optimizes the generated text rather than the latent trajectory itself. Consequently, answer-level feedback provides no direct signal indicating which positions need stronger visual supervision or what evidence they should preserve."

This missing connection is termed the **latent evidence-credit gap**. The paper separates three distinct questions about a generated latent token:

1. **Readout**: Does the answer decoder attend to it?
2. **Grounding**: Does it represent the evidence relevant to the image–question pair?
3. **Utility**: Does intervening on the token change the answer?

These questions require distinct measurements, as a token can receive attention without representing relevant evidence or affecting the answer.

## Methodology

### LVR Formulation

In LVR, the decoder inserts $K$ continuous latent states between the input and the textual answer:

$$x \rightarrow \langle|\text{lvr start}|\rangle, z_1, \ldots, z_K, \langle|\text{lvr end}|\rangle, \text{answer}$$

where $z_t \in \mathbb{R}^d$. A free-running rollout is $\tau = (z_{1:K}, o) \sim \pi_\theta(\cdot|x)$.

### Two-Stage Training

**Stage 1: Target-conditioned visual supervision.** The decoder hidden states $z^{\text{TF}}_{1:T_v}$ are trained to reconstruct a target sequence of visual embeddings $v^\star_{1:T_v}$:

$$\mathcal{L}_{\text{rec}}(\theta) = \frac{1}{T_v} \sum_{t=1}^{T_v} \left\| z^{\text{TF}}_t - v^\star_t \right\|_2^2$$

**Stage 2: Free-running outcome optimization.** Standard GRPO minimizes:

$$\mathcal{L}^{\text{LVR}}_{S2}(\theta) = -\mathcal{J}_{\text{clip}}(\theta; \theta_{\text{old}}, \hat{A}) + \beta \hat{\mathcal{L}}^{\text{text}}_{\text{KL}}(\theta; \theta_0)$$

Both terms score generated text positions, leaving free-running latent states without direct visual-evidence supervision.

### ReaLVR: Outcome-Contrastive Evidence Credit

ReaLVR regenerates a current-model trajectory by recursively feeding back its own hidden states:

$$z^\theta_t = T_\theta\left(x, \langle|\text{lvr start}|\rangle, z^\theta_{1:t-1}\right), \quad t = 1, \ldots, K$$

**Visual contrast (what to preserve)**. The positive prototype is a masked mean over visual tokens: $p^+ = \text{Pool}(V; a^+) = \frac{\sum_n a^+_n v_n}{\sum_n a^+_n + \varepsilon}$. Negative prototypes $\mathcal{N} = \{p^-_s\}_{s=1}^{N^-}$ are pooled from mismatched examples. The margin at latent position $t$ is:

$$g_t = \text{sim}(z^\theta_t, p^+) - \max_{p^- \in \mathcal{N}} \text{sim}(z^\theta_t, p^-)$$

**Answer contrast (where to supervise)**. Correct and wrong answers are teacher-forced after the shared latent span. The readout is:

$$r_t(y) = \text{mean}_{\ell,h,j} A^{(\ell,h)}_{j,t}(y)$$

The selective credit is the positive difference between correct-answer and mean wrong-answer readouts:

$$\gamma_t = [r^+_t - r^-_t]_+$$

**Combined objective**. The per-example evidence loss and Stage 2 objective are:

$$\ell_{\text{ev}} = \sum_{t=1}^{K} \text{sg}(w_t)[m_{\text{ev}} - g_t]_+$$

$$\mathcal{L}^{\text{ReaLVR}}_{S2}(\theta) = \mathcal{L}^{\text{LVR}}_{S2}(\theta) + \lambda_{\text{ev}} \hat{\mathcal{L}}_{\text{ev}}(\theta)$$

where $w_t = \eta/K + (1-\eta)\gamma_t$ combines selective credit with a uniform baseline, and $\text{sg}$ denotes stop-gradient (preventing a shortcut through reducing the weight itself).

## Empirical Validation / Results

### Main Results (Qwen2.5-VL-7B)

| Model | Training | MMVP | BLINK | HR-4K | HR-8K | MME-RW | Average |
|---|---|---|---|---|---|---|---|
| Pixel Reasoner | RL | 66.8 | 53.3 | 69.8 | 64.1 | 49.7 | 60.7 |
| Vision-R1 | RL | 51.5 | 52.7 | 62.9 | 58.6 | 44.2 | 54.0 |
| LVR-SFT | SFT | 63.6 | 53.2 | 69.0 | 63.3 | 49.5 | 59.7 |
| LVR-RL | RL | 64.2 | 53.6 | 69.6 | 64.4 | 50.1 | 60.4 |
| ILVR-Stage2 | S2 | 69.4 | 56.8 | 71.0 | 66.9 | 50.3 | 62.9 |
| Monet-RL | RL | 69.9 | 52.4 | 71.3 | 66.0 | 51.5 | 62.2 |
| **ReaLVR** | **RL** | **72.0** | **55.8** | **71.8** | **66.6** | **52.2** | **63.7** |

ReaLVR achieves **+4.0 points over LVR-SFT**, **+3.3 over LVR-RL**, and **+0.8 over ILVR** (the strongest competing latent-reasoning baseline).

### Scaling Results

| Model | MMVP | BLINK | HR-4K | HR-8K | MME-RW | Average |
|---|---|---|---|---|---|---|
| Qwen3-VL-8B ReaLVR | 68.7 | 54.9 | 70.4 | 62.9 | 49.3 | 61.2 |
| Qwen3-VL-30B ReaLVR | 77.7 | 51.4 | 74.1 | 67.9 | 54.7 | 65.2 |
| Qwen3-VL-235B ReaLVR | 81.9 | 75.4 | – | – | 71.0 | – |
| InternVL3-8B ReaLVR | 72.3 | 52.4 | 64.4 | 55.4 | 45.9 | 58.1 |
| Gemma-3-12B ReaLVR | 60.0 | 40.6 | 37.4 | 38.0 | 31.8 | 41.6 |

### Counterfactual Sensitivity

Vanilla LVR shows weak counterfactual sensitivity: although the correct answer changes in 81.45%–86.33% of edited pairs, LVR changes its prediction in only 5.66%–13.09%. ReaLVR updates its answer much more reliably, with strict paired correct-flip rates of 68.3%–73.1%.

### Component Ablations

| Variant | Five-task Avg. |
|---|---|
| LVR-RL (no evidence loss) | 60.4 |
| Uniform routing ($\eta=1$) | 62.4 |
| Raw attention ($\gamma_t = r^+_t$) | 62.9 |
| No negatives ($\mathcal{N}=\emptyset$) | 62.1 |
| Undetached weights | 62.6 |
| Off-policy targets | 61.9 |
| **ReaLVR (full)** | **63.7** |

### Mechanism Analyses

- **Fixed-context dependence**: Replacing the top-8 answer-attended latent tokens decreases correct-answer probability by **11 percentage points** for ReaLVR (from 0.70 to 0.59), versus only 4 points for LVR.
- **Target-region enrichment**: ReaLVR reaches approximately 2× attention on target regions in middle layers, versus ~1.6 for Monet and ≤1.3 for LVR.
- **Latent variation**: ReaLVR concentrates variation at particular latent positions (top-token variation gap of 0.07, versus 0.02 for Monet and 0.01 for LVR variants).

## Theoretical and Practical Implications

- **Theoretical contribution**: The paper provides the first systematic diagnosis of the *latent evidence-credit gap*, separating readout, grounding, and utility as distinct properties of latent tokens. This framework enables principled evaluation of latent reasoning beyond final-answer accuracy.
- **Methodological insight**: The ablation study reveals that *on-policy regeneration* is the largest contributor to performance (1.8-point drop when removed), confirming that supervision must reach the process that produces latents at inference. The stop-gradient on weights is also crucial, preventing the router from lowering weights on hard positions instead of improving their evidence margin.
- **Practical scalability**: ReaLVR is the **first demonstration of continuous latent visual reasoning trained at 235B parameters**, showing that the approach remains effective at frontier model scales. The method requires no architectural changes and preserves the original LVR inference procedure.
- **Interpretability**: The paper shows that latent states need not be verbalizable — all 376 latent readouts from 47 questions project to the closing-tag token "⟨", yet a linear probe recovers task labels with 99.9% accuracy from the mean latent, demonstrating that continuous states support answers without expressing a rationale.

## Conclusion

This work addresses a fundamental challenge in latent visual reasoning: a correct final answer does not guarantee that preceding latent tokens have learned to preserve the visual evidence necessary to produce it. The authors identify this missing link as the **latent evidence-credit gap** and introduce **ReaLVR**, which:

1. **Contrasts visual evidence** to teach latent tokens *what* to preserve (relevant vs. mismatched prototypes)
2. **Compares correct and wrong answer readouts** to decide *where* supervision is most critical

Across backbones up to 235B parameters, evidence supervision improves over available LVR baselines without changing model architectures or inference procedures. The results support visual-evidence supervision as a way to improve the grounding and use of continuous latent states.

**Future directions** identified by the authors include:
- Deriving finer evidence targets from weak supervision when region annotations are unavailable
- Adapting the latent-token budget to each question (task-dependent budgets showed promise, with $K=8$ giving the highest mean accuracy but task-specific optima varying from $K=4$ to $K=16$)

The paper calls for more attention toward demystifying the internal dynamics of continuous latent reasoning beyond benchmark accuracy, laying a grounded foundation for robust and faithful multimodal systems.

---

_Markdown view of https://picx.dev/p/LKXrs2, served by PicX — AI-generated visual whiteboard summaries of research papers._
