# PanoVLN: Towards Effective Panoramic Vision-and-Language Navigation

> PanoVLN achieves state-of-the-art vision-and-language navigation by pairing panoramic 360-degree inputs with longer action horizons, confidence-guided execution, and geometry-aware visual fusion.

- **Source:** [arXiv](https://arxiv.org/abs/2609.34759)
- **Published:** 2026-10-01
- **Permalink:** https://picx.dev/p/SGefbV
- **Whiteboard:** https://picx.dev/p/SGefbV/image

## Summary

## Summary (Overview)

- **PanoVLN** is a novel vision-and-language navigation (VLN) framework that fully exploits panoramic (360°) observations to improve navigation accuracy and efficiency, achieving state-of-the-art results on R2R-CE and RxR-CE benchmarks.
- The paper identifies that simply replacing perspective images with panoramas yields limited gains, requiring three key adaptations: **longer action-sequence supervision**, **decision-centric training data construction**, and **geometry-aware visual representation**.
- PanoVLN introduces **Confidence-Guided Execution (CGE)**, which dynamically determines how many predicted actions to execute based on action uncertainty, balancing longer-horizon planning with prediction reliability.
- With a 4B backbone and RGB-only input, PanoVLN surpasses previous SOTA by **11.9% and 8.7% in success rate** on R2R-CE and RxR-CE Val-Unseen, respectively.
- Real-world experiments on a quadruped robot demonstrate **faster navigation with fewer pauses** than prior VLN methods, showing successful transfer from simulation to physical environments.

---

## Introduction and Theoretical Foundation

### Background and Motivation

Vision-and-language navigation (VLN) requires an agent to navigate through an environment by following natural-language instructions (Anderson et al., 2018; Krantz et al., 2020). Recent vision-language models (VLMs) have advanced this task by predicting navigation actions from visual observations and instructions.

The core motivation of PanoVLN is straightforward: **an equirectangular panorama (ERP) provides a 360° view**, exposing passages, landmarks, and route alternatives across different viewing directions, providing a more complete visual basis for navigation decisions. For example, a panorama can reveal a passage outside a perspective camera's field of view, allowing the model to identify the intended route without additional exploration.

### Key Insight

The authors find that **simply replacing perspective images with ERPs under the same training and inference setup does not improve performance**. This observation motivates a detailed diagnostic analysis, leading to three targeted adaptations:

1. **Action Prediction**: Wider visibility supports longer-horizon action planning
2. **Training Supervision**: Branching points are relatively sparse in existing VLN training data, requiring decision-centric data construction
3. **Visual Representation**: Panoramic navigation requires understanding spatial relationships across viewing directions, beyond recognizing individual landmarks

---

## Methodology

### 3.1 Preliminaries

**Perspective baseline**: At step $t$, the policy receives a natural-language instruction $x$, the current perspective RGB image $I_t$, and sampled visual history $\mathcal{T}_{<t}$. The language model autoregressively predicts a sequence of actions from $\mathcal{A} = \{\text{forward}, \text{left}, \text{right}, \text{stop}\}$.

**Panoramic baseline**: Replaces perspective images with RGB equirectangular panoramas (ERPs). The agent's heading is centered, with left and right image boundaries adjacent across a seam behind the agent. The policy input becomes $\mathcal{O}_t = (x, \mathcal{P}_{<t}, P_t)$.

### 3.2 Longer Action-Sequence Supervision and Execution

**Action-sequence supervision**: At training state $t$, the target $\mathbf{A}_t^* = (a_{t,1}^*, \ldots, a_{t,H}^*)$ contains the next $H$ expert actions. Using teacher forcing, the loss is:

$$
\mathcal{L}_{\mathrm{act}} = -\frac{1}{H} \sum_{i=1}^{H} \log p_\theta(a_{t,i}^* \mid \mathcal{O}_t, \mathbf{A}_{t,<i}^*),\tag{1}
$$

**Confidence-Guided Execution (CGE)**: Given the logit $z_{t,i}(a)$ for action $a$ at position $i$, uncertainty is defined as:

$$
\begin{array}{c} 
q_{t,i}(a) = \frac{\exp z_{t,i}(a)}{\sum_{b \in \mathcal{A}} \exp z_{t,i}(b)}, \\
u_{t,i} = -\log q_{t,i}(\hat{a}_{t,i}). 
\end{array}\tag{2}
$$

Let $U_t(k) = \sum_{i=1}^{k} u_{t,i}$, with $U_t(0) = 0$. CGE extends the prefix while $U_t(k) \leq B$ for an uncertainty budget $B$, selecting at least $E_{\min}$ actions:

$$
E_t = \max\left\{k \in \{1, \dots, H\}: k \leq E_{\min} \text{ or } U_t(k) \leq B\right\}.\tag{3}
$$

### 3.3 Decision-Centric Data Construction

- **98K navigation trajectories** across 800 HM3D scenes
- Routes constructed to contain **frequent branching points** (at least two visible, traversable paths to different areas)
- Instructions generated by Qwen3.8-27B, verified against route observations through:
  - Motion consistency checks
  - Choice grounding checks  
  - Stop grounding checks
- Training states sampled with **stride-six grid** (for H=18) plus added states at sustained-turn onsets and near termination

### 3.4 Geometry-Aware Visual Representation

- **Visual context allocation**: $N_c$ tokens for current ERP, $N_h < N_c$ tokens per history frame
- **Spatially aligned geometric fusion**: A frozen PanoVGGT encoder extracts geometric features from the current RGB panorama, resampled in ERP coordinates and grouped to match the VLM's merged current tokens

$$
\bar{V}_t = V_t + \alpha f_\psi(G_t),\tag{4}
$$

where $\alpha$ is a fixed residual scale, $f_\psi$ is a trainable MLP, $V_t$ is the VLM's visual tokens, and $G_t$ is the aligned geometric groups.

---

## Empirical Validation / Results

### Simulation Results

**Table 1: Simulation benchmark comparison** — PanoVLN achieves the highest SR and SPL on both benchmarks:

| Method | Pano. | Odo. | Depth | S.RGB | R2R NE↓ | R2R OS↑ | R2R SR↑ | R2R SPL↑ | RxR NE↓ | RxR SR↑ | RxR SPL↑ | RxR nDTW↑ |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| CorrectNav† | | | | ✓ | 4.24 | 67.5 | 65.1 | 62.3 | 4.09 | 69.3 | 63.3 | 75.2 |
| AwareVLN | | | | ✓ | 4.02 | 73.5 | 65.4 | 55.1 | 3.95 | 67.6 | 56.1 | 65.7 |
| DualVLN | | | | ✓ | 4.05 | 70.7 | 64.3 | 58.5 | 4.58 | 61.4 | 51.8 | 70.0 |
| **PanoVLN† (Ours)** | ✓ | | | | **3.10** | **79.7** | **73.9** | **67.9** | **3.14** | **74.1** | **63.9** | **71.4** |
| **PanoVLN (Ours)** | ✓ | | | | **2.83** | **83.1** | **77.3** | **70.6** | **2.85** | **78.0** | **65.9** | **73.3** |

### Ablation Studies

**Prediction horizon**: ERP policies trail perspective policies at short horizons but overtake them as targets lengthen. Perspective performance peaks at H=4, while ERP favors H=18.

**Training-state sampling** (Table 3):

| Sampling | NE↓ | OS↑ | SR↑ | SPL↑ |
|---|---|---|---|---|
| Random | 5.37 | 70.7 | 56.4 | 49.5 |
| **Ours** | **4.40** | 69.9 | **62.8** | **57.2** |

**Execution strategy** (Table 4) — CGE outperforms all fixed and random strategies:

| Execution strategy | R2R NE↓ | R2R SR↑ | R2R SPL↑ | RxR NE↓ | RxR SR↑ |
|---|---|---|---|---|---|
| Fixed: 1 action | 4.42 | 64.6 | 60.5 | 4.11 | 65.7 |
| Fixed: 6 actions | 4.43 | 64.2 | 58.7 | 4.25 | 65.4 |
| Fixed: 18 actions | 4.89 | 57.2 | 50.9 | 5.64 | 54.9 |
| Random: 1–18 actions | 4.62 | 61.3 | 55.6 | 5.07 | 59.9 |
| **CGE (Ours)** | **4.10** | **66.6** | **61.5** | **4.06** | **66.5** |

**Geometry encoder comparison** (Table 5) — PanoVGGT yields the highest success rates:

| Extra encoder | R2R NE↓ | R2R OS↑ | R2R SR↑ | R2R SPL↑ | RxR NE↓ | RxR SR↑ |
|---|---|---|---|---|---|---|
| None | 4.10 | 71.0 | 66.6 | 61.5 | 4.06 | 66.5 |
| UniK3D | 3.77 | 73.2 | 67.8 | 63.3 | 3.92 | 65.4 |
| **PanoVGGT (Ours)** | **3.91** | **73.6** | **68.6** | **63.7** | **4.02** | **67.1** |

### Real-World Results

**Table 2: Real-world execution efficiency** — PanoVLN achieves the best overall navigation efficiency:

| Method | Time (s) ↓ | Speed (cm/s) ↑ | Wait (%) ↓ | Pauses ↓ | Calls ↓ | Latency (s) ↓ |
|---|---|---|---|---|---|---|
| NaVid | 135.2 | 12.9 | 23.2 | 15.7 | 29.4 | 0.90 |
| NaVILA | 194.0 | 8.2 | 24.7 | 29.6 | 41.3 | 1.05 |
| StreamVLN | 117.8 | 13.3 | 27.1 | 8.3 | 34.3 | 0.58 |
| JanusVLN | 412.9 | 5.3 | 39.2 | 95.3 | 101.7 | 1.32 |
| **PanoVLN (Ours)** | **86.4** | **25.7** | **13.6** | **5.4** | **7.4** | **1.08** |

---

## Theoretical and Practical Implications

### Theoretical Implications

1. **Panoramic perception changes policy design**: The finding that panoramic policies favor much longer action horizons (H=18) than perspective policies (H=4) suggests that the *information horizon* of observations fundamentally shapes optimal planning depth in navigation.

2. **Uncertainty-aware execution**: CGE demonstrates that adaptive execution based on prediction confidence is superior to both fixed-length and random execution, providing a principled framework for balancing planning horizon and reliability.

3. **Semantic-geometric fusion**: The combination of semantic (VLM) and geometric (PanoVGGT) features via residual fusion shows that spatial layout understanding is complementary to scene semantics for panoramic navigation, without requiring additional visual tokens.

### Practical Implications

1. **RGB-only navigation**: PanoVLN achieves SOTA results without depth input or odometry, simplifying sensor requirements for real-world deployment.

2. **Data efficiency**: The decision-centric data construction pipeline provides a scalable approach to generating high-quality training trajectories with clear instructions, improving performance across training scales.

3. **Real-world transfer**: Successful deployment on a quadruped robot demonstrates that simulation-trained panoramic policies can transfer to physical environments, with faster navigation (86.4s vs. 117.8s for the next best) and dramatically fewer policy calls (7.4 vs. 29.4+).

---

## Conclusion

### Main Takeaways

PanoVLN demonstrates that fully exploiting panoramic observations in VLN requires coordinated adaptations across action prediction, training supervision, and visual representation:

1. **Longer action-sequence supervision** with H=18 enables panoramic policies to leverage wider visibility for longer-horizon planning
2. **Confidence-Guided Execution** dynamically adapts execution length based on prediction uncertainty, avoiding the pitfalls of both too-short and too-long fixed execution
3. **Decision-centric data construction** with frequent branching points provides targeted supervision for route selection
4. **Geometry-aware visual representation** via PanoVGGT fusion captures spatial layout without adding visual tokens

### Future Directions

The authors' findings suggest that broader visibility should shape how navigation policies learn and act. Future work could explore:
- Extending panoramic VLN to more diverse environments and instruction types
- Investigating whether even longer horizons or multi-scale planning could further improve performance
- Applying the CGE framework to other sequential decision-making tasks beyond navigation
- Exploring adaptive visual-token allocation based on scene complexity or task demands

---

_Markdown view of https://picx.dev/p/SGefbV, served by PicX — AI-generated visual whiteboard summaries of research papers._
