Summary (Overview)

  • PanoVLN is a novel vision-and-language navigation (VLN) framework that fully exploits panoramic (360°) observations to improve navigation accuracy and efficiency, achieving state-of-the-art results on R2R-CE and RxR-CE benchmarks.
  • The paper identifies that simply replacing perspective images with panoramas yields limited gains, requiring three key adaptations: longer action-sequence supervision, decision-centric training data construction, and geometry-aware visual representation.
  • PanoVLN introduces Confidence-Guided Execution (CGE), which dynamically determines how many predicted actions to execute based on action uncertainty, balancing longer-horizon planning with prediction reliability.
  • With a 4B backbone and RGB-only input, PanoVLN surpasses previous SOTA by 11.9% and 8.7% in success rate on R2R-CE and RxR-CE Val-Unseen, respectively.
  • Real-world experiments on a quadruped robot demonstrate faster navigation with fewer pauses than prior VLN methods, showing successful transfer from simulation to physical environments.

Introduction and Theoretical Foundation

Background and Motivation

Vision-and-language navigation (VLN) requires an agent to navigate through an environment by following natural-language instructions (Anderson et al., 2018; Krantz et al., 2020). Recent vision-language models (VLMs) have advanced this task by predicting navigation actions from visual observations and instructions.

The core motivation of PanoVLN is straightforward: an equirectangular panorama (ERP) provides a 360° view, exposing passages, landmarks, and route alternatives across different viewing directions, providing a more complete visual basis for navigation decisions. For example, a panorama can reveal a passage outside a perspective camera's field of view, allowing the model to identify the intended route without additional exploration.

Key Insight

The authors find that simply replacing perspective images with ERPs under the same training and inference setup does not improve performance. This observation motivates a detailed diagnostic analysis, leading to three targeted adaptations:

  1. Action Prediction: Wider visibility supports longer-horizon action planning
  2. Training Supervision: Branching points are relatively sparse in existing VLN training data, requiring decision-centric data construction
  3. Visual Representation: Panoramic navigation requires understanding spatial relationships across viewing directions, beyond recognizing individual landmarks

Methodology

3.1 Preliminaries

Perspective baseline: At step tt, the policy receives a natural-language instruction xx, the current perspective RGB image ItI_t, and sampled visual history T<t\mathcal{T}_{<t}. The language model autoregressively predicts a sequence of actions from A={forward,left,right,stop}\mathcal{A} = \{\text{forward}, \text{left}, \text{right}, \text{stop}\}.

Panoramic baseline: Replaces perspective images with RGB equirectangular panoramas (ERPs). The agent's heading is centered, with left and right image boundaries adjacent across a seam behind the agent. The policy input becomes Ot=(x,P<t,Pt)\mathcal{O}_t = (x, \mathcal{P}_{<t}, P_t).

3.2 Longer Action-Sequence Supervision and Execution

Action-sequence supervision: At training state tt, the target At∗=(at,1∗,…,at,H∗)\mathbf{A}_t^* = (a_{t,1}^*, \ldots, a_{t,H}^*) contains the next HH expert actions. Using teacher forcing, the loss is:

Lact=−1H∑i=1Hlog⁡pθ(at,i∗∣Ot,At,<i∗),(1)\mathcal{L}_{\mathrm{act}} = -\frac{1}{H} \sum_{i=1}^{H} \log p_\theta(a_{t,i}^* \mid \mathcal{O}_t, \mathbf{A}_{t,<i}^*),\tag{1}

Confidence-Guided Execution (CGE): Given the logit zt,i(a)z_{t,i}(a) for action aa at position ii, uncertainty is defined as:

qt,i(a)=exp⁡zt,i(a)∑b∈Aexp⁡zt,i(b),ut,i=−log⁡qt,i(a^t,i).(2)\begin{array}{c} q_{t,i}(a) = \frac{\exp z_{t,i}(a)}{\sum_{b \in \mathcal{A}} \exp z_{t,i}(b)}, \\ u_{t,i} = -\log q_{t,i}(\hat{a}_{t,i}). \end{array}\tag{2}

Let Ut(k)=∑i=1kut,iU_t(k) = \sum_{i=1}^{k} u_{t,i}, with Ut(0)=0U_t(0) = 0. CGE extends the prefix while Ut(k)≤BU_t(k) \leq B for an uncertainty budget BB, selecting at least Emin⁡E_{\min} actions:

Et=max⁡{k∈{1,…,H}:k≤Emin⁡ or Ut(k)≤B}.(3)E_t = \max\left\{k \in \{1, \dots, H\}: k \leq E_{\min} \text{ or } U_t(k) \leq B\right\}.\tag{3}

3.3 Decision-Centric Data Construction

  • 98K navigation trajectories across 800 HM3D scenes
  • Routes constructed to contain frequent branching points (at least two visible, traversable paths to different areas)
  • Instructions generated by Qwen3.8-27B, verified against route observations through:
    • Motion consistency checks
    • Choice grounding checks
    • Stop grounding checks
  • Training states sampled with stride-six grid (for H=18) plus added states at sustained-turn onsets and near termination

3.4 Geometry-Aware Visual Representation

  • Visual context allocation: NcN_c tokens for current ERP, Nh<NcN_h < N_c tokens per history frame
  • Spatially aligned geometric fusion: A frozen PanoVGGT encoder extracts geometric features from the current RGB panorama, resampled in ERP coordinates and grouped to match the VLM's merged current tokens
Vˉt=Vt+αfψ(Gt),(4)\bar{V}_t = V_t + \alpha f_\psi(G_t),\tag{4}

where α\alpha is a fixed residual scale, fψf_\psi is a trainable MLP, VtV_t is the VLM's visual tokens, and GtG_t is the aligned geometric groups.


Empirical Validation / Results

Simulation Results

Table 1: Simulation benchmark comparison — PanoVLN achieves the highest SR and SPL on both benchmarks:

MethodPano.Odo.DepthS.RGBR2R NE↓R2R OS↑R2R SR↑R2R SPL↑RxR NE↓RxR SR↑RxR SPL↑RxR nDTW↑
CorrectNav†✓4.2467.565.162.34.0969.363.375.2
AwareVLN✓4.0273.565.455.13.9567.656.165.7
DualVLN✓4.0570.764.358.54.5861.451.870.0
PanoVLN† (Ours)✓3.1079.773.967.93.1474.163.971.4
PanoVLN (Ours)✓2.8383.177.370.62.8578.065.973.3

Ablation Studies

Prediction horizon: ERP policies trail perspective policies at short horizons but overtake them as targets lengthen. Perspective performance peaks at H=4, while ERP favors H=18.

Training-state sampling (Table 3):

SamplingNE↓OS↑SR↑SPL↑
Random5.3770.756.449.5
Ours4.4069.962.857.2

Execution strategy (Table 4) — CGE outperforms all fixed and random strategies:

Execution strategyR2R NE↓R2R SR↑R2R SPL↑RxR NE↓RxR SR↑
Fixed: 1 action4.4264.660.54.1165.7
Fixed: 6 actions4.4364.258.74.2565.4
Fixed: 18 actions4.8957.250.95.6454.9
Random: 1–18 actions4.6261.355.65.0759.9
CGE (Ours)4.1066.661.54.0666.5

Geometry encoder comparison (Table 5) — PanoVGGT yields the highest success rates:

Extra encoderR2R NE↓R2R OS↑R2R SR↑R2R SPL↑RxR NE↓RxR SR↑
None4.1071.066.661.54.0666.5
UniK3D3.7773.267.863.33.9265.4
PanoVGGT (Ours)3.9173.668.663.74.0267.1

Real-World Results

Table 2: Real-world execution efficiency — PanoVLN achieves the best overall navigation efficiency:

MethodTime (s) ↓Speed (cm/s) ↑Wait (%) ↓Pauses ↓Calls ↓Latency (s) ↓
NaVid135.212.923.215.729.40.90
NaVILA194.08.224.729.641.31.05
StreamVLN117.813.327.18.334.30.58
JanusVLN412.95.339.295.3101.71.32
PanoVLN (Ours)86.425.713.65.47.41.08

Theoretical and Practical Implications

Theoretical Implications

  1. Panoramic perception changes policy design: The finding that panoramic policies favor much longer action horizons (H=18) than perspective policies (H=4) suggests that the information horizon of observations fundamentally shapes optimal planning depth in navigation.

  2. Uncertainty-aware execution: CGE demonstrates that adaptive execution based on prediction confidence is superior to both fixed-length and random execution, providing a principled framework for balancing planning horizon and reliability.

  3. Semantic-geometric fusion: The combination of semantic (VLM) and geometric (PanoVGGT) features via residual fusion shows that spatial layout understanding is complementary to scene semantics for panoramic navigation, without requiring additional visual tokens.

Practical Implications

  1. RGB-only navigation: PanoVLN achieves SOTA results without depth input or odometry, simplifying sensor requirements for real-world deployment.

  2. Data efficiency: The decision-centric data construction pipeline provides a scalable approach to generating high-quality training trajectories with clear instructions, improving performance across training scales.

  3. Real-world transfer: Successful deployment on a quadruped robot demonstrates that simulation-trained panoramic policies can transfer to physical environments, with faster navigation (86.4s vs. 117.8s for the next best) and dramatically fewer policy calls (7.4 vs. 29.4+).


Conclusion

Main Takeaways

PanoVLN demonstrates that fully exploiting panoramic observations in VLN requires coordinated adaptations across action prediction, training supervision, and visual representation:

  1. Longer action-sequence supervision with H=18 enables panoramic policies to leverage wider visibility for longer-horizon planning
  2. Confidence-Guided Execution dynamically adapts execution length based on prediction uncertainty, avoiding the pitfalls of both too-short and too-long fixed execution
  3. Decision-centric data construction with frequent branching points provides targeted supervision for route selection
  4. Geometry-aware visual representation via PanoVGGT fusion captures spatial layout without adding visual tokens

Future Directions

The authors' findings suggest that broader visibility should shape how navigation policies learn and act. Future work could explore:

  • Extending panoramic VLN to more diverse environments and instruction types
  • Investigating whether even longer horizons or multi-scale planning could further improve performance
  • Applying the CGE framework to other sequential decision-making tasks beyond navigation
  • Exploring adaptive visual-token allocation based on scene complexity or task demands

Related papers