Summary (Overview)
- PanoVLN is a novel vision-and-language navigation (VLN) framework that fully exploits panoramic (360°) observations to improve navigation accuracy and efficiency, achieving state-of-the-art results on R2R-CE and RxR-CE benchmarks.
- The paper identifies that simply replacing perspective images with panoramas yields limited gains, requiring three key adaptations: longer action-sequence supervision, decision-centric training data construction, and geometry-aware visual representation.
- PanoVLN introduces Confidence-Guided Execution (CGE), which dynamically determines how many predicted actions to execute based on action uncertainty, balancing longer-horizon planning with prediction reliability.
- With a 4B backbone and RGB-only input, PanoVLN surpasses previous SOTA by 11.9% and 8.7% in success rate on R2R-CE and RxR-CE Val-Unseen, respectively.
- Real-world experiments on a quadruped robot demonstrate faster navigation with fewer pauses than prior VLN methods, showing successful transfer from simulation to physical environments.
Introduction and Theoretical Foundation
Background and Motivation
Vision-and-language navigation (VLN) requires an agent to navigate through an environment by following natural-language instructions (Anderson et al., 2018; Krantz et al., 2020). Recent vision-language models (VLMs) have advanced this task by predicting navigation actions from visual observations and instructions.
The core motivation of PanoVLN is straightforward: an equirectangular panorama (ERP) provides a 360° view, exposing passages, landmarks, and route alternatives across different viewing directions, providing a more complete visual basis for navigation decisions. For example, a panorama can reveal a passage outside a perspective camera's field of view, allowing the model to identify the intended route without additional exploration.
Key Insight
The authors find that simply replacing perspective images with ERPs under the same training and inference setup does not improve performance. This observation motivates a detailed diagnostic analysis, leading to three targeted adaptations:
- Action Prediction: Wider visibility supports longer-horizon action planning
- Training Supervision: Branching points are relatively sparse in existing VLN training data, requiring decision-centric data construction
- Visual Representation: Panoramic navigation requires understanding spatial relationships across viewing directions, beyond recognizing individual landmarks
Methodology
3.1 Preliminaries
Perspective baseline: At step , the policy receives a natural-language instruction , the current perspective RGB image , and sampled visual history . The language model autoregressively predicts a sequence of actions from .
Panoramic baseline: Replaces perspective images with RGB equirectangular panoramas (ERPs). The agent's heading is centered, with left and right image boundaries adjacent across a seam behind the agent. The policy input becomes .
3.2 Longer Action-Sequence Supervision and Execution
Action-sequence supervision: At training state , the target contains the next expert actions. Using teacher forcing, the loss is:
Confidence-Guided Execution (CGE): Given the logit for action at position , uncertainty is defined as:
Let , with . CGE extends the prefix while for an uncertainty budget , selecting at least actions:
3.3 Decision-Centric Data Construction
- 98K navigation trajectories across 800 HM3D scenes
- Routes constructed to contain frequent branching points (at least two visible, traversable paths to different areas)
- Instructions generated by Qwen3.8-27B, verified against route observations through:
- Motion consistency checks
- Choice grounding checks
- Stop grounding checks
- Training states sampled with stride-six grid (for H=18) plus added states at sustained-turn onsets and near termination
3.4 Geometry-Aware Visual Representation
- Visual context allocation: tokens for current ERP, tokens per history frame
- Spatially aligned geometric fusion: A frozen PanoVGGT encoder extracts geometric features from the current RGB panorama, resampled in ERP coordinates and grouped to match the VLM's merged current tokens
where is a fixed residual scale, is a trainable MLP, is the VLM's visual tokens, and is the aligned geometric groups.
Empirical Validation / Results
Simulation Results
Table 1: Simulation benchmark comparison — PanoVLN achieves the highest SR and SPL on both benchmarks:
| Method | Pano. | Odo. | Depth | S.RGB | R2R NE↓ | R2R OS↑ | R2R SR↑ | R2R SPL↑ | RxR NE↓ | RxR SR↑ | RxR SPL↑ | RxR nDTW↑ |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| CorrectNav† | ✓ | 4.24 | 67.5 | 65.1 | 62.3 | 4.09 | 69.3 | 63.3 | 75.2 | |||
| AwareVLN | ✓ | 4.02 | 73.5 | 65.4 | 55.1 | 3.95 | 67.6 | 56.1 | 65.7 | |||
| DualVLN | ✓ | 4.05 | 70.7 | 64.3 | 58.5 | 4.58 | 61.4 | 51.8 | 70.0 | |||
| PanoVLN† (Ours) | ✓ | 3.10 | 79.7 | 73.9 | 67.9 | 3.14 | 74.1 | 63.9 | 71.4 | |||
| PanoVLN (Ours) | ✓ | 2.83 | 83.1 | 77.3 | 70.6 | 2.85 | 78.0 | 65.9 | 73.3 |
Ablation Studies
Prediction horizon: ERP policies trail perspective policies at short horizons but overtake them as targets lengthen. Perspective performance peaks at H=4, while ERP favors H=18.
Training-state sampling (Table 3):
| Sampling | NE↓ | OS↑ | SR↑ | SPL↑ |
|---|---|---|---|---|
| Random | 5.37 | 70.7 | 56.4 | 49.5 |
| Ours | 4.40 | 69.9 | 62.8 | 57.2 |
Execution strategy (Table 4) — CGE outperforms all fixed and random strategies:
| Execution strategy | R2R NE↓ | R2R SR↑ | R2R SPL↑ | RxR NE↓ | RxR SR↑ |
|---|---|---|---|---|---|
| Fixed: 1 action | 4.42 | 64.6 | 60.5 | 4.11 | 65.7 |
| Fixed: 6 actions | 4.43 | 64.2 | 58.7 | 4.25 | 65.4 |
| Fixed: 18 actions | 4.89 | 57.2 | 50.9 | 5.64 | 54.9 |
| Random: 1–18 actions | 4.62 | 61.3 | 55.6 | 5.07 | 59.9 |
| CGE (Ours) | 4.10 | 66.6 | 61.5 | 4.06 | 66.5 |
Geometry encoder comparison (Table 5) — PanoVGGT yields the highest success rates:
| Extra encoder | R2R NE↓ | R2R OS↑ | R2R SR↑ | R2R SPL↑ | RxR NE↓ | RxR SR↑ |
|---|---|---|---|---|---|---|
| None | 4.10 | 71.0 | 66.6 | 61.5 | 4.06 | 66.5 |
| UniK3D | 3.77 | 73.2 | 67.8 | 63.3 | 3.92 | 65.4 |
| PanoVGGT (Ours) | 3.91 | 73.6 | 68.6 | 63.7 | 4.02 | 67.1 |
Real-World Results
Table 2: Real-world execution efficiency — PanoVLN achieves the best overall navigation efficiency:
| Method | Time (s) ↓ | Speed (cm/s) ↑ | Wait (%) ↓ | Pauses ↓ | Calls ↓ | Latency (s) ↓ |
|---|---|---|---|---|---|---|
| NaVid | 135.2 | 12.9 | 23.2 | 15.7 | 29.4 | 0.90 |
| NaVILA | 194.0 | 8.2 | 24.7 | 29.6 | 41.3 | 1.05 |
| StreamVLN | 117.8 | 13.3 | 27.1 | 8.3 | 34.3 | 0.58 |
| JanusVLN | 412.9 | 5.3 | 39.2 | 95.3 | 101.7 | 1.32 |
| PanoVLN (Ours) | 86.4 | 25.7 | 13.6 | 5.4 | 7.4 | 1.08 |
Theoretical and Practical Implications
Theoretical Implications
-
Panoramic perception changes policy design: The finding that panoramic policies favor much longer action horizons (H=18) than perspective policies (H=4) suggests that the information horizon of observations fundamentally shapes optimal planning depth in navigation.
-
Uncertainty-aware execution: CGE demonstrates that adaptive execution based on prediction confidence is superior to both fixed-length and random execution, providing a principled framework for balancing planning horizon and reliability.
-
Semantic-geometric fusion: The combination of semantic (VLM) and geometric (PanoVGGT) features via residual fusion shows that spatial layout understanding is complementary to scene semantics for panoramic navigation, without requiring additional visual tokens.
Practical Implications
-
RGB-only navigation: PanoVLN achieves SOTA results without depth input or odometry, simplifying sensor requirements for real-world deployment.
-
Data efficiency: The decision-centric data construction pipeline provides a scalable approach to generating high-quality training trajectories with clear instructions, improving performance across training scales.
-
Real-world transfer: Successful deployment on a quadruped robot demonstrates that simulation-trained panoramic policies can transfer to physical environments, with faster navigation (86.4s vs. 117.8s for the next best) and dramatically fewer policy calls (7.4 vs. 29.4+).
Conclusion
Main Takeaways
PanoVLN demonstrates that fully exploiting panoramic observations in VLN requires coordinated adaptations across action prediction, training supervision, and visual representation:
- Longer action-sequence supervision with H=18 enables panoramic policies to leverage wider visibility for longer-horizon planning
- Confidence-Guided Execution dynamically adapts execution length based on prediction uncertainty, avoiding the pitfalls of both too-short and too-long fixed execution
- Decision-centric data construction with frequent branching points provides targeted supervision for route selection
- Geometry-aware visual representation via PanoVGGT fusion captures spatial layout without adding visual tokens
Future Directions
The authors' findings suggest that broader visibility should shape how navigation policies learn and act. Future work could explore:
- Extending panoramic VLN to more diverse environments and instruction types
- Investigating whether even longer horizons or multi-scale planning could further improve performance
- Applying the CGE framework to other sequential decision-making tasks beyond navigation
- Exploring adaptive visual-token allocation based on scene complexity or task demands
Related papers
- What Makes World Action Models Generalize? An Empirical Study of Test-Time Future Modeling
Simple-WAM shows a single forward pass over noised future tokens, not iterative denoising, recovers nearly all of explicit world model generalization at latent-level efficiency.
- LEGO-Anything: Coding Agents for 3D Scene Reconstruction
LEGO-Anything turns single images into editable 3D scene programs via coding agents, but current models achieve only about half the fidelity of specialist vision systems.
- VoxMem: Benchmarking Multimodal Memory in Large Audio Language Models
VoxMem reveals large audio-language models remember speech content far better than speaker identity, paralinguistic cues, or environmental sounds, with accuracy collapsing under temporal tracking.