Summary (Overview)
- GST-Bench is a new video-based VQA benchmark for evaluating global spatial awareness in Vision-Language Models (VLMs), requiring models to integrate long-horizon egocentric video observations, reason from novel viewpoints unseen in the input, and map observations onto explicit top-down scene representations.
- The benchmark comprises 12 subtasks across three core competencies (self localization, object localization, scene structure understanding) with 2,762 human-verified questions derived from 6,790 minutes of synthetically generated video in 50 diverse indoor scenes.
- A comprehensive evaluation of 22 state-of-the-art VLMs reveals a striking gap: the strongest zero-shot model (Gemini-3-Pro) achieves only 42.68, far below the human baseline of 79.08.
- Controlled local-variant experiments show that proprietary models fail primarily at cross-frame spatial integration (not single-image perception), while open-source models struggle at both local perception and global integration.
- The authors release GST-Train, a large-scale training dataset; fine-tuning Qwen3-VL-8B on it improves the score from 25.89 to 53.52, surpassing all zero-shot proprietary models.
Introduction and Theoretical Foundation
Background and Motivation
Spatial intelligence is fundamental for embodied agents (e.g., household robots) that must:
- Incrementally build scene representations from egocentric exploration
- Remember where objects were encountered
- Infer relative locations even when objects leave the field of view
While VLMs have made rapid progress in visual perception and reasoning, existing spatial benchmarks are confined to:
- Single-image settings (e.g., CV-Bench, 3DSRBench, Spatial457) — testing local spatial properties like object size and relative depth
- Multi-image settings (e.g., MM-Spatial, MMSI-Bench) — typically 2–3 views, still limited to localized understanding
Gaps in Existing Video-Based Benchmarks
Recent video benchmarks (VSI-Bench, MMSI-Video-Bench, OST-Bench) begin to explore global spatial understanding but have three key limitations:
- No disentanglement of single-frame solvable tasks from those requiring cross-frame global reasoning
- Coarse-grained spatial metrics (e.g., front/back/left/right) rather than precise angles
- Lack of explicit global scene representations (e.g., top-down images) to evaluate global-local correspondence
Design Principles of GST-Bench
- Global reasoning by construction: Target objects are guaranteed invisible from the query viewpoint; single-frame answerable questions are excluded
- Precise numerical answers: Metrics computed over exact distances, angles, and point locations rather than coarse categorical judgments
- Top-down images: Probe the model's ability to map cross-view video onto a global scene representation
- Independent query viewpoints: Query images sampled from off-trajectory camera positions
Methodology
Task Inputs
GST-Bench includes five types of visual inputs:
| Input Type | Description |
|---|---|
| Exploration Video | Long-horizon egocentric traversal of the entire scene |
| Object-Annotated Video | Exploration video with red bounding boxes overlaid on target object frames |
| Short Trajectory Video | Clip from an independently collected trajectory for trajectory matching |
| Top-Down Images | Three abstraction levels: easy (photo-realistic bird's-eye view), medium (occupancy-style map), hard (bare floor plan) |
| Current View | Novel egocentric image from a viewpoint absent from the input video trajectory |
Task Taxonomy
Self Localization (2 tasks):
- Global Orientation: Infer camera orientation of current view on top-down image
- Global Position: Infer camera position of current view on top-down image
Object Localization (6 tasks, with semantic/visual modalities):
- Egocentric Direction: Infer direction from current viewpoint to target object (semantic: by category name; visual: by bounding box)
- Egocentric Distance: Estimate metric distance from current viewpoint to target object
- Global Position: Predict target object's position on top-down image
Scene Structure Understanding (4 tasks):
- Top-Down Selection (easy/medium/hard): Select the matching top-down image among four candidates
- Trajectory Selection: Match egocentric motion to one of four candidate trajectories on a top-down image
Data Generation Pipeline
The pipeline (built on OmniGibson simulation with BEHAVIOR-1K, HyperSim, ArtVIP) operates in six stages:
- Scene Preparation: 50 diverse indoor scenes; evaluation scenes held out from GST-Train
- Exploration Video Generation: Spatially distributed viewpoints connected into efficient traversal trajectories
- Object-Annotated Video Generation: Instance-level bounding boxes from rendered segmentation masks
- Top-Down Image Generation: Three abstraction levels rendered from scene geometry
- Current View Generation: Off-trajectory viewpoints with independent camera orientations; filtered for sufficient visual overlap
- QA Generation & Filtering: Template-based QA pairs with answers computed directly from 3D geometry; automatic filtering removes invalid samples
Human Verification: Annotators verify answerability (target recognizability in video, current-view localizability) — yielding 2,762 human-verified questions.
Evaluation Metrics
For distance prediction, Mean Relative Accuracy (MRA) with thresholds :
For angle prediction, circular error with tolerance thresholds :
For point prediction, Euclidean distance with thresholds pixels:
Overall score is the arithmetic mean across all 12 subtasks:
Empirical Validation / Results
Main Results
Key findings from evaluating 22 VLMs:
-
Huge gap between VLMs and humans: Human average = 79.08 vs. best model (Gemini-3-Pro) = 42.68 (gap of 36.4 points). The gap is most pronounced on orientation estimation (21.52 vs. 85.00) and global position estimation (42.23 vs. 93.00).
-
Open-source models struggle: Best open-source models (Qwen3-VL-32B: 30.43, InternVL3.5-38B: 30.71) remain near random on 10 of 12 subtasks. Four models score below random guessing.
-
Embodied-tuned models inherit the blind spot: RoboBrain2.5-8B (24.61) and Cosmos-Reason2-8B (21.64) fall below general-purpose Qwen3-VL-8B (25.89) at the 8B scale.
Table 1: Key Results on GST-Bench
| Method | Avg. Score | Notable Subtask Scores |
|---|---|---|
| Human Level | 79.08 | Ori: 85.00, GP: 93.00, TDS_E: 100.00 |
| Gemini-3-Pro | 42.68 | TDS_E: 98.15, TDS_M: 82.73, GP_v: 42.23 |
| Gemini-2.5-Pro | 40.95 | EDist_v: 42.00, TDS_E: 97.69 |
| GPT-5 | 40.85 | ED_s: 31.09, GP_s: 43.23 |
| Seed1.8 | 34.04 | TDS_E: 99.07 |
| GPT-4o | 30.33 | TDS_E: 81.02 |
| InternVL3.5-38B | 30.71 | TDS_E: 75.00, TDS_M: 68.64 |
| Qwen3-VL-32B | 30.43 | TDS_E: 94.44, TDS_M: 60.45 |
| Random Guessing | 20.01 | — |
| Qwen3-VL-8B (fine-tuned) | 53.52 | ED_v: 53.72, GP_s: 57.64, TDS_E: 99.07 |
Disentangling Global vs. Local Reasoning
Three controlled settings were constructed: Global (target absent from current view), Local-Video (target visible in current view, video redundant), and Local-Image (video removed, single-image task).
Proprietary models see locally but fail to integrate globally:
- Gemini-3-Pro on ED: 22.11 → 61.20 (Local-Image), nearly tripling
- GPT-5 on ED: 31.09 → 65.61 (Local-Image), +34.52 points
Open-source models fail at both perception and integration:
- Gains are inconsistent; some models degrade in easier settings (e.g., Qwen3-VL-8B on ED: 12.50 → 2.82 in Local-Video)
- Even in Local-Image, open-source scores remain far below proprietary counterparts
Bridging the Gap with Targeted Training
Fine-tuning Qwen3-VL-8B on GST-Train (with general-purpose instruction data mixing):
- Score improvement: 25.89 → 53.52 (+27.63 points)
- Surpasses all zero-shot proprietary models (Gemini-3-Pro: 42.68)
- Notable gains: ED_v: 18.75 → 53.72, GP_s: 13.01 → 57.64, Ori: 17.64 → 44.39
Theoretical and Practical Implications
Diagnostic Insights
-
Two qualitatively different bottlenecks exist between model families:
- Proprietary models have strong single-image spatial perception but fail at cross-frame spatial integration — the bottleneck is memory and integration
- Open-source models lack both local perception and global integration — the bottleneck is more fundamental
-
Egocentric distance estimation is intrinsically difficult: Human baseline itself is low (~41 MRA), suggesting metric distance recovery is a hard problem for both humans and machines.
-
Embodied training recipes are insufficient: Despite explicit tuning for grounded spatial reasoning, embodied models (RoboBrain, Robix, Cosmos-Reason2) show no advantage — current embodied post-training emphasizes local affordances rather than long-horizon spatial memory.
Practical Implications
- GST-Train demonstrates tractability: Targeted supervision can significantly narrow the gap (+27.63 points), suggesting the bottleneck is addressable with appropriate training data.
- Benchmark design matters: The strict "target-invisible" constraint and precise numerical metrics provide a sharper diagnostic tool than coarse categorical evaluations.
- Simulation-based generation scales: The automatic pipeline (with human verification) enables large-scale, controllable benchmark construction.
Conclusion
GST-Bench reveals that global spatial awareness — the ability to build, maintain, and reason over globally consistent spatial representations from long-horizon video — remains a major unsolved challenge for current VLMs. Even the strongest proprietary models achieve less than 55% of human performance, while most open-source models perform near random.
The authors' three main contributions are:
- GST-Bench: A rigorously constructed benchmark that enforces global reasoning by design (target-invisible constraint, off-trajectory viewpoints, precise metrics, top-down representations)
- Diagnostic analysis: Controlled local variants that disentangle perception failures from integration failures, revealing distinct bottlenecks for proprietary vs. open-source models
- GST-Train: A complementary training dataset demonstrating that targeted supervision can substantially improve global spatial reasoning
Future directions include: developing architectures with explicit spatial memory mechanisms, exploring better training recipes for cross-frame integration, and extending the benchmark to more complex, real-world scenarios.
"We hope GST-Bench and GST-Train will facilitate future research toward VLMs that can build, maintain, and reason over globally consistent spatial representations."
Related papers
- Hardening Agent Benchmarks with Adversarial Hacker-Fixer Loops
The hacker-fixer loop, using verifier access and a shared defense pool, automates hardening of agent benchmarks, eliminating all tested reward-hacking attacks on KernelBench and reducing them on Terminal Bench.
- Phantom Gains: Auditing Self-Improvement Against a Measured Null
Transition-level auditing of LLM self-improvement requires measured nulls for every statistic; without them, a frozen model falsely appears to expand at 0.280.
- FormalTCS: Benchmarking End-to-End Frontier Formal Theoretical Computer Science Research of Large Language Models
FORMALTCS reveals LLMs' autoformalization is the critical bottleneck in frontier TCS research, with top models scoring only 11.5 versus 28.6 on proving human-provided formal statements.