Full text not available for this paper
Summary (Overview)
- OneStreamer is a streaming video LLM that unifies query-independent evidence recording and task response through a single proactive generation process, jointly learning when to record/respond and what to generate.
- Proactive Hierarchical Caption Memory (PHCM) generates time-grounded local-detail captions (</Observe>) and semantic summaries (</Summary>) that complement a recent visual window, providing reusable factual context without revisiting historical visual features.
- Proactive State Transition Learning (PSTL) addresses state-label imbalance by preserving supervision at all output anchors and selecting representative state-change and state-persistence tokens, supervising only 27.5% of annotated state tokens.
- OneStreamer-1M is a broad-coverage streaming video interaction dataset (over 1 million records) built via a reusable streaming data synthesis pipeline that aligns output content and timing with available evidence.
- The 4B model achieves state-of-the-art results across all eight evaluated streaming video understanding benchmarks, with an average relative improvement of 25.0% over the Qwen3-VL baseline and 6.1% over the strongest competing method.
Introduction and Theoretical Foundation
Real-world video arrives as an open-ended stream of observations. In live-stream assistants, wearable agents, and security monitoring systems, a model must interpret incoming visual evidence without knowing which observations will matter to future questions. Key challenges:
- Evidence retention: An event that initially appears unimportant may become relevant later, after its source frames have left the model's limited visual context.
- Timely response: A user-visible response must be grounded in available evidence and delivered before the relevant moment passes.
Existing approaches face fundamental limitations:
- Memory compression/retrieval methods extend accessible history but retained historical visual tokens compete with recent observations for limited context capacity.
- Streaming thinking methods maintain evolving reasoning, but reasoning traces do not necessarily serve as reusable time-grounded factual records.
- State-token-based response control enables proactive outputs, but dense supervision overemphasizes repeated silence labels relative to sparse response decisions.
The central research question: How can a streaming video LLM turn incoming observations into reusable factual memory for timely responses that draw on both past and current evidence?
The key insight is to rethink the role of proactive generation: beyond producing user-visible responses, a streaming model can learn to turn observed evidence into reusable factual records, connecting perception learning with query-independent memory formation through a shared causal generation process.
Methodology
3.1 Overall Architecture
OneStreamer processes a live video stream incrementally as temporally ordered clips:
- A vision encoder extracts visual tokens from each incoming clip; a projector maps them into the LLM's embedding space.
- A Recent-N FIFO sliding window retains only visual tokens from the latest observed frames.
- Visual tokens are interleaved with a text history containing accumulated caption records and dialogue.
- The LLM predicts task-specific control tokens and generates associated text when the predicted state initiates an output.
Control tokens:
</Observe>: introduces a local-detail caption of observed objects, actions, scenes, and state changes</Summary>: introduces a semantic summary of a completed event or segment</Standby>: relevant evidence is emerging but the model is not yet ready to answer</Response>: initiates user-visible answer generation</Silence>: continued observation without generating caption or response
3.2 Proactive Hierarchical Caption Memory (PHCM)
PHCM organizes observed evidence into two types of time-aligned records:
| Record Type | Control Token | Granularity | Content |
|---|---|---|---|
| Local-detail captions | </Observe> | Short intervals | Directly observed objects, actions, scenes, state changes |
| Semantic summaries | </Summary> | Coarse temporal granularity | Completed events or segments |
At inference, OneStreamer incrementally builds PHCM using only currently available visual-language context. Each generated record is appended to the text history in generation order, providing long-range context without retrieving or revisiting historical visual features.
3.3 Proactive State Transition Learning (PSTL)
PSTL addresses state-label imbalance where repeated silence tokens outnumber output-initiating tokens. The method:
- Groups state tokens by the transition from the preceding control state to the current one.
- Defines a supervision quota shared across all transition groups in a sequence:
where is the output-anchor set for task , and is the number of occurrences of transition between consecutive annotated control states.
- From each transition group , samples state-token indices uniformly without replacement to form of size:
- Only state tokens indexed by contribute to the state loss.
This preserves supervision for all output-initiating tokens while selecting examples of both state changes and state persistence. Unselected state tokens remain in the causal sequence but are excluded from the state loss.
4 Dataset Construction: OneStreamer-1M
Streaming Caption Synthesis (3 stages):
- High-quality video curation: semantic retrieval, scene detection, visual-richness assessment
- Multi-granularity caption annotation: Gemini and Seed generate timestamp-grounded captions at frame/clip, segment, and video levels, verified for factual and temporal consistency
- Causal streaming sequence construction: frame/clip captions become
</Observe>targets at interval ends; segment captions become</Summary>targets at segment boundaries
Streaming QA Synthesis (3 stages):
- Task-directed streaming QA design: task-specific templates for streaming interaction capabilities
- Coarse evidence interval localization: VLM localizes temporal interval containing supporting evidence; verification via clip-only answering
- Fine-grained response-time calibration: candidate timestamps scored by conditional likelihood (reasoning QA) or visual-text similarity (perception tasks); earliest reliable timestamp selected as response time
OneStreamer-1M task composition includes: hierarchical captioning (52,500), flat captioning (7,500), alerts & reminders (173,930), continuous description (76,220), visual interaction (9,048), active counting (7,633), contextual feedback (4,642), procedural assistance (97), general video QA (162,866 + 97,229), gesture & expression QA (20,782 + 3,606), video dialogue QA (10,000), temporal QA (1,824), procedural QA (1,263 + 12,877), anomaly QA (849), spatial QA (360,603), temporal grounding (92,283), counting QA (26,149), future prediction (17,193), risk reasoning (13,081), answerability judgment (6,515), person identification (1,314).
Empirical Validation / Results
5.1 Main Results
Table 1: Online video benchmark results — OneStreamer achieves the highest score on all eight benchmarks:
| Benchmark | OneStreamer (4B) | Qwen3-VL (4B) | MOSS-VL-Realtime (11B) | Best Competitor |
|---|---|---|---|---|
| OVOBench (Overall) | 72.1 | 58.8 | 70.2 | 70.2 |
| StreamingBench (Real-Time) | 86.9 | 81.8 | 82.9 | 83.2 |
| OVBench (Avg) | 66.8 | 55.4 | 53.7 | 65.4 |
| ODVBench (Overall) | 71.3 | 57.6 | 63.9 | 70.8 |
| ProactiveVideoQA (Avg) | 48.7 | 34.3 | 47.2 | 47.2 |
| OmniMMI (Avg) | 36.6 | 29.4 | 32.7 | 32.7 |
| OVO-Timing (Avg. F1) | 41.6 | 29.4 | 38.5 | 38.5 |
| ViSpeak (Overall) | 2.87 | 2.41 | 2.48 | 2.48 |
5.2 Ablations
Finding 1: Proactively generated captions provide reusable temporal memory beyond the recent visual window.
Table 2: Effects of PHCM — Retaining generated captions improves historical QA:
| Variant | Visual Context | Caption Memory | OVOBench Backward (ASI) | OVOBench Backward (EPM) | OVOBench Real-Time (Avg) | OVOBench Overall | StreamingBench Real-Time (Avg) |
|---|---|---|---|---|---|---|---|
| Full | Full History | × | 67.6 | 63.0 | 70.9 | 70.9 | 80.5 |
| FIFO | Recent-16 | × | 63.5 | 62.0 | 80.9 | 70.1 | 86.3 |
| PHCM (Ours) | Recent-16 | ✓ | 71.6 | 62.6 | 81.4 | 72.1 | 86.9 |
Finding 2: Combining recent visual context with distant caption memory jointly improves perception and memory. PHCM exceeds FIFO's real-time perception (81.4 vs. 80.9) and Full's memory performance (ASI: 71.6 vs. 67.6) within a single representation.
Finding 3: PSTL outperforms dense state-token supervision.
Table 3: Ablation of state-token supervision:
| State Selection | Loss | Supervision Ratio | ProactiveVQA (Avg) | OmniMMI (Avg) | OVO-Timing (Avg F1) |
|---|---|---|---|---|---|
| All state tokens | CE | 100.0% | 26.1 | 30.8 | 1.5 |
| All state tokens | Focal | 100.0% | 47.0 | 32.8 | 26.5 |
| Random Sparse | CE | 27.5% | 26.6 | 27.4 | 1.4 |
| Transition Only | CE | 18.1% | 46.7 | 27.4 | 15.3 |
| PSTL (Ours) | CE | 27.5% | 48.7 | 36.6 | 41.6 |
5.3 Efficiency and Latency
Table 4: Answer-stage efficiency (360s OVOBench sample, single NVIDIA H200):
| Strategy | Visual Context | Caption Memory | GPU Memory (GB) ↓ | Context Tokens ↓ | TTFT (s) ↓ |
|---|---|---|---|---|---|
| Full | Full History | × | 25.18 | 62,094 | 4.560 |
| FIFO | Recent-16 | × | 9.69 | 3,036 | 0.094 |
| PHCM (Ours) | Recent-16 | ✓ | 9.98 | 4,308 | 0.124 |
PHCM reduces context length by 93.1% and GPU memory by 60.4% relative to Full, while reducing TTFT from 4.560s to 0.124s. Compared to FIFO, it adds only 1,272 context tokens, 0.29 GB GPU memory, and 0.030s TTFT.
Theoretical and Practical Implications
-
Proactive generation as a shared learning interface: The results support proactive generation as a unified mechanism connecting visual perception, reusable factual memory, and timely task response within a single streaming model—challenging the view that memory and perception require separate architectural components.
-
Heterogeneous memory representation: PHCM demonstrates that text-based factual records can effectively replace retained visual history for long-range context, providing a more efficient alternative to visual token retention under limited context capacity.
-
Selective supervision beats dense supervision: PSTL shows that selectively supervising representative state transitions and persistence tokens outperforms both dense supervision and frequency-balanced focal loss, with only 27.5% of state tokens supervised—suggesting that what you supervise matters more than how much.
-
Practical deployment efficiency: The 93.1% context reduction and 36.8× TTFT improvement over full-history retention make streaming video interaction feasible for real-time applications with strict latency requirements.
-
Data synthesis as a reusable resource: The streaming data synthesis pipeline provides a template for converting offline video resources into evidence-aligned streaming training sequences, addressing the scarcity of high-quality streaming interaction data.
Conclusion
OneStreamer presents a unified framework for streaming video interaction where proactive generation serves as a shared learning mechanism for:
- Perception: interpreting observed video prefixes through causal streaming caption targets
- Memory formation: turning observed content into time-grounded, reusable factual records via PHCM
- Timely response: learning when available evidence warrants a response through PSTL
Key contributions:
- PHCM complements a recent visual window with hierarchical caption memory, improving historical QA without degrading real-time perception
- PSTL outperforms dense state supervision while supervising fewer state tokens, demonstrating that selective supervision of state changes and persistence is more effective
- OneStreamer-1M provides a broad-coverage training corpus with evidence-aligned content and timing
The 4B model achieves state-of-the-art results across all eight streaming video understanding benchmarks, with an average relative improvement of 25.0% over the Qwen3-VL baseline.
Future directions may include scaling to larger models, extending to additional task families, exploring adaptive memory granularity, and investigating reinforcement learning approaches for response timing policies.
Related papers
- On-Demand Attention: Language Models Know When to Recall
On-demand attention uses a lightweight recall head to predict when global attention helps, recovering most quality with up to 2.65x decoding throughput.
- KV-Kaizen: Learning Context-Adaptive Cache Compression Choices
KV-Kaizen learns per-layer cache compression choices across depth, rank, and precision, achieving 4x compression with no accuracy loss on models 7B and larger.
- LLM sequential decision making under uncertainty in biochemical domains
LLMs overreact to new data and fail to explore in Bayesian optimization due to context-stickiness, a competence gap, not an intention gap, fixable by removing in-context history.