# OneStreamer: Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction

> OneStreamer unifies streaming video perception, memory, and proactive response via hierarchical caption memory and selective state supervision, achieving state-of-the-art results across all eight benchmarks.

- **Source:** [arXiv](https://arxiv.org/abs/2610.01762)
- **Published:** 2026-10-03
- **Permalink:** https://picx.dev/p/v5Ihy2
- **Whiteboard:** https://picx.dev/p/v5Ihy2/image

## Summary

## Summary (Overview)

- **OneStreamer** is a streaming video LLM that unifies query-independent evidence recording and task response through a single proactive generation process, jointly learning *when* to record/respond and *what* to generate.
- **Proactive Hierarchical Caption Memory (PHCM)** generates time-grounded local-detail captions (</Observe>) and semantic summaries (</Summary>) that complement a recent visual window, providing reusable factual context without revisiting historical visual features.
- **Proactive State Transition Learning (PSTL)** addresses state-label imbalance by preserving supervision at all output anchors and selecting representative state-change and state-persistence tokens, supervising only **27.5%** of annotated state tokens.
- **OneStreamer-1M** is a broad-coverage streaming video interaction dataset (over 1 million records) built via a reusable streaming data synthesis pipeline that aligns output content and timing with available evidence.
- The **4B model achieves state-of-the-art results across all eight evaluated streaming video understanding benchmarks**, with an average relative improvement of **25.0%** over the Qwen3-VL baseline and **6.1%** over the strongest competing method.

---

## Introduction and Theoretical Foundation

Real-world video arrives as an open-ended stream of observations. In live-stream assistants, wearable agents, and security monitoring systems, a model must interpret incoming visual evidence without knowing which observations will matter to future questions. Key challenges:

1. **Evidence retention**: An event that initially appears unimportant may become relevant later, after its source frames have left the model's limited visual context.
2. **Timely response**: A user-visible response must be grounded in available evidence and delivered before the relevant moment passes.

Existing approaches face fundamental limitations:

- **Memory compression/retrieval** methods extend accessible history but retained historical visual tokens compete with recent observations for limited context capacity.
- **Streaming thinking** methods maintain evolving reasoning, but reasoning traces do not necessarily serve as reusable time-grounded factual records.
- **State-token-based response control** enables proactive outputs, but dense supervision overemphasizes repeated silence labels relative to sparse response decisions.

The central research question: *How can a streaming video LLM turn incoming observations into reusable factual memory for timely responses that draw on both past and current evidence?*

The key insight is to **rethink the role of proactive generation**: beyond producing user-visible responses, a streaming model can learn to turn observed evidence into reusable factual records, connecting perception learning with query-independent memory formation through a shared causal generation process.

---

## Methodology

### 3.1 Overall Architecture

OneStreamer processes a live video stream incrementally as temporally ordered clips:

- A **vision encoder** extracts visual tokens from each incoming clip; a **projector** maps them into the LLM's embedding space.
- A **Recent-N FIFO sliding window** retains only visual tokens from the latest $N$ observed frames.
- Visual tokens are interleaved with a **text history** containing accumulated caption records and dialogue.
- The LLM predicts **task-specific control tokens** and generates associated text when the predicted state initiates an output.

**Control tokens**:
- `</Observe>`: introduces a local-detail caption of observed objects, actions, scenes, and state changes
- `</Summary>`: introduces a semantic summary of a completed event or segment
- `</Standby>`: relevant evidence is emerging but the model is not yet ready to answer
- `</Response>`: initiates user-visible answer generation
- `</Silence>`: continued observation without generating caption or response

### 3.2 Proactive Hierarchical Caption Memory (PHCM)

PHCM organizes observed evidence into two types of time-aligned records:

| Record Type | Control Token | Granularity | Content |
|---|---|---|---|
| Local-detail captions | `</Observe>` | Short intervals | Directly observed objects, actions, scenes, state changes |
| Semantic summaries | `</Summary>` | Coarse temporal granularity | Completed events or segments |

At inference, OneStreamer incrementally builds PHCM using only currently available visual-language context. Each generated record is appended to the text history in generation order, providing long-range context **without retrieving or revisiting historical visual features**.

### 3.3 Proactive State Transition Learning (PSTL)

PSTL addresses state-label imbalance where repeated silence tokens outnumber output-initiating tokens. The method:

1. Groups state tokens by the **transition** from the preceding control state to the current one.
2. Defines a **supervision quota** $q$ shared across all transition groups in a sequence:

$$q = \max_{X} \max_{a \in A_\tau} n_{X \to a}$$

where $A_\tau$ is the output-anchor set for task $\tau$, and $n_{X \to Y}$ is the number of occurrences of transition $X \to Y$ between consecutive annotated control states.

3. From each transition group $X \to Y$, samples state-token indices uniformly without replacement to form $I_{X \to Y}$ of size:

$$|I_{X \to Y}| = \min(n_{X \to Y}, q)$$

4. Only state tokens indexed by $I = \bigcup_{X,Y} I_{X \to Y}$ contribute to the state loss.

This preserves supervision for all output-initiating tokens while selecting examples of both state changes and state persistence. Unselected state tokens remain in the causal sequence but are excluded from the state loss.

### 4 Dataset Construction: OneStreamer-1M

**Streaming Caption Synthesis** (3 stages):
1. **High-quality video curation**: semantic retrieval, scene detection, visual-richness assessment
2. **Multi-granularity caption annotation**: Gemini and Seed generate timestamp-grounded captions at frame/clip, segment, and video levels, verified for factual and temporal consistency
3. **Causal streaming sequence construction**: frame/clip captions become `</Observe>` targets at interval ends; segment captions become `</Summary>` targets at segment boundaries

**Streaming QA Synthesis** (3 stages):
1. **Task-directed streaming QA design**: task-specific templates for streaming interaction capabilities
2. **Coarse evidence interval localization**: VLM localizes temporal interval containing supporting evidence; verification via clip-only answering
3. **Fine-grained response-time calibration**: candidate timestamps scored by conditional likelihood (reasoning QA) or visual-text similarity (perception tasks); earliest reliable timestamp selected as response time

**OneStreamer-1M task composition** includes: hierarchical captioning (52,500), flat captioning (7,500), alerts & reminders (173,930), continuous description (76,220), visual interaction (9,048), active counting (7,633), contextual feedback (4,642), procedural assistance (97), general video QA (162,866 + 97,229), gesture & expression QA (20,782 + 3,606), video dialogue QA (10,000), temporal QA (1,824), procedural QA (1,263 + 12,877), anomaly QA (849), spatial QA (360,603), temporal grounding (92,283), counting QA (26,149), future prediction (17,193), risk reasoning (13,081), answerability judgment (6,515), person identification (1,314).

---

## Empirical Validation / Results

### 5.1 Main Results

**Table 1: Online video benchmark results** — OneStreamer achieves the highest score on all eight benchmarks:

| Benchmark | OneStreamer (4B) | Qwen3-VL (4B) | MOSS-VL-Realtime (11B) | Best Competitor |
|---|---|---|---|---|
| OVOBench (Overall) | **72.1** | 58.8 | 70.2 | 70.2 |
| StreamingBench (Real-Time) | **86.9** | 81.8 | 82.9 | 83.2 |
| OVBench (Avg) | **66.8** | 55.4 | 53.7 | 65.4 |
| ODVBench (Overall) | **71.3** | 57.6 | 63.9 | 70.8 |
| ProactiveVideoQA (Avg) | **48.7** | 34.3 | 47.2 | 47.2 |
| OmniMMI (Avg) | **36.6** | 29.4 | 32.7 | 32.7 |
| OVO-Timing (Avg. F1) | **41.6** | 29.4 | 38.5 | 38.5 |
| ViSpeak (Overall) | **2.87** | 2.41 | 2.48 | 2.48 |

### 5.2 Ablations

**Finding 1: Proactively generated captions provide reusable temporal memory beyond the recent visual window.**

**Table 2: Effects of PHCM** — Retaining generated captions improves historical QA:

| Variant | Visual Context | Caption Memory | OVOBench Backward (ASI) | OVOBench Backward (EPM) | OVOBench Real-Time (Avg) | OVOBench Overall | StreamingBench Real-Time (Avg) |
|---|---|---|---|---|---|---|---|
| Full | Full History | × | 67.6 | 63.0 | 70.9 | 70.9 | 80.5 |
| FIFO | Recent-16 | × | 63.5 | 62.0 | 80.9 | 70.1 | 86.3 |
| **PHCM (Ours)** | Recent-16 | ✓ | **71.6** | 62.6 | **81.4** | **72.1** | **86.9** |

**Finding 2: Combining recent visual context with distant caption memory jointly improves perception and memory.** PHCM exceeds FIFO's real-time perception (81.4 vs. 80.9) and Full's memory performance (ASI: 71.6 vs. 67.6) within a single representation.

**Finding 3: PSTL outperforms dense state-token supervision.**

**Table 3: Ablation of state-token supervision**:

| State Selection | Loss | Supervision Ratio | ProactiveVQA (Avg) | OmniMMI (Avg) | OVO-Timing (Avg F1) |
|---|---|---|---|---|---|
| All state tokens | CE | 100.0% | 26.1 | 30.8 | 1.5 |
| All state tokens | Focal | 100.0% | 47.0 | 32.8 | 26.5 |
| Random Sparse | CE | 27.5% | 26.6 | 27.4 | 1.4 |
| Transition Only | CE | 18.1% | 46.7 | 27.4 | 15.3 |
| **PSTL (Ours)** | CE | **27.5%** | **48.7** | **36.6** | **41.6** |

### 5.3 Efficiency and Latency

**Table 4: Answer-stage efficiency** (360s OVOBench sample, single NVIDIA H200):

| Strategy | Visual Context | Caption Memory | GPU Memory (GB) ↓ | Context Tokens ↓ | TTFT (s) ↓ |
|---|---|---|---|---|---|
| Full | Full History | × | 25.18 | 62,094 | 4.560 |
| FIFO | Recent-16 | × | 9.69 | 3,036 | 0.094 |
| **PHCM (Ours)** | Recent-16 | ✓ | **9.98** | **4,308** | **0.124** |

PHCM reduces context length by **93.1%** and GPU memory by **60.4%** relative to Full, while reducing TTFT from 4.560s to 0.124s. Compared to FIFO, it adds only 1,272 context tokens, 0.29 GB GPU memory, and 0.030s TTFT.

---

## Theoretical and Practical Implications

1. **Proactive generation as a shared learning interface**: The results support proactive generation as a unified mechanism connecting visual perception, reusable factual memory, and timely task response within a single streaming model—challenging the view that memory and perception require separate architectural components.

2. **Heterogeneous memory representation**: PHCM demonstrates that text-based factual records can effectively replace retained visual history for long-range context, providing a more efficient alternative to visual token retention under limited context capacity.

3. **Selective supervision beats dense supervision**: PSTL shows that selectively supervising representative state transitions and persistence tokens outperforms both dense supervision and frequency-balanced focal loss, with only 27.5% of state tokens supervised—suggesting that *what* you supervise matters more than *how much*.

4. **Practical deployment efficiency**: The 93.1% context reduction and 36.8× TTFT improvement over full-history retention make streaming video interaction feasible for real-time applications with strict latency requirements.

5. **Data synthesis as a reusable resource**: The streaming data synthesis pipeline provides a template for converting offline video resources into evidence-aligned streaming training sequences, addressing the scarcity of high-quality streaming interaction data.

---

## Conclusion

OneStreamer presents a unified framework for streaming video interaction where proactive generation serves as a shared learning mechanism for:

- **Perception**: interpreting observed video prefixes through causal streaming caption targets
- **Memory formation**: turning observed content into time-grounded, reusable factual records via PHCM
- **Timely response**: learning when available evidence warrants a response through PSTL

Key contributions:
1. **PHCM** complements a recent visual window with hierarchical caption memory, improving historical QA without degrading real-time perception
2. **PSTL** outperforms dense state supervision while supervising fewer state tokens, demonstrating that selective supervision of state changes and persistence is more effective
3. **OneStreamer-1M** provides a broad-coverage training corpus with evidence-aligned content and timing

The 4B model achieves state-of-the-art results across all eight streaming video understanding benchmarks, with an average relative improvement of 25.0% over the Qwen3-VL baseline.

**Future directions** may include scaling to larger models, extending to additional task families, exploring adaptive memory granularity, and investigating reinforcement learning approaches for response timing policies.

---

_Markdown view of https://picx.dev/p/v5Ihy2, served by PicX — AI-generated visual whiteboard summaries of research papers._
