Summary (Overview)

  • Declarative Attention (DA) is a novel zero-shot protocol that enables off-the-shelf LLMs to explicitly declare their attention scope via structured tags (<global>, <focus>, <local>) within their chain-of-thought, allowing the inference engine to skip KV cache reads dynamically.
  • Across 15 long-context tasks, DA reduces average attended tokens by 52.0% (Gemma-4-31B) and 31.1% (Qwen-3.6-27B) with modest accuracy drops of 1.27pp and 2.75pp respectively.
  • The approach eliminates the O(N)O(N) per-step selection cost of prior sparse attention methods by deriving the mask directly from the model's own text output rather than approximating it via proxy scores.
  • DA shows positive scaling: accuracy gaps close as model size increases (from 4B to 31B), and absolute token savings grow sharply with context length (up to 21M tokens saved per response).
  • A vLLM integration with block-aligned KV cache masking projects decode wall-clock time reductions to 0.71× (Gemma) and 0.77× (Qwen) of vanilla baselines on optimized serving stacks.

Introduction and Theoretical Foundation

Background and Motivation

Transformers compute attention over every preceding token at each decoding step, making them computationally expensive for long-context tasks. The Key-Value (KV) cache memory access latency dominates decoding time—for instance, Qwen-3.5-397B-A17B requires roughly 15 GB of KV cache loaded per sequence at every decoding step for a 1M-token context.

However, attention scores naturally concentrate on a small subset of context tokens, diverging from the exhaustive mechanism's assumption. The fundamental challenge: true attention scores are unknown a priori—they only become available after computing the full attention matrix.

Prior Approaches and Limitations

ApproachMethodLimitation
Static heuristicsRecency, historical attention magnitudeCannot anticipate query-specific needs
Lightweight scansProxy scores over KV cacheStill incurs O(N)O(N) per-step cost

Key Insight

Recent studies show LLMs encode information about future tokens in hidden states, and chain-of-thought (CoT) surfaces this latent computation as interpretable text. DA extends this principle from dictating what to think to where to attend—eliciting the model to declare its attention plan explicitly.

Methodology

The DA Protocol

DA partitions generation into three attention modes:

  • <global>: Attends to all context segments—used for navigation and surveying the full context
  • <focus magic_chunks="K">: Attends only to named context segments—used for reasoning over a specific region
  • <local>: Attends to none of the context segments—used for self-contained reasoning over the response so far

Example response structure:

<global>
I need the founding year and the IPO year. The company history in Magic Chunk 2 should state the founding.
</global>
<focus magic_chunks="2">"Acme Corp was founded in 2003 in San Jose.</focus>
<global>
The IPO year is still missing. Magic Chunk 7 covers Acme's financial milestones.
</global>
<focus magic_chunks="7">"Acme Corp went public on the NYSE in 2011.</focus>
<local>2011 - 2003 = 8 years.</local>
<answer>8 years</answer>

Prompt Structure

  • Scaffold (always visible): System instruction, question, and DA instruction
  • Context (variable visibility): Long input delivered as addressable "magic chunks" (~2K tokens each) in a simulated tool-use transcript

Decode-Time Interventions

  1. State machine parsing: Watches for tag transitions in the output stream
  2. Block-aligned mask construction: Applies masks at block granularity (16–32 tokens) so existing kernels like FlashAttention run unchanged
  3. vLLM integration: Hooks into attention metadata builder to rewrite KV-cache block tables, with no kernel modifications

Roofline Wall-Time Analysis

DA trades more decode steps for lower per-step attention cost. The roofline model charges each operation at its hardware ceiling:

TFFN=FLOPsPeak FLOPs×MFUT_{FFN} = \frac{FLOPs}{Peak\ FLOPs \times MFU} Tattn=KV bytesPeak BW×MBUT_{attn} = \frac{KV\ bytes}{Peak\ BW \times MBU}

Experimental Setup

  • Models: Gemma-4-{31B, 12B, E4B}, Qwen-3.6-27B, Qwen-3.5-{9B, 4B}
  • Benchmarks: 15 long-context sources from RULER, LongBench v1/v2, LooGLE, ZeroScrolls
  • Baselines: Vanilla (raw context), DA-nm (DA prompt without masking)
  • Evaluation: LLM judge with Gemini-3-Flash rubrics, Qwen-3.5-4B as judge

Empirical Validation / Results

Main Results

MetricGemma-4-31BQwen-3.6-27B
Accuracy drop1.27pp (87.01% → 85.74%)2.75pp (85.31% → 82.56%)
Attended token reduction52.0% (13.43M → 6.45M)31.1% (22.54M → 15.52M)
Per-step attention ratio~0.5~0.5

Key Findings

  1. Mask is the source of efficiency: DA-nm attends 66.2% more tokens than vanilla (due to longer generations), but DA's mask converts this to 52.0% below vanilla—a 71.1% reduction relative to DA-nm on Gemma.

  2. Positive scaling: Relative accuracy rises monotonically with model size—from 29% at Gemma-4-E4B to 99% at Gemma-4-31B.

  3. Context-length scaling: Token savings grow from ~1M at short context to ~21M in the longest bin, while accuracy holds within ~1pp up to 32K tokens.

Wall-Clock Time Estimates

ModelArmMatmul (ms)Global Memory (ms)Local Memory (ms)Total (ms)
Gemma-4-31BVanilla22.9196.549.7269.1
Gemma-4-31BDA30.994.467.1192.3
Qwen-3.6-27BVanilla34.4263.88.0306.2
Qwen-3.6-27BDA45.1181.610.5237.3

Protocol Adherence

  • Focus success rate rises with model size: 58% (Gemma-4-E4B) → 99% (Gemma-4-31B)
  • Focus attempts per response stay in a narrow band (1.4–1.9) regardless of model size
  • <global> accounts for only ~27% of generated tokens; <focus> and <local> save 76–99% per token

Theoretical and Practical Implications

Architectural Relevance

DA's relevance depends on whether global attention remains a major decode cost. Analysis shows:

  • Attention accounts for 94% of decode wall-time at 1M-token contexts for Kimi-K3
  • 56–97% for indexer-based sparse designs (DeepSeek-V4-Pro, GLM-5.3-Flash, etc.)

Agentic Systems

DA addresses a complementary problem to retrieval: retrieval decides what enters context, DA decides what the model attends to among what's already there. Agentic contexts—where tool results accumulate—fall in the regime where DA's savings are largest.

Synergies with Existing Techniques

  • Lightweight scans: Can serve global steps where DA retains full costs
  • Speculative decoding: Collapses sequential steps, offsetting DA's extra decode steps
  • KV cache offloading: DA's span-based attention changes enable prefetching and reversible context compaction

Conclusion

DA demonstrates that off-the-shelf models can control their own attention span through text declarations, turning selective attention from a pattern inferred inside the network into one the model states explicitly. Key future directions include:

  1. Post-training optimization: RL pipelines rewarding both accuracy and attention efficiency could elicit more optimal DA usage
  2. Natural segments: Agentic contexts provide naturally addressable segments (tool calls, user turns)
  3. Thinking-mode integration: Exposing DA operations as standard tool declarations could extend savings to reasoning traces
  4. Global mode reduction: In-context indexes could replace full-fidelity navigation
  5. System-2 sparse attention: DA's legible attention plans serve oversight as much as efficiency

The results represent a lower bound—with training-based methods, DA's potential is expected to grow significantly.

Related papers