# Raven: The Harness of Harnesses for Composable Agentic Intelligence

> Raven's Host Agent orchestrates model-harness pairs as composable units via DAG-based planning, beating Claude Code and Hermes Agent across coding, research, and design benchmarks.

- **Source:** [arXiv](https://arxiv.org/abs/2609.33439)
- **Published:** 2026-10-01
- **Permalink:** https://picx.dev/p/83EjuD
- **Whiteboard:** https://picx.dev/p/83EjuD/image

## Summary

# Summary of "Raven: The Harness of Harnesses for Composable Agentic Intelligence"

## Summary (Overview)

- **Novel Architecture**: Introduces Raven, an open-source multi-agent ecosystem that treats each executable model–harness pair as a composable unit, enabling automatic construction, evolution, and orchestration of specialized agents across domains.
- **Host Agent Orchestration**: A central Host Agent decomposes goals, matches subtasks to registered capabilities, and represents execution dependencies as a directed acyclic graph (DAG), enabling flexible multi-agent collaboration.
- **Harness Self-Evolution**: Raven implements modular harness evolution that improves execution policies around frozen language models, with empirical evidence showing gains transfer beyond training tasks.
- **State-of-the-Art Performance**: Outperforms strong baselines (Claude Code, Hermes Agent) on multi-agent orchestration (MAOB benchmark), research (FRAMES, xBench), coding (SWE-bench), design (PresentBench), and on-call tasks.
- **Theoretical Foundation**: Provides formal analysis of composable agentic intelligence, characterizing sufficient conditions for reliable composition of model–harness pairs under resource constraints.

## Introduction and Theoretical Foundation

The paper addresses two central challenges in AI agent development:

1. **Harness Complexity**: As LLMs advance, the harness (tool interfaces, context management, skills, execution policies, recovery mechanisms) surrounding a model becomes increasingly complex to design manually.

2. **Domain Specificity**: Tighter coupling to specific domains limits the generality of a single harness, making it difficult to scale across diverse tasks.

The authors formalize **composable agentic intelligence** as reliable task coverage under a common resource budget. A model–harness pair is treated as a callable execution unit, and composition is analyzed through formal conditions:

- A plan is **well-typed** when its input, output, and message records have declared types and all identifier references resolve.
- A well-typed plan is compatible with task $t$ and initial set $I$ if $I_0$ holds throughout and specific implications hold for every well-typed state and ledger transition.

The theoretical framework characterizes when composition can extend the capabilities of available agents, providing sufficient conditions for reliable composition. The acceptance function is defined as:

$$\text{ok}_{\cdot D,b}(c_{cfg}, o) = \begin{cases} 1 & \text{if accepted completion in finite time satisfying transition predicate } D \text{ at cost} \leq b \\ 0 & \text{otherwise} \end{cases}$$

## Methodology

### System Architecture

**Host Agent and Execution Adapters**: The ecosystem coordinates native specialists (Raven-Research, Raven-Code, Raven-Design, Raven-Oncall) alongside third-party agents (Claude Code, Codex, Hermes Agent, OpenClaw) through shared execution adapters.

**DAG-Based Orchestration**: The Host Agent:
- Decomposes goals into subtasks
- Matches subtasks to registered capabilities
- Represents dependencies as a directed acyclic graph
- Integrates results through dependency-scoped provenance

**Memory and Experience**: 
- **EverOS**: Organizes memory through extraction, consolidation, and retrieval processes for long-horizon reasoning
- **Skill Forge**: Exposes reusable procedures with sixteen task-domain labels for organization
- Persistent context preserves user experience across sessions

**Harness Self-Evolution**: The system improves execution policies surrounding frozen language models through:
- Repeated adaptive screening on training tasks
- Archive competition for candidate selection
- Held-out comparison for validation

### Evaluation Framework

The **Multi-Agent Orchestration Benchmark (MAOB)** contains 140 requests modeled on occupational tasks, each paired with a reference DAG over a fixed roster of four specialist sub-agents. Four domains yield exactly 11 subsets containing at least two domains.

## Empirical Validation / Results

### Multi-Agent Orchestration (MAOB)

Raven outperformed both baseline systems (Claude Code and Hermes Agent) on all four metrics with each backbone (Qwen and DeepSeek), demonstrating effective specialist selection and dependency planning across model families.

### Raven-Research Results

| Backbone | Raven-Research | Strongest Baseline | Gain |
|----------|---------------|-------------------|------|
| Qwen | Highest pooled accuracy | - | +6.6 pts |
| DeepSeek | Highest pooled accuracy | - | +3.3 pts |
| Third backbone | Highest pooled accuracy | - | +7.6 pts |

All harnesses exceeded 86% on FRAMES, with consistent gains across different backbones.

### Raven-Code Results

With Claude Opus 5 as the main model:
- **Pass@1**: 0.8762
- **Pass@5**: 0.9097

Raven-Code remained competitive across tested development settings with larger advantages on migration tasks.

### Raven-Design Results

On PresentBench (238 slide-generation tasks from five domains), Raven-Design showed gains of **1.9 and 20.5 points** over alternatives, with particularly large improvements for GPT-5.6 Luna.

### Comparison with Strongest Alternatives

Raven led the strongest alternative (Hermes Agent in four settings, Claude Code in two) by **0.3 to 2.8 points** across various benchmarks.

## Theoretical and Practical Implications

### Theoretical Contributions

The paper provides a formal framework for understanding when and how model–harness pairs can be reliably composed. Key insights include:
- Composition reliability depends on well-typed plans, clear contracts, and resource bounds
- The framework distinguishes between execution structure and output quality
- Sufficient conditions for reliable composition are identified, with practical applicability depending on implementation

### Practical Implications

1. **Open Ecosystem**: The architecture supports both native and third-party agents, promoting a collaborative ecosystem rather than a closed system.
2. **Cross-Domain Orchestration**: The DAG-based approach enables flexible composition across research, coding, design, and on-call domains.
3. **Experience Reuse**: Skill Forge and EverOS enable continuous improvement through accumulated experience.
4. **Benchmark Contributions**: MAOB provides a new evaluation standard for multi-agent orchestration.

## Conclusion

Raven addresses the growing complexity of harness design and the difficulty of composing specialized capabilities across domains by treating each executable model–harness pair as a unit of intelligence. The Host Agent organizes collaboration through explicit task dependencies, while modular harness evolution and persistent memory support adaptation.

**Future Work**: The authors envision an **All-Domain Collaboration Network**—an interconnected multi-agent system spanning everyday devices and environments. Each device would have a local Raven instance managing agents within its scope, with instances coordinating collaboration across devices. This network would function as a persistent AI assistant maintaining continuity across devices, with end-to-end evaluations measuring task success, user experience continuity, and coordination cost.

---

_Markdown view of https://picx.dev/p/83EjuD, served by PicX — AI-generated visual whiteboard summaries of research papers._
