Summary (Overview)
- Omni-IO Skills is a plug-and-play Agent Harness that makes existing general-purpose agents (GPT-5.6 Sol, Claude Sonnet 5) "omni-native" without retraining their reasoning cores, enabling them to handle text, images, audio, video, documents, 3D assets, and code.
- The system uses hierarchical Skills (Atomic, Expert, Scenario), a standardized multimodal execution interface (MCP Tool Service), dependency-aware orchestration via Declare Execution Graphs (DEGs), and a persistent Asset Registry for cross-turn reuse.
- It comprises 27 Skills covering 38 representative tasks across four capability families: understanding, generation, reasoning, and retrieval.
- On the UniM-90 benchmark, the harness raises input-support rates from ~40% to 100% for both host agents, with relative Semantic–Quality Coupled Score (SQCS) improving from ~27 to 74.94–77.78 and Strict Structure Score reaching 100.00/99.78.
Introduction and Theoretical Foundation
The paper addresses a critical gap in multimodal AI: while general-purpose agents (e.g., Codex, Claude Code) excel at planning, reasoning, and long-horizon execution, their production capabilities remain fragmented across modalities. Extending foundation models to new modalities is costly (requiring retraining), while assembling specialist models leaves unresolved questions about procedure coordination, dependencies, intermediate assets, and cross-turn revisions.
The authors position Omni-IO Skills as a complementary route to Omni capability—a system layer between agent reasoning and heterogeneous execution backends that turns scattered models, tools, and procedures into selectable, composable, executable, and traceable capabilities. The central research question is:
Can a plug-and-play harness make an existing general-purpose agent omni-native while preserving its reasoning and planning core?
The theoretical foundation builds on two research lines:
- Omni foundation models (unified autoregressive vs. hybrid discrete–continuous designs) which couple capability growth to model updates.
- Agent Harnesses and Skills (e.g., ReAct, MCP, SkillsBench, CUA-Skill, MMSkills) which provide operational substrates but lack multi-asset workflow coordination with persistent artifact state.
Methodology
Four-Layer Architecture
The system positions itself between the host agent and external multimodal tools:
-
Skill Entry Layer: Organizes procedural knowledge into three levels:
- Atomic Skills (A1–A19): Single invocable operations (e.g., image understanding, video generation, code generation).
- Expert Skills (E1–E2): Complete workflows for concrete deliverables (Poster Design, Complex Video Production).
- Scenario Skills (S1–S6): Application-level tasks (Social-Media Post, Office Documents, Job Application, Education Sharing, Event Material, Game Asset).
-
MCP Tool Service Layer: Standardizes external capabilities into understanding, generation, and utility tool groups, converting provider responses into common results.
-
Provider and Configuration Layer: Separates tool capabilities from service implementations, enabling provider/model substitution without rewriting Skills.
-
Asset Registry Layer: Normalizes artifacts with a persistent, append-only JSON registry using atomic writes with file locking.
Declarative Skill Representation
A Skill is formally represented as:
where = applicability conditions, = required inputs, = execution procedure, = expected outputs, and = relationships to other Skills.
Declare Execution Graph (DEG)
A DEG is defined as where each node is:
Registered assets follow:
Runtime Orchestration
- Wave scheduling: Independent nodes execute concurrently; dependencies determine wave assignment.
- Failure handling: Failed nodes cancel pending descendants but independent branches continue.
- Asset propagation: Successful outputs are registered and reused within turns and across turns via
asset_ref.
Empirical Validation / Results
Experimental Setup
- Benchmark: UniM-90, a fixed 90-instance subset of UniM covering all seven modalities.
- Host agents: GPT-5.6 Sol and Claude Sonnet 5, evaluated as Base Agent vs. Agent + Omni-IO Skills.
- Metrics: Input-support rate (τ), absolute and relative variants of SQCS, ICS, StS, LeS.
Main Results (Table 3)
| Base Agent | τ | SQCS (Abs) | SQCS (Rel) | ICS (Rel) | StS | LeS |
|---|---|---|---|---|---|---|
| GPT-5.6 Sol | 40.00% | 67.49 | 26.99 | 34.61 | 47.84 | 72.22 |
| + Omni-IO Skills | 100% | 74.94 | 74.94 | 93.98 | 100.00 | 100.00 |
| Claude Sonnet 5 | 38.89% | 71.53 | 27.82 | 32.00 | 52.21 | 68.57 |
| + Omni-IO Skills | 100% | 77.78 | 77.78 | 83.28 | 99.78 | 100.00 |
Key findings:
- Input-support rates rise to 100% for both agents.
- Relative SQCS gains: +47.95 (GPT-5.6 Sol) and +49.96 (Claude Sonnet 5) percentage points.
- Absolute SQCS with harness exceeds Base Agent scores even on their narrower supported subsets (74.94 vs. 67.49; 77.78 vs. 71.53).
- Strict Structure Scores reach 100.00 and 99.78, validating output-structure control.
Case Studies
Qualitative evaluations demonstrated:
- Art tutorial generation: S4 Education Sharing coordinates video understanding (A2), audio understanding (A3), and image generation (A6) to produce an 11-image tutorial.
- Product promotion: A1 Image Understanding + E1 Poster Design + E2 Complex Video Production + A16 Code Generation produce poster, video, and landing page with consistent product identity.
Theoretical and Practical Implications
Theoretical implications:
- Demonstrates that harness-level capability composition is a viable third route to Omni systems, complementing unified autoregressive and hybrid discrete–continuous model designs.
- Separates task knowledge, tool access, provider binding, and artifact state into distinct layers, enabling independent evolution of each.
- Provides formal representations (DEG, asset records) for multi-asset workflow coordination with persistent provenance.
Practical implications:
- Enables existing agents to handle all seven artifact modalities without retraining, dramatically reducing the cost of modality expansion.
- Supports provider/model substitution without rewriting procedural knowledge, future-proofing against model improvements.
- Enables cross-turn asset reuse and localized revision (only affected components regenerated), improving efficiency in multi-turn production workflows.
- The plug-and-play nature means organizations can extend their existing agent deployments with Omni capabilities immediately.
Conclusion
Omni-IO Skills establishes that a plug-and-play Agent Harness can make general-purpose agents omni-native while preserving their reasoning cores. Key achievements:
- 100% input-support rate on UniM-90 for both GPT-5.6 Sol and Claude Sonnet 5.
- Substantial gains in semantic quality (SQCS +47.95/+49.96 relative points) and structural completeness (StS 100.00/99.78).
- 27 Skills covering 38 tasks across seven modalities and four capability families.
Future directions implied by the work include:
- Expanding the Skill library to additional application domains and task types.
- Further improving the quality of generated outputs through better Expert Skill procedures.
- Exploring integration with additional host agents and execution backends.
- Potential extensions to streaming, real-time, and full-duplex interaction scenarios.
The authors conclude that harness-level composition offers a practical, evolvable route to Omni capability, keeping procedures, providers, and assets independently extensible—providing an application-oriented foundation for agents coordinating heterogeneous media and reusable outputs across multi-turn production workflows.
Related papers
- Recursive self-improvement of AI research agents
AIDE² autonomously discovered seven recursive self-improvements in eight days, yielding an AI research agent that matches or exceeds a human-engineered agent across all held-out benchmarks while reducing reward hacking.
- CliffCompaction: Cost-Efficient Compaction for Long-Horizon Coding Agents
CLIFFCOMPACTION, a rule-based autocompaction method that discards stale context verbatim, cuts inference costs by up to 50% while improving coding agent performance and enabling state-of-the-art continual learning.
- LEGO-Anything: Coding Agents for 3D Scene Reconstruction
LEGO-Anything turns single images into editable 3D scene programs via coding agents, but current models achieve only about half the fidelity of specialist vision systems.