Skip to main content
Back to timeline
arXivSource publication:

Omni-IO Skills lifts GPT-5.6 Sol and Claude Sonnet 5 multimodal input support from about 40% to 100% with 27 skills

Synopsis

The work presents Omni-IO Skills, a plug-and-play Agent Harness that makes existing agents omni-native through hierarchical Skills, a standardized multimodal execution interface, dependency-aware orchestration, and a persistent Asset Registry, representing multi-asset workflows as Declare Execution Graphs; on UniM-90 it raises the input-support rates of GPT-5.6 Sol and Claude Sonnet 5 from 40.00% and 38.89% to 100%, increases relative Semantic–Quality Coupled Score from 26.99 to 74.94 and from 27.82 to 77.78, and reaches Strict Structure Scores of 100.00 and 99.78.

AI-generated editorial illustration: Omni-IO Skills: Harnessing Your Agent Omni-Native

Interpretation

It formulates and implements a plug-and-play Omni-modal Agent Harness that extends a general-purpose agent into an omni-native system without retraining the host model. Prior routes either tie capability growth to costly foundation-model updates or assemble specialist models and tools while leaving unresolved how procedures, dependencies, intermediate assets, and cross-turn revisions are coordinated; this work places unification at the level of task execution and artifact flow. The paper supports this with a four-layer architecture (Skill Entry, MCP Tool Service, Provider and Configuration, Asset Registry) and controlled experiments with two host agents.

It organizes multimodal capability as hierarchical Skills: 19 Atomic, 2 Expert, and 6 Scenario Skills, 27 in total, covering 38 representative tasks across seven artifact modalities and four capability families of understanding, generation, reasoning, and retrieval. Existing Agent Skills work spans general expert tasks, computer use, and visual agents, while Omni-agent research separately establishes the value of coordinating modality experts; their intersection remains underdeveloped for multi-asset execution with persistent artifact state. The paper provides a complete skill catalog (A1–A19, E1–E2, S1–S6) plus task-to-Skill mapping notes.

It expresses control and data dependencies as Declare Execution Graphs, schedules independent nodes concurrently in Waves, cancels pending descendants when a node fails while independent branches continue, and registers successful outputs in the Asset Registry for downstream and cross-turn reuse. Dependency-aware orchestration, failure isolation, and persistent artifact state are brought into one runtime so a workflow is not bound to a particular service or file path, and providers or models can be replaced while Skill descriptions and graph structure stay unchanged. The paper details graph representation, validation (resolvable dependencies and acyclicity), Wave scheduling rules, failure handling, and atomic file-locked registry writes, and illustrates four Waves plus a cross-turn rerun of only the landing page in a product-promotion case.

On UniM-90, both host agents reach 100% input-support rate, relative SQCS rises by 47.95 and 49.96 points, and structure metrics approach full marks. The authors frame the gains as within-agent comparisons before and after loading the harness, and note that post-harness absolute SQCS (74.94 and 77.78) exceeds the Base Agents' 67.49 and 71.53 measured only on their narrower supported subsets. The evaluation uses a controlled 90-instance subset covering text, image, audio, video, document, code, and 3D modalities and their interleaved combinations; the authors explicitly do not interpret cross-agent score differences as a ranking of model capabilities.

Perspective

The results target application settings that require cross-modal input and output plus multi-asset deliverables, such as education sharing, office documents, job applications, event materials, and game assets; the premise is that the host agent keeps its reasoning and planning core while the harness supplies procedural knowledge, an execution interface, and an asset substrate. For practitioners this means underlying models or services can be updated by swapping the Provider and Configuration layer without retraining the foundation model, while Skill descriptions and execution graphs stay unchanged; for researchers it offers a system-layer reference for decoupling multimodal capability growth from the model update cycle.

The evaluation centers on UniM-90, a controlled 90-instance subset covering seven modalities and interleaved combinations, but relative and absolute metrics use different denominators, so readers should note where relative and absolute SQCS diverge. The authors explicitly do not interpret differences between the two hosts as a ranking of model capabilities, so cross-agent comparisons warrant care. The case studies are qualitative, showing consistency across tutorial steps, product identity, and deliverables without case-level quantification. In addition, post-harness absolute SQCS is compared against Base Agent values measured only on their narrower supported subsets, so the two cover different instance ranges. The skill catalog, task-to-Skill mappings, and some figures and tables sit in the appendix, so reading only the main text limits how much of the 38-task composition and selection protocol is visible.

Sources