Skip to main content
Back to timeline
arXivSource publication:

Raven composes model–harness pairs into a multi-agent ecosystem and leads planning, four specialist domains, and skill reuse over Claude Code and other baselines

Synopsis

Raven is an open-source multi-agent ecosystem that treats each executable model–harness pair as a composable unit of intelligence: a Host Agent decomposes goals into a directed acyclic graph and matches subtasks to registered specialists, HarnessBank-style evolution, EverOS memory, and Skill Forge reuse experience across tasks, and the paper gives sufficient conditions under which composition expands reliable task coverage under a shared resource budget, reporting gains over Claude Code, Hermes Agent, OpenClaw and other baselines on the MAOB planning benchmark, four native specialist domains, frozen-backbone harness evolution, and fixed skill-library reuse.

AI-generated editorial illustration: Raven: The Harness of Harnesses for Composable Agentic Intelligence

Interpretation

Raven reframes orchestration from engineering a stronger harness for one domain to autonomously constructing specialized harnesses, improving them through experience, and orchestrating them across domains, treating each executable model–harness pair as a composable unit of intelligence; the Host Agent decomposes goals, matches subtasks to registered capabilities, represents dependencies as a directed acyclic graph, and integrates results. Compared with multi-agent frameworks that organize collaboration through conversation or fixed role workflows such as AutoGen and MetaGPT, Raven connects heterogeneous harnesses including Claude Code, Codex, Hermes Agent, and OpenClaw through execution adapters, and validates each submitted graph through format, structure, capability, status, and environment checks before dispatch, so differing execution capabilities can be composed along explicit dependencies. The paper presents the architecture, the agent registry, the admission-validation rules, the node lifecycle, and the completion-adjudication flow, and states that the Host Agent selects executors from capability tags without a separately trained routing classifier; these are design-level claims, with end-to-end benefit assessed separately in the evaluation sections.

The paper formalizes composable agentic intelligence as reliable task coverage under a common resource budget and establishes sufficient conditions under which complementary local capabilities, compatible handoffs, and bounded planning and execution errors let the composed system solve tasks that the available individual agents cannot reliably solve alone under the same budget. Where prior work reports average success rates or routing effects, this analysis charges planning, coordination, handoffs, verification, and retries to the same budget and gives an explicit complementary construction (two agents each able to read only one bit, plus a node that computes parity) to show composition can yield new task coverage. The results are stated as a lemma, a theorem, propositions, and corollaries, and are explicitly sufficient conditions whose practical applicability depends on whether the implemented agents and host satisfy the premises; the authors also note the result does not establish strict set inclusion or an improvement in average performance and requires matched comparisons against strong individual-agent baselines with comparable models, tools, initial information, and total cost.

On the introduced Multi-Agent Orchestration Benchmark (MAOB), Raven ranks first among the compared systems on all four graph metrics under both tested backbones, with Exact Match gains of several percentage points over the strongest baseline, and Edge F1 remains lower than Node F1 for every system under both backbones, indicating dependency prediction is the harder part of the benchmark. MAOB fixes the reference DAG before generating the request and applies a leakage filter that rejects requests enumerating steps or naming a domain outright, so specialist selection and dependency prediction are measured before any worker is dispatched rather than being confounded with downstream execution quality. The benchmark comprises 140 tasks covering all 11 domain subsets of at least two domains under a fixed four-specialist roster, with reference graphs containing several nodes and edges on average; three systems are compared on identical request text, the same delegation instruction, and matched backbones, with planning measured in isolation.

Four native specialists and the reuse mechanisms each report gains: Raven-Research achieves the highest pooled accuracy with all three shared backbones, Raven-Code achieves the highest score in six of eight settings and leads on whole-repository migration, Raven-Design scores highest in all six benchmark-and-backbone settings, Raven-Oncall reaches lower BPB at lower token use and cost on autoresearch, and a fixed skill library improves performance in all twelve cell–benchmark comparisons with larger gains for Raven than OpenClaw. The evaluation measures orchestration, harness evolution, specialist execution, and skill reuse as separate capability levels, holding the backbone and harness fixed to isolate the component under study; the harness-evolution results come from the published HarnessBank experiments and the skill-reuse results from the published SkillCorpus experiments. Concrete numbers and statistical support are reported, such as held-out Pass@1 gains of 15.4 points on AppWorld, 13.9 on BrowseComp+, and 13.7 on LiveCodeBench with a frozen Qwen3.6-27B backbone, with six comparisons passing the source's paired-gain criterion; the skill experiments report pooled gains with z-scores of 3.2, 3.1, and 4.0 and note that individual-cell GDPval gains are not statistically distinguishable from zero given per-task reward standard deviations of 20 to 27 points.

Perspective

This work addresses builders of agent systems that need long-horizon, cross-domain collaboration: it provides a runnable set of orchestration, evolution, memory, and skill-reuse mechanisms, together with sufficient conditions for when composition expands reliable task coverage. It applies to deployments that have several heterogeneous harnesses and are willing to pay explicit budgets for planning, validation, handoffs, and retries; MAOB conclusions are limited to the four-specialist roster and tested tool configurations, and the theory is conditional on the implemented agents and host satisfying its premises. For a reader, the directly reusable design principles are treating model–harness pairs as units of composition, validating the dependency graph before dispatch, keeping memory records distinct from graph edges, and keeping skill selection distinct from tool permission.

The theoretical results are sufficient conditions whose practical applicability depends on whether the implemented agents and host satisfy compatibility, contract realization, and error-bound premises, and the paper itself notes this does not imply strict set inclusion or improved average performance. On the evaluation side, MAOB reference graphs are authored by a model and reviewed, and the leakage filter excludes explicit planning cues but cannot rule out paraphrased ones; individual-cell GDPval gains are not statistically distinguishable from zero given large per-task reward standard deviations, and the 26-task SWE-bench Verified split does not pass the paired-gain criterion; shared group memory is disabled by default and not evaluated in this report, so its effect on orchestration and task performance remains to be established. In addition, the loaded text has numeric gaps in several places (for example the MAOB Exact Match percentage-point margins and some table cells), so those specific figures cannot be restated here and would need to be checked against the original.

Sources