Skip to main content
Back to timeline
arXivSource publication:

Relic turns recurring collaboration failures into executable protocols, lifting complete-contract delivery from 14.06% to 19.76% across 360 controlled runs

Synopsis

Relic lets multi-agent organizations turn recurring collaboration failures into organization-owned, executable, revisable protocols, raising complete-contract delivery on mainline from 14.06% to 19.76% (+5.71 percentage points) across 360 controlled runs over ten software workloads and three models, and raising behavioral correctness under fresh-member transfer from 25.4% with no inherited protocol and 34.6% with text-only rules to 41.2% with executable bindings.

AI-generated editorial illustration: Relic: From Multi-Agent Collaboration to Persistent Organizational Capability

Interpretation

The paper formulates organizational capability as learned, governed, executable protocol state held outside member-local state, where protocols bind triggers, responsible roles, required evidence, and execution consequences, and remain open to revision and retirement. Prior work (dialogue and workflows in ChatDev, MetaGPT, AutoGen; memory in Reflexion, ExpeL, G-Memory; executable skills in Voyager) preserves knowledge or procedures, but an agreement reached in a conversation does not automatically bind a new member who takes over the role; Relic makes the obligation attach to the role rather than to the individual. The paper gives a formal specification (organizational state O_t=(W_t,Q_t,R_t)) and a governance process (proposal, validation, approval, compilation into bindings, amendment, retirement), and traces one interface-review rule through proposal, adoption, enforcement, and two amendments in a matched case.

Across 360 controlled runs, B3 with the governed protocol lifecycle improves all four verified endpoints over B2, which shares the same SDL backbone but has the lifecycle disabled: complete contracts +5.71 pp (95% CI [2.97, 8.55]), held-out cases +7.72, exposed cases +6.58, and evaluator-confirmed seeded issues +6.96. B2 and B3 share model, tools, task-visible information, SDL parameters, member learning, and execution code, with the governed protocol lifecycle as the only switch, so the difference is attributable to that mechanism rather than to a stronger model or scaffold. Three models (GPT-5.6 Terra, Claude Opus 4.6, DeepSeek-V4-Flash), ten workloads, three seeds, 360 runs; every model stratum has a positive B3−B2 point estimate on all four endpoints, and leave-one-workload-out analyses keep all three reported endpoints positive.

Executable binding itself adds value: the prose-only variant B3-text loses 7.18 pp on complete contracts (95% CI [+3.54, +10.91]) during online rule creation, and under content-matched fresh-member transfer executable bindings beat the same readable text by 6.5 pp (95% CI [0.7, 15.6]) and no inherited protocol by 15.8 pp. Text and Exec use the same frozen rule content, wording, and order, differing only in whether the rules are compiled into runtime bindings, which separates the value of readable rules from that of executable organizational binding. B3 versus B3-text uses 30 matched Claude Opus 4.6 runs; Text and Exec each use 30 GPT-5.6 Terra runs, with Fresh reusing the corresponding 30 B2 main-study runs.

External benchmarks show portability: on the full CooperBench benchmark Relic reaches 367/477 (76.9%) versus 282/477 (59.1%) for the strongest released peer reference and 354/477 (74.2%) for the strongest hierarchical Team reference; on the fixed 48-pair same-model subset Relic reaches 29/48, above Solo's 26/48, while the official Peer baseline is 13/48; on ProgramBench a frozen protocol package raises mini-SWE-agent's mean behavioral pass rate from 64.164% to 70.916%. These results extend the protocol layer from constructed workloads to external software benchmarks and show improved coordination in a peer structure with no permanent lead. CooperBench results exclude broken benchmark pairs under the same exclusion set for every compared system; ProgramBench uses the same 25 tasks with a fixed protocol package and no SDL.

Perspective

The work targets software-production settings: ten frozen workloads (five built from skeletons, five repairing or extending frozen repository versions), a 336-step scheduling horizon, and three models. It supports two deployment modes: adapting rules during work, or reusing a frozen rule package at initialization. For a reader, the directly usable part is writing recurring collaboration friction into rules with triggers, responsible roles, required evidence, and execution consequences, and letting those rules remain in force after member turnover; the transfer experiment indicates that compiling the same rules into runtime bindings is more effective than providing readable text alone.

Rule quality, revision, and retirement over time are listed as open directions; evaluation is concentrated in software production, and broader domains, selective transfer and forgetting, rule-quality diagnosis, alternative decision layers, and human-organization interaction remain to be studied. In addition, this reading covers the paper body and appendix text with figures rendered as descriptions, so checking exact curve shapes or item-level values would still require the original figures.

Sources