Opus 5.5's agent-designed libraries cut downstream code and beat the human production library by 2.3 points, yet 11 of 15 tasks merely reproduce human abstractions
Synopsis
The authors introduce LibraryDesignBench, a two-phase benchmark in which a designer agent implements a full-featured library from a specification that lists required capabilities without prescribing interfaces, and three user agents from different model families then solve problems with it, scored by pass rate squared times simplicity relative to reference solutions written with the real production library; across 15 library-design tasks, 242 expert-validated problems and four languages, Opus 5.5 scores 48.9, 2.3 points above the production library, designers reproduce production-library abstractions on 11 of 15 tasks, 64% of audited excess code traces to rigid or hard-to-use interfaces rather than missing capabilities, and adding prescriptive agent-first guidance raises the score to 46.
Interpretation
The paper introduces LibraryDesignBench, which defines a library's quality by the correctness and simplicity of the programs downstream agents write with it, rather than by the library's own test pass rate or conformity to a prescribed interface. Prior evaluations either test whether a library is correct or compare its interface against a specification, which prescribes the very design under measurement; this work instead observes agents actually using the library and benchmarks simplicity against reference solutions written with the real production library. The benchmark spans 15 library-design tasks in four languages and 242 expert-validated downstream problems; each designer setup produces three libraries per task used by three implementer agents, giving 2,178 evaluated problems per designer, with no-library and production-library baselines run over the same 2,178 trials.
Agents can indeed design libraries that help other agents: Opus 5.5 scores 48.9, 2.3 points (4.9% relative) above the production-library baseline of 46.6, with implementers passing just as many tests while writing simpler code. This is the first measurement that varies the designer and grades library design by downstream value, yielding an initial design baseline; earlier work fixed the library and scored only whether models call it correctly. Table 1 reports scores with 95% confidence intervals for 11 designer setups; pass rates cluster between 84.1% and 86.6% while the no-library condition reaches 86.4%, so ranking is driven mainly by simplicity rather than correctness.
Agent-designed libraries reproduce the abstractions of the human-written production library on 11 of 15 tasks, yet downstream agents still underuse them and reimplement capabilities the library already provides. The paper separates 'the library is correct' from 'the library gets used': even with the production library, implementers reach a simplicity of only 61.5, so part of the gap comes from the user rather than the design. An audit of 810 partially passing, over-length solutions sampled from six designer configurations attributes 82% of primary excess-code classifications to library limitations, with rigidity and verbosity at 64% versus 14% for coverage; in a separate author analysis, 41% of 13,562 hand-labeled rebuilt lines across 180 solutions redo functionality the library shipped but hid or broke.
Prescriptive agent-first guidance changes the resulting designs: having the designer sketch consumer programs, ship runnable usage examples and test with subagents cuts exported-name overlap with production from 19.4% to 13.4% and raises the score from 44.1 to 46.4, mostly through simplicity. This offers an actionable intervention and shows that designing libraries for agents differs from designing them for humans; the result still falls just short of the production library at 46.6 and of standard-prompt Opus 5.5 at 48.9. The intervention is a paired comparison on GPT-6 Astra (Codex, high reasoning) and improves all three implementers and three of four languages; guidance nearly doubles design cost from $4.51 to $8.24 with unchanged per-problem cost.
Perspective
The benchmark targets researchers and engineering teams who design libraries meant to be used by other agents, and applies within the paper's task set, implementer configurations and execution budgets; it measures tested correctness plus the static size and complexity of consumer programs, so it can compare library-design practices across designers, prompts and reasoning efforts, and can serve as a testbed for studying agent-facing interfaces, abstractions and documentation. The guided intervention reported in the paper (consumer-first API sketches, runnable usage examples, testing with subagents) lifts GPT-6 Astra's score from 44.1 to 46.4, offering a starting point for those seeking better downstream reuse.
The paper itself notes that scores characterize utility under specified tasks, consumer configurations and budgets rather than a consumer-independent library ordering, and that confidence intervals cover reruns of the fixed task set rather than generalization to all library domains. The audit covers only partially passing, over-length solutions, is performed once by a single model without human labels or agreement checks, so figures such as 64% and 14% describe that sample rather than all downstream cells. The guidance intervention is evaluated as a combined package without separating component contributions, and it nearly doubles design cost while per-problem cost stays flat, a trade-off whose behavior under other budgets remains to be seen. In addition, this load contains the full paper and the external story but not the complete leaderboards and task repositories on the benchmark site.
