Evidence-Grounded Agentic Formulation Development in an Autonomous Laboratory
Synopsis
The study presents Andromeda 2, an agentic system that reasons over structured in-house experimental evidence and invokes computational and experimental tools, achieving a 50% high-AUC hit rate for paclitaxel self-emulsifying drug delivery system (SEDDS) development at a matched budget of 96 formulations (versus 17% for Andromeda 1 and 2% for wet-lab DoE) and identifying 12 formulations meeting all four target product profile (TPP) objectives (versus 6 and 0), while an ablation showed that access to structured in-house evidence increased mean AUC by 34%.
Figure 1: Paclitaxel formulation performance at a matched experimental budget. Formulation-level distributions of the area under the curve (AUC) of apparently solubilized paclitaxel concentration in FaSSIF from 10 to 240 min achieved using the evidence-grounded agentic platform ( Andromeda 2 ), the optimization framework ( Andromeda 1 ) and a design of experiment (DoE) approach. Each strategy evaluated 96 unique formulations. Andromeda 2 produced an upward-shifted AUC distribution and a greater frequency of higher performing formulations, while both Andromeda platforms generated formulations with higher AUC values than experimental DoE.
arXivInterpretation
Andromeda 2 produced high-performing SEDDS formulations more frequently than probabilistic optimization and wet-lab DoE at a matched budget. Prior autonomous laboratories and sequential optimization typically initialize each campaign largely de novo and learn mainly from data generated within that campaign; this work couples agentic decision-making with the laboratory's existing structured evidence and compares arms under the same formulation space, TPP, assay, and automated platform at matched budget. 96 unique formulations per strategy across six 16-formulation batches on the same miniaturized platform and FaSSIF assay; median AUC was 70.1, 12.0, and 3.5 mg·min/mL, high-AUC hit rates were 50%, 17%, and 2%, and full-TPP passes were 12, 6, and 0; the authors label the formulation-level Mann–Whitney comparison (p=9.6×10⁻⁸) as descriptive rather than inferential.
Andromeda 2's main advantage was producing more high-performing formulations rather than finding a higher isolated optimum. The result separates peak performance from distributional performance: Andromeda 2 and Andromeda 1 reached comparable maximum AUC (157.7 vs 156.4 mg·min/mL), while median AUC and full-TPP counts differed substantially. The highest-AUC formulation met only two of the four TPP objectives, so maximum AUC is not equivalent to a balanced formulation; Andromeda 2's best formulation maintained apparent solubilized paclitaxel through 240 min and reached an AUC about 18 times the unformulated control drug.
Access to structured in-house experimental evidence improved Andromeda 2 performance, and its effect appeared in later batches rather than only the initial batch. The ablation withheld historical in-house evidence while holding constant the formulation objective, executable design space, TPP, feasibility constraints, batch size, and wet-lab workflow, isolating that evidence's contribution. Mean AUC was 62.0 versus 46.2 mg·min/mL with full versus withheld evidence (34% higher), and high-AUC hit rates were 50% versus 31%; the reduced-evidence configuration matched or exceeded full evidence in early batches but subsequently regressed, whereas full-evidence Andromeda 2 improved across successive batches; both configurations continued to generate valid, experimentally executable formulations.
Andromeda 2 both concentrated its experimental budget in the productive region and selected better candidates within that region. Analysis using the Lipid Formulation Classification System (LFCS) showed high-performing formulations concentrated in the Type IIIB/IV region, and Andromeda 2's advantage was not explained solely by identifying that region. High-AUC hit rates by LFCS type were 0% for Type II, 6% for Type IIIA, 35% for Type IIIB, and 32% for Type IV; within Type IIIB/IV, Andromeda 2 achieved a 49% hit rate (bootstrap 95% CI [0.39, 0.59]) versus 29% for Andromeda 1 and 3% for experimental DoE; Andromeda 2 sampled 12 oil-surfactant families in its first batch and 16 overall, fewer than DoE's 90.
Perspective
The results apply to SEDDS formulation screening at a matched 96-formulation budget with endpoints of apparent solubilized paclitaxel AUC in FaSSIF and a four-attribute TPP, using the same miniaturized automated platform, assay, and workflow; they support physicochemical performance of the formulations and provide candidate formulations, including high-performing compositions featuring vitamin E TPGS, for subsequent digestion-coupled testing and pharmacokinetic studies.
The FaSSIF assay does not capture digestion, intestinal permeability or metabolism, or in vivo performance, so whether the observed advantages translate to oral exposure requires digestion-coupled testing and ultimately pharmacokinetic studies; each strategy was evaluated in a single independently initialized campaign, limiting campaign-level statistical inference and motivating replicate autonomous campaigns to establish reproducibility; the experimental DoE explored the shared composition space directly without expert pre-screening and should be interpreted in that context; prospective evaluation across chemically distinct APIs remains an open question; additionally, this is a full-text parse in which specific numerical details in figures and supplementary tables are not itemized, which may affect the precision of restating individual results.
