Designing and Characterizing Antibodies with AI: Three Efforts Spanning Epitope Selection, Affinity Ranking, and Developability Prediction
Synopsis
This article presents three efforts from Amazon Bio Discovery: MochiBind, a sequence-only predictor that reframes binding affinity as pairwise comparison aggregated by TrueSkill into a global ranking, achieving higher pairwise accuracy than every structure-based baseline on four held-out antigens and scoring 200,000 antibody pairs in roughly 13 seconds on a CPU; CA-MAP, a context-aware multi-property predictor that uses example antibodies in the prompt to absorb batch offsets, holding a 0.99 correlation under a simulated batch effect where standard fine-tuning falls to 0.
Interpretation
MochiBind reframes binding affinity prediction from regressing an absolute value to deciding which of two antibodies against the same antigen binds more tightly, then aggregates pairwise comparisons into a global ranking over the candidate pool using TrueSkill, with no structural input at any stage. The text notes that most affinity predictors are evaluated on predicting absolute binding affinity for antigens seen in training, against test sets containing few or no nonbinders; after surveying seven prior studies the authors found none satisfied all conditions needed to train a reliable universal predictor, so they reframed the task. Uses the AlphaBind dataset covering four antigen systems (TIGIT, PD-1, HER2, and the SARS-CoV-1 RBD) with roughly 30,000 experimentally characterized variants each and pairwise sequence similarity between antigens close to zero, under a strictly cross-antigen protocol that trains on two antigens, validates on a third, and tests on the fourth, rotating so each is held out once. MochiBind achieved higher pairwise accuracy than every structure-based baseline on all four held-out antigens, outperforming the closest competitor by almost 10% on average; it had the highest retrieval accuracy on all four antigens and the highest retrieval precision on three of four; and it scored 200,000 antibody pairs in roughly 13 seconds on a CPU, a more than 100-fold inference speedup.
CA-MAP places example antibodies with their measured properties into the prompt context, letting the model adjust for batch effects at inference time without retraining. The text notes that a model fine-tuned on one lab's data quietly inherits that lab's batch offsets, and that a model trained on data from a single source can learn to ignore the examples and rely on the query sequence alone. The authors introduce the AB-context-aware training strategy, which applies a hidden random transformation to both the context properties and the expected answer, resampled for every prompt, so the transformation can be recovered only from the context and the model must use it. On a fine-tuned domain-specific multimodal LLM, TxGemma, predicting hydrophobicity: without batch effects, standard fine-tuning and AB-context-aware training perform comparably at a Spearman correlation of 0.99 with ground truth; with a simulated additive batch effect in the 0-0.3 range, standard fine-tuning falls to 0.58 while the context-aware model remains at 0.99. Trained on a synthetic dataset of 876,898 antibody-heavy chains covering six developability properties, CA-MAP achieves a Spearman correlation (denoted rho) greater than 0.8 on several properties and outperforms the fine-tuned TxGemma baseline across all four properties tested jointly; it has roughly 182,000 trainable parameters versus TxGemma's 40 million and is about 200 times as fast per prompt at inference.
CA-MAP can be queried for properties absent from its training data and draws on correlations between properties to improve prediction. Because properties are specified as text, the model can be asked about properties not seen in training; the text reports that with only the two target properties as context, immunogenicity prediction reached rho = 0.25 and positive-charge heterogeneity (PosCh) reached rho = 0.08, while with all six correlated properties in the context both reached rho = 0.73. The experiment trained CA-MAP on only four of the dataset's six developability properties and tested it on the other two (PosCh and immunogenicity), a held-out-property generalization test; the authors read the gains as the model drawing on correlations between developability properties, suggesting expensive assays could be estimated in part from cheaper ones.
The third effort chains predictive and generative models into an end-to-end design process and yields experimentally validated binders against a novel target with no experimental structure and no public antibody information. The target was identified by collaborators at Dr. Nai-Kong V. Cheung's Lab at Memorial Sloan Kettering Cancer Center by sequencing patient tumor specimens for proteins that sit on the tumor cell surface, are driven by a specific genetic error, and are largely absent from healthy tissue, for desmoplastic small round-cell tumors, a rare and aggressive pediatric cancer; with no template to graft, no prior campaign to affinity-mature from, and no possibility the design models encountered this antigen during training. The hotspot recommendation agent orchestrates seven bioinformatics tools (solvent-accessible surface area, secondary structure, hydrophobicity, sequence uniqueness against user-specified negative targets, matching against 500,000 entries in NIAID's Immune Epitope Database, and Pfam domain annotation, among others); on antibody-antigen complexes from the SAbDab benchmark it recovered at least one true epitope residue within its top five proposed regions about 80% of the time on a diverse holdout set, and it proposed eight hotspot regions for the campaign target. Three generative models, RFantibody, IgGM, and mBER, each produced 96,000 designs; after multi-objective Pareto filtering the candidate selection agent prioritized 100,000 candidates for experimental screening; after yeast surface display and two rounds of sorting and filtering, none of the 116 surviving candidates bound an unrelated control protein, and all 116 were individually measured, with 46 identified as strong binders.
Perspective
These results are aimed at antibody and nanobody discovery settings: MochiBind fits cases where a candidate pool must be ranked and the target may be unseen in training, with evaluation built on four antigen systems and a cross-antigen rotating protocol; CA-MAP fits cases where example antibodies from the same laboratory are available, batch offsets need correction, and multiple developability properties are predicted, with training on a synthetic dataset; the end-to-end design process fits novel targets with no experimental structure and no public antibody information, with validation relying on yeast surface display and two rounds of sorting. For teams wanting to plug predictive models into their own experimental workflows, these efforts offer a reference evaluation protocol, training strategy, and agent orchestration pattern.
Several open questions remain for a careful reader: whether MochiBind's cross-antigen advantage holds beyond the four antigen systems and how relative orderings perform in real screening pools; whether CA-MAP stays stable under measured batch variation beyond simulated batch effects, and whether gains on held-out properties translate into practical substitution for expensive assays; and what affinity levels, developability, and room for further optimization the 46 strong binders have, plus how repeatable the pipeline is on other targets. In addition, this piece is a narrative overview and does not reproduce the figures or full experimental detail of each paper, so assessing specific statistical settings and ablation results would require consulting the original papers.
