LLMsFold: Integrating Large Language Models and Biophysical Simulations for De Novo Drug Design
Synopsis
This work presents LLMsFold, a computational framework that identifies binding pockets geometrically, has Llama-3.3-70B generate candidate small molecules as SMILES strings, evaluates each with the Boltz-2 co-folding model for bound pose and binding affinity, and iteratively refines candidates through a feedback loop, yielding molecules for ACVR1 and CD19 that pass drug-likeness, synthetic accessibility, and novelty filters, with the ACVR1 candidate reaching a predicted affinity probability of 0.953 and predicted pIC50 of about 10.72 and the CD19 Pocket 1 candidate reaching a predicted pIC50 of about 7.73.
Interpretation
It builds a pipeline that embeds structural awareness into every generation stage: geometric pocket detection with the Convex Hull Pocket Finder comes first, then the pocket residue description is written directly into the prompt so the LLM generates molecules within a specified region rather than enumerating a library and filtering afterwards. Compared with the conventional virtual-screening paradigm of building a library and filtering post hoc, spatial information is moved forward into the generation stage, constraining generation to a biologically meaningful region of the protein surface. The methodology is described in full, with deterministic pocket-filtering parameters (minimum dimension 8 Å, isotropic margin +5 Å, maximum box size 30 Å, proximity radius 8 Å), and is demonstrated on both ACVR1 and CD19.
It relies purely on in-context learning rather than fine-tuning, giving Llama-3.3-70B clinically relevant molecules in the prompt as 'inspired by' leads while enforcing Lipinski's Rule of Five and avoidance of PAINS substructures. Compared with frameworks such as REINVENT and DrugGPT that require task-specific training or fine-tuning, this approach needs no dedicated GPU cluster or large chemical corpus, and switching to a new target requires only reformulating the prompt. The paper states explicitly that no retraining or fine-tuning was performed, and notes that increasing the number of reference molecules in context can introduce noise and that the absence of gradient-based optimization limits deep internalization of structure-activity relationships for a specific target.
Boltz-2 serves a dual role of filtering and ranking, outputting binding probability, pIC50, ipTM, and pLDDT, with results fed back into the generative loop to form an iterative optimization trajectory. Here Boltz-2 not only predicts how a ligand binds but also provides an affinity estimate and rapid virtual screening, and its diffusion module implicitly accounts for protein flexibility, capturing induced-fit effects that rigid docking might miss. ACVR1 Molecule 1 shows ipTM = 0.983, ligand pLDDT = 0.958, and affinity probability 0.953; AutoDock Vina on ACVR1 produced a pose similar to the Boltz-2 top pose with RMSD < 1.5 Å, providing one cross-check.
It produces candidates for ACVR1 and CD19 that pass all filters, with final leads returning no PubChem match, suggesting likely novel chemical entities. The ACVR1 candidates are polycyclic heterocycle scaffolds with QED 0.55–0.64 and SA scores of 2.64 and 2.81, satisfying all four Lipinski criteria with no PAINS alerts; the three CD19 pockets show a predicted affinity gradient, highest for Pocket 1 and lowest for Pocket 3. Fifty unique molecules were generated per target, with 15 of roughly 50 ACVR1 proposals judged strong binders by Boltz-2; for CD19, 7 molecules were identified across three pockets. All affinity and structural values are computational predictions.
Perspective
The framework targets the early lead-discovery stage and applies to target proteins with experimentally resolved structures in PDB format, requiring a pocket description and reference inhibitors; the authors position it as a tool to accelerate early ideation rather than a replacement for medicinal chemistry expertise. Its design intent covers both kinase-like deep pockets (such as the ACVR1 ATP-binding site) and shallow protein-protein interaction surfaces (such as the three CD19 pockets), the latter historically resistant to small-molecule intervention. Because generation and co-folding run on remote API endpoints, the method is accessible to academic groups, patient foundations, and small biotech companies with limited computing resources, and the authors specifically note its disproportionate relevance to rare conditions such as ACVR1-related FOP, which affects approximately one in two million individuals worldwide.
All binding affinities and structural predictions are computational estimates; although Boltz-2 approaches AlphaFold3-level accuracy on established benchmarks, its outputs still require confirmation through biophysical assays such as surface plasmon resonance or isothermal titration calorimetry. The pipeline does not currently incorporate absorption, distribution, metabolism, excretion, and toxicity predictions, and QED and SA scores do not substitute for dedicated pharmacokinetic modeling. Language models are susceptible to chemical hallucinations, potentially generating SMILES strings that parse correctly but encode chemically implausible or unstable structures; multi-stage filtering mitigates but cannot eliminate this risk. Generation quality depends on the choice of reference compounds in the prompt, and a poorly chosen set of examples may bias the model toward a narrow region of chemical space. Boltz-2 affinity predictions are useful for ranking but should not be interpreted as quantitative IC50 values without calibration against experimental data for the specific targets examined. In addition, the main text does not disclose the specific structures of the generated molecules, and the appendix benchmark was conducted only on the single ACVR1 target, so readers assessing cross-target generalization will need to await validation on further targets.
