Expansion of DNA-Encoded Library Hits Using Generative Chemistry and Ultra-Large Compound Catalogs
Synopsis
This work initialized and biased the HIDDEN GEM structure-guided generative virtual screening workflow with screening data from a focused DNA-encoded library against the 53BP1 tandem Tudor domain (UNCDEL003, 58,080 compounds), nominated 57 purchasable compounds from the roughly 37-billion-compound Enamine REAL Space, and validated 14 as active hits by TR-FRET displacement (3 with IC50 ≤50 µM and 11 with IC50 ≤100 µM), with the AI-nominated hits showing greater chemical diversity, improved drug-likeness, and off-the-shelf purchasability relative to the initial DEL hits.
Figure 1. Workflow for the 53BP1 hit-finding campaign. The schematic illustrates a dual-pathway approach
bioRxiv · Page 8Interpretation
It presents and experimentally validates a hit-expansion workflow coupling empirical DEL screening data with a generative chemistry model: a similarity search from UNC8531 into Enamine REAL Space supplied the 1 million most similar compounds to initialize HIDDEN GEM, and two iterative cycles of docking, generation, similarity expansion, and re-docking nominated 57 compounds for testing. Prior machine-learning hit nomination from DEL data relied largely on supervised models whose generalization is constrained by training-data quality and scope; here a DEL hit instead initializes a structure-guided generative workflow, allowing exploration of chemical space beyond the training data within an ultra-large purchasable catalog. The workflow is described in detail (Glide SP docking, a ChEMBL-pretrained generative model, 100K molecules generated per cycle, similarity searches returning up to 10K purchasable compounds), and the docking model was validated by a UNC8531 docked pose within <1.0 Å RMSD of the crystal ligand pose.
Across two HIDDEN GEM cycles, 57 compounds were nominated and 14 were confirmed active by TR-FRET displacement, with 3 showing IC50 ≤50 µM and 11 showing IC50 ≤100 µM; the most potent, UNC10413788A, gave an IC50 of 22 ± 3.6 µM, comparable to the DEL-derived UNC10413788A at 22 ± 3 µM. The result extends hit finding from a roughly 58K-member focused library into the roughly 37-billion-compound Enamine REAL Space at the time of screening, and reports measured activity for purchasable compounds rather than computational scores alone. Activity was measured by TR-FRET displacement, with IC50 values reported as mean ± SD of at least three replicates, and a parallel traditional DEL analysis path based on aggregated enrichment from two UNCDEL003 replicates served as a benchmark.
The AI-nominated hits occupy chemical space clearly distinct from the reference compound, with maximum Tanimoto similarity to UNC8531 of only 0.23 (cycle 1) and 0.19 (cycle 2), and t-SNE shows substituents such as benzoisothiazole, naphthalene, and indole that were not captured by the DEL. This indicates that generative expansion moved into chemical regions not covered by the physical DEL library rather than making minor modifications around an existing scaffold, and the compounds are readily purchasable without extensive medicinal chemistry follow-up or synthon aggregation. Similarity was computed as Tanimoto similarity on Morgan fingerprints (radius=3, 2048 bits), and chemical space was visualized with scikit-learn t-SNE using perplexity 30, learning rate 200, and 1000 iterations.
Iterative generation improved drug-like properties while retaining activity: statistically significant shifts between cycles appeared in molecular weight (p = 0.0176), heteroatom count (p = 0.0249), and rotatable bonds (p = 0.0304), and FastSolv-predicted logS for the generative hits was higher across 250–350 K than for both the baseline control and the traditional DEL-derived hit. This suggests iterative generation can shift the physicochemical property distribution toward more favorable values without sacrificing activity, and adds a solubility dimension beyond QED and SA benchmarks. Property differences were assessed with two-tailed t-tests, and solubility was predicted by the FastSolv deep learning model across a 250–350 K gradient in DMSO, making it a computational prediction rather than an experimental measurement.
Perspective
The workflow applies where the protein binding site is known and at least one validated DEL hit is available for initialization, as in this study's starting point of the 53BP1 tandem Tudor domain with the UNC8531 co-crystal structure (PDB: 8SWJ). It is aimed at academic and industrial drug discovery teams seeking low-cost expansion of purchasable hits from focused or commercial DEL data, particularly CRO-run DEL campaigns where clients receive only a few disclosed binders. The chemical space addressed is the roughly 37-billion-compound Enamine REAL Space lead-like library at the time of screening, and the reported activities fall in the low- to mid-micromolar range.
A careful reader would still watch several things: the generative hits sit in the low- to mid-micromolar range, so the path toward higher affinity remains a question for follow-up optimization; the solubility and drug-likeness conclusions come mainly from FastSolv, QED, and SA predictions, with experimental physicochemical measurements not presented here; the significant property shifts between cycles rest on distribution comparisons of the nominated sets, and their reproducibility across other targets and initialization conditions remains to be tested; and this is a preprint that has not yet been peer reviewed, with specific numerical details in figures and supporting materials not itemized within this reading scope.
