Multi-ISM maps 5,000 genes with 45-fold fewer model evaluations and improves rare-variant expression prediction
Synopsis
The work introduces Multi-ISM, a framework that reformulates in silico saturation mutagenesis (ISM) as a sparse recovery problem, generating mutational maps with 45-fold fewer model evaluations than exhaustive single-variant ISM while matching or exceeding its accuracy on variant-effect benchmarks, and applies it to 5,000 protein-coding genes (including 3,317 OMIM disease genes) to produce base-pair-resolution, tissue-resolved attribution maps across 500-kb windows for enhancer-gene prioritization and cell-type-specific regulatory element identification, with gene-level summaries showing more constrained genes have smaller predicted mutational effects and with aggregation into gene-level rare-variant burdens improving personalized expression prediction over a common-variant elastic net, wi
Figure 1: Multi-ISM makes efficient and accurate variant-effect predictions. A, Schematic of base Multi-ISM. Multiple spatially separated mutations are introduced into each sequence, and individual variant-effect estimates are recovered from the resulting multi-mutant measurements by sparse regression. B, Schematic of adaptive Multi-ISM. Variant effects are iteratively estimated via regression, high-effect vari- ants are re-evaluated and pruned from subsequent designs, and the remaining variants are re-estimated using additional Multi-ISM measurements. C, Pearson correlation between Multi-ISM and single-ISM predictions for HBG1 averaged across output tracks, as a function of the number of forward passes. The comparison includes variants within the central 10-kb window with absolute single-ISM effects greater than 0.01. Blue, base Multi-ISM; green, adaptive Multi-ISM. Inset, representative comparison of adaptive Multi-ISM and single-ISM scores averaged across K562 tracks. D, Top, adaptive Multi-ISM attribution scores across the 500 kb region surrounding HBG1 in K562. In gene tracks, orange and blue denote plus- and minus-strand genes. Bottom, CAGI5 MPRA measurements, single-ISM, adaptive Multi-ISM and Input × Gradient attri- bution scores near HBG1 promoter. Dashed boxes highlight regulatory motifs. Nucleotide-level substitution effects are shown at right for the indicated region; for gradient track, nucleotide-level gradient scores are shown. E, Performance of base Multi-ISM (blue) and adaptive Multi-ISM (green) on the CAGI5 MPRA saturation mutagenesis benchmark as a function of computational budget. Pearson and Spearman correla- tions quantify agreement with measured variant effects, whereas AUROC and AUPRC quantify classification of over-expression versus under-expression variants as defined previously (Methods). Dashed orange lines indicate single-ISM performance, and orange stars indicate the estimated cost of single-ISM across a 500 kb sequence. N indicates the number of variants used for regression or classification.
bioRxiv · Page 4Interpretation
Multi-ISM reformulates ISM as a sparse recovery problem, generating mutational maps in one pass rather than scoring mutations one at a time. Conventional ISM requires millions of model evaluations for a single gene and billions to trillions at genome scales; Multi-ISM reduces evaluations 45-fold. The abstract reports the 45-fold reduction and states it matches or exceeds exhaustive single-variant ISM accuracy on variant-effect benchmarks; specific benchmarks and values are not given.
Multi-ISM is architecture-agnostic and transfers across large-scale sequence-to-function models. This indicates the method is not tied to one model and can be reused as new architectures appear. Stated in the abstract as 'architecture-agnostic' and 'transfers across large-scale sequence-to-function models'; the set of models tested is not listed.
Applied to 5,000 protein-coding genes, including 3,317 OMIM disease genes, it generates base-pair-resolution, tissue-resolved attribution maps across 500-kb windows. Extends nucleotide-resolution interpretation from dedicated single-gene efforts to large-scale, cross-gene and cross-tissue maps. The abstract gives gene counts, OMIM disease gene count, window size, and resolution, but not the number of maps or the tissue list.
The maps supported enhancer-gene prioritization and cell-type-specific regulatory element identification; gene-level summaries showed more constrained genes had smaller predicted mutational effects; aggregating predictions into gene-level rare-variant burdens improved personalized expression prediction over a common-variant elastic net, with the largest gains at expression outliers. Moves mutational maps from visualization output toward downstream tasks of regulatory element prioritization and expression prediction. The abstract reports these downstream results and the comparison qualitatively, without effect sizes, sample sizes, or statistical metrics.
Perspective
The framework targets researchers and genomic analysis pipelines using long-context sequence-to-function models, in settings that need base-pair-resolution, tissue-resolved regulatory attribution maps across 500-kb windows, such as enhancer-gene prioritization, cell-type-specific regulatory element identification, and aggregating predictions into gene-level rare-variant burdens to improve personalized expression prediction. Its value lies in letting new architectures, functional readouts, and cellular contexts be mapped as they appear, without dedicated per-gene computation.
Only the abstract and the competing interest statement were loaded; the main text, figures, and tables were not, so it is not possible to verify how the 45-fold reduction was measured, which benchmark datasets and accuracy metrics were used, which sequence-to-function models were tested, the selection criteria for the 5,000 genes and 3,317 OMIM genes, the tissue and cell-type coverage, or the effect sizes and statistical significance of the rare-variant burden improvement in expression prediction. The abstract's statements that more constrained genes have smaller predicted mutational effects and that gains are largest at expression outliers are qualitative, and their robustness and reproducibility await the main text. In addition, authors include employees of and a consultant for Calico Life Sciences, so application and translation directions merit independent assessment alongside the main text.
