跳到主要内容
返回时间线
bioRxiv来源发表:

Multi-ISM 以少 45 倍模型评估量完成 5000 个基因的饱和突变扫描,并提升罕见变异表达预测

核心概要

该工作提出 Multi-ISM,将计算机饱和突变(ISM)重构为稀疏恢复问题,用比逐变异穷举 ISM 少 45 倍的模型评估量生成突变图谱,在变异效应基准上达到或超过其精度,并应用于 5000 个蛋白编码基因(含 3317 个 OMIM 疾病基因),在 500-kb 窗口内生成碱基分辨率、组织分辨的归因图谱,用于增强子—基因优先级排序和细胞类型特异调控元件识别,其基因水平汇总显示受约束更强的基因预测突变效应更小,将预测聚合为基因水平罕见变异负荷后相较常见变异弹性网改善了表达预测,在表达离群样本上增益最大。

Source-provided article image: Scalable saturation mutagenesis reveals gene regulatory architecture and rare variant effects
Figure 1 ·

Figure 1: Multi-ISM makes efficient and accurate variant-effect predictions. A, Schematic of base Multi-ISM. Multiple spatially separated mutations are introduced into each sequence, and individual variant-effect estimates are recovered from the resulting multi-mutant measurements by sparse regression. B, Schematic of adaptive Multi-ISM. Variant effects are iteratively estimated via regression, high-effect vari- ants are re-evaluated and pruned from subsequent designs, and the remaining variants are re-estimated using additional Multi-ISM measurements. C, Pearson correlation between Multi-ISM and single-ISM predictions for HBG1 averaged across output tracks, as a function of the number of forward passes. The comparison includes variants within the central 10-kb window with absolute single-ISM effects greater than 0.01. Blue, base Multi-ISM; green, adaptive Multi-ISM. Inset, representative comparison of adaptive Multi-ISM and single-ISM scores averaged across K562 tracks. D, Top, adaptive Multi-ISM attribution scores across the 500 kb region surrounding HBG1 in K562. In gene tracks, orange and blue denote plus- and minus-strand genes. Bottom, CAGI5 MPRA measurements, single-ISM, adaptive Multi-ISM and Input × Gradient attri- bution scores near HBG1 promoter. Dashed boxes highlight regulatory motifs. Nucleotide-level substitution effects are shown at right for the indicated region; for gradient track, nucleotide-level gradient scores are shown. E, Performance of base Multi-ISM (blue) and adaptive Multi-ISM (green) on the CAGI5 MPRA saturation mutagenesis benchmark as a function of computational budget. Pearson and Spearman correla- tions quantify agreement with measured variant effects, whereas AUROC and AUPRC quantify classification of over-expression versus under-expression variants as defined previously (Methods). Dashed orange lines indicate single-ISM performance, and orange stars indicate the estimated cost of single-ISM across a 500 kb sequence. N indicates the number of variants used for regression or classification.

bioRxiv · 第 4 页

深度剖析

Multi-ISM 把 ISM 重新表述为稀疏恢复问题,从而一次性生成突变图谱,而不是逐个变异评分。 传统 ISM 对单个基因需数百万次模型评估、基因组尺度需数十亿至数万亿次;Multi-ISM 将评估量降低 45 倍。 摘要报告了 45 倍评估量缩减,并称其在变异效应基准上达到或超过穷举单变异 ISM 的精度;未给出具体基准名称或数值。

Multi-ISM 与模型架构无关,并可跨大规模序列到功能模型迁移。 说明该方法不绑定某一特定模型,可随新架构出现而复用。 摘要以“architecture-agnostic”和“transfers across large-scale sequence-to-function models”陈述,未列出所测试的模型清单。

在 5000 个蛋白编码基因(含 3317 个 OMIM 疾病基因)上生成碱基分辨率、组织分辨的归因图谱,覆盖 500-kb 窗口。 把核苷酸分辨率解释从单基因专用工作扩展到跨基因、跨组织的规模化图谱。 摘要给出基因数量、OMIM 疾病基因数量、窗口大小与分辨率,但未提供图谱数量或组织清单。

图谱支持增强子—基因优先级排序与细胞类型特异调控元件识别;基因水平汇总显示受约束更强的基因预测突变效应更小;将预测聚合为基因水平罕见变异负荷后,相较常见变异弹性网改善了个体化表达预测,在表达离群样本上增益最大。 把突变图谱从可视化产物推进到可用于调控元件优先排序和表达预测的下游任务。 摘要以定性方式报告这些下游结果与比较,未给出效应量、样本量或统计指标。

启示与展望

该框架面向使用长上下文序列到功能模型的研究者与基因组分析流程,适用于需要在 500-kb 窗口内获得碱基分辨率、组织分辨调控归因图谱的场景,例如增强子—基因优先级排序、细胞类型特异调控元件识别,以及把预测聚合为基因水平罕见变异负荷以改进个体化表达预测。其价值在于让新架构、新功能读出与新细胞背景出现时即可被映射,而不必为每个基因单独投入专用计算。

当前仅读到摘要与利益冲突声明,正文、图与表均未加载,因此无法核对 45 倍缩减的测量方式、基准数据集与精度指标、所测试的序列到功能模型清单、5000 个基因与 3317 个 OMIM 基因的筛选标准、组织与细胞类型覆盖范围,以及罕见变异负荷改善表达预测的效应量与统计显著性。摘要中“受约束更强的基因预测突变效应更小”与“表达离群样本增益最大”均为定性表述,其稳健性与可复现性有待正文验证。此外,作者中有 Calico Life Sciences 雇员与顾问,应用与转化方向值得结合正文独立评估。

来源