Skip to main content
Back to timeline
发表出处待核验Source publication:

functional-standard-atlas: an attenuation-corrected, territory-resolved benchmark of variant effect predictors against saturation genome editing

Synopsis

This work builds a frozen, content-hashed data asset and a uniform scoring harness that maps seven MaveDB saturation genome editing (SGE) score sets to GRCh38, harmonises orientation and freezes them into immutable matrices, then evaluates nineteen variant effect predictors across sixteen strata using per-gene Spearman rho pooled by DerSimonian-Laird random-effects meta-analysis with per-stratum measurement-reliability estimates, attenuation correction, paired dependent-correlation tests and leave-one-gene-out validation, covering 64,178 variants and seven cancer susceptibility genes.

AI-generated editorial illustration: functional-standard-atlas: attenuation-corrected, territory-resolved benchmarking of variant effect predictors against saturation genome editing

Interpretation

It establishes an auditable, reproducible benchmark asset: seven MaveDB SGE score sets are mapped to GRCh38 and orientation-harmonised, then frozen into immutable matrices with a SHA-256 manifest, with one directory per predictor carrying a pinned environment and a score.py entry point. Relative to prior practice of using MaveDB scores without freezing, hashing or harmonising orientation, this treats the data asset itself as a first-class deliverable that can be verified byte by byte. Supported by the repository layout, the config/assays.yaml registry, the SHA-256 manifest under data/frozen/, and the CONVENTIONS.md rules that frozen data are immutable and dependencies pinned exactly; at the test level it reports 202 collected tests, 178 passing from a clean extract of the release archive, with a 179th verifying mapped reference bases and skipping where the Ensembl REST service is unreachable.

On the evaluation side it introduces per-stratum measurement-reliability estimates and attenuation correction, pooling per-gene Spearman rho across sixteen strata by DerSimonian-Laird random-effects meta-analysis, alongside paired dependent-correlation tests and leave-one-gene-out validation. Relative to reporting raw correlation coefficients alone, this explicitly estimates measurement reliability and corrects for attenuation, separating predictor performance from the ceiling imposed by the SGE measurements themselves. Supported by the analysis modules atlas.reliability (attenuation ceilings), atlas.evaluate_ext (the 19 x 16 grid), atlas.robustness (ties, power, paired tests) and atlas.simulate_attenuation (a 151,200-trial validation), with every number read from files under results/ rather than transcribed.

It provides a tiered reproduction path that distinguishes what can be re-run from what can be re-derived. Relative to simply stating that code is open, it uses four tiers to state what each level can do, its wall clock and disk footprint, and notes that Tier 2 (adding the Zenodo archive) regenerates every figure and table in the paper. Tier 1 runs 130 code-only guardrails after a git clone (72 skip without data, about 25 s, 20 MB); Tier 2 with the archive takes under 1 min and 60 MB; Tier 3 rebuilds the frozen matrix from MaveDB in about 20 min and 2.2 GB; Tier 4 re-scores every predictor from scratch in about 40 h and 6 GB, needing an API key and a GPU.

It distinguishes and discloses licence terms per data layer and per predictor column. Relative to labelling a whole dataset as openly licensed, it notes that the frozen matrix (assay measurements only, no predictor columns) is CC BY 4.0 without qualification, while the score matrices under results/ and everything derived from them carry nineteen predictor columns with mixed terms, eight more restrictive than CC BY 4.0 and three with no licence terms that could be located. Per-column licences with source URLs are recorded in results/predictor_resources_v1.tsv as the single source of truth, and LICENSE-DATA restates the full summary so it does not depend on that file being present.

Perspective

The benchmark targets SGE measurements for seven cancer susceptibility genes and nineteen predictors, and is meant for evaluation settings where predictor performance must be separated from the reliability ceiling of the measurements; its frozen matrix and scoring harness can be used directly to reproduce the paper's figures and tables (Tier 2), and can be extended to new predictors or new MaveDB assays via the template and assay registry. Re-deriving predictor scores from scratch requires an AlphaGenome API key and one NVIDIA A6000 (48 GB) GPU, with Evo2-7B alone taking 25 h for 46,392 SNVs; readers without a key or GPU can still reproduce every figure and number, but cannot re-derive those scores from the API.

Readers should still watch for the following: this repository text does not give the specific correlation values or rankings per gene and per predictor, which must be read from the tables under results/; attenuation correction depends on per-stratum measurement-reliability estimates whose derivation and uncertainty intervals need to be checked in the manuscript; the nineteen predictor columns carry mixed licence terms, three of which could not be located, so reuse of the score matrices and derived figures requires per-column confirmation; and the companion manuscript on ranking versus calibration of splice-region predictors uses an overlapping subset of the same MaveDB assays but estimates no measurement reliability and applies no attenuation correction, sharing no results section, figure or table, so the two sets of conclusions should not be interchanged directly.

Sources