SHERLOCK models perturbation effects as structured interventions on a latent state, recovering pathway relationships and predicting combinatorial effects across CRISPR, chemical, and spatial perturbation data
Synopsis
The authors present SHERLOCK, an interpretable deep generative framework that represents genetic and pharmacological perturbations as structured interventions on a latent baseline cellular state, learns correlated and sparse perturbation representations, enables counterfactual estimation of downstream transcriptional effects under explicit identifiability assumptions, quantifies condition-dependent responses, and compositionally models combinatorial perturbations; across genome-scale CRISPR, chemical, and spatial perturbation datasets, it recovers perturbation relationships concordant with known biological pathways and pharmacological properties, identifies condition-dependent responses, and predicts combinatorial perturbation effects.
Figure 1: SHERLOCK: a generative framework for disentangling perturbation effects in single- cell data. (a) Overview of SHERLOCK. Single-cell gene expression data, together with perturbation and condition labels, are modeled using a structured variational autoencoder (VAE) to disentangle perturbation effects, context-dependent responses, and combinatorial interactions. (b) Plate diagram of the SHERLOCK generative process. Uncolored circles denote inferred variables, colored circles denote observed variables, and squares denote observed constants. (c) Main functionalities of SHERLOCK. (d) Left: UMAP of the original gene expression space from the Replogle dataset, colored by pathway-level annotations. Middle: UMAP of the original gene expression space, colored by Leiden clusters. Right: SHERLOCK latent representation of the same cells, colored by pathway-level annotations. (e) Correlation matrix of the learned perturbation embedding. Reference pathway annotations were annotated by Replogle et al., whereas mapped pathways were inferred from SHERLOCK groupings. (f) Gene set enrichment analysis (GSEA) of the inferred down- stream target genes for each perturbation group, highlighting pathway-relevant functional signatures. Color indicates −log10-transformed adjusted p-values, and point size represents the number of overlapping genes between the inferred target genes and enriched pathways. (g) Benchmarking of SHERLOCK on the Replogle
bioRxiv · Page 6Interpretation
SHERLOCK represents perturbation effects as structured interventions on a latent baseline cellular state, learning correlated and sparse perturbation representations that organize genetic and pharmacological perturbations by shared transcriptional responses. Existing computational approaches are typically designed for individual tasks such as predicting perturbation responses, characterizing downstream transcriptional effects, or modeling perturbation combinations, lacking a unified framework; this work places interpretable perturbation representation learning alongside downstream causal analysis and condition-dependent response characterization in one framework. The abstract states that the representation recovers perturbation relationships concordant with known biological pathways and pharmacological properties across genome-scale CRISPR, chemical, and spatial perturbation datasets; specific model details and quantitative metrics are not given in the loaded text.
By formulating perturbations as interventions within a structural causal model, SHERLOCK enables counterfactual estimation of downstream transcriptional effects under explicit identifiability assumptions. This extends perturbation analysis from correlation-based prediction toward counterfactual causal inference, with the identifiability assumptions stated explicitly. The abstract states the causal formulation and identifiability assumptions; the loaded text does not include the specific form of those assumptions or validation experiment details.
The same framework quantifies how perturbation responses vary across conditions and compositionally models combinatorial perturbations, supporting prediction of held-out combinations and classification of genetic interactions. Condition-dependent response characterization and combinatorial perturbation modeling were previously handled by separate methods; here they are unified within a single generative framework. The abstract reports identifying condition-dependent responses and predicting combinatorial perturbation effects across CRISPR, chemical, and spatial perturbation data; specific dataset sizes and evaluation metrics are not given in the loaded text.
Perspective
The work targets single-cell perturbation experiment settings and is meant for researchers who need interpretable perturbation representations, causal analysis of downstream effects, and characterization of condition-dependent responses together, for example teams working with genome-scale CRISPR screens, chemical perturbations, or spatial perturbation data. The applicability of counterfactual estimation rests on the explicit identifiability assumptions stated in the abstract; combinatorial perturbation prediction targets held-out combinations, and genetic interaction classification targets perturbation pairs amenable to compositional modeling.
The loaded text is an incomplete abstract and does not include figures, specific dataset sizes, baseline comparisons, or quantitative results, so the accuracy of counterfactual estimation on real data cannot be judged, nor can the conditions under which the identifiability assumptions hold be confirmed. Readers may watch for: the specific form of the structured intervention representation and sparsity constraints, the stability of condition-dependent responses across different condition settings, and the generalization of combinatorial perturbation prediction on held-out combinations. In addition, the authors declare a provisional patent application related to the described methods, which readers may note when assessing application prospects.
