Skip to main content
Back to timeline
bioRxivSource publication:

Speciesformer: Learning Conserved Cellular States for Cross-Species Generative Virtual Cell Modeling

Synopsis

This work presents Speciesformer, a cross-species generative single-cell foundation model pretrained on SpeciesCorpus (about 131 million cells from 11 species, 154 tissues, and more than 923 cell types) that maps species-specific genes into a shared evolution-informed gene space via a macrogene vocabulary, whose encoder learns transferable cell and gene representations and whose unified generative decoder supports bidirectional generation between transcriptomic states and biological text as well as prediction of post-perturbation transcriptomes from initial cell states and intervention descriptions, reporting performance above existing baselines across multiple benchmarks.

Source-provided article image: Speciesformer learns conserved cellular states for cross-species generative virtual cell modeling
Fig. 1d ·

specific outputs (Fig. 1d). Full details of Speciesformer architecture are provided in Methods and Supplementary Notes.

bioRxiv · Page 4

Interpretation

It builds a shared cross-species gene space: using ESM-2 protein embeddings together with kernel archetypal analysis and an inverse evolutionary-tree strategy, it constructs a Deuterostomia-level macrogene vocabulary yielding 862 macrogenes shared across 11 species, aligning species-specific genes without relying on one-to-one ortholog mapping. Unlike prior cross-species models that treat protein embeddings directly as gene tokens, this work uses them to construct an evolutionarily layered macrogene vocabulary, preserving species-specific transcriptional information within a shared gene space. The method section describes the macrogene construction procedure and the count of 862 shared macrogenes, with pretraining on SpeciesCorpus (131M cells, 11 species, 462 datasets, 154 tissues).

The encoder provides cell and gene representations that generalize across tasks: in cell type annotation on five human tissue datasets, Speciesformer improves overall accuracy over scGPT, Geneformer, UCE, and CellFM and achieves the highest macro-F1 on four of five datasets; on the Adamson and Norman CRISPR Perturb-seq datasets it reaches global Pearson correlations of 0.98 and 0.99 on held-out perturbations. It applies cross-species pretrained representations to annotation and genetic perturbation prediction and evaluates generalization on held-out perturbations rather than interpolating within seen perturbations. Annotation covers three intra-dataset and two inter-dataset settings; perturbation tasks are evaluated on held-out perturbations in Adamson (86 single-gene perturbations, 68,603 cells) and Norman (105 single-gene and 131 double-gene perturbations, 91,205 cells) across six metrics.

Gene embeddings encode interpretable functional structure before fine-tuning: on a human immune dataset, gene embeddings group functionally related genes into 25 modules, where M10 is enriched in membrane transport and the ESCRT-III-mediated nuclear membrane closure process and M11 involves translation, nonsense-mediated mRNA degradation, and rRNA processing, with within-module average correlation higher than random gene sets matched for expression mean and variance (999 permutations, all with BH q = 0.0125). It shows that pretrained gene embeddings themselves carry regulatory and functional relationships, not only as a product of downstream fine-tuning. Based on Reactome AUROC analysis across 25 modules and 35 cell types, with permutation testing and BH q values, and target cell type AUROC of 0.913 and 0.859 for M10 and M11.

The unified generative decoder enables bidirectional cell-state/text generation and perturbation-conditioned generation: on a human cross-tissue immune dataset (29,790 cells, 34 cell types), expression-to-text generation achieves 73.4% exact label match (2,185 of 2,978 test cells correct), and text-to-expression generation outperforms scVI, scDiffusion, scGPT, C2S, scMMGPT, and InstructCell in both zero-shot and fine-tuned settings; on Tahoe-100M, SciPlex3, and Parse-PBMC, perturbation-conditioned generation achieves leading or near-leading performance across multiple metrics. It unifies comparative analysis, representation probing, cell-state generation, and perturbation prediction within one architecture and one orthology-aware macrogene space rather than treating them as disconnected problems. Text generation is evaluated with BLEU, ROUGE, METEOR, MMD, EMD, and label accuracy; pseudo-cell generation with an external kNN classifier and marker-gene patterns; perturbation prediction reports PCC(Delta), PDS(L2), DE overlap, AUROC, and AUPRC under a unified Cell-Eval protocol.

Perspective

The results are aimed at cross-species single-cell transcriptomics research, applicable where a shared cellular state space, cross-species knowledge transfer, text-conditioned cell-state generation, and cross-context perturbation response prediction are needed, for example transferring from data-rich to data-poor species or predicting drug or cytokine responses in unobserved cellular contexts. The design intends comparisons across species, tissues, and assays within one orthology-aware macrogene space, so applicability to human-only analyses or analyses requiring spatial organization depends on future extensions.

A careful reader would still watch how batch effects and dataset-specific biases are handled within the unified framework, which the text lists as a shared open issue for current foundation models; spatial transcriptomics is not yet incorporated, leaving how tissue architecture and local neighborhoods shape cellular states open; perturbation prediction performance across broader perturbation modalities, datasets, and biological contexts needs continued evaluation; and because this is a preprint, some results depend on supplementary figures and tables, so verifying specific metrics and module details is limited if those materials are not visible.

Sources