Corpus-wide causality: Algorithm design & application for aggregating gene-disease causal evidence
Synopsis
This work develops a method to infer a Corpus-Wide Causal Score (CWCS) for a gene-disease pair by integrating network-based causal signals in a gene regulatory network (CWCS-Net) with corpus-wide literature evidence from PubMed abstracts quantified by a newly developed Truth Discovery algorithm (CWCS-TD), achieving a causal class F1 score of 0.600 across ten diseases using OMIM as an external expert-curated reference, outperforming GPT-4o (0.505) and MMed-Llama 3 (0.522).
Fig. 1. Overview of our work. Note that the same abstract can contain multiple G-D pairs (e.g., Abstract 1) and a G-D pair can be mentioned in different abstracts (e.g., G1-D1). C/NC represents the causal/non-causal prediction from CRED-trained model, where CRED denotes Causal Relation Extraction Dataset. RRF refers to Reciprocal Rank Fusion. CWCS (comb) score of a G-D pair quantifies the evidence for G causing D by combining network-based evidence (CWCS-Net) and literature-based evidence (CWCS-TD) scores. Abs denotes Abstract. G-G denotes a gene-gene interaction.
bioRxiv · Page 2Interpretation
It proposes the CWCS method that integrates network evidence and literature evidence into a corpus-wide causal score for gene-disease pairs. Prior studies mostly extracted causal relations from abstracts but rarely summarized corpus-level evidence; this work combines network causal signals with literature evidence into a unified causal scoring framework. Evaluation used OMIM as an external expert-curated reference, reporting a causal class F1 score of 0.600 across ten diseases.
It develops a new Truth Discovery (CWCS-TD) algorithm that jointly and iteratively estimates causal scores for multiple gene-disease pairs while modeling the reliability of PubMed abstracts co-mentioning them. By incorporating bibliometric features of publications, the algorithm addresses the sparsity of abstracts that assert a gene-disease causal relation, which the authors describe as an advance in the field of Truth Discovery algorithms. The text reports ablation studies supporting the design that integrates network- and literature-based evidence.
On the causal classification task, the CWCS method outperforms two large language model baselines, GPT-4o and MMed-Llama 3. There was a distinct lack of explicit benchmarking comparing generalized LLM-based methods against specialized, domain-aware frameworks for corpus-wide causal inference; this work provides that comparison. CWCS achieved a causal class F1 of 0.600 versus 0.505 for GPT-4o and 0.522 for MMed-Llama 3; the performance trend persists when using area under the precision-recall curve as the evaluation metric. Both LLMs exhibit high recall with comparatively low precision, resulting in many false positive predictions.
Perspective
The result is intended for research and application settings that need to aggregate gene-disease causal evidence from network and literature sources, for example to focus drug discovery; evaluation uses OMIM as an external expert-curated reference across ten diseases, and the method is designed to integrate network- and literature-based evidence, on which basis the authors propose broader applicability.
A careful reader may wonder about generalization beyond the ten diseases, robustness under sparse abstracts, and the relative contribution of network versus literature evidence; because the current reading scope is the abstract level, figures and full ablation details are not included, so specific values on these points remain open questions.
