PlantRegMoD: An Integrative and AI-Driven Multi-Omics Database for Plant Regeneration Research
Synopsis
This work constructed PlantRegMoD, an AI-powered integrated multi-omics database dedicated to plant regeneration, hosting 20.54 TB of standardized multi-omics data from 147 projects across 32 plant species and 2,593 samples, establishing a unified hierarchical classification system covering five major categories and nine regeneration models, curating 236 regeneration genes and their 28,190 homologs across 58 representative plant species, and containing 196,423 single cells and over 8.81 million epigenetic peaks, equipped with nine online omics tools and a RAG-based intelligent Q&A system.
Interpretation
Established a unified hierarchical classification system for plant regeneration covering nine experimental models including root tip regeneration, grafting, de novo root regeneration, callus regeneration, protoplast regeneration, and somatic embryogenesis, providing data summaries, pivotal regulatory genes, core molecular regulatory mechanisms, and brief experimental protocols for each model. Previously, the absence of unified classification criteria hindered cross-study mechanistic comparison and integration; this work revised and optimized this classification system based on prior work. Based on systematic curation of 147 projects covering 32 species and 2,593 samples, with all sequencing data processed through a standardized pipeline to ensure comparability and reliability.
Systematically compiled 236 regeneration-associated candidate genes in Arabidopsis thaliana (102 experimentally validated as core regeneration genes) and identified 75 orthologous groups comprising 28,190 regeneration-related genes across 58 representative plant species covering major evolutionary lineages including bryophytes, ferns, gymnosperms, lycophytes, basal angiosperms, monocots, and dicots. The only previously available regeneration-focused database, REGENOMICS, merely covered transcriptomic data lacking epigenetic resources and cross-species homolog analysis; this module provides phylogenetic trees, gene structure diagrams, multiple sequence alignments, and conserved motif maps for evolutionary analysis. The gene set was verified by GO enrichment analysis for high reliability, and homolog identification covered multiple evolutionary lineages across 58 species.
Integrated 2,221 transcriptomic datasets, 39 high-quality scRNA-seq datasets (covering 196,423 individual cells annotated into 39 cell types), and epigenomic data from 355 samples (including 78 ATAC-seq, 224 ChIP-seq, 39 BS-seq, and 14 m6A-seq datasets, totaling 8.81 million annotated peaks). No existing platform integrated multi-layered omics data with intelligent analytical tools to support user-friendly data mining and hypothesis validation for wet-lab researchers; this work fills this gap. The single-cell module integrates 210 experimentally validated marker genes and 609 newly screened candidate markers identified via rigorous statistical filtering; the epigenome module supports flexible data retrieval via gene-based or genomic locus-based queries.
Developed nine user-friendly bioinformatics utilities (including BLAST homology searching, multiple sequence alignment, phylogenetic visualization, GO and KEGG functional enrichment pipelines, PlantTF-PK for transcription factor and protein kinase prediction, JBrowse genome browser, PCR primer and sgRNA design tools) and a RAG-based intelligent Q&A system powered by GLM-4-Flash large language models. No existing platform integrated multi-layered omics data with intelligent analytical tools to lower bioinformatic barriers for wet-lab researchers; the Q&A system leverages a curated knowledge base and annotated literature to support natural language interaction. The Tools module integrates nine utilities, and the Q&A system is built on GLM-4-Flash large language models and RAG technology.
Perspective
The platform is intended to serve the plant regeneration research field, particularly to support wet-lab researchers in data mining and hypothesis validation. Its classification system covers nine regeneration models, and the gene module uses Arabidopsis thaliana as a reference model extended to 58 representative plant species. The platform plans regular updates of new datasets and functions to continue as a community resource.
This is a preprint that has not yet undergone peer review. The platform's actual usability, data update frequency, and long-term maintenance remain to be verified. The accuracy and coverage of the RAG Q&A system in practical use need user feedback for evaluation. The applicability of the classification system to non-model plants and the reliability of homolog gene function inference are directions worth watching. Additionally, because the current parse does not include supplementary figures and tables, some technical details (such as specific parameters of the standardized pipeline) cannot be presented in this summary.
