Skip to main content
Back to timeline
arXivSource publication:

RAGenome scales a retrieval-based genomic language model to 13,312 nucleotides, lifting gene-finding MCC from 0.45 to 0.60

Synopsis

The authors introduce RAGenome, the first retrieval-based genomic language model: it retrieves homologous sequences from a 100-vertebrate whole-genome alignment, discards the 73% of tokens that are gaps, and scales pretraining context to 13,312 nucleotides, 100 times longer than existing MSA-based gLMs; it raises gene-finding MCC from GPN-Star's 0.45 to 0.60 and reaches 0.87 AUROC on pathogenic-variant prioritization, unifying long-range and evolutionary signals with 168M parameters and 16,032 GPU hours.

Source-provided article image: RAGenome: Scaling Retrieval-Based Genomic Language Models to Long Contexts
Figure 1 ·

Figure 1: Overview of the retrieval mechanism of RAGenome. For each query sequence, we retrieve aligned homologous sequences from a WGA and remove gap tokens, preserving the original alignment column of each retained token as its positional embedding. We then concatenate all sequences, add taxonomy embeddings to distinguish species, and restrict retrieval to a fixed budget B B using the phylogenetic tree. Tokens are then masked per clade, and the resulting sequence is passed to the RAGenome Transformer, which predicts the masked query tokens.

arXiv

Interpretation

RAGenome extends MSA-based genomic language models from windows of roughly 128–256 nucleotides to 13,312 nucleotides while retaining explicit cross-species evolutionary modeling. Earlier MSA-based models (GPN-MSA, GPN-Star) process every token of the alignment, so cost grows with both context length and species count and they remain confined to short windows; RAGenome keeps only aligned nucleotides, discards gap tokens, and selects species under a fixed token budget guided by the phylogenetic tree. The authors report that 73% of tokens in the 100-vertebrate WGA they use are gaps, versus 41% gaps in the top 5% most conserved regions, showing that dropping gaps substantially reduces input tokens for whole-genome pretraining; the retrieval budget is bounded by GPU memory and set to 80,000 tokens in the largest context setting.

Context extension drives large gains in gene finding while variant effect prediction stays flat, indicating long-range capability and evolutionary signal can coexist in one model. Prior models tended to trade one for the other: MSA-based models excel at variant effect prediction but have short contexts, while large-scale gLMs handle long range but are weaker on variant effect prediction. Gene-finding MCC rises from 0.47 to 0.52 when context extends to 4,096 and from 0.54 to 0.57 when it extends to 13,312; the final model reaches MCC 0.60 versus GPN-Star's 0.45. Variant effect prediction remains flat across pretraining stages, ending at 0.87 AUROC. An inference-context ablation shows performance improves monotonically with inference context length at a fixed checkpoint.

RAGenome is competitive on both tasks with 168M parameters and 16,032 GPU hours, narrowing the gap to large-scale gLMs. NT-MS (2.5B parameters, 215,000 GPU hours) reaches 0.66 gene-finding MCC but only 0.57 on variant effect prediction; Evo2 (7B parameters, about 1M GPU hours) reaches 0.81 and 0.80; RAGenome reaches 0.60 and 0.87 at much smaller scale. The comparison draws on Table 1's parameters, maximum length, MSA usage, GPU hours, and both metrics; per-class confusion matrices show RAGenome clearly outperforms NT-MS on donor and acceptor splice sites, while NT-MS is better only on introns, which span far more nucleotide positions in the test set.

Retrieval information is essential for both tasks, but the two tasks depend on the retrieval budget to different degrees. The ablation varies the retrieval budget from 0 to the training budget, directly testing the contribution of the retrieval mechanism rather than context length alone. Without retrieval, performance collapses, as expected since the model was never trained without retrieved context; a small budget already yields substantial improvement; variant effect prediction saturates at a moderate budget while gene finding improves slightly up to the full budget.

Perspective

The work targets genomic tasks that require modeling both cross-species evolutionary relationships and within-species long-range interactions, such as gene finding and variant effect prediction. It assumes a whole-genome alignment is available at inference: the model uses human as the query species and retrieves homologous sequences from a 100-vertebrate alignment, so it currently applies to sequences covered by a human-referenced alignment. The authors note the architecture supports any query species and alignment, but applying it to a new species would likely require additional pretraining. Code and weights are public, and training and evaluation data are publicly available, supporting reproduction and extension under the same setting.

The conclusions rest on a single 168M-parameter model with human as the query species; scaling in model size, data, and context length remains untested. On gene finding, RAGenome still trails NT-MS (0.66) and Evo2 (0.81), and the authors expect further scaling to close the gap, an expectation that remains to be verified. Per-class results show RAGenome is stronger on splice sites while NT-MS is stronger on introns, and the aggregate MCC is influenced by how many nucleotide positions introns occupy in the test set, so the overall score can mask class-level differences. The model also depends on an alignment being available at inference and does not yet apply to sequences or species without any alignment.

Sources