Evo 2 shows in-context learning on five binary classification tasks with F1 up to 0.902 on short sequences, but collapses at kilobase scale and the 7B model beats the 40B
Synopsis
This study maps the in-context learning operating regime of Evo 2, a nucleotide-level foundation genomic language model, across five binary classification tasks spanning biological and artificial sequences, finding robust performance on shorter natural sequences (F1=0.902 for miRNA, 0.785 for Toxins), degradation with sequence length and collapse at kilobase scale, no benefit from model scaling (the 7B model systematically outperforms the 40B variant), poor prediction of accuracy by perplexity, and mechanistic interpretability via logit-lens and Jacobian Scope suggesting a prediction-generalisation trade-off and that models might track prompt structure rather than signal-carrying content.
Figure 1. Experiment structure. Schematic representation of the main steps composing the experimental structure, exemplified with a single Evo 2 model and microRNA sequences. Step 1, each class pool (class 0, miRNA; class 1, non-miRNA) is split independently into train and test sets. Step 2, per-class splits are merged into balanced train and test sets. Step 3, N examples per class from the train set are concatenated with a test set query, each encoded as a [sequence] →[buffer_run][label_run] pair. Frozen Evo 2 scores both candidate labels, and the lower-perplexity candidate is shown as the predicted class. Step 4, step 3 is repeated for each test query, with different values of N; predictions are then evaluated against ground truth with Macro F1.
bioRxiv · Page 4Interpretation
The study establishes empirical boundaries of Evo 2's in-context learning: robust performance on shorter natural sequences with F1=0.902 for miRNA and 0.785 for Toxins, degrading with sequence length and collapsing at kilobase scale. In-context learning has been extensively studied in general large language models, while its capability boundaries and underlying mechanisms remain poorly characterised in genomic language models; this work systematically maps Evo 2's operating regime across five binary classification tasks spanning biological and artificial sequences. Evidence comes from F1 metrics reported across five binary classification tasks covering biological and artificial sequences; the visible text is an abstract and does not provide sample sizes, data splits, or statistical tests.
Model scaling brings no benefit: the 7B variant systematically outperforms the 40B variant. This runs counter to the usual scaling intuition that larger models perform better, suggesting that scale is not a reliable driver of performance in in-context learning for genomic language models. Evidence is a systematic comparison of two model sizes (7B and 40B) across five tasks, described in the abstract as 'systematically outperforms', without per-task values or significance tests.
Perplexity, a widely used proxy for genomic language model performance, poorly predicts classification accuracy. This finding offers an empirical caution against the common practice of using perplexity for model selection and evaluation, indicating a weak link to downstream classification accuracy. Evidence comes from a comparison between perplexity and accuracy; the abstract reports no correlation coefficients or specific statistics.
Mechanistic interpretability analyses offer a consistent explanation: logit-lens profiling suggests a prediction-generalisation trade-off, while Jacobian Scope indicates models might track prompts' structure rather than the signal-carrying content. The work connects behavioural performance observations with internal mechanism analyses, offering mechanistic clues for why in-context learning degrades with length and why scale does not help in genomic language models. Evidence comes from two interpretability methods (logit-lens and Jacobian Scope), presented in the abstract with hedged language such as 'suggests' and 'might track', making these suggestive rather than confirmatory conclusions.
Perspective
The results define the usable regime of Evo 2 for binary classification on shorter natural sequences, applicable to settings that use prompt-injected examples for inference with relatively short sequences, such as miRNA- and toxin-related classification; for kilobase-scale long-sequence tasks, the current evidence suggests not relying on its in-context learning. For practitioners, this means that when selecting genomic language models, one should not assume that larger scale or lower perplexity implies better downstream classification performance, but should evaluate directly on the specific task and sequence length. The mechanistic logit-lens and Jacobian Scope analyses offer testable hypotheses for follow-up work, namely whether models actually use signal-carrying sequence content.
The currently visible text is only an abstract and does not provide the specific datasets, sample sizes, sequence length distributions, or training and evaluation splits for the five tasks, nor per-task values and significance tests for the 7B versus 40B comparison, so the magnitude and robustness of 'systematically outperforms' remain to be confirmed. The perplexity-accuracy relationship lacks quantitative descriptions such as correlation coefficients. The logit-lens and Jacobian Scope conclusions are phrased with 'suggests' and 'might track', making them suggestive explanations whose causal status and reproducibility await verification in the full text. In addition, all five tasks are binary classification, so extrapolation to multi-class, regression, or generative tasks is unclear; the specific threshold for the kilobase-scale collapse is also not given in the abstract.
