BioEvidence lets frozen LLMs match strong specialized protein predictors in zero-shot variant ranking, with performance rising as evidence gets richer
Synopsis
The authors introduce BioEvidence, a training-free and model-agnostic interface that converts structural and evolutionary information from standard biological tools into compact evidence for frozen large language models, and on the ProteinGym benchmark they observe evidence scaling: performance improves as evidence becomes richer, structural and evolutionary evidence each help and combining them yields further gains, mismatching the same evidence to the wrong variants degrades performance below the no-evidence baseline, and the interface enables zero-shot ranking to reach strong specialized protein predictors on matched evaluations while the improvement persists on data released after the model's knowledge cutoff.
Figure 1: BioEvidence pipeline. Structural features and evolutionary statistics are computed for the candidate variants and rendered as compact evidence blocks for a frozen general-purpose LLM. The model parameters remain unchanged; only the biological evidence available at inference time varies.
bioRxiv · Page 2Interpretation
BioEvidence is a training-free and model-agnostic interface that converts structural and evolutionary information from standard biological tools into compact evidence for frozen large language models. Rather than adapting model parameters to close the gap with specialized protein models, this work scales the biological evidence available to a frozen model. The abstract states the interface is training-free and model-agnostic and evaluates it on the ProteinGym benchmark; implementation details and the specific tool list are not given in the abstract.
On ProteinGym the authors observe evidence scaling: performance improves as evidence becomes richer, with structural and evolutionary evidence each improving performance and their combination yielding further gains. This positions external evidence as a complementary scaling axis alongside model capability and inference effort, rather than a single-point method improvement. An empirical observation on the ProteinGym benchmark; the abstract reports directional findings without specific numbers or effect sizes.
Whether evidence matches the variant determines the direction of the effect: mismatching the same evidence to the wrong variants degrades performance below the no-evidence baseline. This indicates the gains do not come from generic value in the evidence text but depend on the correspondence between evidence and target variant. The abstract describes a mismatching condition as a control but does not report the magnitude of the degradation.
BioEvidence enables zero-shot ranking to reach strong specialized protein predictors on matched evaluations, the improvement persists on post-cutoff data released after the model's knowledge cutoff, and evidence interacts with conventional scaling, for example GPT-5.6 Sol with evidence at low reasoning effort outperforms the no-evidence condition at medium effort, while a six-model analysis associates stronger no-evidence performance with larger margins over evolutionary rank fusion. It places evidence scaling in the same frame as inference effort and model capability and adds a cross-model association. The abstract reports matched evaluations, post-cutoff data, and a six-model analysis, all as directional statements without specific metric values.
Perspective
This work speaks to researchers and practitioners who need zero-shot prediction of protein variant effects, in settings where standard biological tools can supply structural and evolutionary information and model parameters stay frozen. It suggests an actionable path: instead of fine-tuning a model, organize structural and evolutionary evidence into a compact input and keep evidence strictly matched to the target variant. It also suggests a trade-off between inference effort and evidence effort, where low reasoning effort with sufficient evidence may outperform simply raising reasoning effort.
The visible text is only the abstract and contains no figures, specific metric values, evidence-construction details, or benchmark subset breakdowns, so the size of the gains, the magnitude of the mismatching degradation, and the robustness of the six-model association cannot be judged. Whether evidence scaling holds beyond ProteinGym or in other scientific prediction domains, and how much evidence quality versus evidence quantity contributes, remain open questions to watch.
