SemVac: A Semantic Vaccinology Paradigm Powered by LLMs for Antigen Discovery
Synopsis
The work introduces semantic vaccinology and implements it as SemVac: publications linked to each protein are retrieved through PaperBLAST, condensed into a structured semantic profile, and an LLM is prompted to return an antigenicity probability; on a curated 246-protein bacterial benchmark the best of 14 general-purpose LLMs matched or exceeded the precision of the specialized predictor PLGDL, with open-weight Kimi K2 0905 offering the strongest performance-cost balance, predictions were robust to masking of vaccine keywords, reproducible across repeated inference, and generalized to a 1,200-protein cross-pathogen dataset; explicit chain-of-thought reasoning increased recall but lowered precision in every model tested; applied to the mpox virus proteome, SemVac recovered the established
Interpretation
It proposes and implements semantic vaccinology, treating the published literature as an explicit, auditable third modality alongside sequence and structure for protective antigen prediction. Existing reverse vaccinology methods rely mainly on sequence and structure and overlook the functional and immunological knowledge recorded in the literature; SemVac retrieves protein-linked publications via PaperBLAST, condenses the evidence into a structured semantic profile, and prompts an LLM for an antigenicity probability, making literature evidence a traceable basis for prediction. Benchmarked against a curated 246-protein bacterial benchmark and the specialized protein-language and geometric-deep-learning predictor PLGDL, the best of 14 general-purpose LLMs matched or exceeded PLGDL's precision; predictions were robust to masking of vaccine keywords, reproducible across repeated inference, and generalized to a 1,200-protein cross-pathogen dataset.
It provides an actionable LLM selection reference for the performance-cost trade-off: the open-weight Kimi K2 0905 offered the strongest balance. The comparison spans 14 general-purpose LLMs rather than a single-model demonstration, giving empirical grounding for deployment trade-offs. Based on a head-to-head evaluation on the same 246-protein bacterial benchmark, with Kimi K2 0905 reported as offering the strongest performance-cost balance.
It finds that explicit chain-of-thought reasoning causes over-reasoning in biological scoring: recall rose while precision fell in every model tested. This counterintuitive result suggests that adding explicit reasoning steps does not necessarily improve discrimination quality on antigenicity scoring, offering directional evidence for prompt design. The pattern of increased recall and lowered precision was observed consistently across all models tested.
On the mpox virus proteome it recovered the established antigen repertoire and prioritized uncharacterized candidates, while showing that confabulation can be detected and corrected. A30L and C19L received independent experimental support in recent orthopoxvirus vaccine development, providing external validation; B20R produced a coherent but false TNF-decoy narrative unsupported by curated annotations, showing that cross-checking reasoning traces against curated resources can expose confabulation. External validation comes from independent experimental support in recent orthopoxvirus vaccine development; confabulation detection is achieved by cross-checking reasoning traces against curated resources.
Perspective
The paradigm targets research and early discovery settings that screen protective antigens from pathogen proteomes, and applies where relevant literature can be retrieved for each protein and condensed into a structured semantic profile; its value lies in turning literature evidence into an auditable third modality for vaccine target prioritization. For proteins with sparse literature coverage, the information content of the semantic profile is correspondingly limited.
A careful reader would still watch how the construction of semantic profiles affects antigenicity probabilities when literature is sparse or descriptions vary in quality; whether the recall-up, precision-down effect of chain-of-thought holds across broader tasks; how B20R-style confabulation can be identified for candidates lacking curated annotations to cross-check against; and how reproducible the external experimental support for A30L and C19L is in larger cohorts.
