Hybrid lexical-semantic retrieval over SNOMED CT: combining two retrieval paradigms to facilitate clinical data entry
Synopsis
The work proposes and implements a hybrid retrieval architecture that lets deterministic lexical matching and learned semantic matching coexist over SNOMED CT's own curated descriptions, combining multi-prefix search, BioLORD-2023-M embeddings, an optional BGE cross-encoder for re-ranking, and a local LLM for query normalization, with Reciprocal Rank Fusion and a hierarchy filter; on a search-only linking evaluation over 542 disease mentions from the DisTEMIST corpus (Spanish, zero-shot), semantic search with re-ranking and no LLM query pre-processing reached accuracy@1 of 0.60 and recall@10 of 0.80, and on 12,897 mentions from English real-EHR discharge notes a field-scoped typeahead placed the concept on the top-10 picker list for 73% of mentions (accuracy@1 0.
Figure 1 summarizes these functions and the flow between them.
medRxiv · Page 7Interpretation
It proposes a reference architecture in which lexical and semantic retrieval coexist in one retrieval flow rather than being chosen between, and provides a reproducible instance over the SNOMED CT International edition. Where the conventional way to avoid adding a semantic channel has been to strengthen the lexical one by curating an additional local interface vocabulary, this work instead computes additional access routes over SNOMED CT's own maintained descriptions. One reproducible instance was implemented over the SNOMED CT International edition, combining deterministic multi-prefix search, BioLORD-2023-M embeddings, an optional BGE cross-encoder, and a local LLM (gemma 4), integrated with Reciprocal Rank Fusion and a hierarchy filter.
On the DisTEMIST Spanish zero-shot search-only linking evaluation, the best and computationally cheapest configuration was semantic search with re-ranking and no LLM query pre-processing. That configuration reached accuracy@1 of 0.60 and recall@10 of 0.80 over 542 disease mentions, and because these figures are computed on gold mention spans they isolate normalization from mention detection and are not directly comparable to the DisTEMIST shared task's end-to-end scores. The external evaluation covered 542 disease mentions, resolved gold codes through SNOMED CT historical associations, and ablated query pre-processing, re-ranking, and each retrieval channel in isolation.
The two channels coexist without a trade-off: on this clean-normalization corpus the semantic channel carried the accuracy, and adding the lexical channel neither helped nor harmed it. Accuracy@1 was 0.60 versus 0.61, because rank fusion followed by re-ranking suppresses whichever channel is off-task; the lexical channel's own strength, partial as-typed input, is a regime this corpus does not test. This came from ablations isolating each retrieval channel, supported by a near-miss analysis showing many remaining cases retrieved a parent or child of the intended concept.
On English real-EHR discharge notes, a field-scoped typeahead placed the concept on the top-10 picker list for most mentions, and a context-aware pre-process worked as a selective rescue. Across 12,897 mentions, 73% reached the top-10 picker list (accuracy@1 0.52; findings 0.54, close to DisTEMIST); most residual errors were soft, the hard misses were dominated by ambiguous EHR shorthand, and a single context-aware pre-processing call roughly doubled recovery of those hard misses, though applying it to every mention slightly lowered overall accuracy by perturbing the easy majority. This was estimated on the English real-EHR discharge notes of the SNOMED CT Entity Linking Challenge, 12,897 mentions across all clinical domains, and tested a context-aware pre-process.
Perspective
The results are meant for settings that build retrieval services over SNOMED CT's maintained descriptions, and for clinical data entry flows that want to support direct search without additional local interface vocabulary; the authors note that curated interface vocabularies remain useful in selected guided workflows. The reported figures are specific to the evaluated configuration, not an architecture-independent estimate.
The lexical channel's own strength, partial as-typed input, is a regime this corpus does not test, so whether the two channels also coexist without a trade-off there remains to be seen; LLM query pre-processing helped dirty, lay input but slightly hurt clean terminology-like queries, and applying context-aware pre-processing to every mention slightly lowered overall accuracy, so the trigger conditions for its selective use deserve attention; and because the search-only figures are computed on gold mention spans, they are not directly comparable to end-to-end scores, leaving end-to-end performance an open question.
