Context-Aware Semantic Similarity Measurement for Unsupervised Word Sense Disambiguation
Synopsis
The paper proposes a context-aware semantic similarity (CASS) method for unsupervised word sense disambiguation: given a target word, its context and an exclusion list, the context and the context with each candidate synonym substituted are encoded as vectors, and the synonym that minimizes the resulting semantic change (maximizes cosine similarity) is chosen as the sense decision; evaluated for accuracy on CoarseWSD-20, a Wikipedia-derived, noun-only benchmark of 20 words with 2 to 5 senses and 10,196 instances, against a random-option baseline (43.73%) and a most-frequent-sense baseline (73.43%), the BERT-embedding configuration reached 77.74% (7,927 hits), the only setting above the strong baseline, while USE (71.94%), ELMo (68.75%) and WMD (60.
Fig. 1 Results obtained for the CoarseWSD-20 dataset using UWSD+BERT. The solid blue bar represents the results obtained. While the black and red colors represent the weak and strong baselines, respectively
· Page 11Interpretation
It defines a CASS decision procedure: given a word w, a context C and an exclusion list E, the method encodes C and the context with each candidate synonym s_i substituted for w, and selects the candidate that minimizes one minus the cosine similarity, i.e. the smallest semantic change. Earlier semantic similarity measures relied mainly on intrinsic word features or corpus co-occurrence rather than on how the surrounding context changes after substitution; this work formalizes that criterion in Eq. 3 and realizes it with pre-trained transformer embeddings for unsupervised sense assignment. The contribution is a formal definition accompanied by comparative experiments across four embedding families (BERT, ELMo, USE, WMD) and several concrete models; the author reports that every configuration beats the weak baseline on CoarseWSD-20, and that only the BERT configuration beats the strong baseline.
On CoarseWSD-20 the BERT-based CASS configuration reaches the highest accuracy of 77.74% (7,927 hits), above the MFS strong baseline of 73.43%, which relies on external knowledge. The author reports this gain without any annotated training data or external knowledge; within the same family, all-mpnet-base-v2 (77.74%), all-MiniLM-L12-v2 (75.05%) and all-MiniLM-L6-v2 (74.63%) all sit above MFS. An accuracy comparison on a single benchmark, reported as hit counts and percentages (Table 1, Table 2); the author also notes being unaware of other published unsupervised WSD results on this task for direct comparison.
Within-family comparisons show embedding choice is a decisive variable: USE peaks at 71.94%, ELMo at 68.75% and WMD at 60.00%, all above the weak baseline of 43.73% but below the MFS 73.43%, with USE's Large model coming closest to the baseline. Earlier work often treats the introduction of context itself as the source of improvement; this paper presents the same CASS pipeline across four embedding families and their individual models. Tables 3, 4 and 5 list concrete hit counts and accuracies per model within each family, giving a comparable within-pipeline contrast; the author also notes that why some considered models rank better than others remains future work.
The evaluation setup uses CoarseWSD-20 (built from Wikipedia, nouns only, 20 words, 2 to 5 coarse senses per word, 10,196 instances), with accuracy reported per use case and globally, and the source code released. Because the method is fully unsupervised, the author deliberately avoids comparison with solutions that use training samples and compares only against the weak (random option) and strong (most frequent sense) baselines, noting that some mapping classes were adapted to ease disambiguation. The dataset and metric are described explicitly in the text, the baselines cover both a weak (random) and a strong (most frequent sense) reference, and per-word figures plus global summary tables accompany the results.
Perspective
This work makes word sense assignment possible under no-annotation conditions by relying on off-the-shelf contextual embeddings: it is directly relevant for low-resource languages with scarce annotation and for settings that depend on semantic similarity such as search, document classification, question answering and text summarization; its conclusions are framed for English nouns, the coarse senses of CoarseWSD-20 (2 to 5 per word) and the accuracy metric. The author suggests future directions of incorporating additional contextual sources or combining the approach with other WSD techniques, and the released code makes it practical to move this synonym-substitution-and-embedding-similarity pipeline onto other embedding models or corpora.
The author notes that the approach depends on context quality: when the context is noisy or too sparse, disambiguation accuracy may be affected, and the reasons why different embedding models rank differently are not explained, with deeper analysis left as future work. As read here, the text presents Figures 1 to 4 and Tables 1 to 5, but the per-word bar values in the figures cannot be read off item by item, so per-word accuracy details and whether significance testing was performed would need checking against the original figures. Two further points are worth watching: the mapping classes for coarse senses were adapted, and the paper does not list other published unsupervised results on the same task for side-by-side comparison.
