Skip to main content
Back to timeline
arXivSource publication:

SignRAG combines hierarchical pretraining with retrieval-augmented reinforcement fine-tuning to make a gloss-free SLT model beat gloss-supervised methods on all CSL-Daily metrics

Synopsis

SignRAG is a unified gloss-free sign language translation framework that first learns linguistically grounded sign representations via fine-grained sign-text alignment, then jointly pretrains the sign encoder with Qwen3-8B, adds a target-domain-only retrieval gallery for instance-specific translation cues, and applies RUG-RFT combining translation-quality and retrieval-utility rewards to suppress harmful reliance; it sets new state-of-the-art results on CSL-Daily, PHOENIX-2014T, How2Sign, and OpenASL, and is reported as the first gloss-free approach to surpass gloss-supervised methods across all reported metrics on CSL-Daily.

Source-provided article image: SignRAG: Unified Retrieval-Augmented Gloss-Free Sign Language Translation
Figure 1 ·

Figure 1: Motivation. Hierarchical pretraining and retrieval augmentation improve sign-to-language learning. (a) , (b) Comparison of different pretraining strategies. The solid and dashed lines denote the previous state-of-the-art performance under the gloss-free and gloss-based settings, respectively, corresponding to Geo-Sign ( Fish and Bowden, 2026 ) and CV-SLT ( Zhao et al., 2024 ) . (c) Comparison of the semantic distance between generated and ground-truth (GT) texts with and without RAG. Semantic distance is measured using the independent sentence encoder paraphrase-multilingual-MiniLM-L12-v2 . Bars report the mean cosine distance over all samples, with error bars denoting ± \pm SEM. Samples are drawn from the CSL-Daily dev set.

arXiv

Interpretation

Hierarchical pretraining splits large-scale sign pretraining into two stages: first learning linguistically grounded sign representations with a bidirectional InfoNCE contrastive objective and a CoCa-style skeleton-to-text objective, then jointly pretraining the aligned skeleton encoder, a GELU-MLP projector, and a LoRA-adapted Qwen3-8B. Prior gloss-free pretraining largely targeted T5-style encoder-decoder models or pretrained only the encoder; this work reorganizes pretraining for a decoder-only LLM via staged alignment, explicitly aiming to mitigate cross-modal optimization imbalance. On CSL-Daily, hierarchical pretraining reaches BLEU-4 29.72 and ROUGE-L 58.78, versus 9.53 and 38.70 for end-to-end pretraining and 3.20 and 22.51 for encoder-only pretraining; ablation shows pretraining raises the SFT baseline from ROUGE-L 55.03 to 58.78 and BLEU-4 from 23.46 to 29.72.

Retrieval-augmented SFT builds a sign-to-text gallery from the target-domain training split only, retrieves semantically related samples and their translations via attention-weighted frame-level sign similarity as instance-specific cues, and applies retrieval dropout during training. Existing retrieval-based SLT mainly targets gloss-based settings; this work extends retrieval augmentation to the gloss-free setting and enforces that retrieved knowledge always comes from the target domain, with validation and test samples used only as queries. On CSL-Daily, adding retrieval raises ROUGE-L from 58.78 to 60.20 and BLEU-4 from 29.72 to 31.22; retrieval-quality analysis shows retrieval F1 and translation performance first improve then decline as the number of retrieved contexts grows, indicating that extra less-relevant contexts can interfere.

RUG-RFT, built on GRPO, combines a sentence-level quality reward with a retrieval-utility reward that measures the teacher-forcing conditional-likelihood gain from retrieval under a frozen reference policy, retaining only positive utility and normalizing reward magnitudes. Standard SFT optimizes token-level likelihood and gives no explicit signal about whether retrieved contexts help; this work turns retrieval utility into an optimizable signal while suppressing reliance on misleading contexts. On CSL-Daily, quality reward alone gives BLEU-4 31.97 and retrieval-utility reward alone 31.66, while combining both reaches 32.91 with ROUGE-L 62.34; chrF rises from 26.93 to 28.42 and BERTSim from 82.10 to 82.96, indicating gains beyond BLEU/ROUGE; a failure case shows SFT misled by incorrect retrieval while RFT recovers the correct translation.

Systematic evaluation across four downstream benchmarks shows SignRAG outperforms existing gloss-free state-of-the-art methods both with and without external pretraining data, and surpasses gloss-supervised methods on CSL-Daily. The authors report this as the first gloss-free SLT approach to outperform gloss-supervised methods across all reported metrics on CSL-Daily, and show BLEU-4 rising from 22.72 to 28.58 and 32.91 as the Qwen3 backbone scales from 0.6B to 4B and 8B. On CSL-Daily, SignRAG-RFT with external pretraining reaches ROUGE-L 62.34 and BLEU-4 32.91 versus Geo-Sign's 57.97 and 27.42; on OpenASL it reaches 46.91 and 26.62 versus Uni-Sign's 43.22 and 23.14; on How2Sign 38.5 and 16.7 versus SSVP-SLT's 38.4 and 15.5; on PHOENIX-2014T, without additional pretraining, 53.07 and 28.63.

Perspective

This work targets gloss-free sign language translation in settings where large-scale sign-text pretraining corpora exist and a retrieval gallery can be built from the target-domain training split; the gallery is built only from that split, validation and test samples serve only as queries, and training excludes the query itself and entries with identical reference translations. The authors note PHOENIX-2014T results use no additional pretraining because no large-scale German sign language pretraining dataset exists, and gains over prior state of the art are more moderate on English datasets than on Chinese ones. On efficiency, online inference latency rises from 42.69 ms/sample to 61.98 ms/sample, throughput falls from 23.42 to 16.13 samples/s, and peak GPU memory rises from 23.28 to 29.03 GiB/GPU; offline retrieval takes 23.88 ms/query on the 18.4K CSL-Daily gallery and 705.85 ms/query on the 95,888-sample OpenASL gallery. The complete training pipeline requires about 413.99 GPU-hours.

Million-scale retrieval remains unverified: the authors state they do not report direct retrieval results for BOBSL (about 1.2 million sentences) and avoid extrapolating current measurements to that scale, and they do not claim coarse-to-fine retrieval can fully preserve exhaustive-retrieval accuracy. The retrieval-utility reward relies on likelihood differences under a frozen reference policy, and its stability across languages and data scales needs further validation. Gains on English datasets are more moderate, which the authors partly attribute to dataset difficulty, evaluation-protocol differences, and incomplete YouTube-ASL pretraining data; how these factors affect extrapolation of the conclusions is worth watching.

Sources