Near-tie restriction lets watermark-fingerprinted models keep text quality at equal detectability
Related research and updatesSynopsis
The authors revisit a watermark-distillation protocol for LLM fingerprinting, show that its utility evaluation understates text-quality degradation in open-ended generation, and propose near-tie restriction, which uses top-1-relative logit gaps to confine the watermark bias to tokens close to the base model's top prediction; across Llama-3.2-3B, Qwen-2.5-3B, Llama-3.1-8B, and Gemma-9B, this improves detection-quality frontiers under sampling, system prompts, quantization, pruning, and fine-tuning, preserves higher text quality across query budgets, and further improves existing schemes when combined with them.
Figure 1: Overview of rethinking watermark teachers for model fingerprinting. Top: watermark-based fingerprinting distills a text watermark into model weights on a semantic domain such as French. The watermark splits tokens into green and red sets and biases green tokens to create a detectable statistical signal. Bottom: prior teachers bias every green token for individual-response detection, whereas fingerprint verification aggregates evidence across responses and permits sparser signals. Our near-tie restriction applies the bias only to green tokens near the top-1 logit.
arXivInterpretation
A re-evaluation finds that a recent watermark-fingerprinting protocol understates text-quality degradation in open-ended generation: in a Llama-3.1-8B-Instruct reproduction, French Benchmark accuracy drops only from 0.64 to 0.61, while PPL rises from 4.63 to 28.27 and the LLM-judge score falls from 7.76 to 3.73. The prior evaluation relied on accuracy-based tasks and reported a metric labeled PPL that is actually mean token entropy rather than standard perplexity, which favors overly strong watermark teachers. Reproduction on 500 French-translated WritingPrompts examples under the baseline setting, reporting accuracy, PPL, and LLM-judge scores together, with degradation consistently visible in open-ended generation.
The work argues that fingerprint verification can aggregate signal across queries, making sparser watermark signals viable, and then analyzes where the sparse signal should be placed: at matched base probability mass and bias strength, placing the bias on tokens the base model finds more plausible yields a smaller surprisal shift. Prior watermark distillation directly inherited a decoding-time KGW teacher designed for per-output detectability and biased every green token; this work identifies the objective difference between fingerprint verification and text provenance and provides a theoretical comparison of signal placement (Proposition 5.1). Proposition 5.1 proves that subsets with equal base probability mass give the same expected watermark evidence and KL divergence from the base distribution, while the lower-average-surprisal subset induces a smaller surprisal shift; a random-placement control matched on teacher green mass and induced KL yields greater surprisal shift and a larger student PPL ratio.
The work proposes near-tie restriction, which uses top-1-relative logit gaps to confine the watermark bias to green tokens close to the base model's top prediction, yielding KGW-NT while keeping the null green-token rate at 1/2 so the standard analytical detector remains applicable. Unlike KGW-TK, which truncates by a fixed top-k rank, near-tie selects tokens by logit gap; it is complementary to SWEET and MorphMark and can further restrict their biased sets to give SWEET-NT and MorphMark-NT. Detection-quality frontiers are built across four models, French and medicine fingerprint domains, and five deployment transformations; KGW-NT achieves higher worst-case z-scores at comparable text quality and outperforms KGW-TK, indicating the gain comes from logit-gap selection rather than sparsity alone.
In query-budget analysis, KGW-NT attains the lowest PPL ratio at each budget while remaining robustly detectable; some 3B-model settings reach PPL ratios below 1, and diversity (distinct-2/3, self-BLEU) stays close to the base model with LLM-judge scores comparable to or above it. The detectability-quality trade-off is extended from single-point comparisons to achievable-quality curves as a function of query budget, showing that aggregating evidence across queries can buy higher text quality. Lowest PPL ratios per budget are reported on Llama-3.2-3B, Qwen-2.5-3B, Llama-3.1-8B, and Gemma-2-9B; a temperature-tuning control shows KGW cannot recover the KGW-NT frontier even with additional tuning, indicating the gain is not merely distribution sharpening.
Perspective
The result targets black-box ownership verification where the owner retains the base model and secret key and can access a suspect model only through queries, with French as the primary fingerprint domain and an additional medicine domain; the method applies to KGW-style red-green watermarks and their extensions and is evaluated under sampling, system prompts, FP8/INT4 quantization, WANDA/SparseGPT pruning, and French WildChat fine-tuning. For model owners, this means choosing a sparser teacher signal closer to top-1 at a given query budget, reducing quality loss within the fingerprint domain while maintaining robust detection; for watermark designers, near-tie offers a way to narrow the biased token set without replacing the original mechanism.
The near-tie threshold is calibrated as a percentile of the top-1-top-2 logit-gap distribution on fingerprint training data, so how that percentile transfers across data distributions remains to be observed; the authors also note that computational constraints prevented exhaustive exploration of the joint hyperparameter space, and SWEET-NT and MorphMark-NT were evaluated only at representative configurations. The detector scores only positions satisfying the key-independent logit-gap criterion, while the teacher still applies bias when top-1 is green, and how this asymmetry behaves under more extreme deployment transformations deserves continued tracking. In addition, text quality is mainly measured by an external reference model's PPL and GPT-5 scores focused on grammar, disfluency, malformed words, and excessive repetition, so broader task-level effects remain an open question.
