Skip to main content
Back to timeline
arXivSource publication:

SWIFT co-serves watermark construction with the vLLM inference engine, reaching 4.87 utility and 99.65% detection at 0.905 s latency on text generation

Related research and updates

Synopsis

The authors propose SWIFT, a framework that asynchronously co-serves text generation and watermark construction on the same vLLM backend, LLM, and GPU: it uses instruction-guided candidate generation with entity protection to produce context-aware substitutions, embeds watermarks via key-conditioned Tournament selection, and adapts watermark request admission with a scheduler driven by queue length and waiting-time pressure; on C4 text generation it attains the highest utility score of 4.87, 99.65% detection accuracy, and 0.905 s latency, retains 97.7% detection under substitution attacks, preserves task accuracy on GSM8K and MedMCQA, and reaches 528.7 tokens/s throughput with a 96.8% prefix-cache hit rate.

Source-provided article image: Adaptive Co-Serving LLM Watermarking on Modern Inference Engines
Figure 1 ·

Figure 1: SWIFT example. Multiple prompts are processed concurrently through a shared vLLM backend. For a Kyoto guide, SWIFT generates context-aware substitutions (e.g., “beautiful” → \rightarrow “wonderful”) while preserving key information such as locations and times. Watermark construction runs asynchronously alongside text generation, enabling streaming with minimal latency overhead.

arXiv

Interpretation

SWIFT turns watermark construction from an auxiliary pipeline into an asynchronous co-serving stream that shares one LLM, one vLLM engine, and one GPU with text generation, submitting completed segments as watermark requests at sentence boundaries while generation continues without waiting. Existing watermarking methods largely focus on the algorithm and often rely on external NLP tools such as spaCy, Stanza, or NLTK, or on auxiliary models, with limited use of asynchronous generation, KV-cache reuse, prefix sharing, and scheduling; SWIFT applies those serving optimizations directly to the watermark workload. Ablations show that disabling continuous batching and request prioritization (SWIFT-FB) raises latency by 20.67% with no utility change; replacing LLM-based entity handling with an external NER tool (SWIFT-NER) raises latency by 87.90% and lowers utility by 5.36%; removing entity-protection instructions (SWIFT-EN) improves latency by only 0.55% while utility drops by 50.95%.

Instruction-guided candidate generation lets the same LLM identify watermarkable words, protect named entities and sensitive spans, and generate context-aware substitutions under semantic and factual constraints, after which key-conditioned Tournament selection fixes the final substitutions. Unlike in-processing methods that act at every decoding step via logit bias or sampling, SWIFT selectively replaces watermarkable words within sentence segments and constrains candidates to a JSON mapping through structured decoding, avoiding extra semantic reranking models or similarity filters. On C4 text generation it reaches 4.87 utility, 99.65% detection, and 0.905 s latency; on CNN/Daily Mail summarization, 4.57 utility, 98.8% detection, and 0.559 s latency; in qualitative comparison KGW changes a tenure from 20 to 25 years and adds unsupported roles and personnel, SynthID changes it to 23 years and adds unsupported policing details, while SWIFT preserves key facts.

An adaptive co-serving scheduler estimates normalized watermark pressure from the watermark queue length and the oldest request's waiting time, periodically updates the watermark admission ratio, and allocates watermark capacity within the shared batch. Rather than reserving a fixed batch fraction for watermarking, the scheduler leaves capacity to text generation under low pressure, raises watermark admission as pressure grows to limit backlog, and returns capacity once pressure falls. Against fixed admission ratios, the adaptive scheduler achieves the lowest latency of 0.905 s; too-low fixed ratios cause watermark request backlogs and higher latency, while an overly high ratio interferes with text generation and also raises latency.

A key-conditioned contrastive detector encodes text with RoBERTa, maps the key to a deterministic embedding, and combines text and key representations via element-wise interaction and absolute difference, yielding 99.2–99.7% detection with the correct key and near-zero with incorrect keys. The detector learns watermark patterns from text-key pairs instead of predefined rules or statistical scores that attacks may disrupt, and a single detector generalizes across LLMs without per-model retraining. Across Gemma-3, Qwen-2.5, and LLaMA-3.1, correct-key detection is 99.2–99.7% while incorrect-key detection is 0–1.9%; on the unseen Ministral model only 3.7% of outputs are detected with the Gemma-3 key; ROC analysis gives AUC 0.9999 with 99.8% TPR at 1% FPR.

Perspective

The result targets a proprietary LLM service provider exposing an API, under a black-box adversary that can query the API and observe outputs but cannot access model parameters, architecture, or internal states. The method builds on vLLM and depends on prefix caching, continuous batching, structured decoding, and chunked prefill, with experiments on a single NVIDIA A100 80GB GPU, batch size 32, and KV-cache block size 16. Watermark quality depends on the capability of the shared LLM, since the same model performs text generation, watermarkable-word identification, and candidate substitution. Outputs are currently streamed at the sentence level rather than token by token. The detector is trained once on C4 samples and reused without retraining for summarization, mathematical reasoning, and medical QA.

Watermark density tracks detectability: reasoning text at 22.23% density yields 90.70% detection while C4 at 40.51% density yields 99.65%, so detection strength falls on tasks with constrained substitution choices. The SIRA attack reduces detection to 26.4%, but its attacked outputs achieve only a 0.145 pairwise win rate against the original watermarked outputs and it takes 8.933 s, so the trade-off between attack strength, utility, and overhead still warrants observation under more attacks. KV-cache reuse for segment-specific suffixes remains unimplemented because text generation uses the original prompt and history while candidate generation uses a separate watermarking context. Token-by-token streaming is not yet supported. The ethics statement notes that watermark detection is probabilistic and should not be treated as sole evidence of authorship, misuse, or policy violations, and that real-world deployment also requires secure key management and consideration of adversarial removal or forgery, privacy, fairness, and distribution shift.

Sources