Public articles linked to the same research event.
arXiv The authors propose SWIFT, a framework that asynchronously co-serves text generation and watermark construction on the same vLLM backend, LLM, and GPU: it uses instruction-guided candidate generation with entity protection to produce context-aware substitutions, embeds watermarks via key-conditioned Tournament selection, and adapts watermark request admission with a scheduler driven by queue length and waiting-time pressure; on C4 text generation it attains the highest utility score of 4.87, 99.65% detection accuracy, and 0.905 s latency, retains 97.7% detection under substitution attacks, preserves task accuracy on GSM8K and MedMCQA, and reaches 528.7 tokens/s throughput with a 96.8% prefix-cache hit rate.
The authors propose SWIFT, a framework that asynchronously co-serves text generation and watermark construction on the same vLLM backend, LLM, and GPU: it uses instruction-guided candidate generation with entity protection to produce context-aware substitutions, embeds watermarks via key-conditioned Tournament selection, and adapts watermark request admission with a scheduler driven by queue length and waiting-time pressure; on C4 text generation it attains the highest utility score of 4.87, 99.65% detection accuracy, and 0.905 s latency, retains 97.7% detection under substitution attacks, preserves task accuracy on GSM8K and MedMCQA, and reaches 528.7 tokens/s throughput with a 96.8% prefix-cache hit rate.
The authors propose SWIFT, a framework that asynchronously co-serves text generation and watermark construction on the same vLLM backend, LLM, and GPU: it uses instruction-guided candidate generation with entity protection to produce context-aware substitutions, embeds watermarks via key-conditioned Tournament selection, and adapts watermark request admission with a scheduler driven by queue length and waiting-time pressure; on C4 text generation it attains the highest utility score of 4.87, 99.65% detection accuracy, and 0.905 s latency, retains 97.7% detection under substitution attacks, preserves task accuracy on GSM8K and MedMCQA, and reaches 528.7 tokens/s throughput with a 96.8% prefix-cache hit rate.
The authors propose SWIFT, a framework that asynchronously co-serves text generation and watermark construction on the same vLLM backend, LLM, and GPU: it uses instruction-guided candidate generation with entity protection to produce context-aware substitutions, embeds watermarks via key-conditioned Tournament selection, and adapts watermark request admission with a scheduler driven by queue length and waiting-time pressure; on C4 text generation it attains the highest utility score of 4.87, 99.65% detection accuracy, and 0.905 s latency, retains 97.7% detection under substitution attacks, preserves task accuracy on GSM8K and MedMCQA, and reaches 528.7 tokens/s throughput with a 96.8% prefix-cache hit rate.