Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

Randomized Key Selection Cuts Blind-Attacker Watermark Forgery Success from 87% to 1% Without Further Utility Loss

The work proposes a defense that is agnostic to the underlying watermarking method: it randomizes watermark key selection per query and accepts content as genuine only if a watermark is detected by exactly one key, yielding a sample-count-independent upper bound on forgery success for blind attackers conditional on key-symmetric, independent detector outcomes without further degrading model utility; against the adaptive blind attackers evaluated, harmful-text forgery success at r=4 keys drops from as high as 87% with a single key to as low as 1%, and a preliminary image study shows a reduction from 100% to 2%.