NUS team introduces a residual-transferability metric and CoverLock, cutting forgery success below 0.25% across three high-RT watermarking systems
Synopsis
The work formalizes neural image watermark forgery as residual transferability (RT), a metric that uses ground-truth residuals to measure how well watermark evidence stays decodable after transfer across unrelated images, finds that common training-side variations cannot explain the large RT gaps while architectural design plays a central role, identifies Broadcast–GAP and content-adaptive embedding as two mechanisms that strengthen watermark dependence on the cover image, and introduces CoverLock, a plug-and-play strategy that drives average forgery success below 0.25% under two existing residual-based attacks across the high-RT systems CIN, MBRS, and VINE while attaining higher robustness.
Interpretation
The paper introduces residual transferability (RT) as a system-level property, transferring ground-truth residuals from source images to a disjoint set of target images and measuring normalized average transfer accuracy, thereby characterizing susceptibility to residual-based forgery independently of any specific residual-estimation procedure. Prior work focused on constructing increasingly effective forgery attacks; this work makes the question of why residuals transfer across images the object of study and provides a metric decoupled from attack method. RT values are reported for five representative watermarking schemes: CIN 1.00, VINE 0.99, MBRS 0.98, versus HiDDeN 0.28 and RivaGAN 0.08; Appendix D shows the qualitative high/low-RT ordering holds across RT@1, RT@10, RT@100, and RT@1000.
By comparing encoder and decoder behavior in high- and low-RT systems, the paper finds that high-RT systems encode and decode watermark signals with limited reliance on the host image, showing more consistent residual patterns after cross-image averaging and a smaller relative contribution of cover-induced variance at the decoder's final feature layer. This supplies behavioral evidence for a cover-agnostic shortcut, moving the RT gap from an attack phenomenon to a difference in internal representations. Residuals are averaged over groups of five cover images to suppress cover-dependent components, and relative feature variance is compared under fixed-message/varying-cover and fixed-cover/varying-message conditions at the decoder's final layer.
Through controlled interventions the paper identifies two architectural mechanisms that strengthen coupling between watermark evidence and the cover image: the coupled Broadcast–GAP design in HiDDeN, where the message is spatially broadcast during embedding and the decoder preserves spatial representations until late global average pooling, and the content-adaptive attention in RivaGAN's encoder. Ablation further shows Broadcast and GAP do not act independently: removing broadcast alone leaves RT@100 nearly unchanged, removing GAP alone raises it moderately, and removing both sharply increases RT@100, indicating a coupled design. Original and modified models are retrained under the same distortion-free setting for controlled comparison; introducing BG into high-RT watermarks substantially reduces RT, while directly introducing it into VINE did not converge reliably given VINE's already challenging optimization; replacing RivaGAN's adaptive mechanism with a fixed embedding pattern causes a large RT increase.
The paper proposes CoverLock, a watermark-agnostic plug-in that derives an image-conditioned binary code from frozen DINOv2 patch-feature channel-wise mean and standard deviation, wraps the message by bitwise XOR, and recomputes the code from the received image at detection, strengthening verification's dependence on image content without modifying the original watermark encoder or decoder architecture or parameters. Unlike architectural redesign that is method-specific and requires retraining, CoverLock trains only a lightweight projection head and a learnable logit scale, using only MSCOCO natural images and their transformed views, without involving the watermark encoder, decoder, or any forgery samples. Evaluated on CIN, MBRS, and VINE across DIV2K, ImageNet, and MS-COCO under GT@1, Yang et al. (2024a) with 100 references, and WMForger with a single reference, CoverLock keeps even its highest average ASR below 0.25%, with average ASR of 0.243%, 0.015%, and 0.059%, outperforming the ConvNeXt classifier baseline and MHDW; average bit accuracy under noise, brightness, blur, JPEG, contrast, and crop distortions exceeds MHDW.
Perspective
The work targets post-processing neural image watermarking. The threat model assumes an adversary with access to one or a few released watermarked images carrying the same message, who may use external datasets or pretrained models to estimate transferable watermark evidence, but who does not know the deployed watermarking method or any protection mechanism, has no access to encoder or decoder parameters, and cannot query the encoder for additional watermarked images. The RT metric uses ground-truth residuals to remove the confounding effect of any particular residual-estimation procedure; the architectural conclusions come from the two low-RT positive cases HiDDeN and RivaGAN, while CoverLock is validated on the three high-RT systems CIN, MBRS, and VINE. For practitioners hardening deployed watermarking systems, CoverLock offers a path that requires neither redesign nor retraining of the underlying watermark model; for researchers designing new watermarking architectures, Broadcast–GAP and content-adaptive embedding point to concrete directions for strengthening cover dependence.
CoverLock achieves low forgery success and a favorable security–robustness trade-off within the distortions considered, but whether that trade-off holds under more severe or unseen image transformations, broader distribution shifts, and adaptive attacks specifically targeting the image-binding mechanism remains an open question. CoverLock introduces an additional image-binding module and associated training and deployment overhead. The paper identifies two mechanisms associated with low RT, but other architectures may achieve cover dependence through different design principles. In addition, introducing Broadcast–GAP into VINE did not converge reliably, indicating the design is not equally feasible for all architectures. This summary is based on the full paper text and does not verify every figure and table value in the appendices.
