Yutong Wu, Wenyue Li, Hewang Nie, Ziqi Zhou, Jin Li, Yixiao Gong, Songfeng Lu · Cybersecurity 2026 · 2026
DOI: 10.1186/s42400-026-00654-8
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Model watermarking supports intellectual-property claims by verifying a model’s responses to a secret key set, but this behavior-only interface is vulnerable to fabricated evidence. This work presents FakeMark , a gradient-guided false-claim attack for image classifiers that uses a white-box surrogate but never queries or accesses the victim model during attack construction. Under a simplified linear decision-boundary model, targeted perturbations can acquire a nonzero component along the watermark-trigger direction; experiments on deep networks provide only conditional, setting-dependent support for this intuition. FakeMark caches selected convolutional and fully connected layer outputs from clean surrogate batches and injects them through stochastic multi-layer, channel-wise interpolation to improve transfer. Across 16 distinct architectures and an additional adversarially trained ResNet-50 checkpoint variant, over eight evaluated watermark variants, retrospective best-case behavioral target-label accuracy reaches 1.00 on CIFAR-10 and 0.99 on ImageNet. ImageNet transfer varies substantially across checkpoint and surrogate settings, ranging from near zero to 0.99. Matched baselines and detector analyses motivate provenance-aware, multi-factor ownership protocols.
No comments yet — start the discussion below.