FakeMark: gradient-guided false watermark claims via robust feature fusion

Abstract Model watermarking supports intellectual-property claims by verifying a model’s responses to a secret key set, but this behavior-only interface is vulnerable to fabricated evidence. This work presents FakeMark , a gradient-guided false-claim attack for image classifiers that uses a white-box surrogate but never queries or accesses the victim model during attack construction. Under a simplified linear decision-boundary model, targeted perturbations can acquire a nonzero component along the watermark-trigger direction; experiments on deep networks provide only conditional, setting-dependent support for this intuition. FakeMark caches selected convolutional and fully connected layer outputs from clean surrogate batches and injects them through stochastic multi-layer, channel-wise interpolation to improve transfer. Across 16 distinct architectures and an additional adversarially trained ResNet-50 checkpoint variant, over eight evaluated watermark variants, retrospective best-case behavioral target-label accuracy reaches 1.00 on CIFAR-10 and 0.99 on ImageNet. ImageNet transfer varies substantially across checkpoint and surrogate settings, ranging from near zero to 0.99. Matched baselines and detector analyses motivate provenance-aware, multi-factor ownership protocols.

Authors

Institutions

Publication Details

Journal
Cybersecurity
Published
2026-09-17
DOI
https://doi.org/10.1186/s42400-026-00654-8
Primary Topic
Adversarial Robustness in Machine Learning
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

FakeMark: gradient-guided false watermark claims via robust feature fusion

Songfeng Lu, Hewang Nie, Jin Li, Yixiao Gong et al.
Cybersecurity
Adversarial Robustness in Machine Learning
article

FakeMark: gradient-guided false watermark claims via robust feature fusion

Songfeng Lu, Hewang Nie, Jin Li, Yixiao Gong, Ziqi Zhou, Wenyue Li, Yutong Wu
article en

Abstract

Abstract Model watermarking supports intellectual-property claims by verifying a model’s responses to a secret key set, but this behavior-only interface is vulnerable to fabricated evidence. This work presents FakeMark , a gradient-guided false-claim attack for image classifiers that uses a white-box surrogate but never queries or accesses the victim model during attack construction. Under a simplified linear decision-boundary model, targeted perturbations can acquire a nonzero component along the watermark-trigger direction; experiments on deep networks provide only conditional, setting-dependent support for this intuition. FakeMark caches selected convolutional and fully connected layer outputs from clean surrogate batches and injects them through stochastic multi-layer, channel-wise interpolation to improve transfer. Across 16 distinct architectures and an additional adversarially trained ResNet-50 checkpoint variant, over eight evaluated watermark variants, retrospective best-case behavioral target-label accuracy reaches 1.00 on CIFAR-10 and 0.99 on ImageNet. ImageNet transfer varies substantially across checkpoint and surrogate settings, ranging from near zero to 0.99. Matched baselines and detector analyses motivate provenance-aware, multi-factor ownership protocols.

CybersecurityVol. 9(1)
Chongqing University (CN), Guangxi Normal University (CN), Second Hospital of Yichang (CN), Shenzhen Technology University (CN), Tongji Hospital (CN), Huazhong University of Science and Technology (CN), Guilin University of Electronic Technology (CN)
National Natural Science Foundation of China, Science, Technology and Innovation Commission of Shenzhen Municipality, Natural Science Foundation for Young Scientists of Shanxi Province, Young Scientists Fund, Changsha Science and Technology Project
Peace, Justice and strong institutions
Openalex Percentile: Top 9%
Adversarial Robustness in Machine Learning
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.