
Trung Dang Thanh, Thi Thanh Thuy Pham, Huong-Giang Doan · Engineering Technology & Applied Science Research 2026 · 2026
DOI: 10.48084/etasr.20412
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
In this paper, we propose SPARTA (Sparse Patch Attack for Revealing Transformer Attention vulnerabilities), an attention-guided white-box adversarial attack framework for investigating attention-based weaknesses in Vision Transformer (ViT)-based medical image classifiers. SPARTA employs attention rollout to identify attention-critical patches and restricts adversarial perturbations to these highly influential regions, generating sparse yet effective adversarial examples while preserving visual fidelity. SPARTA is evaluated on four medical imaging datasets, namely ISIC2018, ChestX-ray14, OCT2017, and PathMNIST, using ViT-B/16, DeiT-S, and ResNet50 architectures. The experimental results show that SPARTA consistently outperforms conventional adversarial attacks, achieving attack success rates of up to 89.5% while perturbing only a small fraction of the input image. Visual quality analysis further demonstrates that SPARTA preserves image fidelity, yielding higher SSIM and PSNR values and lower LPIPS scores than conventional full-image attacks. Ablation studies indicate that rollout-guided patch selection is the primary factor underlying the effectiveness of the proposed framework. Overall, the results provide strong empirical evidence that transformer attention mechanisms are closely associated with a practical attack surface and that predictive importance in ViTs is highly concentrated within a limited set of attention-dominant patches. These findings provide new insights into the structural vulnerabilities of transformer-based medical imaging systems and highlight the need for attention-aware robustness evaluation and defense strategies.
No comments yet — start the discussion below.