Yixiang Pan, Zhouyan He, Ting Luo, Wenpeng Xing, Yufeng Li, Meng Han · Neurocomputing 2026 · 2026
DOI: 10.1016/j.neucom.2026.135262
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Text-to-Image (T2I) models demonstrate powerful image generation capabilities, yet their misuse for generating harmful content poses serious societal risks. Common defense strategies such as concept erasure and content moderation suffer from notable limitations: the former often fails to completely remove harmful concepts and may impair generation quality, while the latter is susceptible to evasion and lacks fine-grained semantic understanding. To address these challenges, we propose EmbedGuard, a concept-aware safety framework for T2I models. EmbedGuard constructs a multi-level semantic anchor system within the text encoder, leveraging its hierarchical semantic properties to precisely characterize various harmful concepts. By computing the strength of the semantic association between token embeddings and anchors at different semantic depths, EmbedGuard can accurately quantify and identify a broad spectrum of risks, including explicit harmful terms, implicit malicious expressions, and adversarial perturbations. Once risks are identified, a learnable embedding transformation module uses a nonlinear mapping to smoothly guide high-risk embeddings toward semantically safer representations, helping suppress harmful content generation while preserving benign semantics and image quality. Extensive experiments demonstrate that EmbedGuard significantly improves T2I safety by reducing Not-Safe-For-Work (NSFW) content generation while maintaining high image quality. EmbedGuard also exhibits strong robustness against adversarial attacks, with the Max configuration achieving an NSFW generation rate of only 0.2% under the challenging Ring-A-Bell attack.
No comments yet — start the discussion below.