Wenjin Hu, Qiheng Sun · Journal of King Saud University - Computer and Information Sciences 2026 · 2026
DOI: 10.1007/s44443-026-00998-8
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Thangka, a distinct cultural heritage of Tibet, is characterized by rigorous symmetric layouts, concentrated core semantics, and a unique "centric-radiant" composition. However, existing Mamba-based image captioning models lack domain-specific prior knowledge, making it difficult to generate precise and culturally compliant descriptions. To address this, we propose GCP-Mamba (Geometric and Compositional Priors Mamba), a state space model that integrates geometric and compositional priors. Specifically, we design a Symmetry-Aware Fusion (SAF) module to explicitly model geometric symmetry, thereby calibrating spatial orientation logic and enhancing feature robustness. Furthermore, we introduce a Compositional Priors State Space Model (CP-SSM), which concurrently disentangles the fine-grained entity associations within the Principal Deity region and the macro-level "inside-out" layout. Experimental results on the D-Thangka dataset demonstrate that GCP-Mamba achieves a CIDEr score of 895.7 and a BLEU-4 of 94.4, outperforming baseline models by 34.18% and 3.20%, respectively. These results validate the effectiveness of our proposed modules in capturing the complex structures and microscopic details of Thangka imagery, providing a novel technical path for the intelligent protection of digital cultural heritage.
No comments yet — start the discussion below.