Hossam Magdy Balaha, Ahmed Sharafeldeen, Magdy Hassan Balaha · Information 2026 · 2026
DOI: 10.3390/info17100973
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Background: Vision Transformers (ViTs) have demonstrated efficacy in visual recognition tasks through global dependency modeling via self-attention. However, standard ViT architectures lack native mechanisms for integrating external clinical metadata (e.g., patient demographics, lesion characteristics, image quality indicators) into the visual processing pipeline, limiting their utility in multi-modal diagnostic scenarios. Methods: We propose two novel attention mechanisms: Q-Conditioning and Q-Gating, designed for efficient metadata fusion within the transformer framework. In Q-Conditioning, metadata is projected into the query space and added to the query vector, guiding attention toward contextually relevant regions during early computation. In Q-Gating, a learned gate modulates computed attention scores, enabling soft, differentiable control over token interactions. We further introduce a hybrid attention framework wherein standard, Q-Conditioning, and Q-Gating heads coexist within the same model. Results: We evaluated our approach on four medical imaging benchmarks (PAD-UFES-20, TissueNet, PH2, and NDB-UFES). Specific configurations, particularly Q-Gating and hybrid variants (e.g., QC_QG, QG_S), achieved >95.5% accuracy on PAD-UFES-20 and >92% average score on PH2, significantly outperforming image-only baseline models (p<0.001). On TissueNet, where auxiliary inputs consist of PhikonV2 foundation-model embeddings rather than raw clinical metadata, all configurations exceeded 99% accuracy; however, we note that these gains are substantially attributable to the pretrained encoder and are reported separately from the primary clinical-metadata claims. Attention concentration metrics, encompassing both theoretical conditional entropy reduction and empirical absolute entropy dispersion, correlated strongly with diagnostic performance gains (ρ=0.89), indicating improved focus on diagnostically salient regions while avoiding spurious attention collapse. Conclusions: Metadata-aware attention mechanisms enhance both performance and interpretability in domains where auxiliary information is critical. This work provides a scalable extension to ViTs that supports multi-input modeling while preserving architectural integrity and computational efficiency.
No comments yet — start the discussion below.