Shiyu Wang, Dajun Liu · International Journal of Pattern Recognition and Artificial Intelligence 2026 · 2026
DOI: 10.1142/s0218001426510146
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Fine-grained neural discrepancy learning is an important problem in multimodal representation learning because heterogeneous modalities often contain partially aligned, weakly correlated, or locally contradictory semantic information. Existing multimodal learning methods mainly rely on deterministic feature fusion or cross-modal attention, but they often fail to model hidden topic-level uncertainty and fine-grained discrepancy between textual, visual, and auxiliary representations. To address this problem, this paper proposes VTL-FND, a Variational Topic Latent Modeling Algorithm for Finegrained Neural Discrepancy Learning. VTL-FND first encodes textual, visual, and metadata information into a shared representation space, then infers a latent topic variable through variational inference, and finally combines topic-level semantic abstraction with contradiction-aware multimodal evidence for semantic discrepancy prediction. The proposed framework integrates ELBO-based latent regularization, cross-modal consistency learning, and supervised classification into a unified optimization objective. Experiments on multiple multimodal semantic consistency benchmarks show that VTL-FND consistently outperforms representative baselines in Accuracy, Precision, Recall, F1-score, and AUC. The results demonstrate that latent topic modeling and contradiction-aware learning provide robust and interpretable improvements for fine-grained multimodal discrepancy learning.
No comments yet — start the discussion below.