Langping Wang, Bin Fang · International Journal of Pattern Recognition and Artificial Intelligence 2026 · 2026
DOI: 10.1142/s0218001426590421
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Video captioning faces significant challenges in balancing global semantic consistency with fine-grained syntactic accuracy, particularly in complex multi-event scenarios. Existing global alignment methods often suffer from semantic dilution, while fine-grained approaches based on Part-of-Speech tagging frequently encounter compositional failures. To bridge these gaps, we propose EAS-Cap (Event-Aware Semantic-Enhanced Captioning), a novel framework that redefines the fundamental learning unit as Atomic Events—indivisible Subject-Verb-Object triplets extracted via syntactic analysis. EAS-Cap decomposes the captioning task into three synergistic objectives: atomic event extraction, cross-modal semantic distillation, and global-local consistency optimization. Specifically, our Event Module employs learnable queries to distill textual priors into visual encoders, enhancing sensitivity to specific actions, while a Visual-to-Event constraint ensures local features remain consistent with global context. Extensive experiments on the MSVD and MSR-VTT datasets demonstrate that EAS-Cap achieves excellent performance, notably leading in CIDEr scores. These results confirm our method's superior capability in preserving syntactic integrity and accurately describing complex temporal dynamics without relying on noisy global signals or explicit grammatical constraints.
No comments yet — start the discussion below.