
Timothy Dillan, Sani Muhamad Isa, Abba Suganda Girsang, Derwin Suhartono · Engineering Technology & Applied Science Research 2026 · 2026
DOI: 10.48084/etasr.20178
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
The rapid growth of digital media has increased the availability of posts that combine spoken narration, visual demonstrations, embedded images, captions, and video sequences within a single communicative unit. Although existing information extraction systems perform well on textual data, their effectiveness remains limited when entities and relations are distributed across synchronized text, image, audio, and video modalities. This study proposes an LLM-guided framework for multimodal entity and relation extraction from sequential audio-visual sources. The framework integrates temporal decomposition, context injection, schema-guided verification, and cross-chunk entity resolution to support structured knowledge extraction across modalities. Evaluation on 200 annotated multimodal documents shows an overall entity extraction F1 score of 82.4%, with modality-specific F1 scores of 87.3% for text, 83.1% for video, 81.6% for embedded images, and 78.9% for audio. Temporal decomposition with context injection improves cross-chunk entity linking accuracy by 9.7%, while schema-guided verification reduces incomplete and unsupported outputs and increases overall F1 by 5.5%. The framework achieves this result at an estimated cost of $0.64 per item under the evaluated deployment scenario, indicating that prompt-based multimodal extraction can be implemented without supervised fine-tuning. This study contributes an integrated approach for extracting structured knowledge from four-modal sequential content units and clarifies the methodological trade-offs involved in applying multimodal large language models to real-world information extraction tasks.
No comments yet — start the discussion below.