Mohaned Zkaria Salem, Ali Hussein Khalaf AL-Sammarraie · Bilad Alrafidain Journal for Engineering Science and Technology 2026 · 2026
DOI: 10.56990/bajest/2026.050218
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Vision-Language Transformers (VLTs) support strong video-text alignment. Still, they are trained with quadratic complexity, making them impractical for deployment on resource-constrained edge devices. In this study, we present Dynamic Token Pruning for Cross-Modal Alignment (DTP-CMA). This framework achieves faster inference speeds while maintaining semantic fidelity. Specifically, DTP-CMA employs a Cross-Modal Saliency Estimator to identify query-relevant tokens, develops a Dynamic Pruning Scheduler to tailor compression ratios at each layer, and integrates a Context Preservation Mechanism to prevent the erasure of semantic variations. Evaluated on UCF-Crime, ActivityNet, and a newly defined Arabic video subset, DTP-CMA reduces FLOPs by 52% and Latency by 41% on NVIDIA Jetson AGX Orin while retaining 98.8% of baseline retrieval performance. Zero-shot cross-lingual performance improves by +4.5% R@1 with a lightweight Arabic adapter. Energy consumption reductions of 41.3% enable ~1.70× battery-life extension and near real-time Inference (7 FPS). We find that performance falls off at high power constraints (< 5W) and for ambiguous queries, but demonstrate how uncertainty-aware pruning can alleviate these limitations; our results position DTP-CMA as a strong candidate framework across diverse edge-deployable multimodal video analytics challenges
No comments yet — start the discussion below.