
Muhammad Arifuddin Aljufri, Diana Purwitasari, Hilmil Pradana · Engineering Technology & Applied Science Research 2026 · 2026
DOI: 10.48084/etasr.20893
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Automating cargo Harmonized System (HS) Code classification via non-intrusive imaging is a critical objective for modern port logistics and customs enforcement, particularly as manual physical inspections account for less than 2% of global trade. While multimodal contrastive frameworks have successfully advanced Zero-Shot Learning (ZSL) in computer vision, their application to radiographic cargo imaging remains limited. This study introduces a Hierarchical Contrastive Learning framework, which is a multi-task optimization approach that structures a joint vision-language latent space according to the explicit three-level hierarchical topology of the HS nomenclature: Chapter ( ), Heading ( ), and Subheading ( ). Unlike conventional multimodal pipelines that treat textual descriptors as flat strings, the proposed formulation inherently preserves the taxonomic relationships of the nomenclature. To evaluate this approach, this study introduced a large-scale multimodal dataset pairing cargo gamma-ray scans with HS codes from real-world export transactions. The proposed framework was benchmarked under low-resource (10%) and full-scale (100%) training regimes across 5,675 unique HS code pseudo-class descriptors. Empirical results demonstrate that the proposed hierarchical formulation consistently outperforms the flat-text baseline. Specifically, under full-scale training, the subheading-dominant set (0-1-9) secures a Top-10 accuracy of 0.852 and an mAcc of 0.724 on seen categories, while the balanced set (1-4-5) achieves Top-10 accuracy of 0.249 on unseen categories.
No comments yet — start the discussion below.