
Hoang-Tu Vo, Vinh Dinh Nguyen, Huu-Hoa Nguyen · Engineering Research Express 2026 · 2026
DOI: 10.1088/2631-8695/ae9b62
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Visual inspection in degraded field imagery remains difficult when disease cues are subtle, obscured, or distorted. In aquatic disease screening, these challenges are amplified by turbidity, unstable illumination, blur, and background interference. Symptom descriptions add semantic context, but may introduce text-based shortcuts when label-like information is uncontrolled. Reliable assessment therefore requires a unified protocol covering accuracy, robustness, textual leakage, and reproducibility. This study proposes JITCA, a Joint Image-Text Contrastive Attention framework for reliability-oriented aquatic disease screening. JITCA uses cross-modal attention to integrate leakage-aware symptom descriptions into image-region representations, allowing local evidence to be interpreted with symptom-level context rather than relying on visual appearance alone. This fusion pathway is trained with paired image-text alignment and training-only visual consistency learning, improving cross-modal agreement and visual stability. JITCA restricts the auxiliary image view to training, preserving a single-view prediction pathway. The symptom-text protocol confines descriptions to observable manifestations, while shortcut risk is addressed through filtering, expert audit, and controlled text tests. We evaluate the framework on a curated aquatic field-image dataset against strong baselines under matched conditions. JITCA achieves 98.51% accuracy and 98.08% macro F1 Score, outperforming the primary image-only baseline by 4.06 and 4.93 percentage points. Ablations show that symptom-conditioned fusion, image-text alignment, and visual consistency learning each improve performance, while sensitivity tests indicate stability under moderate loss reweighting. Shortcut checks show that the standardized descriptions retain class-related information. However, performance decreases when image-text pairing is broken or symptom content is replaced, suggesting that correct correspondence contributes beyond the text-only signal. Under severe synthetic random occlusion, JITCA retains 88.43% macro F1 Score, compared with 76.49% for the primary image-only baseline. These findings support improved in-domain performance and robustness under controlled perturbations. Independent cross-site and cross-device validation is required before broader generalization or deployment claims can be made.
No comments yet — start the discussion below.