Qunyang Zuo, Zifan Wang, Jiayi Li, Yifan Guan, Jianjun Chen, Chuanhong Yang, Zhihui Fu · Array 2026 · 2026
DOI: 10.1016/j.array.2026.101139
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Complete two-section lumbar MRI report generation from paired radiologist-selected sagittal T1w/T2w slices is a constrained radiology-assistance task in which the image inputs must support precise level-wise localization and pathology description. Medical visual language models (VLMs) are increasingly used for radiology assistance, yet their performance remains uneven in anatomically specific and data-scarce settings. Public datasets rarely provide paired sagittal inputs and free-text expert reports for this setting. This paper presents AVP2CD , a hierarchical post-training framework that adapts a compact general VLM from A natomical V isual P erception to C linical D iagnosis for complete two-section lumbar MRI report generation from paired sagittal images. AVP2CD decomposes report learning into segmentation-derived anatomy-grounding VQA, diagnostic VQA from structured lumbar findings, and report-generation refinement using private paired sagittal T1- and T2-weighted images. Reinforcement learning is guided by a rule-based clinical-semantic-format reward combining entity-level matching, BERTScore similarity, and report-structure compliance, with DAPO-style asymmetric clipping and truncated-completion masking for long-form refinement. Experiments use 219 private lumbar MRI cases and Spider annotations, with 175 private report pairs used for final report training. Starting from Qwen2.5-VL-3B-Instruct, AVP2CD increases the secondary average report-generation score from 0.1443 to 0.3463 on the private held-out test set and reports higher automatic metric values than the evaluated SFT-only, GRPO, API-based, and medical VLM baselines under the same selected-slice evaluation protocol. Entity-level analysis further shows fewer omissions and unsupported predicted entities than SFT-only training, although residual localization errors remain. These findings support structured anatomy-aware post-training as a practical research strategy for low-resource lumbar MRI report generation.
No comments yet — start the discussion below.