Xin Li, H. Liang, Xu Wang, Fan Yang, Feiyang Xu · Journal on Image and Video Processing 2026 · 2026
DOI: 10.1186/s13640-026-00704-8
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
The development of large language models for the chemical domain relies heavily on high-quality structured data. However, key experimental information in chemical literature is often scattered across PDFs in multimodal forms, such as reaction schemes, experimental tables, figure captions, and footnotes. This makes structured information extraction much more challenging. To address problems such as the difficulty of finely segmenting reaction schemes in complex layouts, the challenge of unified semantic parsing across multiple modalities, and the lack of a workflow for local deployment and continuous optimization, this paper proposes CREST, a multimodal framework for reaction information extraction from chemical literature. CREST combines VisualHeist and MinerU to perform fine-grained segmentation of PDF content. It also extracts titles, descriptive text, and footnotes as auxiliary text prompts. Based on this, we build an end-to-end information extraction pipeline. The pipeline uses the Qwen3 multimodal model together with LoRA fine-tuning to jointly parse reaction schemes, experimental tables, and related text. We also design an evaluation mechanism that integrates automatic evaluation with expert verification. Experimental results show that CREST can effectively improve multimodal reaction information extraction from chemical literature on a self-built dataset. It also achieves performance comparable to some closed-source models, even with a relatively small model size. This demonstrates its practical value and application potential. Finally, we have launched an online information extraction system based on CREST, available at: https://ai4s.iflytek.com/1024/extract .
No comments yet — start the discussion below.