Dinghao Pan, Yuanyuan Sun, Jiru Li, Ling Luo, Jian Wang, Hongfei Lin · Information Processing & Management 2026 · 2026
DOI: 10.1016/j.ipm.2026.105203
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
In recent years, Large Language Models (LLMs) have demonstrated superior performance in information extraction tasks. Leveraging these models for Document-Level Relation Extraction (DocRE) can benefit from their powerful generative capabilities. However, we observe that LLMs still face challenges in DocRE tasks: Document Structure Parsing Error, Relation Definition Ambiguity, and Entity Boundary Recognition Error. To address these issues, we propose SDB-DRE, an LLM-based DocRE model that does not rely on pre-labeled entities during inference. To tackle the Document Structure Parsing Error, we introduce a novel Structure-Aware QA training approach, enabling LLMs to learn coreference relationships and entity types within the document. To address relation definition ambiguity and entity boundary recognition errors, we introduce relation definition learning and mention boundary learning in the second stage of relation extraction training. These components improve the internal document representation of the LLM, ensuring the output triples are consistent with the relation definitions and have more accurate entity boundaries. Experimental results demonstrate that SDB-DRE surpasses LLM-based methods employing multi-stage reasoning in terms of performance under a single-stage reasoning setup, while also achieving higher inference efficiency. The code will be made publicly available upon acceptance.
No comments yet — start the discussion below.