Simon Pietro Romano, Giancarlo Sperlí, Mario Varlese, Andrea Vignali · Information Processing & Management 2026 · 2026
DOI: 10.1016/j.ipm.2026.105132
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
General-purpose Named Entity Recognition (NER) models struggle with domain-specific challenges, particularly in the legal sector, where complex syntax, specialized terminology, and privacy concerns pose significant obstacles. This paper presents a novel data-driven two-stage framework that leverages Large Language Models (LLMs) fine-tuned for legal applications to enhance NER for criminal process documents. Our contributions include: (i) the design of a novel framework for structured entity extraction in legal texts, (ii) the definition of an enhanced ontology adapted to criminal law, and (iii) a comprehensive evaluation of the framework on a real-world dataset. As a side contribution of the paper, we extend the renowned OntoNotes5 ontology by integrating new domain-specific entities tailored to criminal law. To evaluate our approach, we construct and manually annotate a real-world dataset comprising over 100 criminal judgments from four Italian courts, sourced from the Italian Antimafia and Anti-terrorism National Directorate (Direzione Nazionale Antimafia e Antiterrorismo — DNAA). Experimental results demonstrate the effectiveness of fine-tuned LLMs in accurately identifying legal entities. Among the evaluated models, LLaMA-3.2-1B demonstrates the lowest training time, while LLaMA-2.7B achieves the fastest inference time. In terms of predictive performance, LLaMA-2.7B and Vicuna-7B consistently yield the highest accuracy scores across evaluation metrics.
No comments yet — start the discussion below.