
Айгерим Айтим, Anel Auyezova, Б.К. Синчев · Engineering Technology & Applied Science Research 2026 · 2026
DOI: 10.48084/etasr.20152
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Legal decision support often depends on multiple conditions distributed across policy and regulatory documents. Retrieval-only systems can identify individually relevant passages, but may return redundant or collectively insufficient evidence. This study proposes a Legal NLP framework that combines MPNet-based semantic retrieval with controlled-cardinality evidence selection. Retrieved clauses are organized into compact subsets by jointly optimizing condition coverage, semantic relevance, redundancy, and subset size. The evidence-selection stage is formulated as a constrained set-cover and subset-selection problem. Experiments were conducted on LexIR-PolicyEvidence, a clause-level corpus comprising 850 policy and regulatory documents, approximately 14,000 clauses, and 1,800 query or case-description instances. The evaluation compared BM25, MPNet top-k retrieval, MPNet with greedy set cover, and the proposed method. MPNet improved Recall@5 from 0.68 to 0.79, MRR from 0.61 to 0.73, and nDCG@10 from 0.66 to 0.78. The proposed selection method achieved coverage of 0.86, sufficiency of 0.84, and redundancy of 0.24. The results demonstrate that retrieval quality alone does not ensure decision-ready legal support and that explicit evidence construction is required to produce compact and interpretable clause bundles.
No comments yet — start the discussion below.