Venkata Sai Aditya Kondru, Venkata Rao Kasukurthi · International journal of intelligent engineering and systems 2026 · 2026
DOI: 10.22266/ijies2026.0930.53
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
The rapid growth of unstructured textual data in domains such as healthcare, education, and enterprise systems has increased the risk of inadvertent exposure of Personally Identifiable Information (PII).Effective deidentification is essential to enable secure data sharing while preserving privacy and regulatory compliance.Existing approaches, including rule-based and deep learning-based Named Entity Recognition (NER) methods, primarily focus on detection and often lack contextual validation and robust replacement mechanisms, leading to information loss or residual privacy risks.This work proposes DLGNet, a pipeline-based integrated framework PII de-identification framework that integrates DeBERTa-based token-level detection, LLM-based semantic verification, and type-aware replacement.A BIO-tagged DeBERTa model identifies candidate PII spans, which are refined using a TinyLlamabased semantic verification module for contextual validation.Detected entities are then replaced using a hybrid generative and rule-based strategy to ensure irreversible transformation while preserving textual coherence.Experimental results demonstrate strong recall-oriented performance, achieving precision of 90.3%, recall of 94.7%, F1-score of 92.4%, and Fβ-score (β = 5) of 94.52%.Utility evaluation shows a length ratio of 0.96 and token overlap of 0.73, indicating effective preservation of textual structure and semantics.The proposed framework provides a unified and practical solution for privacy-preserving text processing by jointly improving detection accuracy, contextual reliability, and data utility.
No comments yet — start the discussion below.