Pramod Prakash · International Journal of Intelligent Systems and Data Science 2026 · 2026
DOI: 10.67231/822qz209
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Machine learning systems use sensitive personal data to train and infer models, but developers are largely unaware of privacy attack vectors that evade conventional data security controls. This paper surveys the privacy threat landscape inherent to ML workflows, including membership inference attacks, model inversion attacks, reconstruction attacks, and re-identification attacks, which can occur even when data is transmitted over encrypted channels or is explicitly de-identified. It then catalogues cryptographic and perturbation-based defense mechanisms, including homomorphic encryption protocols, garbled circuits, secret-sharing architectures, secure processors, and differential privacy frameworks, that enable privacy-preserving machine learning (PPML) without sacrificing model utility. The extensive study of commercial systems (ShareMind, RAPPOR, local DP in Chrome, etc.) and academic systems shows what is being deployed in practice and what limitations remain. Finally, the paper provides an overview of the important barriers to adoption, namely algorithmic rigidity, the trade-off between computational scalability and privacy, non-collusion assumptions, gaps in policy enforcement, and human-data-interaction principles (legibility, agency, negotiability) for realizing the right to be forgotten and granting data owners control over the processes of inference.
No comments yet — start the discussion below.