Usama Mehboob · Preprints.org 2026 · 2026
DOI: 10.20944/preprints202609.2305.v1
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Anonymizing sensitive healthcare data while preserving data utility is always a tradeoff between suppression and generalization. In this experiment, we employ a genetic algorithm to search the anonymization policy space using synthetic healthcare-style data. Each policy candidate specifies different levels of generalization for quasi-identifiers such as age, a five-digit numeric location code, and sex, while any equivalence class that does not satisfy the constraints of k-anonymity or distinct ℓ-diversity is suppressed. The GA experiment is carried out under constraints of k = 5, ℓ = 2, and a 30% suppression limit. There are 32 possible anonymization policies, and the search space is intentionally kept small so that the results of the GA-based search can be compared directly with exhaustive search to identify its accuracy, gaps, and limitations. Across 60 seeded runs using three different table sizes, the GA search recovered the exact Pareto set in 59 runs while evaluating a median of 28–29 policies. These results show that the GA was capable of reliably recovering the Pareto-optimal solutions in this small synthetic setting, but GA still ended up evaluating most of the policies. The efficiency advantage of GA has yet to be evaluated in the future on large real world healthcare data with substantially larger policy spaces to establish if GA indeed offers more clinical utility with less compute compared to other traditional methods.
No comments yet — start the discussion below.