Sachin Sudesh Desh, Prakash MG Achary, Saurav Nayak · bioRxiv (Cold Spring Harbor Laboratory) 2026 · 2026
DOI: 10.64898/2026.08.11.741471
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Background Synthetic data generation is increasingly proposed as a strategy to support privacy-preserving data sharing, augmentation of small or restricted biomedical datasets, and benchmarking of artificial intelligence tools in laboratory medicine. However, model selection remains difficult because synthetic data generators differ in fidelity, privacy risk, stability, and generalisability. Existing evaluations have rarely examined performance jointly across conditioning signal strength, synthetic output scale, and train–test generalisation. Methods We developed the Synthetic Fidelity–Stability Framework (SFSF), a systematic benchmark of 17 synthetic tabular data generation models using NHANES as a complex biomedical reference dataset. Models included statistical, copula-based, resampling, variational autoencoder, generative adversarial network, and diffusion-based approaches. Synthetic datasets were generated across 11 seed sizes, from 0 to 500 real conditioning observations, and six output scales, from 50 to 5,000 rows, yielding 1,122 synthetic datasets per run. Each dataset was evaluated against the full original dataset, the training subset, and a held-out test subset across five tiers: univariate distributional fidelity, moment agreement, tail behaviour, multivariate dependency structure, and privacy/memorisation risk. Composite rankings and seed-versus-output stability profiles were derived. Results Univariate fidelity was broadly recovered across model classes and was the least discriminating tier. Resampling-based methods ranked highest overall but showed the greatest privacy risk, reflecting proximity to real observations rather than true generative novelty. VAE-family models reproduced moment statistics relatively well but consistently failed on tail and shape fidelity. GAN-family models showed substantial moment-level instability, while VineCopula demonstrated severe multivariate dependency failure. Diffusion-based models, particularly ForestDiffusion, provided the most favourable privacy–utility balance, combining competitive fidelity with the lowest privacy risk and the smallest train–test gap. Conclusions No single synthetic data generator dominated across fidelity, stability, and privacy dimensions. The SFSF framework provides a practical, multi-criterion approach for selecting synthetic tabular data generators according to intended clinical laboratory use, balancing statistical realism, dependency preservation, privacy risk, and robustness to seed and output scale.
No comments yet — start the discussion below.