
Seongjung Na, Seunghyun Gwak, Suhwan Kim · Scientific Reports 2026 · 2026
DOI: 10.1038/s41598-026-71346-z
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Reusable text representations are routinely cached and compressed, yet downstream utility and source-content exposure are rarely measured together. We evaluate token sequences and pooled vectors across six languages, five encoders, classification and dense retrieval, matched attacker budgets, and transformation-aware attacks. In a five-seed mBERT experiment on parallel MASSIVE data, mean token F1 was $$0.9370\pm 0.0030$$ from full token states and $$0.7211\pm 0.0333$$ from one mean-pooled vector; the full-sequence advantage occurred in English, Korean, Arabic, Japanese, Swahili, and Telugu. On six-language Belebele retrieval, clean macro nDCG@10 was 0.9027 for multilingual-E5-base and 0.9220 for BGE-M3; int8 retained 0.9025/0.9217, whereas PCA-64 fell to 0.7984/0.8043. Five-seed, content-disjoint attack curves show that rankings depend on paired-data and compute budgets. Fully transformation-aware strong-mT5 and Vec2Text attacks find little protection from int8 or masking and only partial reductions from PCA-256 or Gaussian noise. A condition-blinded audit by the three study authors supports higher factual recoverability from full Korean token sequences and higher sensitive-content recovery from full English token sequences, while other contrasts remain uncertain. Actual payload accounting and Pareto analysis further show that storage reduction, utility, and leakage are not interchangeable. Granularity remains an important determinant under matched conditions, but exposure depends jointly on language, task, encoder, attacker access and adaptation, and representation budget.
No comments yet — start the discussion below.