Spyridon Manolidis · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.23171338
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
We consider the current landscape in AI red-teaming and assess how fine-tuning plays an increasingly critical role.We provide an overview of the etymological origins of the English lexicon, and outline how existing literature and data have motivated us to create an encoding scheme based on etymology.We then formally define EtymoStega: a novel etymological steganographic scheme, and a simpler scheme inspired by EtymoStega where the etymological class of the first verb in a response carries a binary bit, which we then train two 7B-8B open-weight models on to empirically demonstrate the feasibility of such schemes in the research landscape of today.The fine-tuned models adopt our modified steganographic scheme with over 96% accuracy on one model and 90% accuracy on another model, after a conservative adjustment for evaluation database contamination.We then evaluate the trained models, highlighting both the advantages and the disadvantages of our results.Lastly, we provide a comprehensive review of the limitations, assumptions, and uncertainties of our research.The source code and research artifacts are also provided.
No comments yet — start the discussion below.