Satya Sri Rajiteswari Nimmagadda, Ethan Young, Niladri Sengupta, Ananya Jana, Aniruddha Maiti · Big Data and Cognitive Computing 2026 · 2026
DOI: 10.3390/bdcc10090288
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
This paper studies whether structured representations can retain the meaning contained in scientific sentences. A lightweight LLM is fine-tuned to generate JSON structures from scientific sentences. The objective was to capture information contained in the sentence using a hierarchical structure. The structured representations are then used to reconstruct sentences with a generative model to test whether such representations are useful to retain the meaning of the sentence. The evaluation compares the reconstructed sentences with the originals using semantic and lexical similarity. Our results show that hierarchical structured formats preserve semantic information with an average cosine similarity of 0.87.
No comments yet — start the discussion below.