Sheyasree Ghosh, Akanksha Srivastava, Soumen Bhowmik, Subhajyoti Maity · International Journal of Innovative Science and Research Technology (IJISRT) 2026 · 2026
DOI: 10.38124/ijisrt/26sep333
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Large Language Models (LLMs) find increasing usage in applications that involve reasoning over large spans of text such as document comprehension, question-answering across multiple documents, and large scale code understanding. However, even current state-of-the-art designs struggle to efficiently capture long-range dependencies. Despite Transformerbased designs being highly effective in handling tasks with short contexts, they suffer from scalability issues as well as degradation of contextual effects caused by the high complexity of the underlying attention computations. In this paper, we conduct a comparative study of Transformer-based and State-Space Model designs when it comes to long-context reasoning. A selection of representative models from both classes is systematically studied through carefully conducted experiments involving various types of long-sequence generation and comprehension tasks. In order to ensure comparability, a new metric called Context Retention Score (CRS) is proposed which quantifies the effect of distant tokens on model decisions. As evidenced in our results, Transformer-based models have a pronounced tendency to forget distant information while StateSpace Model-based designs retain information more stably and yield better results for tasks that involve long context.
No comments yet — start the discussion below.