Chukwuemeka Christiantus Ndubuisi · INTERNATIONAL JOURNAL OF COMPUTER SCIENCE AND MATHEMATICAL THEORY E-ISSN 2026 · 2026
DOI: 10.56201/ijcsmt.vol.12.no4.2026.pg239.255
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Large Language Models (LLMs) have enabled the emergence of autonomous AI agents capable of reasoning, planning, tool use, and iterative decision-making. Despite rapid development, the field remains architecturally fragmented, with limited conceptual clarity regarding memory integration, planning mechanisms, and operational reliability. This study presents a systematic review and critical synthesis of LLM-based autonomous agents, focusing on architectural paradigms, memory models, planning strategies, and real-world deployment constraints. Using a structured review approach, this study examines existing LLM based agent systems across key design components to uncover common patterns, differences in implementation, and recurring structural weaknesses. The review reveals persistent and structurally significant challenges across all four dimensions: long-horizon reasoning stability degrades as task length increases; memory consistency is undermined by retrieval noise, embedding drift, and summarisation errors; tool alignment failures propagate errors across modular pipelines; and evaluation standardisation remains insufficient to support reliable cross-paper comparison. A consistent cross-paradigm finding emerges: autonomy and reliability trade off systematically as agent complexity increases, with current systems achieving capability gains through heuristic design rather than principled theoretical foundations. Based on this synthesis, the review proposes a consolidated analytical framework that maps common structural elements and trade-offs across reviewed systems, and outlines a research agenda directed toward formalised agent architectures, memory consistency guarantees, verified planning algorithms, standardised reliability metrics, and benchmark frameworks adequate for long-horizon, real-world evaluation conditions.
No comments yet — start the discussion below.