Bibek Ghimire · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22998787
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Large language models (LLMs) are increasingly being used in decision-support applications, but their reliability remains limited by hallucination, reasoning errors, prompt sensitivity, uncertainty, inconsistent tool use, distribution shift, and inadequate human oversight. This review examines major failure modes affecting LLM-based decision-support systems and synthesizes current approaches for evaluating and mitigating these risks. Particular attention is given to factual reliability, reasoning robustness, confidence and calibration, retrieval-augmented generation, tool use, explainability, human-in-the-loop oversight, and deployment monitoring. The paper proposes a practical failure-layer framework covering input, reasoning, external information and tool interaction, recommendation generation, and human decision stages. It also discusses research gaps relevant to increasingly autonomous and agentic AI systems. The review emphasizes that reliable deployment requires more than benchmark accuracy and should combine technical safeguards, transparent evaluation, uncertainty awareness, and appropriate human supervision.
No comments yet — start the discussion below.