Tuukka Hyytiäinen, Pasi Luukka · International Journal of Human-Computer Interaction 2026 · 2026
DOI: 10.1080/10447318.2026.2721889
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
In monthly financial cost variance analysis, AI is valuable when its outputs can be interpreted, challenged, and incorporated into accountable decisions, not only when it predicts accurately. Following a Design Science Research approach, we developed and evaluated a human-in-the-loop workflow combining TimeGPT-based anomaly screening, bounded local evidence retrieval, LLM-generated explanations, and blinded expert evaluation. Across three monthly runs, screening processed 445,633 data points and flagged 417 anomalies in 371 series. Eight finance professionals completed 144 paired case assessments and 288 explanation ratings. In the curated evaluation set, screening achieved precision 0.991, recall 0.941, and F1 0.965, indicating strong expert concordance rather than production-level performance. ChatGPT was rated higher for reliability and relevance, whereas Gemini was rated slightly higher for clarity. Qualitative feedback identified failures involving hallucinated or misattributed evidence, incomplete grounding, and contextual mismatches. We derive design implications for grounding, selective disclosure, and expert oversight in high-accountability financial work.
No comments yet — start the discussion below.