
Anne Stockem Novo · Scientific Reports 2026 · 2026
DOI: 10.1038/s41598-026-74237-5
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Traditional system testing relies on predefined scenarios and handcrafted rules, whereby the system output is tested against requirements. However, this approach is not scalable and therefore not feasible when testing complex AI systems. Human understandable explanations of the model decisions are required and must be generated by automated pipelines. This work explores the applicability of a scalable, adaptable concept bottleneck model which is trained with the system under test and provides explainable concept representations alongside model decisions. Moreover, the generated explanations can be fed back to the model as auxiliary information to improve model performance. Both aspects are studied using pedestrian intention detection as a use case: first, it is evaluated how much the incorporation of concepts into the prediction process enhances performance, despite the challenges posed by the subjective and ambiguous nature of behavioral concepts. Second, interpretability is quantified with the technique of Integrated Gradients. While concept classification accuracy is mediocre, CBMs outperform baseline models that do not explicitly use concepts. Moreover, CBMs focus attention more within pedestrian bounding boxes, unlike baselines that attend more to irrelevant background. These findings support the integration of explanations into decision-making systems while highlighting the complexity of applying subjective concepts.
No comments yet — start the discussion below.