Simona Moldovanu, Sanda Florentina Mihalache, Cătălin Anghel, Dan Munteanu, Ioana Diana Moldovanu · AI 2026 · 2026
DOI: 10.3390/ai7100401
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Deep learning has achieved strong predictive performance in facial emotion recognition (FER), but interpreting the internal representations underlying model decisions remains challenging. In this study, we propose a cross-representation concept-based explainability framework that extends TCAV analysis beyond a single visual representation by evaluating the same human-interpretable facial concepts in both latent image features and explicit landmark-derived geometric features. AffectNet and CK+ facial image databases were analyzed using a frozen pretrained EmoNet backbone with a task-specific classification head, and landmark-derived geometric features were classified with FLAML AutoML. We analyzed seven human-interpretable concepts, such as mouth width, mouth opening, bilateral eye openness, bilateral eyebrow raising, and facial asymmetry, through visual TCAV and P25/P75-based tabular TCAV. The image-based model outperforms the geometric model, with reported accuracies of 0.7111 and 0.9329 on AffectNet and CK+, respectively, versus 0.6640 and 0.8265 on tabular data. TCAV found strong concept–emotion relationships, but they depended on the concept representation. Visual concepts tended to have stronger sensitivities, while geometric concepts had more moderate relationships that were directly interpretable. Visual and geometric representations provide complementary explanations for FER decisions. Concept stability, dataset dependence, and concept-class confounding are still important considerations for reliable TCAV interpretation.
No comments yet — start the discussion below.