Gilson Shimizu, Rafael Izbicki, Fernando Rezende Zagatti, Filipe Loyola Lopes, André Gomes Regino, Rodrigo Bonacin, André C. P. L. F. de Carvalho · Journal of the Brazilian Computer Society 2026 · 2026
DOI: 10.5753/jbcs.2026.6073
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
One of the main limitations for the trust in the use of machine learning models is the understanding of how they produce their predictions. This limitation is related to an increasing demand for transparency in decision-making. Interpretability refers to how well a user can understand the model’s decision-making process without necessarily knowing its internal mechanisms. Several interpretability methods have emerged in the last years, such as LIME and Shapley values, which are valued for their flexibility, intuitive appeal, and strong theoretical foundations. While these methods have significantly contributed to better model transparency, they face challenges in several model deployment scenarios, such as handling irrelevant features or maintaining stability under small data perturbations. This article introduces two new agnostic interpretability methods, namely VarImp and SupClus, which overcome these issues by using local regressions fits with a weighted distance that takes into account variable importance. Whereas VarImp generates interpretations for each instance and can be applied to datasets with more complex relationships, SupClus interprets data clusters of instances with similar interpretations and can be applied to simpler datasets where data clusters can be found. In this paper, we compare these proposed methods with state-of-the-art methods and show that the proposed methods generate either equal or better interpretations, according to several proposed quantitative metrics (mean square error of coefficients, effect correlation, prediction correlation and ICE effect correlation), particularly in high-dimensional problems with irrelevant features and when the relationship between features and target is non-linear.
No comments yet — start the discussion below.