Suvendu Barai · Cologne Open Science (TH Köln) 2026 · 2026
DOI: 10.57684/cos-1497
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Tree-based models remain highly competitive on tabular learning tasks despite the rapid progress of deep learning. In this work, we present a systematic comparison of tree-based, neural, and recent pretrained tabular models across a diverse benchmark of classification and regression datasets. Our study goes beyond clean-data evaluation by combining three perspectives: predictive performance after hyperparameter optimization, robustness under controlled data challenges, and stability under covariate shift. We evaluate classical tree ensembles (Random Forest, XGBoost, LightGBM, and CatBoost), neural baselines (MLP and TabNet), and recent tabular foundation-style models (TabPFN and APT) on six datasets, including two regression variants of the Diamonds dataset with different correlation filtering thresholds. To assess robustness, we introduce missingness and synthetic high-cardinality categorical features, and we analyze clean versus shifted test distributions using both predictive degradation and SHAP-based stability measures. Across the selected benchmark, boosted tree methods show the most consistent empirical pattern: they usually remain close to the best clean-data performance while also preserving stronger robustness and feature-attribution stability than the neural baselines. Neural models remain competitive in selected cases, but in this benchmark they are generally more sensitive to featurespace changes, hyperparameter choice, and attribution instability. Pretrained tabular models reduce part of this gap, yet they do not consistently surpass boosted tree methods. These findings suggest that the continued strength of tree-based models in this study is linked not only to predictive accuracy, but also to robustness and stability under controlled tabular stress conditions. The results should therefore be interpreted as evidence from a controlled benchmark rather than as a universal claim.
No comments yet — start the discussion below.