Evan Solecki, Peifen Zhu · Computational Materials Science 2026 · 2026
DOI: 10.1016/j.commatsci.2026.115106
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Band-gap datasets often contain many zero-gap entries and closely related compositions, both of which affect reported model errors. We quantify these effects across six public datasets containing 1306–106,113 materials with experimental, computed, or surrogate gaps. Aggregate tolerance accuracy can conceal error on non-zero-gap materials: on the Castelli set it reaches 0.93 while the non-zero-material mean absolute error (MAE) is about 1 eV. Reporting MAE and root-mean-square error (RMSE) on the true non-zero subset resolves this ambiguity. A classify-then-regress model with enriched descriptors lowers non-zero MAE by 23–42% relative to the original single-stage baselines. With features held fixed, the evaluated pipeline bundle reduces MAE by 4.8–14.7% in nested comparisons; the nominal paired tests support the differences for the experimental and Castelli sets, but not Wolverton. A separate fixed-configuration Materials Project analysis gives a descriptive 7.8% reduction. Holding out whole element-set chemistries raises MAE by 55% on the experimental set but barely changes it on several datasets of discrete stoichiometric compounds. In an exploratory six-dataset comparison, the two datasets with non-zero fractional-formula share have the largest grouped/random MAE ratios, but the rank association is uncertain (Spearman ρ = 0.85, exact p = 0.067); a within-Tol thinning analysis is also confounded by sample size. Across five Materials Project splits, graph networks using crystal structure reduce non-zero MAE by 31–42% relative to composition-only XGBoost. We therefore recommend reporting non-zero metrics, labeling each evaluation protocol, and using grouped splits when the intended task is extrapolation beyond the chemistries represented in training.
No comments yet — start the discussion below.