Linfang Ding · Zenodo (CERN European Organization for Nuclear Research) 2026 · 2026
DOI: 10.5281/zenodo.22912418
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Background. Supervised machine learning (ML) studies on tabular data frequently proceed without any assessment of whether the planned task is feasible with the available or obtainable sample. Consequences include research waste, publication bias toward silently failed projects, and misapplied one-size-fits-all sample size heuristics imported from clinical prediction modeling.Objectives. To specify a reusable feasibility screening framework that (i) integrates a technical validation domain (discrimination, AUC) with an interpretability domain (SHAP-based evidence of pattern learning), (ii) replaces universal decision thresholds with explicit domain-calibration principles, and (iii) defines a mandatory validation protocol including positive controls, permutation-based negative controls, synthetic signal injection, and leak-free simulation.Methods. The framework was specified axiomatically: signal strength is defined as the strength of the feature-outcome relationship; feasibility is a two-dimensional construct requiring evidence in both domains; and thresholds are local conventions that must be stated and justified before primary analysis. A simulation validation protocol was designed in which synthetic binary classification tasks span a grid of signal strengths (class separation) and label noise levels, and framework classifications are compared against ground-truth conditions.Results. The specification defines a GREEN/YELLOW/RED classification core with documented domain-specific refinements (YELLOW-A/YELLOW-B subtypes for marginal discrimination; GREEN* for screening-instrument contexts where labels share a construction pathway with features). The simulation protocol yields ground-truth-verifiable classification accuracy as a ceiling estimate; companion simulation code is released with this note. Real-world application of the framework across six domains (sport monitoring, respiratory medicine, educational data mining, autism screening, developmental screening, and cross-domain computational evaluation) is reported separately in companion manuscripts.Conclusions. Feasibility screening should be a standard pre-commitment step in ML study design. This note provides a public, citable specification intended to decouple the framework's design decisions from any single application, and to make threshold choices auditable rather than implicit.
No comments yet — start the discussion below.