
Sunil Kumar Vytla · Engineering Technology & Applied Science Research 2026 · 2026
DOI: 10.48084/etasr.21503
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Production machine learning systems can remain infrastructure-healthy while degrading through data drift, feature-quality loss, calibration shift, dependency change, or prediction instability. This paper presents OBSERVE-ML, an auditable decision layer that converts heterogeneous feature, model, trace, and incident evidence into bounded trust components, noncompensatory gates, denominator-safe metrics, and retained decision logs. Its novelty is an admissibility contract: a production incident cannot receive a confirmed root-cause label unless telemetry completeness, feature-service lineage, posterior confidence, denominator membership, and auditability requirements are satisfied. A sanitized 90-day enterprise case study covered 1.8 TB of authorized telemetry, 14 production workloads, and 118 confirmed machine-learning-relevant incidents. In the 30-day blind window, the recall was 94.9% (112/118; 95% CI 89.3-97.6) and attribution precision was 87.5% (98/112; 80.1-92.4). Median triage fell from 5 to 2 steps, abrupt-failure alerting from 17.6 to 3.5 min, gradual-drift detection from 51.3 to 2.1 h, and the false-positive group rate from 31.0% to 10.8% (65.0% relative reduction). These results are bounded single-site deployment evidence, not universal causal proof.
No comments yet — start the discussion below.