P. T. Manasa Visakai, Sadhana Venkatraghavan · Iconic Research and Engineering Journals 2026 · 2026
DOI: 10.64388/irev10i2-1722663
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Companies of every size now collect substantial volumes of customer data, but converting that data into segments, forecasts, and explainable decisions remains difficult, particularly for student researchers and small organisations without access to large, professionally curated datasets. This study builds and transparently evaluates a complete customer-analytics pipeline on a dataset of 238 individual customers, covering data cleaning, feature engineering, exploratory analysis, K-Means segmentation, engagement-tier prediction, SHAP-based explainability, and a rule-based personalised marketing framework. K-Means clustering (k=4, validated using the Elbow and Silhouette methods) produced four interpretable segments: High-Value Loyal, Established, Growing/Potential, and New/Low-Engagement. Four classifiers — Logistic Regression, Decision Tree, Random Forest, and Gradient Boosting — were trained on demographic predictors alone (age, annual income, region) to forecast engagement tier, achieving test accuracies between 96.7% and 100%. Diagnostic analysis traces this near-perfect performance to severe multicollinearity in the dataset (Variance Inflation Factors of 37–197; a single principal component explaining 98.8% of variance) rather than to genuine predictive strength, a conclusion corroborated by SHAP, which identifies annual income as the dominant predictor, age as a weaker secondary predictor, and region as largely irrelevant. Rather than overstating predictive novelty, the study advances a narrower and more defensible claim: that a segmentation–prediction–explainability framework already validated on large industrial datasets can be meaningfully applied to, and honestly evaluated on, the small, highly collinear data typical of student and SME research. The pipeline is operationalised as an interactive dashboard application.
No comments yet — start the discussion below.