Jain Vasu Raj · World Journal of Advanced Engineering Technology and Sciences 2026 · 2026
DOI: 10.30574/wjaets.2026.20.3.0457
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Click-through rate (CTR) prediction underpins bidding, ranking, and budget allocation in programmatic advertising, where predicted probabilities, not merely rankings, drive economic decisions. We present an interpretable machine learning framework for CTR prediction designed to avoid two failure modes that inflate reported performance in much of the applied literature: target leakage through label-derived aggregate features, and evaluation on randomly shuffled temporal data. The framework constructs historical, past-only behavioral features (per-device, per-campaign, per-app, and per-site click-through aggregates with Bayesian smoothing) and evaluates all models under a chronological train/validation/test split. On a 5.39-million-impression sample of the public Avazu benchmark, we compare gradient-boosted ensembles, Random Forest, and an embedding-based deep neural network against logistic regression and three deep baselines (DeepFM, DCN-v2, FinalMLP) trained on identical data and splits. XGBoost achieves the best five-seed mean ROC-AUC (0.7374 ± 0.0006), and the embedding DNN reaches 0.7402 in its best run, significant over every alternative by DeLong tests (all p < 10⁻²⁴). A controlled leakage experiment shows that naive label-derived features combined with a random split inflate ROC-AUC to 0.8145, a gap of +7.7 points, quantifying a pitfall practitioners should actively guard against. Feature ablations show the leakage-safe historical CTR features provide the largest single gain (+1.08 ROC-AUC points), and we show that class-weighted probabilities require a closed-form prior correction before use in bidding. Code and the full experimental protocol are released for reproducibility.
No comments yet — start the discussion below.