Yashar Mousavi, Rashin Mousavi, Arash Mousavi, Mahsa Tavasoli, İbrahim Beklan Küçükdemiral, Afef Fekih, Ümit Cali · Neurocomputing 2026 · 2026
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Temporal credit assignment becomes intractable when reward signals are displaced hundreds of steps from causal actions, limiting learning in robotic manipulation, autonomous navigation, and long-horizon planning. Conventional exponentially-decaying eligibility traces allocate negligible credit beyond moderate horizons, as their geometric forgetting renders causal attribution effectively impossible at extended lags. This paper introduces FracTrace, a Fractional Trace Actor-Critic framework that replaces exponential decay with power-law memory kernels derived from fractional calculus. Variable-order Mittag-Leffler functions are incorporated at two architectural levels: hierarchical traces maintain multi-timescale state representations across fast, medium, and slow dynamics, while fractional eligibility traces induce algebraic decay that sustains meaningful credit assignment on tasks with reward delays of 500–800 steps. A key design principle retains the standard distributional Bellman operator for the critic, preserving -contraction in Wasserstein metric, while fractionalization is applied exclusively to policy-side advantages. For an idealized policy-gradient variant of the algorithm, almost-sure convergence of the critic to the true value function and of the policy to a stationary point neighborhood of radius , a truncation-dependent bias term, are established under two-timescale stochastic approximation, where vanishes as the truncation horizon ; the practical algorithm uses the PPO clipped surrogate and is evaluated empirically. The critic achieves sample complexity , matching classical actor-critic rates while capturing substantially longer dependencies. Empirical evaluation across six long-horizon tasks demonstrates substantial improvements over PPO, PPO-LSTM, R2D2, and GTrXL, with FracTrace uniquely solving extreme-horizon tasks (800 steps) at per-step complexity. The kernel-swap ablation, which isolates kernel choice within an otherwise identical architecture, supports the conclusion that these gains stem from power-law memory structure; the PPO-LSTM comparison provides converging evidence at the level of a full baseline architecture rather than a single-factor control.
No comments yet — start the discussion below.