John Goodman · · 2026
DOI: 10.31224/8438
Counts differ because each database indexes a different set of publications. We treat OpenAlex as the canonical count; Google Scholar is not shown (no API, and crawling it violates its ToS).
Reinforcement-learning (RL) schedulers are now flying on real spacecraft: NASA’s Carruthers Geocorona Observatory (launched 24 September 2025, per NASA’s own mission status page — not itself a LEO mission; it orbits near the Sun–Earth L1 point) uses deep RL as its default operational scheduler for long-horizon operations under power, thermal, and instrument constraints. We cite it only as evidence that RL scheduling is flight-proven in general, not as a LEO example. For low-Earth-orbit (LEO) platforms specifically, the battery is frequently the lifetime-limiting consumable, and an RL scheduler’s exploration-exploitation setting — how often it tries a novel action instead of its current best estimate — is usually tuned purely as a learning-speed knob. We show that the exploration coefficient is also an independent driver of physical wear, separate from and in addition to its effect on how quickly mission value is delivered. Holding delivered mission value fixed by construction (every scheduler setting runs until it reaches the same cumulative reward, not a fixed wall-clock duration), full exploration costs 1.47–1.71× the wear or aging of a purely exploitative scheduler, depending on the physical model - battery depth-of-discharge cycling (two models: a power-law damage law and an independently-published log-linear cycle-life law) or, in a third model, mechanical actuator wear on a documented robot platform (Archard’s law). A zero-initialised greedy baseline is unfairly handicapped: it locks onto an arbitrary first action and, without a brief warm start, reverses the effect in 13–19 of 30 independent task draws (median ratios 0.73–1.04×, none significant). Every arm here therefore uses a warm start, applied identically at every exploration rate, and results are reported over 30 independent environments per model. With it, the battery result is fully robust (0/30 reversals in both models, Wilcoxon p = 9.3×10−10 each); the actuator-wear generalization is real but less uniform (2/30 reversals, ratio range 0.83–3.30×). We report both the effect and its exception rate, rather than the effect alone. What transfers across all three models is the qualitative claim: exploration costs physical wear independently of task throughput, through both more total actions and (usually, to a lesser degree) higher wear per action. We frame this as a design implication for anyone tuning exploration schedules on hardware with a finite-cycle-life component, not as a validated property of any specific deployed system - and as a demonstration that a single-seed simulation result, however clean it looks, is not evidence until it survives being replayed on environments it was not tuned on.
No comments yet — start the discussion below.