Loading…
When Sparse Reward Meets Dense Distillation: Training Dynamics of On-Policy Distillation · Researchar