Loading…
An RL View of OPD: Least Square Policy Distillation for Sample-Efficient LLM Reasoning · Researchar