Github repo to learn the OPD/OPSD and how they perform compared to GRPO, on a consumer grade GPU [P]
<!-- SC_OFF --><div class="md"><p>I am trying to learn concepts like On Policy Distillation (OPD), On Policy Self Distillation (OPSD) and how do they compare to RL algorithms like GRPO.</p> <p>There are a lot of papers on this, but because of limited compute I cannot try these papers out and learn them by implementing them myself.</p> <p>If someone here has worked with these algorithms and their implementation on SLMs (something that can fit a consumer grade GPU like Nvidia RTX 4090 or 5090), ca
















