Dr. OPD: Learning What to Follow for Optimal On-Policy Distillation of Large Language Models
Preprint · 2026
OPD provides dense token-level supervision. We formulate optimal OPD as bilevel optimization, learning how to weight teacher signals to improve the student’s performance.