Read the original at HF Daily Papers
Researchers find that restricting reverse KL to a student-selected top-16 vocabulary subset achieves accuracy comparable to full shared-vocabulary on-policy distillation, outperforming evaluated cross-tokenizer baselines.
Carried by: HF Daily Papers. First seen: .