Live page · Day archive

Rethinking Cross-Tokenizer On-Policy Distillation: From Alignment Coverage to Supervision Reliability

Read the original at HF Daily Papers

Summary

Researchers find that restricting reverse KL to a student-selected top-16 vocabulary subset achieves accuracy comparable to full shared-vocabulary on-policy distillation, outperforming evaluated cross-tokenizer baselines.

Carried by: HF Daily Papers. First seen: .