Read the original at HF Daily Papers
Researchers introduce SCAPO, a variant of GRPO that uses semifactual stability to adjust token-level credit assignment, improving reasoning accuracy without updating model weights.
Carried by: HF Daily Papers. First seen: .