Live page · Day archive

Semifactual Credit-Augmented Policy Optimization

Read the original at HF Daily Papers

Summary

Researchers introduce SCAPO, a variant of GRPO that uses semifactual stability to adjust token-level credit assignment, improving reasoning accuracy without updating model weights.

Carried by: HF Daily Papers. First seen: .