Self-Align to Explain: Comparing Post-Training Methods for Counterfactual Generation
-
Updated
Aug 21, 2026 - Python
Self-Align to Explain: Comparing Post-Training Methods for Counterfactual Generation
Controlled GRPO/GDPO post-training for tool calling: Format SFT, KL-path audit, matched no-KL ablation, and held-out evaluation on Qwen2.5-1.5B.
To associate your repository with the gdpo topic, visit your repo's landing page and select "manage topics."