Hi authors,
Thanks for the interesting TIP paper. That's really inspiring! I have a few clarification questions about the experimental setup.
In Section 7.1, the paper lists the math reasoning model pair as Qwen3-8B (GRPO) → Qwen3-4B. Could you clarify whether Qwen3-8B (GRPO) refers to a publicly available GRPO-tuned checkpoint, or whether the teacher was further trained with GRPO by the authors? If it is public, could you share the exact checkpoint name or release source? If it was trained by the authors, could you share whether you plan to release the checkpoint, and if not, could you provide the training data and main GRPO hyperparameters?
Also, for the OPD/TIP training on the mathematical reasoning tasks, could you share the total number of training steps or epochs? I found the training prompts, learning rate, rollout number, max response length, and other hyperparameters, but could not locate the total training length in the paper or appendix.
Thanks again for the helpful paper!
Hi authors,
Thanks for the interesting TIP paper. That's really inspiring! I have a few clarification questions about the experimental setup.
In Section 7.1, the paper lists the math reasoning model pair as Qwen3-8B (GRPO) → Qwen3-4B. Could you clarify whether Qwen3-8B (GRPO) refers to a publicly available GRPO-tuned checkpoint, or whether the teacher was further trained with GRPO by the authors? If it is public, could you share the exact checkpoint name or release source? If it was trained by the authors, could you share whether you plan to release the checkpoint, and if not, could you provide the training data and main GRPO hyperparameters?
Also, for the OPD/TIP training on the mathematical reasoning tasks, could you share the total number of training steps or epochs? I found the training prompts, learning rate, rollout number, max response length, and other hyperparameters, but could not locate the total training length in the paper or appendix.
Thanks again for the helpful paper!