Skip to content

Clarification on Qwen3-8B (GRPO) teacher and training steps #6

Description

@SeanZh30

Hi authors,

Thanks for the interesting TIP paper. That's really inspiring! I have a few clarification questions about the experimental setup.

In Section 7.1, the paper lists the math reasoning model pair as Qwen3-8B (GRPO) → Qwen3-4B. Could you clarify whether Qwen3-8B (GRPO) refers to a publicly available GRPO-tuned checkpoint, or whether the teacher was further trained with GRPO by the authors? If it is public, could you share the exact checkpoint name or release source? If it was trained by the authors, could you share whether you plan to release the checkpoint, and if not, could you provide the training data and main GRPO hyperparameters?

Also, for the OPD/TIP training on the mathematical reasoning tasks, could you share the total number of training steps or epochs? I found the training prompts, learning rate, rollout number, max response length, and other hyperparameters, but could not locate the total training length in the paper or appendix.

Thanks again for the helpful paper!

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions