Dear authors,
Thank you for releasing the code for Prot2Text-V2. I've been comparing the paper (https://arxiv.org/pdf/2505.11194) with the implementation and found several differences in the hyperparameters described in Section 4 (Experimental Setup).
Main Issue: It doesn't clearly separate which hyperparameters apply to Stage 1 (contrastive learning) vs Stage 2 (supervised fine-tuning). The hyperparameters section describes them together, making it ambiguous which settings are used for each stage.
Specific mismatches in SFT training code (scripts/train_instruct.py):
-
Optimizer:
- Paper (p.7): AdamW with ϵ = 1×10⁻⁶, β₁ = 0.9, β₂ = 0.999
- Code (line 405):
Adam(model.parameters(), lr=args["learning_rate"]) (no weight decay, default params)
-
Learning Rate Scheduler:
- Paper (p.7): Cosine scheduler with 6% warmup
- Code (line 406):
StepLR(optimizer, step_size=1, gamma=args["scheduler_gamma"]) (no warmup)
-
LoRA Target Modules:
- Paper (p.7): "apply it to the self-attention modules in both the ESM encoder and LLaMA decoder"
- Code (lines 141-150): Only applied to LLaMA decoder (self-attention + MLP), ESM encoder has no LoRA
-
Batch Size:
- Paper (p.7): "batch size per device is 1024" (for contrastive), then "batch size is set to 4 per GPU" (unclear which stage)
- Unclear if these apply to both stages or differ between them
Could you clarify:
- Which hyperparameters were actually used for the reported results in each stage?
- Can you provide a clear breakdown of hyperparameters per training stage (Stage 1 vs Stage 2)?
- Does the code or paper reflect the actual trained model?
Best regards,
Dear authors,
Thank you for releasing the code for Prot2Text-V2. I've been comparing the paper (https://arxiv.org/pdf/2505.11194) with the implementation and found several differences in the hyperparameters described in Section 4 (Experimental Setup).
Main Issue: It doesn't clearly separate which hyperparameters apply to Stage 1 (contrastive learning) vs Stage 2 (supervised fine-tuning). The hyperparameters section describes them together, making it ambiguous which settings are used for each stage.
Specific mismatches in SFT training code (
scripts/train_instruct.py):Optimizer:
Adam(model.parameters(), lr=args["learning_rate"])(no weight decay, default params)Learning Rate Scheduler:
StepLR(optimizer, step_size=1, gamma=args["scheduler_gamma"])(no warmup)LoRA Target Modules:
Batch Size:
Could you clarify:
Best regards,