Hi, thank you for releasing Prot2Text-V2. I am trying to reproduce the H-SCALE results and noticed three implementation differences between the paper and the current code. Could you clarify which version was used for the reported results?
1. H-SCALE loss
Equation 4 defines two separate losses:
L_align = 0.7 * L_InfoNCE(mean_protein, mean_text)
+ 0.3 * L_InfoNCE(std_protein, std_text)
The code instead concatenates mean and std and applies one InfoNCE loss. Which objective should be used?
2. ESM training in Stage 1
The paper says Stage 1 trains the projector and part of ESM, and reports LoRA on ESM self-attention. The released script is marked Without LoRA and fully freezes ESM.
Was ESM LoRA used for the reported H-SCALE model? If so, which attention modules were targeted?
An exact Stage 1 config or checkpoint would be very helpful. Thank you!
Hi, thank you for releasing Prot2Text-V2. I am trying to reproduce the H-SCALE results and noticed three implementation differences between the paper and the current code. Could you clarify which version was used for the reported results?
1. H-SCALE loss
Equation 4 defines two separate losses:
The code instead concatenates mean and std and applies one InfoNCE loss. Which objective should be used?
2. ESM training in Stage 1
The paper says Stage 1 trains the projector and part of ESM, and reports LoRA on ESM self-attention. The released script is marked
Without LoRAand fully freezes ESM.Was ESM LoRA used for the reported H-SCALE model? If so, which attention modules were targeted?
An exact Stage 1 config or checkpoint would be very helpful. Thank you!