The public training scripts (train_surface.py / main_drivaerml_surface.py) appear to run in FP32 (model.float(), no autocast / GradScaler). The paper mentions training in fp16 or bf16.
I tried running Transolver-3 on DrivAerML in fp16 and training blows up (NaNs). Could you share the mixed-precision recipe you used (fp16 vs bf16, AMP settings, any extra stability tricks)?
The public training scripts (
train_surface.py/main_drivaerml_surface.py) appear to run in FP32 (model.float(), noautocast/ GradScaler). The paper mentions training in fp16 or bf16.I tried running Transolver-3 on DrivAerML in fp16 and training blows up (NaNs). Could you share the mixed-precision recipe you used (fp16 vs bf16, AMP settings, any extra stability tricks)?