Skip to content

[Klaud Cold] Update qwen3.5-fp4-mi355x-sglang SGLang image to v0.5.20-rocm720-mi35x-20260919 / 将 qwen3.5-fp4-mi355x-sglang 的 SGLang 镜像更新至 v0.5.20-rocm720-mi35x-20260919 - #3311

Draft
adibarra wants to merge 1 commit into
mainfrom
klaud/auto-450cf9e2a8690947-ba4ad4932f9ef993
Draft

adibarra wants to merge 1 commit into
mainfrom
klaud/auto-450cf9e2a8690947-ba4ad4932f9ef993

Conversation

@adibarra

@adibarra adibarra commented Sep 20, 2026

Copy link
Copy Markdown
Collaborator

Goal: Update SGLang image from lmsysorg/sglang-rocm:v0.5.19-rocm720-mi35x-20260913 to lmsysorg/sglang-rocm:v0.5.20-rocm720-mi35x-20260919.
Baseline: 2026-09-14 · lmsysorg/sglang-rocm:v0.5.19-rocm720-mi35x-20260913
Mean latency · Sources: API 1, API 2, API 3

Point Total tok/s/GPU ↑ Output tok/s/GPU ↑ TTFT ms ↓ TPOT ms ↓
8k/1k c4 TP2 EP1 8d071a 2,001.03 222.6 343.27 8.42
8k/1k c4 TP4 EP1 f4bc3f 1,077.53 119.87 273.61 7.86
8k/1k c8 TP2 EP1 a6851d 3,158.79 355.1 513.18 10.38
8k/1k c8 TP4 EP1 a3a77e 1,797.94 202.12 424.11 9.11
8k/1k c16 TP2 EP1 031729 4,580.15 507.51 571.66 14.57
8k/1k c16 TP4 EP1 8a13b4 2,880.95 319.23 490.04 11.43
8k/1k c32 TP2 EP1 dde6f7 6,031.28 674.47 719.09 22.38
8k/1k c64 TP2 EP1 e22360 8,271.69 917.67 1,078.54 33.16
8k/1k c128 TP2 EP1 0cc9d4 10,186.9 1,128.36 1,773.84 54.06
8k/1k c256 TP2 EP1 7f1993 12,026.57 1,336.67 3,081.46 91.41
Eval Score ↑ Samples
gsm8k/em_strict · c64 96.74% 1,319
gsm8k/em_strict · c256 96.74% 1,319

Status: Klaud Cold stopped before any GPU dispatch. The baseline above was frozen locally by the reporting helper from the public API (2026-09-14, producer run 34758656500, head 4796505be3e4) and rendered with the canonical renderer, but the helper could not publish it or any typed report: the session credential is not the Klaud-Cold account, so the lifecycle helper rejects this PR as not owned. See the first Klaud comment for the source comparison, the blocker and the next step.

AI model disclosure

Prepared by Claude Fable 5.1 (claude-fable-5-1) running as the autonomous Klaud Cold candidate session; no other models or delegated agents were used.

中文

**目标:**将 SGLang 镜像从 lmsysorg/sglang-rocm:v0.5.19-rocm720-mi35x-20260913 更新为 lmsysorg/sglang-rocm:v0.5.20-rocm720-mi35x-20260919
**基线:**2026-09-14 · lmsysorg/sglang-rocm:v0.5.19-rocm720-mi35x-20260913
平均延迟 · 来源: API 1, API 2, API 3;数值及异常说明见上表。

**状态:**Klaud Cold 在任何 GPU 调度之前停止。上表基线由报告助手从公开 API 本地冻结(2026-09-14,生产运行 34758656500,head 4796505be3e4)并用规范渲染器渲染,但助手无法将其发布到 PR 或发布任何类型化报告:会话凭据不是 Klaud-Cold 账号,生命周期助手将此 PR 判定为非自有。来源比较、阻塞原因与下一步见第一条 Klaud 评论。

**AI 模型披露:**由 Claude Fable 5.1(claude-fable-5-1)作为自主 Klaud Cold 候选会话准备;未使用其他模型或委派代理。

🤖 Generated with Claude Code

…0-rocm720-mi35x-20260919

Bump the qwen3.5-fp4-mi355x-sglang master image from
lmsysorg/sglang-rocm:v0.5.19-rocm720-mi35x-20260913 (SGLang main 14b647cf27,
2026-09-13) to lmsysorg/sglang-rocm:v0.5.20-rocm720-mi35x-20260919 (SGLang main
7fac84b639, 2026-09-19, includes the v0.5.20 tag commit 94602c9c; digest
sha256:2bd66396d60d01aef2245fbbf143a8ef941b48b0bbc233c4f4a9c172adf2c994).
AITER 4ad99832, Triton 42270451, ROCm 7.2 and PyTorch 2.9.1 are unchanged.
The recipe script, model, precision, topology and search space are unchanged.

将 qwen3.5-fp4-mi355x-sglang 的主镜像从
lmsysorg/sglang-rocm:v0.5.19-rocm720-mi35x-20260913(SGLang main 14b647cf27,
2026-09-13)更新为 lmsysorg/sglang-rocm:v0.5.20-rocm720-mi35x-20260919(SGLang
main 7fac84b639,2026-09-19,包含 v0.5.20 tag 提交 94602c9c;digest
sha256:2bd66396d60d01aef2245fbbf143a8ef941b48b0bbc233c4f4a9c172adf2c994)。
AITER 4ad99832、Triton 42270451、ROCm 7.2 与 PyTorch 2.9.1 均未变化。
配方脚本、模型、精度、拓扑与搜索空间保持不变。

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution!

  • Review: If this PR changes files owned by someone other than a repository admin or @SemiAnalysisAI/core, ask one eligible CODEOWNER to complete the latest PR_REVIEW_CHECKLIST.md before contacting a core maintainer on Slack. Follow the template exactly, including As a PR reviewer and CODEOWNER, I have reviewed this and have, so sign-off verification triggers.
  • PR verification: Sweeps only run on labeled PRs. Add full-sweep-fail-fast (strongly recommended); use full-sweep-enabled only when matrix jobs should continue after a failure.
  • After merging: PR authors must ensure all GitHub Actions jobs pass. Transient failures often pass on rerun; see how to rerun failed jobs.
中文

感谢你的贡献!

  • **审阅:**如果 PR 修改的文件归属于仓库管理员及 @SemiAnalysisAI/core 之外的 CODEOWNER,请先联系一位有资格的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,再通过 Slack 联系核心维护者。必须严格遵循模板,并保留 As a PR reviewer and CODEOWNER, I have reviewed this and have,才能触发签核验证。
  • **PR 验证:**扫描仅在带有标签的 PR 上运行。强烈建议添加 full-sweep-fail-fast;仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled
  • **合并后:**PR 作者必须确保所有 GitHub Actions 任务通过。临时性失败通常可以通过重新运行恢复;参见重新运行失败任务的说明

@adibarra

Copy link
Copy Markdown
Collaborator Author

Initial update 0/5 · Not dispatched · no run · 2026-09-20 06:47 UTC
lmsysorg/sglang-rocm:v0.5.20-rocm720-mi35x-20260919 · c2cb8aa99bc0 · 8k/1k · TP2/EP1, TP4/EP1 · Mean latency
Change: Bump only the qwen3.5-fp4-mi355x-sglang master image from the 2026-09-13 ROCm 7.2 MI35x daily (SGLang main 14b647cf27, digest sha256:d1a7c4cae3870711a5548cfe38e0b33db4c61c8977b04f66b7ab9b3bc1769a64) to the 2026-09-19 daily (Docker Hub tag, digest sha256:2bd66396d60d01aef2245fbbf143a8ef941b48b0bbc233c4f4a9c172adf2c994, SGLang main 7fac84b639). Both image configs record the exact bundled commit in SETUPTOOLS_SCM_PRETEND_VERSION (0.5.19.dev20260913+g14b647cf270.5.20.dev20260919+g7fac84b639), so the release-string-mismatch / unstable-image hints resolve to: dated nightlies built from main, and the new image contains the v0.5.20 tag commit 94602c9c (tagged 2026-09-18) as an ancestor. Coupled dependencies are unchanged between the two images: AITER 4ad99832, Triton 42270451, ROCm 7.2.70200, PyTorch 2.9.1, gfx950-rocm720; sgl-kernel moves 0.4.6.post1 → 0.4.7 (built from source in the ROCm image). Across the 318 upstream commits, every flag used by benchmarks/single_node/fixed_seq_len/qwen3.5_fp4_mi355x.sh (--attention-backend, --mem-fraction-static, --model-loader-extra-config, --watchdog-timeout, --disable-radix-cache, --max-running-requests, --page-size, --kv-cache-dtype, --trust-remote-code) is still defined with unchanged semantics in python/sglang/srt/arg_groups/fields/, and SGLANG_USE_AITER, SGLANG_USE_AITER_UNIFIED_ATTN, SGLANG_MAMBA_SSM_DTYPE and ROCM_QUICK_REDUCE_QUANTIZATION keep their definitions and defaults in environ.py / mamba_utils.py / quick_all_reduce.py. Qwen3.5 model files change only a parallel-state accessor; the AITER attention backend adds a spec-decode plan-cache reset (#40325 area) that is inert without speculative decoding; the Quark W8A8 FP8 fused-MoE runner selection changed, which should not apply to this checkpoint's MXFP4 routed experts. No recipe flag or environment change is needed and no engine patch exists on the selected launch path (runners/launch_mi355x-amds.sh → recipe fallback qwen3.5_fp4_mi355x.sh). Weights amd/Qwen3.5-397B-A17B-MXFP4-AttnFP8-V2 (HF revision e17e5f0e) are the same as the published 2026-09-14 baseline and are fetched through the launcher's shared HF cache mount. Not runtime-verified: no benchmark was dispatched.

Blocker · session credential identity: the GitHub credential supplied to this candidate session authenticates as a login other than Klaud-Cold, so this PR is authored by that other login and the lifecycle helper rejects it as not owned (report, check-final, finish and the Stop-hook check-stop all fail with Candidate ownership mismatch; runs it dispatched would be filtered out of ownership discovery the same way). The complete 2026-09-14 public baseline (10 throughput points plus the two published gsm8k evals from producer run 34758656500) was frozen locally and is rendered in the PR body with the canonical renderer, but the typed record could not be embedded and no verified completion receipt can exist. No e2e-tests.yml run was dispatched, so there is nothing to cancel. Capacity checks for mi355x passed at every gate (before edits, before branch creation and before the intended smoke dispatch). Same failure as #3279 on 2026-09-19.

Next: Maintainer decision: restore the Klaud-Cold identity for the candidate credential, then either adopt this PR (the change is source-verified; a full-sweep-fail-fast label runs the full family of 10 points and 2 default gsm8k evals under maintainer ownership) or close it. This PR stays draft with no sweep labels so the family is not re-selected every six hours while the credential is wrong.

中文

初始更新 0/5 · 未调度 · 无运行 · 2026-09-20 06:47 UTC
lmsysorg/sglang-rocm:v0.5.20-rocm720-mi35x-20260919 · c2cb8aa99bc0 · 8k/1k · TP2/EP1, TP4/EP1 · 平均延迟
**变更:**仅将 qwen3.5-fp4-mi355x-sglang 的主镜像从 2026-09-13 的 ROCm 7.2 MI35x 每日构建(SGLang main 14b647cf27)更新为 2026-09-19 的每日构建(SGLang main 7fac84b639,digest sha256:2bd66396…,包含 v0.5.20 tag 提交 94602c9c)。两个镜像均在 SETUPTOOLS_SCM_PRETEND_VERSION 中记录了精确的 SGLang 提交;AITER、Triton、ROCm 7.2 与 PyTorch 2.9.1 未变化。配方脚本使用的全部启动参数与环境变量在新旧两个提交中定义与语义一致,无需修改配方,所选启动路径上不存在引擎补丁。权重与 2026-09-14 公开基线相同。未经运行时验证:未调度任何基准测试。来源链接见上文英文部分。

**阻塞 · 会话凭据身份:**本候选会话使用的 GitHub 凭据并非 Klaud-Cold 账号,因此本 PR 的作者为其他登录名,生命周期助手将其判定为非自有(reportcheck-finalfinish 与 Stop 钩子 check-stop 均报 Candidate ownership mismatch)。完整的 2026-09-14 公开基线(10 个吞吐点与 2 个已发布的 gsm8k 评测)已在本地冻结并用规范渲染器渲染到 PR 正文,但类型化记录无法嵌入,也无法生成经验证的完成回执。未调度任何 e2e-tests.yml 运行,无需取消。各关口的 mi355x 容量检查均通过。与 2026-09-19 的 #3279 为同一故障。

**下一步:**维护者决定:为候选凭据恢复 Klaud-Cold 身份,然后采用本 PR(变更已经过源码比对;添加 full-sweep-fail-fast 标签即可在维护者名下运行 10 个点与 2 个默认 gsm8k 评测的完整配置族)或关闭它。本 PR 保持草稿且不带 sweep 标签,以免在凭据修复前每六小时重复选择该配置族。

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

2 participants