Skip to content

[Klaud Cold] Update qwen3.8next-fp4-b200-sglang-agentic-mtp SGLang image to v0.5.20-cu130 / 将 qwen3.8next-fp4-b200-sglang-agentic-mtp 的 SGLang 镜像更新至 v0.5.20-cu130 - #3325

Closed
adibarra wants to merge 1 commit into
mainfrom
klaud/auto-8066badf6f119f59-cb3f7f5ee450245b
Closed

adibarra wants to merge 1 commit into
mainfrom
klaud/auto-8066badf6f119f59-cb3f7f5ee450245b

Conversation

@adibarra

@adibarra adibarra commented Sep 21, 2026

Copy link
Copy Markdown
Collaborator

Goal: Update SGLang image from lmsysorg/sglang:qwen38flashnext to lmsysorg/sglang:v0.5.20-cu130.
Baseline: 2026-09-15 · lmsysorg/sglang:qwen38flashnext
AgentX · TP1/EP1 · Mean latency · Sources: API 1, API 2, API 3, API 4

Concurrency Total tok/s/GPU ↑ Output tok/s/GPU ↑ TTFT ms ↓ TPOT ms ↓
1 14,445.12 159.68 665.82 3.57
4 23,191.29 246.97 595.02 4.3
8 39,051.57 370.67 858.57 5.43
16 71,333.43 637.9 1,570.11 12.65
Eval Score ↑ Samples
gsm8k/em_strict · c16 97.95% 1,319
中文

**目标:**将 SGLang 镜像从 lmsysorg/sglang:qwen38flashnext 更新为 lmsysorg/sglang:v0.5.20-cu130
**基线:**2026-09-15 · lmsysorg/sglang:qwen38flashnext
AgentX · TP1/EP1 · 平均延迟 · 来源: API 1, API 2, API 3, API 4;数值及异常说明见上表。

…ge to v0.5.20-cu130

Move the family from the model-branch tag lmsysorg/sglang:qwen38flashnext
(commit 593134d1, 2026-09-03) to the official lmsysorg/sglang:v0.5.20-cu130
release (commit 94602c9c), which ships upstream Qwen3.8-Flash-Next support
(sgl-project/sglang#37500). v0.5.20 removed the deprecated --cuda-graph-max-bs
alias (sgl-project/sglang#38375), so the unshared recipe now passes the same
value through --cuda-graph-max-bs-decode.

将该配置从模型分支标签 lmsysorg/sglang:qwen38flashnext(提交 593134d1,2026-09-03)
迁移到官方发布镜像 lmsysorg/sglang:v0.5.20-cu130(提交 94602c9c),该版本已包含
上游 Qwen3.8-Flash-Next 支持(sgl-project/sglang#37500)。v0.5.20 移除了已弃用的
--cuda-graph-max-bs 别名(sgl-project/sglang#38375),因此该配方独有的脚本改用
--cuda-graph-max-bs-decode 传递相同的值。

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution!

  • Review: If this PR changes files owned by someone other than a repository admin or @SemiAnalysisAI/core, ask one eligible CODEOWNER to complete the latest PR_REVIEW_CHECKLIST.md before contacting a core maintainer on Slack. Follow the template exactly, including As a PR reviewer and CODEOWNER, I have reviewed this and have, so sign-off verification triggers.
  • PR verification: Sweeps only run on labeled PRs. Add full-sweep-fail-fast (strongly recommended); use full-sweep-enabled only when matrix jobs should continue after a failure.
  • After merging: PR authors must ensure all GitHub Actions jobs pass. Transient failures often pass on rerun; see how to rerun failed jobs.
中文

感谢你的贡献!

  • **审阅:**如果 PR 修改的文件归属于仓库管理员及 @SemiAnalysisAI/core 之外的 CODEOWNER,请先联系一位有资格的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,再通过 Slack 联系核心维护者。必须严格遵循模板,并保留 As a PR reviewer and CODEOWNER, I have reviewed this and have,才能触发签核验证。
  • **PR 验证:**扫描仅在带有标签的 PR 上运行。强烈建议添加 full-sweep-fail-fast;仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled
  • **合并后:**PR 作者必须确保所有 GitHub Actions 任务通过。临时性失败通常可以通过重新运行恢复;参见重新运行失败任务的说明

@adibarra

adibarra commented Sep 21, 2026

Copy link
Copy Markdown
Collaborator Author

Initial attempt · Failed · Run 35550052895 / attempt 1 · 2026-09-21 01:15 UTC
lmsysorg/sglang:v0.5.20-cu130 · 0dc6546247f1 · AgentX · TP1/EP1 · Mean latency
Change: Move the master image from the model-branch tag lmsysorg/sglang:qwen38flashnext (commit 593134d1, 2026-09-03) to the official release lmsysorg/sglang:v0.5.20-cu130 (tag commit 94602c9c), which ships upstream Qwen3.8-Flash-Next support (#37500), and rename the recipe's --cuda-graph-max-bs to --cuda-graph-max-bs-decode because v0.5.20 deleted the deprecated alias (#38375).

Concurrency Output tok/s/GPU ↑ TTFT ms ↓ TPOT ms ↓
1 N/A N/A N/A

Note: All rows: failed.
Note: All rows: request errors unavailable.
Note: All rows: Δ N/A: no matched baseline.

Eval Score ↑ Samples
gsm8k/em_strict · c16 N/A 1,319/N/A (old/new)

Note: All rows: failed.
Note: All rows: Δ N/A: eval score unavailable.

Next: Maintainer fix of the shared B200 Nscale launcher (out of Klaud scope): check_env_vars MODEL_PATH at runners/launch_b200-nscale-slurm.sh:163 exits before the staged-path fallback, so every qwen3.8next fp4 job on this pool, including main sweeps, fails at launch.

Provenance: old image index digest sha256:5ae5816783d58e2e56e84d2e863f5441425056f500b7fbd7448c4aae017a2521; new image index digest sha256:06e4f2ed21afde4ff513cda65070124e727ba23ccaeff7712b8c40e1097d611f (amd64 sha256:b27fce60bc5494c118c4910702812bcfa8cee67abcdd1ff8b0902f21647552f4). The old tag is a locally built model-branch image (BRANCH_TYPE=local, sgl-kernel 0.4.6.post1, FlashInfer 0.6.18, CUDA 13.0.3) whose 15 branch commits diverge from the #37500 merge, so the exact old→new model-code delta is not fully reconstructible; v0.5.20 is a release build (sgl-kernel 0.4.7, sgl-deep-gemm 0.2.0, FlashInfer 0.6.18, CUDA 13.0.3) that also carries the Qwen3.8 follow-ups #39126 and #39474. At both commits arg_groups/fields/* define every other recipe flag (linear_attn_{prefill,decode}_backend, mamba_ssm_dtype, max_mamba_cache_size, tokenizer_worker_num, NEXTN 3/1/4, reasoning_parser auto, max_running_requests); the SM100 --linear-attn-decode-backend flashinfer--mamba-ssm-dtype bfloat16 gate in attention_hook.py is unchanged; SGLANG_SIMULATE_ACC_* and SGLANG_TIMEOUT_KEEP_ALIVE keep their definitions; models/qwen4_exp{,_mtp}.py differ only in PLE-offload config plumbing and the draft-extend-v2 hidden-state selection. No engine patch exists on runners/launch_b200-nscale-slurm.shqwen3.8next_fp4_b200_sglang_mtp.sh and none is added; weights stay at the launcher's staged /scratch/models/Qwen3.8-Flash-Next-NVFP4. The digest is recorded here rather than in the config because the sibling B300 PR's @sha256 reference failed at enroot import.

Blocker · session credential identity: the GitHub credential supplied to this candidate session is not the Klaud-Cold login, so the lifecycle helper rejects this PR as not owned (report, check-final and finish fail with Candidate ownership mismatch, and runs dispatched here are invisible to its ownership discovery). Same fault as #3279, #3311 and #3318. The baseline and this comment were rendered with the canonical renderer directly; no verified completion receipt can exist for this session. Capacity for b200-nscale passed before edits, before branch/PR creation and before this dispatch.

Failure: First error in both jobs (agentic c1 job 106183052128 on b200-nscale-slurm_00, agentic eval c16 job 106183052112 on b200-nscale-slurm_01), ~35 s after job start, on the runner host before enroot import or any server launch: check_env_vars MODEL_PATHError: The following required environment variables are not set: - MODEL_PATH → exit 1, from runners/launch_b200-nscale-slurm.sh line 163 in the qwen3.8next/fp4 branch. The launcher's own fallback to /scratch/models/Qwen3.8-Flash-Next-NVFP4 (lines 164-168, #3270) is unreachable because check_env_vars exits first; the benchmark workflow exports no MODEL_PATH. Independent of the image; not a recipe or master-config setting. Log excerpt (job 106183052128): + check_env_vars MODEL_PATH · Error: The following required environment variables are not set: · - MODEL_PATH · + exit 1. Job links: c1 throughput, c16 eval. Deterministic launcher failure, not an infrastructure retry; no in-scope repair exists (single-node master configs have no environment passthrough and the recipe runs only after the launcher), so no repair budget was spent. Outcome requested: readiness-blocked (phase targeted, repairs 0). The lifecycle finish cannot verify or clean up this session (credential blocker above), so the PR stays draft without sweep labels for the maintainer.

中文

初次尝试 · 失败 · Run 35550052895 / attempt 1 · 2026-09-21 01:15 UTC
lmsysorg/sglang:v0.5.20-cu130 · 0dc6546247f1 · AgentX · TP1/EP1 · 平均延迟
**变更:**将主镜像从模型分支标签 lmsysorg/sglang:qwen38flashnext(提交 593134d1,2026-09-03)更新为官方发布镜像 lmsysorg/sglang:v0.5.20-cu130(tag 提交 94602c9c,已包含上游 Qwen3.8-Flash-Next 支持 #37500);并因 v0.5.20 删除了弃用别名(#38375),将配方中的 --cuda-graph-max-bs 改为 --cuda-graph-max-bs-decode。实测数值及异常说明见上表。
**下一步:**需由维护者修复共享的 B200 Nscale 启动器(超出 Klaud 编辑范围):runners/launch_b200-nscale-slurm.sh:163check_env_vars MODEL_PATH 在回退到预置路径之前就退出,导致该池上所有 qwen3.8next fp4 作业(含 main 分支 sweep)在启动阶段失败。

**来源说明:**旧标签为本地构建的模型分支镜像(15 个分支提交与 #37500 合并存在分歧,旧→新的模型代码差异无法完全重建);v0.5.20 为官方发布构建,并包含 #39126、#39474 两项 Qwen3.8 后续修复。两个提交中配方使用的其他全部参数均有定义,SM100 上 flashinfer 线性注意力解码需 --mamba-ssm-dtype bfloat16 的门控未变,所选启动路径上无引擎补丁且未新增。因 B300 兄弟 PR 的 @sha256 引用在 enroot import 时失败,digest 记录于上文而非配置中。

**阻塞 · 会话凭据身份:**本会话的 GitHub 凭据并非 Klaud-Cold 登录名,生命周期助手判定本 PR 非自有(reportcheck-finalfinish 均报 Candidate ownership mismatch)。与 #3279#3311#3318 为同一故障。基线与本评论由规范渲染器直接渲染;本会话无法生成经验证的完成回执。

**失败:**两个作业均在作业启动约 35 秒后、容器导入与服务启动之前,于运行器主机上因 runners/launch_b200-nscale-slurm.sh:163check_env_vars MODEL_PATHMODEL_PATH 未设置而退出(exit 1);#3270 新增的预置路径回退位于其后而无法执行,工作流也不导出 MODEL_PATH。与镜像无关,亦非配方或主配置可设置项。属确定性启动器故障而非基础设施重试;不存在范围内的修复,因此未消耗修复预算。请求的结果为 readiness-blocked(阶段 targeted,修复 0 次)。因上述凭据阻塞,finish 无法验证或清理本会话,PR 保持草稿且不带 sweep 标签,待维护者处理。

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

1 participant