diff --git a/docs/eval-agentx-procedures.md b/docs/eval-agentx-procedures.md index 06c9ef6cf9..ba58b91743 100644 --- a/docs/eval-agentx-procedures.md +++ b/docs/eval-agentx-procedures.md @@ -344,6 +344,45 @@ gh run cancel --repo SemiAnalysisAI/InferenceX Use `scancel` or process termination only with explicit approval and a concrete reason. They can bypass cleanup or strand the runner. After a recipe fix, dispatch one targeted fast e2e point, inspect it live, then reserve a canonical run/full sweep for the candidate that passed. +## 11. Lessons from the September 2026 DSpark and GLM-5.2 AgentX sweeps + +Evidence-backed rules from the DeepSeek-V4.1-Flash DSpark arms (#3239-#3247, #3216) and the GLM-5.2 nightly bumps (#3275-#3277). Cited run ids are GitHub Actions runs in this repository. + +**Memory on Blackwell SGLang DSpark arms** + +- Keep `--max-running-requests` inside the decode CUDA-graph batch (64). A DSpark verify step for a batch above the captured tier runs eagerly and allocates its attention workspace on the fly; the H200 eval OOMed that way at 128 running requests with 2 GiB free (run 35306704553). +- Use `--mem-fraction-static 0.70` and `--chunked-prefill-size 4096` on every Blackwell arm. GB300 kept 0.75/8192/128 and c128 OOMed 82 minutes into warmup when the TVM sparse-attention prefill kernel asked for 20 GiB with 12.5 GiB free (run 35307635202); B300 completed c1-c128 on 0.70/64. +- Per-request pool caps scale with the DSpark block: the indexer and verify buffers grow with chunk × context, so the prefill chunk, not the static fraction, is the first lever when a long prompt OOMs. + +**MI355X (gfx950) SGLang preview** + +- The Engram tables must stay on the GPU. `SGLANG_ENABLE_DSV41_ENGRAM_HOST_TABLE=1` raised `hipErrorIllegalAddress` on every rank during decode graph capture (run 35306715045). With the tables resident the weights take 117-129 GB of each 288 GB card, so the static fraction has to drop (0.60 shipped). +- The ROCm image runs the torch caching allocator with fixed segments (aiter logs `expandable_segments=False`). Eager chunked prefill of one 126k-token prompt hoarded every non-static byte and RCCL aborted with `HSA_STATUS_ERROR_OUT_OF_RESOURCES` at 0 MB free with a single request running (runs 35376928227, 35460273555). `PYTORCH_HIP_ALLOC_CONF=expandable_segments:True` plus a 2048-token chunk cleared c1, c2, c4 and the eval (run 35468635539). `GPU_MAX_HW_QUEUES=2` and `HSA_NO_SCRATCH_RECLAIM=1` on MEC firmware below 177 are the other RCCL levers on this SKU. +- When a concurrency point exceeds what the eager 1M-context prefill can hold, publish the points that fit and trim the conc-list, as #2661 did for B300 HiCache; 13 concurrent sessions did not fit at any static fraction. + +**vLLM DSpark on B200 TP2** + +- The eval path (block rejection with adaptive verification) died with `cudaErrorIllegalAddress` twice on TP2, first in the eager draft `lm_head` GEMM and then at the adaptive verifier's `record_confidences` sync even with lm-eval pinned to 32 requests inside the graph tier (runs 35399984613, 35403890075). The identical config passes on TP4. Turn adaptive verification off for TP2 evals; block rejection alone is exact. +- An AgentX warmup that aborts with `ClientOSError: Can not write request body` and `ServerDisconnectedError` while the server log shows only `auto-aborting request due to dropped stream` is a client-side abort during a server stall (autotune, first long prefill), not a crash. Rerun the point; raise `AIPERF_HTTP_TCP_USER_TIMEOUT` if it recurs. + +**Detecting hung jobs** + +- A GitHub job can sit "in progress" for hours with no Slurm allocation behind it (B300 c4 ran 5.5 hours while siblings finished in 2; GLM-5.2 B300 evals ran 4 hours against an 18-minute B200 baseline). Compare against the durations of finished siblings on the same or a sibling SKU, then check `squeue` for the runner name. `PrivateData=jobs` hides other users' jobs on some clusters, so an empty queue is not proof; the duration gap is. Cancel the run and `gh run rerun --failed` to keep the green points. +- Server logs vanish minutes after a job ends because the next job reuses the runner workspace. For crash hunting on a shared jumpbox, run a detached capture loop (`nohup`, pid file) that copies `results/server.log` when it matches a crash signature; SSH-tethered loops die with the connection. + +**Queue and lease mechanics** + +- The lease controller admits jobs by publishing a `ci-job-*` label onto a leased runner. Idle runners tagged `ci-slurm-unavailable`, or a pool held by multi-node `ci-lease-*` tokens and `ci-job-1.000` priority work, mean priority-0 single-node jobs never dispatch even with runners online. +- GitHub cancels a job that has been queued for 24 hours (GB300 run 35307635202 attempt 1, B300 c4 attempt 2). `gh run rerun --failed` re-queues only the cancelled jobs and preserves successes and the reuse authorization. +- A force-push while the previous head's sweep is still running leaves the new run `pending` behind the pull-request concurrency group; `gh run cancel` can lag, so use the `force-cancel` API when the old run lingers. + +**Repository mechanics** + +- `check-changelog` validates that every config key named by a new entry exists in the master configs (run 35403589805 rejected a misspelled key). Use the exact key. +- The reuse gate only requires the green run's head commit to remain in the PR. Merge `origin/main` (append-only changelog resolution) to keep reuse valid; rebase to a single fresh commit only when a new sweep is wanted, and diff against the merge-base, not `origin/main`, or the rebase silently reverts work that landed on main in between. +- Launcher refactors change conflict shapes: #3270 folded `launch_b200-nscale-compat.sh` into `launch_b200-nscale-slurm.sh`, so a PR editing the old file must port its hunk into the merged launcher and accept the deletion. +- SGLang nightlies from 2026-09-15 reject `--cuda-graph-max-bs` as an ambiguous prefix; pass `--cuda-graph-max-bs-decode`. ROCm `vllm-openai-rocm` nightly tags disappear from Docker Hub within days, so verify the tag on the day you pin it and expect a re-pin before merge. + ## Completion checklist - Matrix preview matches intended scenario, topology, eval mode, and concurrency. diff --git a/docs/eval-agentx-procedures_zh.md b/docs/eval-agentx-procedures_zh.md index 03c0580439..b252fef378 100644 --- a/docs/eval-agentx-procedures_zh.md +++ b/docs/eval-agentx-procedures_zh.md @@ -342,6 +342,45 @@ gh run cancel --repo SemiAnalysisAI/InferenceX 只有在获得明确批准且有具体理由时才使用 `scancel` 或终止进程;否则可能绕过 cleanup 或使 runner 残留。修复 recipe 后,先分派一个目标 fast e2e 点并实时检查,只有通过检查的 candidate 才值得进行 canonical 运行/完整 sweep。 +## 11. 2026 年 9 月 DSpark 与 GLM-5.2 AgentX sweep 的经验教训 + +以下规则来自 DeepSeek-V4.1-Flash DSpark 配方(#3239-#3247、#3216)与 GLM-5.2 nightly 升级(#3275-#3277)的实证,所引用的运行 id 均为本仓库的 GitHub Actions 运行。 + +**Blackwell SGLang DSpark 配方的显存** + +- 让 `--max-running-requests` 不超过 decode CUDA graph batch(64)。超出已捕获层级的 DSpark verify 步骤会以 eager 方式运行并临时分配 attention 工作区;H200 eval 在 128 个运行请求、仅剩 2 GiB 时因此 OOM(运行 35306704553)。 +- 所有 Blackwell 配方使用 `--mem-fraction-static 0.70` 与 `--chunked-prefill-size 4096`。GB300 保留 0.75/8192/128 时,c128 在 warmup 82 分钟后因 TVM 稀疏注意力 prefill 内核申请 20 GiB 而仅剩 12.5 GiB 而 OOM(运行 35307635202);B300 在 0.70/64 下完成了 c1-c128。 +- 逐请求的池上限随 DSpark block 缩放:indexer 与 verify 缓冲随“分块 × 上下文”增长,因此长提示 OOM 时首先调整 prefill 分块,而非静态比例。 + +**MI355X(gfx950)SGLang 预览版** + +- Engram 表必须驻留 GPU。`SGLANG_ENABLE_DSV41_ENGRAM_HOST_TABLE=1` 会在 decode graph 捕获时于所有 rank 触发 `hipErrorIllegalAddress`(运行 35306715045)。表驻留后权重占用每块 288 GB 显卡的 117-129 GB,静态比例必须下调(最终为 0.60)。 +- ROCm 镜像的 torch 缓存分配器使用固定段(aiter 日志显示 `expandable_segments=False`)。一条 126k token 提示的 eager 分块 prefill 会囤积全部非静态显存,即使只有一个请求,RCCL 也会因 `HSA_STATUS_ERROR_OUT_OF_RESOURCES`(剩余 0 MB)中止(运行 35376928227、35460273555)。`PYTORCH_HIP_ALLOC_CONF=expandable_segments:True` 加 2048 token 分块后 c1、c2、c4 与 eval 全部通过(运行 35468635539)。该 SKU 上其余 RCCL 手段为 `GPU_MAX_HW_QUEUES=2` 以及 MEC 固件低于 177 时的 `HSA_NO_SCRATCH_RECLAIM=1`。 +- 当某个并发点超出 eager 1M 上下文 prefill 的容量时,参照 #2661 对 B300 HiCache 的做法,只发布可容纳的点并裁剪 conc-list;13 个并发会话在任何静态比例下都放不下。 + +**B200 TP2 上的 vLLM DSpark** + +- eval 路径(block rejection 加自适应验证)在 TP2 上两次因 `cudaErrorIllegalAddress` 崩溃:先在 eager 草稿 `lm_head` GEMM,随后即便 lm-eval 已固定为 graph 层级内的 32 个请求,仍在自适应验证器的 `record_confidences` 同步处崩溃(运行 35399984613、35403890075)。相同配置在 TP4 上通过。TP2 eval 应关闭自适应验证;仅 block rejection 即为精确验证。 +- AgentX warmup 若因 `ClientOSError: Can not write request body` 与 `ServerDisconnectedError` 中止,而 server 日志只有 `auto-aborting request due to dropped stream`,说明是服务端停顿(autotune、首个长 prefill)期间的客户端中止,而非崩溃。重跑该点;若复发则提高 `AIPERF_HTTP_TCP_USER_TIMEOUT`。 + +**识别挂起的作业** + +- GitHub 作业可能在没有任何 Slurm allocation 的情况下“运行”数小时(B300 c4 跑了 5.5 小时而同级点 2 小时完成;GLM-5.2 B300 eval 跑了 4 小时而 B200 基线为 18 分钟)。先与同 SKU 或姊妹 SKU 上已完成同级点的时长对比,再用 runner 名称查 `squeue`。部分集群的 `PrivateData=jobs` 会隐藏其他用户的作业,队列为空不算证据,时长差距才是。取消运行并用 `gh run rerun --failed` 保留已通过的点。 +- 作业结束几分钟后 server 日志就会被下一个作业复用工作区时覆盖。在共享 jumpbox 上排查崩溃时,运行脱离终端的捕获循环(`nohup`、pid 文件),在 `results/server.log` 匹配崩溃特征时立即拷贝;随 SSH 连接存活的循环会随连接一起中断。 + +**队列与 lease 机制** + +- lease controller 通过在已 lease 的 runner 上发布 `ci-job-*` 标签来准入作业。空闲 runner 带有 `ci-slurm-unavailable`,或池被多节点 `ci-lease-*` 令牌与 `ci-job-1.000` 高优先级作业占满时,priority 0 的单节点作业即使 runner 在线也永远不会派发。 +- GitHub 会取消排队满 24 小时的作业(GB300 运行 35307635202 第 1 次、B300 c4 第 2 次)。`gh run rerun --failed` 只重新排队被取消的作业,并保留成功点与 reuse 授权。 +- 在上一个 head 的 sweep 仍在运行时 force-push,会让新运行卡在 pull-request 并发组后的 `pending` 状态;`gh run cancel` 可能滞后,旧运行迟迟不退时改用 `force-cancel` API。 + +**仓库机制** + +- `check-changelog` 会校验新条目引用的每个 config key 都存在于 master 配置(运行 35403589805 拒绝了拼错的 key)。使用精确的 key。 +- reuse gate 只要求绿色运行的 head 提交仍在 PR 中。合并 `origin/main`(changelog 仅追加解析)以保持 reuse 有效;仅在需要新 sweep 时才 rebase 为单个新提交,并且要与 merge-base 而非 `origin/main` 做 diff,否则 rebase 会悄悄回退期间已合入 main 的改动。 +- launcher 重构会改变冲突形态:#3270 将 `launch_b200-nscale-compat.sh` 并入 `launch_b200-nscale-slurm.sh`,编辑旧文件的 PR 必须把 hunk 移植到合并后的 launcher 并接受删除。 +- 2026-09-15 起的 SGLang nightly 将 `--cuda-graph-max-bs` 视为歧义前缀而拒绝;改用 `--cuda-graph-max-bs-decode`。ROCm `vllm-openai-rocm` nightly tag 会在数天内从 Docker Hub 消失,务必在固定当天验证 tag,并预期合并前需要重新固定。 + ## 完成检查清单 - 矩阵预览符合预期 scenario、topology、eval mode 与 concurrency。