From 421ae3ef309f0a94500064c3dd6426168fbfb052 Mon Sep 17 00:00:00 2001 From: functionstackx <47992694+functionstackx@users.noreply.github.com> Date: Sun, 20 Sep 2026 14:18:40 -0400 Subject: [PATCH 1/3] =?UTF-8?q?dsv41flash-fp4-gb200-vllm-agentic-dspark:?= =?UTF-8?q?=20add=20TP2=20arm=20with=20Engram=20host=20offload=20/=20?= =?UTF-8?q?=E6=96=B0=E5=A2=9E=20Engram=20=E4=B8=BB=E6=9C=BA=E5=8D=B8?= =?UTF-8?q?=E8=BD=BD=E7=9A=84=20TP2=20=E8=87=82?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Add a TP2 search-space line beside the existing TP4 arm. The Engram tables stay in pinned host DRAM via --engram-config cpu_offload; weights rise to ~133 GiB on each 256 GiB GPU, which still leaves room for the sparse-attention indexer's 32 GiB logits buffer at the upstream batched-token default plus ~47 GiB of KV per GPU. Unlike the 180 GB B200 TP2 arm (#3216) this SKU needs no batched-token or graph-capture caps, so no benchmark script changes. Co-Authored-By: Claude Opus 5 (1M context) --- configs/nvidia-master.yaml | 6 ++++++ perf-changelog.yaml | 11 +++++++++++ 2 files changed, 17 insertions(+) diff --git a/configs/nvidia-master.yaml b/configs/nvidia-master.yaml index c5d5edac68..4cc02d3ceb 100644 --- a/configs/nvidia-master.yaml +++ b/configs/nvidia-master.yaml @@ -8105,6 +8105,12 @@ dsv41flash-fp4-gb200-vllm-agentic-dspark: search-space: # Engram weights use UVA DRAM; the KV cache stays GPU-resident. - { tp: 4, kv-offloading: none, spec-decoding: mtp, conc-list: [1, 2, 4, 8, 16, 32, 64, 128] } + # TP2 halves the GPU count per replica. Weights rise to ~133 GiB on each + # 256 GiB GPU with the Engram tables in pinned host DRAM, which still + # leaves room for the indexer's 32 GiB logits buffer at the upstream + # batched-token default plus ~47 GiB of KV per GPU, so this arm keeps + # the upstream scheduler settings (source run 35307253872). + - { tp: 2, kv-offloading: none, spec-decoding: mtp, conc-list: [1, 2, 4, 8, 16, 32, 64, 128] } # SGLang arm for DeepSeek-V4.1-Flash AgentX on GB200, from the SGLang cookbook # (https://lmsysorg.mintlify.app/cookbook/autoregressive/DeepSeek/DeepSeek-V4_1). diff --git a/perf-changelog.yaml b/perf-changelog.yaml index c5b4309f5f..bd53449a7b 100644 --- a/perf-changelog.yaml +++ b/perf-changelog.yaml @@ -8455,3 +8455,14 @@ description: - "Update B200 vLLM AgentX to DSpark6 and a new image with TP8 and DEP8 configurations." pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3274 + +- config-keys: + - dsv41flash-fp4-gb200-vllm-agentic-dspark + scenario-type: + - agentic-coding + description: + - "Add a TP2 arm to the GB200 vLLM DeepSeek-V4.1-Flash AgentX recipe alongside the existing TP4 arm, keeping the Engram tables in pinned host DRAM via --engram-config cpu_offload; weights rise to ~133 GiB on each 256 GiB GPU, leaving room for the sparse-attention indexer's 32 GiB logits buffer at the upstream batched-token default plus ~47 GiB of KV per GPU, so the arm keeps the upstream scheduler settings" + - "Unlike the 180 GB B200 TP2 arm the GB200 GPUs need no --max-num-batched-tokens or CUDA-graph capture caps, so no benchmark script changes accompany this entry" + - "为 GB200 vLLM DeepSeek-V4.1-Flash AgentX 配方在现有 TP4 臂旁新增 TP2 臂,Engram 表继续通过 --engram-config cpu_offload 放在固定页主机 DRAM;每张 256 GiB GPU 的权重升至约 133 GiB,仍可容纳 indexer 在上游 batched-token 默认值下的 32 GiB logits 缓冲区,并为每张 GPU 留出约 47 GiB KV,因此本臂沿用上游调度器设置" + - "与 180 GB 的 B200 TP2 臂不同,GB200 无需 --max-num-batched-tokens 或 CUDA graph 捕获上限,因此本条目不涉及 benchmark 脚本改动" + pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3320 From 1b10b56ec651020a953fc8a0c4b0d571cfc706bd Mon Sep 17 00:00:00 2001 From: functionstackx <47992694+functionstackx@users.noreply.github.com> Date: Sun, 20 Sep 2026 14:37:52 -0400 Subject: [PATCH 2/3] Cap batched tokens and graph capture on TP2 arms GB200 TP2 c128 died in memory profiling with a -5.34 GiB KV budget (run 35528745995): weights are 145 GiB per rank, the indexer's logits buffer is 32 GiB at the upstream 16384 batched tokens, and graph capture at 1024 adds ~22 GiB. Apply the same caps the B200 TP2 arm uses (#3216): batched tokens 4096, scheduler batch bounded to 2x concurrency (16-256), capture stopped at 512. TP4 and TP8 arms keep the upstream defaults. Co-Authored-By: Claude Opus 5 (1M context) --- .../agentic/dsv41flash_fp4_vllm_mtp.sh | 28 +++++++++++++++++++ configs/nvidia-master.yaml | 10 +++---- perf-changelog.yaml | 8 +++--- 3 files changed, 37 insertions(+), 9 deletions(-) diff --git a/benchmarks/single_node/agentic/dsv41flash_fp4_vllm_mtp.sh b/benchmarks/single_node/agentic/dsv41flash_fp4_vllm_mtp.sh index d5d00eb8a9..d843953e95 100644 --- a/benchmarks/single_node/agentic/dsv41flash_fp4_vllm_mtp.sh +++ b/benchmarks/single_node/agentic/dsv41flash_fp4_vllm_mtp.sh @@ -46,6 +46,33 @@ while (( CAPTURE_SIZE < CONC * (1 + NUM_SPEC_TOKENS) && CAPTURE_SIZE < 2048 )); CAPTURE_SIZE=$((CAPTURE_SIZE * 2)) done +# TP2 halves the rank count, so each GPU holds ~145 GiB of weights on B200 and +# GB200 and ~175 GiB on GB300 even with the Engram tables offloaded. At the +# upstream 16384 batched tokens the sparse-attention indexer's +# [batched-tokens, 1M] fp8 logits buffer is 32 GiB, and graph capture at 1024 +# (reached at c128) adds ~22 GiB, which drove the KV budget to -5.34 GiB on +# GB200 c128 (run 35528745995). Cap batched tokens at 4096 (8 GiB, as the H100 +# arm does), bound the scheduler batch to the AgentX fan-out, and stop +# capturing above 512 tokens; TP4 and TP8 keep the upstream defaults. +TP2_ARGS=() +if (( TP == 2 )); then + MAX_NUM_SEQS=$((2 * CONC)) + if (( MAX_NUM_SEQS > 256 )); then + MAX_NUM_SEQS=256 + fi + # FlashInfer's autotune dummy run batches max-num-seqs requests through the + # DSpark draft head; with 2-8 requests it selected an invalid MXFP8 split-K + # tactic and the engine never started (run 35320655804). 16 is the smallest + # value that has passed. + if (( MAX_NUM_SEQS < 16 )); then + MAX_NUM_SEQS=16 + fi + if (( CAPTURE_SIZE > 512 )); then + CAPTURE_SIZE=512 + fi + TP2_ARGS=(--max-num-batched-tokens 4096 --max-num-seqs "$MAX_NUM_SEQS") +fi + # Pyxis shares the host network; port 8888 can already belong to a host service. select_available_server_port export AIPERF_SERVER_URL="http://localhost:${PORT}" @@ -71,6 +98,7 @@ VLLM_CMD=( --speculative-config "$SPEC_CONFIG" --max-model-len 1048576 --max-cudagraph-capture-size "$CAPTURE_SIZE" + "${TP2_ARGS[@]}" --disable-uvicorn-access-log "${LOAD_ARGS[@]}" ) diff --git a/configs/nvidia-master.yaml b/configs/nvidia-master.yaml index 4cc02d3ceb..468031eb56 100644 --- a/configs/nvidia-master.yaml +++ b/configs/nvidia-master.yaml @@ -8105,11 +8105,11 @@ dsv41flash-fp4-gb200-vllm-agentic-dspark: search-space: # Engram weights use UVA DRAM; the KV cache stays GPU-resident. - { tp: 4, kv-offloading: none, spec-decoding: mtp, conc-list: [1, 2, 4, 8, 16, 32, 64, 128] } - # TP2 halves the GPU count per replica. Weights rise to ~133 GiB on each - # 256 GiB GPU with the Engram tables in pinned host DRAM, which still - # leaves room for the indexer's 32 GiB logits buffer at the upstream - # batched-token default plus ~47 GiB of KV per GPU, so this arm keeps - # the upstream scheduler settings (source run 35307253872). + # TP2 halves the GPU count per replica. Weights rise to ~145 GiB on each + # 256 GiB GPU with the Engram tables in pinned host DRAM, so the arm + # takes the same caps as the B200 TP2 arm: batched tokens 4096 (the + # indexer's logits buffer is 32 GiB at the upstream 16384) and graph + # capture stopped at 512, leaving ~49 GiB of KV per GPU. - { tp: 2, kv-offloading: none, spec-decoding: mtp, conc-list: [1, 2, 4, 8, 16, 32, 64, 128] } # SGLang arm for DeepSeek-V4.1-Flash AgentX on GB200, from the SGLang cookbook diff --git a/perf-changelog.yaml b/perf-changelog.yaml index bd53449a7b..794c117ba5 100644 --- a/perf-changelog.yaml +++ b/perf-changelog.yaml @@ -8461,8 +8461,8 @@ scenario-type: - agentic-coding description: - - "Add a TP2 arm to the GB200 vLLM DeepSeek-V4.1-Flash AgentX recipe alongside the existing TP4 arm, keeping the Engram tables in pinned host DRAM via --engram-config cpu_offload; weights rise to ~133 GiB on each 256 GiB GPU, leaving room for the sparse-attention indexer's 32 GiB logits buffer at the upstream batched-token default plus ~47 GiB of KV per GPU, so the arm keeps the upstream scheduler settings" - - "Unlike the 180 GB B200 TP2 arm the GB200 GPUs need no --max-num-batched-tokens or CUDA-graph capture caps, so no benchmark script changes accompany this entry" - - "为 GB200 vLLM DeepSeek-V4.1-Flash AgentX 配方在现有 TP4 臂旁新增 TP2 臂,Engram 表继续通过 --engram-config cpu_offload 放在固定页主机 DRAM;每张 256 GiB GPU 的权重升至约 133 GiB,仍可容纳 indexer 在上游 batched-token 默认值下的 32 GiB logits 缓冲区,并为每张 GPU 留出约 47 GiB KV,因此本臂沿用上游调度器设置" - - "与 180 GB 的 B200 TP2 臂不同,GB200 无需 --max-num-batched-tokens 或 CUDA graph 捕获上限,因此本条目不涉及 benchmark 脚本改动" + - "Add a TP2 arm to the GB200 vLLM DeepSeek-V4.1-Flash AgentX recipe alongside the existing TP4 arm, keeping the Engram tables in pinned host DRAM via --engram-config cpu_offload; weights rise to ~145 GiB on each 256 GiB GPU" + - "Apply the B200 TP2 caps in dsv41flash_fp4_vllm_mtp.sh to every TP2 arm: --max-num-batched-tokens 4096 (the sparse-attention indexer's [batched-tokens, 1M] fp8 buffer is 32 GiB at the upstream 16384), --max-num-seqs at twice the concurrency (16-256) and CUDA graph capture stopped at 512, which leaves ~49 GiB of KV per GPU; TP4 and TP8 arms keep the upstream defaults" + - "为 GB200 vLLM DeepSeek-V4.1-Flash AgentX 配方在现有 TP4 臂旁新增 TP2 臂,Engram 表继续通过 --engram-config cpu_offload 放在固定页主机 DRAM;每张 256 GiB GPU 的权重升至约 145 GiB" + - "在 dsv41flash_fp4_vllm_mtp.sh 中将 B200 TP2 的上限推广到所有 TP2 臂:--max-num-batched-tokens 4096(上游 16384 时 indexer 的 [batched-tokens, 1M] fp8 缓冲区达 32 GiB)、--max-num-seqs 为并发的两倍(16-256)、CUDA graph 捕获上限 512,为每张 GPU 留出约 49 GiB KV;TP4 与 TP8 臂沿用上游默认值" pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3320 From 975a8375dd10198bc10324f83cb682b392635255 Mon Sep 17 00:00:00 2001 From: functionstackx <47992694+functionstackx@users.noreply.github.com> Date: Sun, 20 Sep 2026 20:26:01 -0400 Subject: [PATCH 3/3] chore: refresh PR #3320 for sweep reuse [skip-sweep] MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Sync with origin/main after the green sweep run 35529685706; the reuse gate authorizes that run on this head. 在绿色 sweep 运行 35529685706 之后与 origin/main 同步;reuse gate 在此 head 上授权该运行。 Co-Authored-By: Claude Fable 5.1