Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
126 changes: 126 additions & 0 deletions benchmarks/single_node/fixed_seq_len/qwen3-0.6b_bf16_h100_trt.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,126 @@
#!/usr/bin/env bash
set -eo pipefail

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 (optional) This script adds set -eo pipefail (line 2), unlike every other single_node/fixed_seq_len benchmark script, so operators lose the GSM8K eval results and clean GPU metrics that siblings still produce after a benchmark hiccup. If run_benchmark_serving (line 106) fails transiently, the script exits immediately and never reaches run_eval (line 120) or stop_gpu_monitor/rm -f "$EXTRA_CONFIG_FILE" (lines 124-125), so the eval stage is skipped entirely and gpu_metrics.csv is left with a truncated tail instead of being finalized. Fix: keep run_benchmark_serving and run_eval failures non-fatal (e.g. capture their exit codes) so cleanup and the eval stage still run on transient failures, consistent with dsr1_fp8_h200_trt.sh and other sibling scripts that intentionally omit set -e.

Extended reasoning...

grep '^set -e' across benchmarks/single_node/fixed_seq_len/*.sh shows this is the only script with errexit; all siblings (e.g. dsr1_fp8_h200_trt.sh) omit it deliberately, relying on falling through to stop_gpu_monitor even when run_benchmark_serving fails. run_benchmark_serving (benchmark_lib.sh:793) returns non-zero on ordinary conditions (missing args, benchmark_exit_code from infx.bench_serving.benchmark_serving, capture failures) not just catastrophic errors. With set -e active here, that non-zero return terminates the script at line 106 before line 119's RUN_EVAL check, so the GSM8K accuracy eval that this PR's own validation section relies on never runs for that job. It also skips stop_gpu_monitor (benchmark_lib.sh:478), which normally appends a final nvidia-smi sample and repairs a truncated trailing CSV row via _repair_truncated_gpu_metrics_tail; skipping it leaves gpu_metrics.csv with a partial/truncated last row. The mktemp'd EXTRA_CONFIG_FILE at line 52 also leaks since rm -f at line 125 is never reached.

Verification: nit. The mechanism is real and reachable but low severity. Line 2 of the new file is set -eo pipefail (the only fixed_seq_len script with errexit; siblings such as dsr1_fp8_h200_trt.sh omit it). run_benchmark_serving (line 106) is invoked as a bare command and returns $benchmark_exit_code (benchmark_lib.sh:1018), which is non-zero when run_server_client fails, when the `server_watch…


source "$(dirname "$0")/../../benchmark_lib.sh"

check_env_vars \
MODEL \
TP \
CONC \
ISL \
OSL \
MAX_MODEL_LEN \
RANDOM_RANGE_RATIO \
RESULT_FILENAME \
EVAL_ONLY \
RUN_EVAL \
PORT \
HF_HUB_CACHE

if [[ -n "$SLURM_JOB_ID" ]]; then
echo "JOB $SLURM_JOB_ID running on $SLURMD_NODENAME"
fi

python3 - <<'PY'
from importlib.metadata import version

expected = "1.3.0rc27"
actual = version("tensorrt_llm")
if actual != expected:
raise SystemExit(
f"Expected TensorRT-LLM {expected} for the pinned NGC image, got {actual}"
)
PY

python3 -m pip install --quiet --disable-pip-version-check \
"modelscope==1.40.1" "modelscope-hub==0.4.3"
python3 "$(dirname "$0")/../../../runners/patch_trtllm_modelscope.py"

export TRTLLM_USE_MODELSCOPE=true
MODELSCOPE_CACHE=$(mktemp -d /tmp/modelscope-cold.XXXXXX)
COLD_HF_HOME=$(mktemp -d /tmp/modelscope-hf-empty.XXXXXX)
export MODELSCOPE_CACHE
SNAPSHOT_HELPER="$(dirname "$0")/../../../runners/modelscope_snapshot.py"
SNAPSHOT_REPORT=/workspace/modelscope_snapshot_report.json
python3 "$SNAPSHOT_HELPER" before --model "$MODEL" --cache "$MODELSCOPE_CACHE" \
--hf-home "$COLD_HF_HOME" --report "$SNAPSHOT_REPORT"

echo "TP: $TP, CONC: $CONC, ISL: $ISL, OSL: $OSL"
nvidia-smi

SERVER_LOG=/workspace/server.log
EXTRA_CONFIG_FILE=$(mktemp --suffix=.yaml)
MAX_BATCH_SIZE=$((CONC > 16 ? CONC : 16))
MAX_NUM_TOKENS=$((((ISL + CONC + 127) / 128) * 128))
MAX_NUM_TOKENS=$((MAX_NUM_TOKENS > 8192 ? MAX_NUM_TOKENS : 8192))

cat > "$EXTRA_CONFIG_FILE" <<EOF
dtype: bfloat16
print_iter_log: true
kv_cache_config:
free_gpu_memory_fraction: 0.9
enable_block_reuse: false
cuda_graph_config:
enable_padding: true
max_batch_size: $MAX_BATCH_SIZE
EOF

if [[ "$EVAL_ONLY" == "true" ]]; then
# The caller supplies the model-specific context ceiling. Avoid a hub
# lookup before the server performs its cold ModelScope download.
export EVAL_MAX_MODEL_LEN="$MAX_MODEL_LEN"
MAX_NUM_TOKENS="$EVAL_MAX_MODEL_LEN"
fi

start_gpu_monitor

set -x
PYTHONNOUSERSITE=1 HF_HUB_OFFLINE=0 HF_HOME="$COLD_HF_HOME" \
HF_HUB_CACHE="$COLD_HF_HOME/hub" HUGGINGFACE_HUB_CACHE="$COLD_HF_HOME/hub" \
TRANSFORMERS_CACHE="$COLD_HF_HOME/hub" mpirun -n 1 --oversubscribe --allow-run-as-root \
trtllm-serve "$MODEL" --port="$PORT" \
--backend=pytorch \
--max_batch_size="$MAX_BATCH_SIZE" \
--max_seq_len="$MAX_MODEL_LEN" \
--max_num_tokens="$MAX_NUM_TOKENS" \
--tp_size="$TP" \
--extra_llm_api_options="$EXTRA_CONFIG_FILE" \
> "$SERVER_LOG" 2>&1 &

SERVER_PID=$!

wait_for_server_ready --port "$PORT" --server-log "$SERVER_LOG" --server-pid "$SERVER_PID"

# Resolve only after readiness; this must reuse the fresh server download.
HF_HUB_OFFLINE=1 python3 "$SNAPSHOT_HELPER" after --model "$MODEL" --cache "$MODELSCOPE_CACHE" \
--hf-home "$COLD_HF_HOME" --report "$SNAPSHOT_REPORT"
MODEL_PATH=$(python3 - "$SNAPSHOT_REPORT" <<'PYCODE'
import json
import sys
from pathlib import Path
print(json.loads(Path(sys.argv[1]).read_text())["snapshot"])
PYCODE
)
export MODEL_PATH

run_benchmark_serving \
--model "$MODEL" \
--tokenizer "$MODEL_PATH" \
--port "$PORT" \
--backend openai \
--input-len "$ISL" \
--output-len "$OSL" \
--random-range-ratio "$RANDOM_RANGE_RATIO" \
--num-prompts "$((CONC * 10))" \
--max-concurrency "$CONC" \
--result-filename "$RESULT_FILENAME" \
--result-dir /workspace/

if [[ "$RUN_EVAL" == "true" ]]; then
run_eval --framework lm-eval --port "$PORT"
append_lm_eval_summary
fi

stop_gpu_monitor
rm -f "$EXTRA_CONFIG_FILE"
set +x
17 changes: 17 additions & 0 deletions configs/nvidia-master.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -3953,6 +3953,23 @@ qwen3.5-fp8-h100-sglang:
- { tp: 8, ep: 1, conc-start: 1, conc-end: 8 }
- { tp: 8, ep: 8, conc-start: 16, conc-end: 256 }

# ModelScope integration coverage for TensorRT-LLM. The image version maps to
# NVIDIA/TensorRT-LLM tag v1.3.0rc27 at commit 6e1cc953c071b8a9055b03ef2ae4ee0bc4c645c4.
qwen3-0.6b-bf16-h100-trt-modelscope:
image: nvcr.io/nvidia/tensorrt-llm/release:1.3.0rc27
model: Qwen/Qwen3-0.6B
model-prefix: qwen3-0.6b
runner: cluster:h100-dgxc
precision: bf16
framework: trt
multinode: false
scenarios:
fixed-seq-len:
- isl: 8192
osl: 1024
search-space:
- { tp: 1, conc-list: [1, 4, 16, 32, 64] }

qwen3.5-fp8-h100-sglang-mtp:
image: lmsysorg/sglang:v0.5.19-cu130
model: Qwen/Qwen3.5-397B-A17B-FP8
Expand Down
53 changes: 53 additions & 0 deletions docs/eval-agentx-procedures.md
Original file line number Diff line number Diff line change
Expand Up @@ -162,6 +162,40 @@ python3 -m infx.evals.validate_scores \

Validation resolves the threshold in this order: `models.<prefix>.<task>`, `default.<task>`, then `--min-score` (default `0.85`). By default it checks numeric, non-stderr metrics beginning with `exact_match,`. It fails when a score is below threshold, no metric matches, a requested concurrency is absent, metadata has duplicates/invalid values, any point is marked failed, or result suffixes do not match the manifest. Current floors are authoritative in [`thresholds.yaml`](../infx/evals/thresholds.yaml). See [threshold resolution](../infx/evals/validate_scores.py#L61-L69) and the [validation flow](../infx/evals/validate_scores.py#L174-L302).

### Qwen3-0.6B GSM8K floor

`models.qwen3-0.6b.gsm8k` is **0.60** for both strict-match and flexible-extract.
This is a conservative integration regression floor for the 0.6B checkpoint, not
an expected leaderboard score. The global 0.90 floor remains unchanged.

Keep the standard five-shot chat evaluation, full 1,319-question test split,
`temperature=0`, `top_p=1`, and 5,376 generated-token limit. The
[initial H100 BF16 run](https://github.com/SemiAnalysisAI/InferenceX/actions/runs/35534330261)
at concurrency 64 scored 886/1,319 (0.6717) strict and 895/1,319 (0.6785)
flexible. All responses were nonempty; 86 lacked the required numeric `####`
answer marker, 87 lacked a closing `</think>`, and 61 repeated an identical
nonempty line of at least 30 characters five or more times. These categories
overlap. The nine-answer extraction gain does not explain most errors; inspected
wrong answers also contained arithmetic and reasoning mistakes.

The [Qwen3 technical report, Table 8 and Section 3.3](https://arxiv.org/html/2505.09388v1)
reports 59.59% GSM8K for **Qwen3-0.6B-Base**, using four-shot chain of thought.
That is scale context only: the checkpoint and prompt differ, and it must not be
presented as a comparable five-shot chat baseline. No directly comparable
published baseline was established. The 0.60 floor is an explicit conservative
policy choice supported by that context and the inspected full-split result;
it leaves 7.17 percentage points below the observed strict score (reported
standard error 1.29 points), rather than rounding the observed score into a gate.
Validate the same fixed floor independently at concurrency 32 and 64.

[Qwen's model guidance](https://huggingface.co/Qwen/Qwen3-0.6B#best-practices)
recommends sampling for thinking mode and warns that greedy decoding can repeat.
This integration keeps InferenceX's deterministic protocol for comparability;
the floor does not establish optimal Qwen quality or excuse request failures.
Changing thinking mode, sampling, prompts, or token budget requires a separately
documented policy and fresh full-split evidence. Preserve failed and passing
artifacts, and do not lower this floor in response to a later regression.

A manual combined throughput+eval recipe uploads eval output but the template's automatic score gate is specific to eval-only jobs. Run the validator explicitly for manual or combined runs.

## 6. Collect and inspect eval artifacts
Expand Down Expand Up @@ -354,3 +388,22 @@ Use `scancel` or process termination only with explicit approval and a concrete
- Every backend/frontend and metrics source is represented in live evidence.
- Fast/smoke results are labeled diagnostic. Only the canonical candidate is used for final comparison.
- Workflow and artifact collection conclude green before success is reported.

### Qwen3-0.6B ModelScope cold-cache coverage

The H100 ModelScope recipe starts `trtllm-serve` with the remote model ID and a
new, verified-empty ModelScope cache for every job. Its Hugging Face home/cache
is independently empty. The server performs the download; the recipe does not
predownload the weights or tokenizer. After readiness, the offline resolver must
return a snapshot within the new ModelScope cache, containing weights and the
required tokenizer/config assets. Unexpected Hugging Face cache files fail the
job; a cache version marker and the empty Transformers scaffolding files
`modules/__init__.py` and `modules/hf_remote_code.lock` are allowed. The benchmark
client uses the same resolved tokenizer path.

`modelscope_snapshot_report.json` records the initial empty caches, resolved
snapshot path, file sizes and SHA256 hashes. Eval jobs upload this report with
their raw results; job logs also contain the report. The matrix supplies the
model-specific context ceiling before startup, avoiding a separate hub lookup.
This uses the source-matched TensorRT-LLM 1.3.0rc27 backport documented in the
engine-patch waiver. It does not build the newer TensorRT-LLM PR branch.
40 changes: 40 additions & 0 deletions docs/eval-agentx-procedures_zh.md
Original file line number Diff line number Diff line change
Expand Up @@ -162,6 +162,32 @@ python3 -m infx.evals.validate_scores \

手动的吞吐量+eval 组合 recipe 会上传 eval 输出,但模板的自动分数 gate 专用于 eval-only 作业。对手动或组合运行必须显式执行 validator。

### Qwen3-0.6B 的 GSM8K 下限

`models.qwen3-0.6b.gsm8k` 对 strict-match 和 flexible-extract 均使用 **0.60**。
这是针对 0.6B 检查点的保守集成回归下限,不是排行榜预期分数;全局 0.90 下限保持不变。

保留标准五样本聊天评测、完整的 1,319 道测试题、`temperature=0`、`top_p=1`
和 5,376 个生成 token 的上限。
[首次 H100 BF16 运行](https://github.com/SemiAnalysisAI/InferenceX/actions/runs/35534330261)
在并发 64 下取得 strict 886/1,319(0.6717)、flexible 895/1,319(0.6785)。
所有响应均非空;86 个响应缺少所需的数字 `####` 答案标记,87 个缺少 `</think>`
结束标记,61 个将同一条至少 30 个字符的非空行重复了五次或更多。这些类别存在重叠。
宽松提取仅多判对九题,无法解释大多数错误;抽查的错误答案也包含算术和推理错误。

[Qwen3 技术报告表 8 和第 3.3 节](https://arxiv.org/html/2505.09388v1)
报告 **Qwen3-0.6B-Base** 在四样本思维链设置下的 GSM8K 分数为 59.59%。
该结果仅用于说明模型规模背景:检查点和提示不同,不能视为可比的五样本聊天基线。
目前未找到直接可比的已发表基线。0.60 是结合该背景和完整测试集响应检查作出的保守策略选择;
它比已观察到的 strict 分数低 7.17 个百分点(报告的标准误为 1.29 个百分点),
并非将单次分数取整后作为门槛。应在并发 32 和 64 下独立验证同一个固定下限。

[Qwen 模型指南](https://huggingface.co/Qwen/Qwen3-0.6B#best-practices)
建议思考模式使用采样,并警告贪心解码可能产生重复。为保持可比性,本集成沿用
InferenceX 的确定性协议;该下限不代表 Qwen 的最佳质量,也不豁免请求失败。
改变思考模式、采样、提示或 token 预算,需要另行记录策略并提供新的完整测试集证据。
保留失败和成功的产物,不应因为后续回归而继续降低此下限。

## 6. 收集并检查 eval artifact

收集工作流会下载 `eval_*`,用 `infx/results/collect_eval_results.py` 聚合原始集合,上传 `eval_results_all/agg_eval_all.json`,并将表格写入 step summary([`collect-evals.yml`](../.github/workflows/collect-evals.yml))。
Expand Down Expand Up @@ -352,3 +378,17 @@ gh run cancel <RUN_ID> --repo SemiAnalysisAI/InferenceX
- 每个 backend/frontend 与 metrics source 都在实时证据中有所体现。
- Fast/smoke 结果明确标为诊断用途;只有 canonical candidate 用于最终比较。
- 在报告成功前,工作流与 artifact collection 均已得出 green 结论。

### Qwen3-0.6B ModelScope 冷缓存覆盖

H100 ModelScope 配方在每个任务中使用远程模型 ID 和新建、确认为空的 ModelScope
缓存启动 `trtllm-serve`,并为服务提供独立的空 Hugging Face home/cache。模型由
服务自行下载,配方不会预下载权重或分词器。服务就绪后,离线解析器必须返回新
ModelScope 缓存内部的快照,且其中包含权重和必需的分词器、配置文件。检测到非预期
Hugging Face 缓存文件时任务失败;允许缓存版本标记,以及 Transformers 创建的空文件
`modules/__init__.py` 和 `modules/hf_remote_code.lock`。基准客户端使用同一快照中的分词器。

`modelscope_snapshot_report.json` 记录初始空缓存、最终快照路径、文件大小和 SHA256。
评测任务将该报告与原始结果一起上传,任务日志中也包含报告。矩阵在启动前提供模型
上下文上限,避免另行查询模型仓库。本配方使用 engine-patch 豁免中记录的、与源码
匹配的 TensorRT-LLM 1.3.0rc27 回移补丁,并不构建较新 TensorRT-LLM PR 分支。
49 changes: 49 additions & 0 deletions docs/waiver/3324.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,49 @@
# Inference-engine patch waiver — PR #3324

Filed per [`docs/PR_REVIEW_CHECKLIST.md`](../PR_REVIEW_CHECKLIST.md): this PR patches the pinned
TensorRT-LLM image before serving because the released image predates ModelScope model loading.

## Config covered

- **Master config entry:** `qwen3-0.6b-bf16-h100-trt-modelscope` in
[`configs/nvidia-master.yaml`](../../configs/nvidia-master.yaml)
- **Pinned image:** `nvcr.io/nvidia/tensorrt-llm/release:1.3.0rc27`
- **Image source:** NVIDIA/TensorRT-LLM tag `v1.3.0rc27`, commit
`6e1cc953c071b8a9055b03ef2ae4ee0bc4c645c4`
- **Patch entrypoint:**
[`runners/patch_trtllm_modelscope.py`](../../runners/patch_trtllm_modelscope.py), invoked by
[`benchmarks/single_node/fixed_seq_len/qwen3-0.6b_bf16_h100_trt.sh`](../../benchmarks/single_node/fixed_seq_len/qwen3-0.6b_bf16_h100_trt.sh)

## What is patched

The patcher backports the ModelScope integration from
[SemiAnalysisAI/TensorRT-LLM#2](https://github.com/SemiAnalysisAI/TensorRT-LLM/pull/2) to the two
installed Python modules that participate in this benchmark:

- `tensorrt_llm/llmapi/utils.py` routes full and partial snapshot downloads through ModelScope when
`TRTLLM_USE_MODELSCOPE=true`, while preserving Hugging Face as the default.
- `tensorrt_llm/llmapi/llm.py` loads the tokenizer, generation config, and model config from the
resolved local snapshot rather than retrying the remote Hugging Face model ID.

The backport is source-matched to `v1.3.0rc27`, exact-anchor gated, and idempotent. It refuses an
unknown or partially patched installed source tree. `modelscope==1.40.1` and
`modelscope-hub==0.4.3` are installed in the H100 container before the patch is applied.

## Why the unmodified upstream image cannot run this benchmark

TensorRT-LLM `1.3.0rc27` resolves remote model IDs exclusively with `huggingface_hub`. It has no
ModelScope switch or downloader and subsequently loads tokenizer and configuration files from the
original remote ID. Therefore the stock image cannot validate TensorRT-LLM model loading from
ModelScope for `Qwen/Qwen3-0.6B`; installing the optional ModelScope dependency alone does not change
that behavior.

## Upstream PR

- https://github.com/SemiAnalysisAI/TensorRT-LLM/pull/2

## Removal plan

Once an NGC TensorRT-LLM release includes the ModelScope integration, update
`qwen3-0.6b-bf16-h100-trt-modelscope` to the first matching release image and verify its source tag.
In the same PR, remove `runners/patch_trtllm_modelscope.py`, remove its invocation and runtime package
installation from the benchmark script, and delete this waiver.
3 changes: 3 additions & 0 deletions infx/evals/thresholds.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -50,6 +50,9 @@
"minimaxm2.5": {
"gsm8k": 0.92
},
"qwen3-0.6b": {
"gsm8k": 0.60
},
"qwen3.5": {
"gsm8k": 0.94
}
Expand Down
14 changes: 14 additions & 0 deletions perf-changelog.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -8455,3 +8455,17 @@
description:
- "Update B200 vLLM AgentX to DSpark6 and a new image with TP8 and DEP8 configurations."
pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3274

- config-keys:
- qwen3-0.6b-bf16-h100-trt-modelscope
description:
- "Add H100 TensorRT-LLM 1.3.0rc27 coverage for Qwen3-0.6B in BF16, resolving the model through ModelScope with a source-matched runtime backport of SemiAnalysisAI/TensorRT-LLM#2."
- "为 Qwen3-0.6B BF16 添加 H100 TensorRT-LLM 1.3.0rc27 覆盖,通过 ModelScope 解析模型,并使用与镜像源码匹配的 SemiAnalysisAI/TensorRT-LLM#2 运行时回移补丁。"
pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3323

- config-keys:
- qwen3-0.6b-bf16-h100-trt-modelscope
description:
- "Exercise ModelScope cold downloads in every H100 Qwen3-0.6B job using empty, isolated hub caches and record snapshot hashes before benchmarking or evaluation."
- "每个 H100 Qwen3-0.6B 任务均使用独立空缓存测试 ModelScope 冷下载,并在基准或评测前记录模型文件哈希。"
pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3324
Loading
Loading