[PowerX] enable NVL72 smart provisioning on measured compute-module power / 为 NVL72 启用基于实测计算模块功耗的智能预配 - #1190
edwingao28 wants to merge 31 commits into
Conversation
…t time PowerX read chip telemetry by downloading and parsing gpu_metrics GitHub artifacts on every page view, and lost the data once GitHub's 90-day artifact retention expired. This moves that work to ingest time. - Migration 016 adds gpu_metric_series (one row per artifact CSV), gpu_metric_samples (full-resolution per-GPU samples), gpu_metric_gpu_stats (per-GPU min/max/mean/median/p95/p99/stddev digest), and benchmark_result_gpu_metrics (point <-> series links). - A pure nvidia-smi/amd-smi CSV parser plus artifact discovery and an idempotent upsert (same CSV hash refreshes links only; a changed CSV replaces samples and digest in one transaction). - CI ingest links gpu_metrics_<suffix> next to bmk_<suffix> using the same pairing rule as server logs. - New backfill CLI: bun run admin:db:backfill-gpu-metrics --all --yes (bounded by GitHub retention; the GCS backup does not mirror gpu_metrics). - /api/gpu-metrics serves the stored digest first and falls back to live GitHub artifacts for runs that are not ingested yet. - New /api/v1/gpu-metrics-point?id=N and a PowerX tab on the per-point detail page showing the telemetry recorded while that point ran. - Shared benchmark-result lookup extracted from the server-log backfill. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…lemetry charts Add a Points / Rolling average control to the PowerX telemetry charts on both the explorer page and the per-point PowerX tab, with a 10/30/60/300 s window selector. The average is a centered time-window mean per chip (pure helper in telemetry-smoothing.ts, unit-tested for empty input, single sample, irregular timestamps and inclusive window edges), so it follows elapsed time rather than sample count. Averaged mode draws smooth lines and keeps the sample circles as invisible hover targets so the tooltip and crosshair still work. The per-point PowerX tab gains the same chip legend as the explorer so individual chips can be hidden and restored, scoped to the selected series. The chart's t=0 now comes from the whole series rather than the visible chips so hiding a chip does not shift the time axis. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Add a Lines control (Per chip / Mean of chips / Both) next to the display-mode toggle on both PowerX surfaces. The mean line averages the currently visible chips at each timestamp of the longest series, aligning the other chips by nearest sample within one estimated poll interval (median gap), so the few ms of skew between nvidia-smi rows and the occasional dropped row do not interpolate or misalign. It is drawn in the foreground color at a heavier stroke, labelled in a key row under the caption, and its tooltip reports how many chips contributed. The rolling-average mode smooths the mean line too. The shared line layer gains an optional per-series getStrokeWidth so one keyed join can hold both the chip lines and the heavier mean line. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Add an "Overlay decode throughput" switch (default off) to the per-point PowerX tab. When on, the point's decodeTps series from the trace server metrics is drawn over the telemetry in violet on its own right-hand axis, with a key-row entry and a row in every tooltip giving the nearest decode rate. The trace's startNs is an epoch-ns wall-clock timestamp, so the two series are aligned by absolute time; when a trace carries no wall-clock start the overlay falls back to the telemetry start and the note under the switch says so. Points without server metrics get an "unavailable" note instead of an empty overlay. The chart takes the overlay as an optional prop and renders it in a custom D3 layer that redraws on zoom and removes its own axis when toggled off. In rolling-average mode the overlay is smoothed with the same time window as the chip lines so both curves describe the same interval. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
At 1 s cadence the raw per-chip points read as noise, so both PowerX surfaces now open in rolling-average mode (30 s window); Points remains one click away in the Display control. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…rlay Replace the decode-only overlay switch on the per-point PowerX tab with a single-select menu of server metrics. The menu lists only the series the point's trace actually reports (decode/prefill throughput, KV and host KV cache utilization, prefix cache hit rate and hits, queue depth), defaults to "None", and is disabled when the trace has no server metrics. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… digest Register GET /api/v1/gpu-metrics-point as a page-owned BFF in the API route catalog and refresh the /api/gpu-metrics digest after the DB-first digest read landed in d5e90c4. Neither route is part of the public API reference, so the OpenAPI document and human reference are unchanged; the exclusion reason for /api/gpu-metrics now describes the stored-digest-then-live-artifact behavior. 中文:在 API 路由目录中登记 GET /api/v1/gpu-metrics-point(页面专用 BFF),并在 d5e90c4 引入 DB 优先读取后刷新 /api/gpu-metrics 的摘要。两条路由都不属于公开 API 参考,OpenAPI 文档与人工参考无需改动;/api/gpu-metrics 的排除说明改为描述 "已入库摘要优先、否则回退实时制品"的行为。
The pipeline doc blamed an artifact-name filter in the GCS backup. The backup keeps every artifact name; it only mirrors scheduled and main-branch runs, so PR sweeps and manual dispatches - nearly every telemetry-bearing run - are never copied, and our own GCS reader ignores non-bmk_/server_logs_ objects anyway. Also note that the multinode template uploads no gpu_metrics_ artifact, so multinode and disaggregated points never get a telemetry series. 中文:修正数据管线文档中 gpu_metrics telemetry 回填受 90 天限制的原因说明。 GCS 备份并不按 artifact 名过滤,而是只镜像定时任务和 main 分支的运行,PR sweep 与手动触发的运行从不被复制;应用自身的 GCS reader 也只读取 bmk_/server_logs_ 对象。同时注明多节点模板不上传 gpu_metrics_ artifact,多节点与 disagg 数据点 不会有 telemetry 序列。
30 PGlite-backed tests over the PowerX telemetry ingest code that had none. - queries/gpu-metrics: null on unknown run and on a run with zero series, full run payload shape, includeSamples=false, latest-attempt selection, multinode point fan-out, availability map, and a characterization of the `?? 0` sample defaulting (a dropped reading reports 0 W). - lib/benchmark-result-lookup: natural-key match scoped to run attempt and config, offload_mode exact match vs. unique fallback vs. ambiguous drop, agentic null isl/osl, and artifact JSON mapping. - backfill-gpu-metrics: candidate selection driven through the real CLI main against a migrated PGlite database, with GitHub listing stubbed. Pins the 90-day default --since, the resume hazard (runs with stored series are skipped without --force), --run bypassing both filters, --from-run, --limit, and that --shard-count/--shard-index do not partition the set. No production source was changed. 中文:为此前完全没有测试的 PowerX 遥测 ingest 代码补充 30 个基于 PGlite 的单元 测试,覆盖 gpu-metrics 读取查询、backfill 候选筛选与 benchmark 点位匹配。要点: 读取查询在 run 不存在或没有 series 时返回 null,多节点点位返回全部关联 series, 并固化了 `?? 0` 默认值行为(采集中断的样本会显示为 0 W,与真实的 0 W 无法区分); 点位匹配按 run attempt 与完整 config 自然键限定,并区分 offload_mode 精确匹配、 唯一回退与歧义丢弃;backfill 测试通过真实 main 驱动候选筛选,固化默认 90 天 --since 窗口、--force 之前会跳过已有 series 的 run(断点续跑隐患)、--run 同时 绕过日期与 series 过滤,以及 --shard-count/--shard-index 实际不分片的现状。 未修改任何生产代码。
…timate Pin inferencex_power_model 963ead8b (feat/gb200-nvl72-rack-model) and emit gb200/gb300 rack profiles plus 204 Python parity cases; the x86 chassis profiles and their 292 cases regenerate byte-identically. Port the rack evaluation to TypeScript with the source rounding order (rack AC rounded before PUE). modelSystemPower admits NVL72 rows as compute trays (one host, four GPUs, two Grace sockets) whose module or GPU-board + Grace-socket watts are measured: cpu_power_valid=1 and the Grace-side keys are required, the Grace side is never modelled, DLC PUE 1.1 applies once at rack AC, a partial tray extrapolates only the GPU-board share, and the result carries topologyBasis 'nvl72-trays', measuredBasis, sensorKind and the new cpu-telemetry reason. x86 rows are unchanged. 中文:固定 inferencex_power_model 963ead8b(feat/gb200-nvl72-rack-model), 生成 gb200/gb300 机架 profile 和 204 条 Python 对照用例,x86 机箱 profile 及 其 292 条用例逐字节不变。TypeScript 按原模型的舍入顺序移植机架计算(机架交流 功率先舍入再乘 PUE)。modelSystemPower 将 NVL72 行按计算 tray(单主机、4 张 GPU、2 个 Grace socket)接纳,输入为实测模块功耗或 GPU 板卡 + Grace socket 功耗:要求 cpu_power_valid=1 及 Grace 侧指标,Grace 侧从不建模,DLC PUE 1.1 在机架交流侧只应用一次,部分分配的 tray 仅外推 GPU 板卡份额,结果新增 topologyBasis 'nvl72-trays'、measuredBasis、sensorKind 和 cpu-telemetry 原因。x86 行为不变。
`gh api …/runs/<id>/artifacts` returns HTTP 404 once a run is deleted or purged, and `retryArtifactOperation` spent the full backoff schedule on it before `backfill-gpu-metrics --all --dry-run` aborted on the first such run (34533943809, 18 runs into a 224-run sweep). Classify that 404 as `WorkflowRunNotFoundError`, a `NonRetryableArtifactError` the retry helper rethrows at once, and let both gpu-metrics and server-log backfills report the run as gone and continue. The gpu-metrics header also drops the wrong "GCS mirrors only bmk_/server_logs_" retention explanation. 中文:`gh api` 在 run 被删除或清理后返回 HTTP 404,之前 `retryArtifactOperation` 会把完整退避重试跑完,`backfill-gpu-metrics --all --dry-run` 在 224 个候选中 第 19 个 run(34533943809)处直接崩溃。现将该 404 归类为不可重试的 `WorkflowRunNotFoundError`,重试助手立即抛出,gpu-metrics 与 server-log 两个 backfill 记录该 run 已不存在并继续;同时修正 gpu-metrics 头注释中错误的 GCS 镜像保留说明。
…nd name the power basis Planning kW/GPU now admits nvl72-trays estimates whose trays are all fully measured next to the eight-GPU single-node chassis; partial trays stay rejected. Two frontier knots must share the same measured basis and sensor kind or the bracket stays unavailable. Modeled rows carry powerSource (topology, measured basis, sensor kind, PUE, pinned profile path, revision and source SHA-256); the bar tooltip, a caption line and three CSV columns name it per row. English strings for x86 rows are byte-identical. 中文:规划 kW/GPU 除八卡单节点机箱外,接受全部 tray 完整实测的 NVL72 估算, 部分 tray 仍被拒绝;两个前沿数据点的实测口径或传感器类型不同时保持不可用。 估算行携带 powerSource(拓扑、实测口径、传感器类型、PUE、固定 profile 路径、 版本和源文件 SHA-256),柱形提示、标注行和三列 CSV 逐行标出;x86 行英文字符串 逐字节不变。
…onstants, ingest and the API reference avg_cpu_socket_power_w, avg_total_cpu_power_w, total_cpu_energy_j, avg_total_module_power_w and total_module_energy_j join MEASURED_POWER_METRIC_KEY_LIST (withheld with power_valid=0 like every measured key); cpu_power_valid joins the contract discriminators and is normalized as a verdict independent of power_valid. extractPowerAudit keeps a bounded power_audit.cpu block matching the consumer's audit_summary. The API registry documents the six keys, the cpu audit schema, and a bilingual measured-power note and example; no route catalog digest changes. 中文:五个 CPU 侧实测指标加入 MEASURED_POWER_METRIC_KEY_LIST(与其他实测键一样 在 power_valid=0 时移除);cpu_power_valid 加入契约字段并按独立于 power_valid 的验证结论归一化;extractPowerAudit 保留与消费端 audit_summary 一致的有界 power_audit.cpu。API 文档新增六个键的说明、cpu 审计 schema 以及中英文 measured-power 说明与示例;路由目录摘要无需刷新。
…ress spec profit-fixtures gains an NVL72 SKU (one four-GPU tray with CPU-side keys and the module sensor) kept out of PROFIT_SKUS so existing bar counts hold. The new case checks the control stays hidden while the gate is locked, then prices the tray on Measured + modeled, names the measured module basis and DLC PUE 1.1 in the caption, and leaves GB300 unavailable. 中文:profit-fixtures 新增 NVL72 SKU(单个四卡 tray,含 CPU 侧指标与模块传感器), 不加入 PROFIT_SKUS 以保持现有柱形数量。新用例验证功能开关锁定时控件隐藏,解锁后 按实测 + 估算为该 tray 定价,标注行写明实测模块口径与液冷 PUE 1.1,GB300 保持不可用。
Measured input and admission rules, the modeled residual table with the UNVERIFIED parameters and their ranges, the shelf overflow bound, and the Profit Estimator gate rules, in English and in the 中文说明 section. The Profit Estimator paragraphs now name the per-hardware PUE policy, the same-basis knot rule and the CSV provenance columns in both languages. 中文:新增 NVL72 机架估算一节:实测输入与接纳条件、含 UNVERIFIED 参数及范围的 建模残差表、电源架容量上限和利润估算器门槛规则,中英文同步;利润估算器段落 双语补充按硬件取值的 PUE 策略、同口径数据点规则和 CSV 出处列。
…yment mean Aggregate multinode producers (Kimi K3 B200 dynamo-vLLM TP8/PP2, H200 vLLM TP16×2) emit no per-worker telemetry, so `modelSystemPower` rejected every such row as `topology` and the smart-provision basis showed no NVIDIA K3 SKU. Add a `uniform-hosts` topology basis for non-disaggregated multinode rows without a worker array: the telemetry GPU count must fill whole eight-GPU hosts and match TP×PP×PCP×replicas, and each chassis is modeled at the deployment mean; rows that carry per-worker telemetry keep the exact worker-hosts path and disaggregated rows still require it. `planningKwPerGpu` now admits every fully measured multi-chassis estimate instead of only one single-node chassis. Tooltip copy (en/zh) names the uniform-hosts assumption and the system-power doc records the admission rule. 中文:聚合多节点的采集端(Kimi K3 B200 dynamo-vLLM TP8/PP2、H200 vLLM TP16×2)不输出逐 worker 功耗,`modelSystemPower` 一律按 `topology` 拒绝, smart provision 里因此没有任何 NVIDIA K3 SKU。新增 `uniform-hosts` 拓扑基准: 非 disagg 多节点且无 worker 数组时,要求 telemetry GPU 数填满整数个八卡主机并 等于 TP×PP×PCP×副本数,每个机箱按部署平均功耗建模;带逐 worker 功耗的行仍走 精确的 worker-hosts 路径,disagg 行仍需逐 worker 数据。`planningKwPerGpu` 改为接受所有机箱均完整实测的估算。tooltip 中英文说明该假设,系统功耗文档同步。
… one rack - ingest withholds the CPU-side keys on cpu_power_valid != 1 and GPU-side keys on power_valid = 0; mapper tests cover all four verdict combinations; the power manifest carries cpu_power_valid - exporter shares defaultSystemPue with the dashboard (1.3 chassis, 1.1 DLC NVL72); NVL72 rows carry the rack profile, measured basis, sensor kind and CPU-side inputs; x86 rows byte-identical - chart tooltip names NVL72 compute trays and the measured basis instead of eight-GPU chassis (en/zh) - heterogeneous measured trays fold into one rack at their mean; the shelf curve is evaluated once at rack DC load, matching gb200_nvl72_rack_power; estimateRackPower and parity cases unchanged 中文:摄取按 cpu_power_valid 独立清除 CPU 侧指标,GPU 侧仍按 power_valid,mapper 测试覆盖四种组合,功耗清单附带 cpu_power_valid;导出器与仪表板共用 defaultSystemPue(机箱 1.3、液冷 NVL72 1.1),NVL72 行补充机架 profile、实测口径、传感器类型与 CPU 侧输入,x86 行逐字节不变;图表提示改用 NVL72 计算 tray 措辞并标出实测口径(中英文);多 tray 按均值折算为整机架,电源架曲线只在机架直流负载处求值一次,与 Python 模型一致,estimateRackPower 与对照用例不变。
Multinode InferenceX jobs upload no `gpu_metrics_` artifact; their per-GPU 1 Hz power lives in `power_audit_<suffix>/LOGS/power/samples.csv`, one deployment-wide CSV from srt-slurm's `dcgm-power` collector. Every Kimi K3 NVIDIA point therefore had an empty PowerX tab while the data sat in GitHub. Accept the bundle as a fallback telemetry artifact: `gpuMetricsArtifactSuffix` recognises `power_audit_`, discovery (CI ingest) and backfill pairing prefer a `gpu_metrics_` sibling and use the bundle only when none exists, and `prepareGpuMetricsArtifact` regroups `samples.csv` by hostname into one power-only series per host (`file_name` `LOGS/power/samples.csv#<host>`, manifest as context sidecar, GPU UUIDs as identity). Power-only series made a latent reader defect visible: `toSampleRow` coerced null clocks, temperature and utilization to 0, so the UI offered and plotted fabricated flat-zero metrics. The five non-power core fields are now optional end to end (reader, `GpuMetricRow`, anomaly checks, chart and correlation points, tooltip, correlation default). The PowerX tab labels the collector from the recorded producer instead of `nvidia-smi`, and takes the point's hardware key for the TDP line because multinode artifact names are hash-truncated. Docs record the adapter. Backfilled into the branch DB: K3 B200 34674595026 (7 points, 14 series), GB300 34873998796 (5 of 11; the six disaggregated jobs report `multinode_power_contract_missing`), H200 34744300699 (10 points, 40 series). Not covered here: the explorer's live GitHub fallback for not-yet-ingested runs still reads `gpu_metrics_` only. 中文:多节点 InferenceX 任务不上传 `gpu_metrics_` artifact,每 GPU 1 Hz 功耗数据 在 `power_audit_<suffix>/LOGS/power/samples.csv`(srt-slurm `dcgm-power` 采集, 整个部署一个 CSV)里,Kimi K3 NVIDIA 数据点的 PowerX 标签页因此一直为空。 本次将该 bundle 作为回退 telemetry artifact:suffix 识别 `power_audit_`, CI 入库发现与 backfill 配对优先使用 `gpu_metrics_`,仅在没有时使用 bundle; `prepareGpuMetricsArtifact` 按 hostname 拆成每主机一条仅含功耗的序列。 同时修正读取端把空的时钟/温度/利用率读成 0 的问题(五个字段全链路改为可选), PowerX 标签页的 Collector 改为显示真实采集器,TDP 参考线改用数据点的硬件 key。 已回填 branch DB:K3 B200、GB300(5/11,disagg 任务无功耗合约)、H200。 未覆盖:explorer 对未入库 run 的 GitHub 实时回退仍只读 `gpu_metrics_`。
…with multinode chassis planning Brings in the base's deleted-run skip for artifact backfills, the `uniform-hosts` topology basis (aggregate multinode chassis modeled at the deployment mean) and the per-host `power_audit_` telemetry ingest, and reconciles them with the NVL72 tray path: - `modelSystemPower`: the `uniform-hosts` branch keeps the base's x86 semantics unchanged and now fills the shared `MeasuredUnit` list; it is chassis-only (`!rack`), so an aggregate multinode NVL72 row without a per-worker array stays `topology`-unavailable rather than inferring a tray count from the GPU total. The supported-estimate union carries `'single-node' | 'worker-hosts' | 'uniform-hosts'` beside `'nvl72-trays'`. - Profit planning gate: every fully measured estimate is admitted (`chassisBasis: 'full'`); `nvl72-trays` estimates keep the measured-basis and sensor-kind source label, every chassis topology is labeled `chassis`. - Tooltip: the uniform-hosts note renders in the same position as on the base, alongside the NVL72 tray copy; English and Chinese x86 strings are unchanged. - Docs and tests from both sides kept; new tests pin the chassis-only uniform-hosts decision and the `chassis` label for multi-host sources. 中文:合并 feat/powerx-db-ingest(跳过 GitHub 已删除的 run、聚合多节点机箱按 部署平均功耗建模的 `uniform-hosts` 拓扑基准、按主机拆分的 `power_audit_` telemetry 入库),并与 NVL72 tray 路径对齐:`uniform-hosts` 分支保持 x86 语义不变、仅对机箱硬件生效,无 worker 数组的 NVL72 聚合多节点行仍按 `topology` 不可用;Profit Estimator 规划闸门接受所有完整实测的估算, NVL72 tray 保留实测基准与传感器标签,机箱拓扑统一标为 `chassis`;tooltip 的 uniform-hosts 说明位置与基线一致,x86 中英文文案不变;两侧文档与测试 均保留,并新增测试固定上述决策。
… array The Kimi K3 GB200 aggregate producer (dynamo-vLLM TP16, sixteen GPUs on four trays) emits no `workers` array, so `modelSystemPower` rejected every such row as `topology` and the tray path never reached the Profit Estimator. Generalize the base's `uniform-hosts` branch from eight-GPU chassis to the unit size of the hardware: on GB200/GB300, a non-disaggregated multinode row without a worker array is `gpuCount / 4` compute trays, each fed the deployment mean (module total per tray when the module keys are present, otherwise GPU-board plus Grace-socket watts per tray, exactly as on the worker-hosts tray path), with `chassisBasis: 'full'`, `topologyBasis: 'nvl72-trays'` and the usual measured basis and sensor kind. A GPU total that does not fill whole trays is `gpu-count`; the tray count must agree with the Grace-socket count recovered from the CPU-side keys and, when the CPU leg recorded it, with `power_audit.cpu.observed_sockets`, otherwise `cpu-telemetry`. The x86 uniform-hosts path, the worker-hosts tray path and every English string are unchanged. Tests cover module and Grace-socket bases, both socket mismatches, the uneven count and the Profit Estimator source label; the system-power doc records the rule in English and Chinese. 中文:Kimi K3 GB200 聚合采集端(dynamo-vLLM TP16,16 张 GPU 分布在 4 个 tray) 不输出 `workers` 数组,`modelSystemPower` 一律按 `topology` 拒绝,tray 路径无法进入 Profit Estimator。现将基线的 `uniform-hosts` 分支从八卡机箱推广到硬件的单元大小: GB200/GB300 上无 worker 数组的非 disagg 多节点行按 GPU 总数 ÷ 4 推算 tray 数,每个 tray 取部署平均值(有模块指标时用模块总功耗 ÷ tray 数,否则用 GPU 板卡 + Grace socket 每 tray 功耗,与 worker-hosts tray 路径一致),`chassisBasis: 'full'`、 `topologyBasis: 'nvl72-trays'`,实测口径与传感器类型不变。GPU 总数无法填满整数个 tray 时为 `gpu-count`;tray 数须与 CPU 侧指标推算的 Grace socket 数一致,且在 CPU 采集记录了 `power_audit.cpu.observed_sockets` 时与之一致,否则为 `cpu-telemetry`。 x86 uniform-hosts 路径、worker-hosts tray 路径及所有英文文案不变。测试覆盖模块与 Grace socket 两种口径、两类 socket 不一致、非整数 tray 数以及 Profit Estimator 的 来源标签;系统功耗文档中英文同步。
…-side comments Three comments referenced a task-tracker ticket that does not exist in this repository; point them at the InferenceX producer contract instead. No code change. 中文:三处注释引用了仓库中不存在的本地工单,改为指向 InferenceX 的生产端契约文档;无代码变更。
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
…timeline / 新增功耗边界、标尺、对比序列与功耗时间线 (#1177) * feat: share placed perf rulers through the i_rulers URL param Rulers placed on the primary Inference chart now live in a provider-owned store and serialize into the share link as `i_rulers` (`isoX|curveA|curveB;...`). Opening such a link restores the rulers once both curves have rendered, switches the tool on, clears them when the x-axis mode changes the rendered chart, and ignores malformed values. Overlay-run curves are addressable too. 中文:图表上放置的 Perf Ruler 现在保存在 provider 级的 store 中,并以 `i_rulers` 参数写入分享链接;打开链接后会在两条曲线都渲染完成时恢复标尺并自动 开启该工具,切换 X 轴模式导致图表更换时清除,格式错误的值将被忽略。非官方运行 的 overlay 曲线同样可被引用。 * feat: derive provisioned and modelled power boundary fields Adds `lib/power-basis.ts` with the B2–B4 boundary maths and wires it into `buildDerivedChartFields`: GPU provisioned (registry TDP), utility provisioned (registry all-in kW per GPU) and utility modelled (measured GPU power carried through the chassis model to the utility meter, PUE applied once). Energy variants count every allocated GPU per output token; values are omitted, never estimated, when a spec, throughput or telemetry input is missing. Registers the six `y_*` metric keys with bilingual labels and axis explanations. 中文:新增 `lib/power-basis.ts`,实现 B2–B4 功耗边界的计算并接入 `buildDerivedChartFields`:GPU 额定(注册表 TDP)、全电源配置(注册表每 GPU all-in kW)以及数据中心建模(GPU 实测功耗经机箱模型推算至市电侧,PUE 仅计入一 次)。能耗指标按每输出 token 计入全部已分配 GPU;缺少规格、吞吐量或遥测输入时 省略数值而不做估算。同时注册六个 `y_*` 指标键及中英文标签与坐标轴说明。 * feat: add a Boundary control to the gated measured power groups The ↑↑↓↓-gated Measured Power / Measured Energy controls gain a Boundary select (GPU measured, GPU provisioned, utility provisioned, utility modelled). The boundary rides on `i_metric`, so existing share links keep working and a shared boundary view renders even while the gate is locked. Boundary watt gauges draw the same upper envelope as measured watts and stay envelope-locked under Optimal Only; tied maxima now stay on the envelope so a flat TDP series spans its tested range. The availability panel explains per point why a boundary value is withheld, and the chart caption states each boundary's formula and assumptions. 中文:↑↑↓↓ 门控的 Measured Power / Measured Energy 控件新增“功耗边界”下拉 (GPU 实测、GPU 额定、全电源配置、数据中心建模)。边界随 `i_metric` 传递,现有 分享链接不受影响,门控锁定时分享的边界视图仍可渲染。边界功率指标与实测功率一 样绘制上包络线,并在“仅最优”下保持包络锁定;包络线现在保留并列最大值,使平坦 的 TDP 序列覆盖完整测试范围。可用性面板逐点说明边界值缺失的原因,图表说明列出 各边界的公式与假设。 * Shorten chip config comparison label / 缩短芯片配置对比文案 (#1168) * fix: shorten chip config comparison label Shorten the comparison selector placeholder in the inference chart and profit estimator, with matching Simplified Chinese copy.\n\n中文:缩短推理图表和利润估算器中的芯片配置对比选择器占位文案,并同步更新简体中文。 * test: update chip config placeholder assertions Update the profit estimator E2E assertions for the shortened English and Simplified Chinese placeholders.\n\n中文:更新利润估算器端到端测试断言,使其匹配缩短后的英文和简体中文占位文案。 * feat: add a measured power timeline display to the gated power group Third value of the Measured Power Display control (`timeline`, `y_measuredPowerTimeline`). The key aliases the measured average so the point set, table view, availability panel and share link are unchanged; ChartDisplay swaps ScatterGraph for the new PowerTimeline, which joins each validated point to its `gpu_metrics_<RESULT_FILENAME>` artifact via `power_audit.source`, fetches one prefixed request per workflow run, and draws one-second per-GPU power over the whole job with the validated window emphasized, TDP references per hardware, an opt-in all-in line, wall-clock / elapsed axes, per-GPU lines, overlay-run colours, and an explicit list of undrawn configs by reason. `/api/gpu-metrics` gains `series=power` (compact one-second buckets, ~1 MB instead of ~27 MB raw for a 25-config run) and `prefix=` to skip other models' artifacts; downloads run four at a time. 中文:在受门控的实测功耗组中新增「时间线」显示方式 (`y_measuredPowerTimeline`)。该指标键复用实测平均功耗,因此数据点集合、 表格视图、可用性面板和分享链接保持不变;ChartDisplay 在该模式下用新的 PowerTimeline 替换散点图:按 `power_audit.source` 将每个有效数据点关联到 对应的 `gpu_metrics_<RESULT_FILENAME>` 产物,按工作流运行各发起一次带前缀的 请求,绘制整个基准测试任务期间每 GPU 的逐秒功耗,突出显示有效测量窗口, 按硬件绘制 TDP 参考线(全电源配置线可选开启),支持实际时刻 / 相对起点两种 时间轴、每 GPU 单独曲线、非官方运行的配色,并按原因列出未绘制的配置。 `/api/gpu-metrics` 新增 `series=power`(一秒分桶的紧凑序列,25 个配置约 1 MB,而原始行约 27 MB)和 `prefix=` 参数以跳过其他模型的产物;下载并发数为 4。 * fix(landing): hide DeepSeek V4 Pro and Qwen3.8 Flash Next / 首页隐藏两个模型 (#1169) * fix(landing): hide DeepSeek V4 Pro and Qwen3.8 Flash Next 中文:从首页隐藏 DeepSeek V4 Pro 和 Qwen3.8 Flash Next,保留比较页及仪表板行为,并添加双语回归测试。 * test(landing): update model row and badge expectations 中文:更新首页模型行数及 NEW 标记数量的测试预期。 * fix: temporarily label DSpark run 35166686551 as MoRI UMBP SGLang (#1170) 中文:为 DSpark 运行 35166686551 添加三周有效的 MoRI UMBP SGLang 显示名称,保留其他运行和原始数据,并测试到期边界。 Co-authored-by: Perplexity Computer <computer@users.noreply.github.com> * docs: require exact AI model disclosure (#1171) 中文:要求 PR 描述列出确切的 AI 模型名称/版本及工作内容,更新 agent 指引并新增双语 PR 模板。 * docs: expand bilingual glossary from Rubin AgentX article (#1172) Add 12 terms, update 10 definitions, and test bilingual coverage and metric boundaries. 中文:根据 Rubin AgentX 文章扩充双语术语表,新增 12 个词条、更新 10 个定义,并补充双语覆盖与指标边界回归测试。 * feat: redirect /xvideo and /xxvideo aliases to /video (#1173) Add a small alias table (`video-alias-redirects.ts`) wired into `next.config.ts` so `/xvideo` and `/xxvideo` (plus the `/zh` siblings) 308-redirect to the canonical `/video` route. Subpaths and query strings carry through via `:path*`. Unit test mirrors the inference-model alias redirect test. 中文:新增 `video-alias-redirects.ts` 别名表并接入 `next.config.ts`,使 `/xvideo` 与 `/xxvideo`(含 `/zh` 版本)以 308 永久重定向到规范路由 `/video`,子路径和查询参数通过 `:path*` 原样透传。单元测试与推理模型 别名重定向测试保持一致。 * fix: refresh the gpu-metrics route digest after formatting The pre-commit formatter rewrote src/app/api/gpu-metrics/route.ts after its catalog digest was computed, so the API route catalog guard failed on a clean checkout. Documentation and classification are unchanged. 中文:提交前的格式化工具在计算摘要之后重写了 gpu-metrics 路由文件,导致 API 路由目录校验在干净检出时失败;此处仅刷新 SHA-256 摘要,文档与分类不变。 * feat: overlay power boundary and role comparison series via i_pcompare Add a Compare control to the gated Measured controls that overlays sibling series on the selected metric without changing it: every power boundary (GPU measured, GPU provisioned, utility provisioned, utility modeled) or the prefill / decode worker pools. useChartData and the ?unofficialrun= overlay processor both append one clone per sibling to every base point (same x, y from the sibling's field, powerVariant set) through utils/power-compare.ts; base points are untouched, so charts without a comparison render byte-identically. ScatterGraph keys series on hwKey + precision + variant, draws siblings in the hardware colour (overlay runs keep the run colour) with a per-variant dash and their own frontier, dims clone points, adds legend rows with line swatches that toggle and highlight a series across hardware, and labels the series in tooltips, the Table and the CSV export. Clones stay out of best-per-SKU ranking, tier counts, the legend points table, the availability panel and the date-comparison GPUGraph. For PowerX Figure 7 the prefill pool's J per input token is carried onto the output-token axis by the served input:output ratio (utils/role-energy.ts, reconstructedPrefillJPerOutputToken) so it stacks against the decode pool's J per output token. The comparison rides on the new i_pcompare URL parameter (boundaries | roles; empty = off) and pauses with a hint on metrics without a common axis instead of being cleared. 中文:在受门控的 Measured 控件中新增“对比”选项,可在所选指标上叠加同源系列而不改变 该指标:四种功耗边界(GPU 实测、GPU 额定、全电源配置、数据中心建模)或预填充 / 解码 worker 池。useChartData 与 ?unofficialrun= overlay 处理路径均通过 utils/power-compare.ts 为每个基础数据点追加对应系列的克隆点(x 相同、y 取自对应字段、标记 powerVariant), 基础点保持不变,未开启对比的图表渲染结果与之前完全一致。ScatterGraph 以 hwKey + 精度 + 变体作为系列键,同色不同虚线绘制,图例新增可切换、可高亮的线型行, tooltip、表格与 CSV 标注系列;克隆点不参与 best-per-SKU、功耗等级计数、图例数据点表、 可用性面板及日期对比 GPUGraph。针对 PowerX 图 7,utils/role-energy.ts 按实际服务的 输入/输出 token 比将预填充池的每输入 token 能耗折算到每输出 token 轴 (reconstructedPrefillJPerOutputToken),与解码池并列。对比模式由新 URL 参数 i_pcompare 承载(boundaries | roles,空为关闭),在没有公共坐标轴的指标上暂停并提示, 而不是被清除。 * Add Qwen3.8-27B / 新增 Qwen3.8-27B (#1176) Register the qwen3.827b and qwen3.827beager DB keys as the experimental Qwen3.8-27B and Qwen3.8-27B-Eager dashboard models (InferenceX#3260, #3263): constants, ETL normalizer path, Model enum and MODEL_CONFIG, compare slugs, compare SSR known models, refreshed constants digest, AGENTS.md parameter row, and the pinned registry counts. Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> * [Cache Reuse] add prefix-cache tier share tab and agentic entry links / 新增前缀缓存复用页及智能体入口链接 (#1175) * feat: add Prefix Cache Reuse tab with agentic entry links Footer-only /cache-reuse (+ /zh) stacks per-concurrency shares of prompt tokens served from HBM, the host tier, or recomputed, for one configuration. Host tier is the CPU-offload rate, falling back to the external rate; never summed. TensorRT-LLM offload rows draw one combined segment. Fixed sequences keep the scenario selector and explain that no cache tiers are recorded. Unofficial runs on the same hardware render as outlined series. c_cfg seeds the configuration from a share link. Agentic chart footer and point summary link into the tab. 中文:新增 footer 级 /cache-reuse(含 /zh)页面,按并发数堆叠展示单一配置的 prompt token 来源占比:HBM 缓存、主机层、未复用。主机层取 CPU offload 命中率, 缺失时改用外部缓存命中率,两者不相加;TensorRT-LLM offload 行绘制为单一合并段。 固定序列保留场景选择器并说明无缓存分层数据。同硬件的非官方运行以描边系列叠加。 c_cfg 参数用于分享链接定位配置。智能体图表页脚和数据点摘要新增入口链接。 * fix: pin point links to one precision and drop the empty official series The point-detail link now writes i_prec from the point, so the tab keys its groups by bare hardware key and c_cfg matches instead of falling back to another sweep. buildCacheReuse lists an official series only when official rows exist, so an overlay-only configuration draws full-width run bars. 中文:数据点详情页链接现在写入该点的 i_prec,页面按纯硬件键分组,c_cfg 能精确 匹配而不会回退到其他配置。buildCacheReuse 仅在存在官方数据行时列出官方系列, 仅存在于非官方运行中的配置以全宽柱形绘制。 * feat: draw worker-pool power traces from power-audit bundles on the timeline PowerX Figure 1 on the gated Timeline display (y_measuredPowerTimeline): - /api/gpu-metrics?series=power also downloads power_audit_<RESULT_FILENAME> bundles (Slurm / Dynamo DCGM collector, up to 256 MiB), cuts LOGS/power/samples.csv into one series per power_validation_*.json window (±60 s) and labels devices <hostname>/<GPU-uuid> with their prefill / decode role; a corrupt archive skips that artifact instead of failing the run. - The timeline joins bundle-cut series by validation file name, keeps the gpu_metrics join for single-node rows, and gains a "Prefill / decode pools" line mode: summed pool watts with pool size × TDP references, pool tooltip, and an all-GPU total for traces without roles. - Pinned scatter tooltips on measured-power metrics offer "View power trace", which opens the timeline focused on that config (official and ?unofficialrun= points); the anchor href carries the canonical chart state. - Docs, catalog digest, unit / component / e2e coverage in en and zh. 中文:在受门控的功耗时间线视图上实现 PowerX 图 1。/api/gpu-metrics?series=power 现在也会下载 Slurm / Dynamo(DCGM)运行上传的 power_audit_* 产物包(上限 256 MiB), 按每个 power_validation_*.json 的测量窗口(前后各 60 s)切出独立序列,并按 prefill / decode 角色标注每个 <hostname>/<GPU-uuid> 设备;损坏的压缩包只跳过该产物, 不再让整个请求失败。时间线通过校验文件名关联这些序列(单节点行仍按 gpu_metrics 关联),新增"预填充 / 解码 GPU 池"线型:按池求和的功耗、池内 GPU 数 × TDP 的参考线、 池级提示框,无角色的曲线则显示全部 GPU 总功耗。实测功耗散点图的固定提示框新增 "查看功耗曲线"操作,可直接跳转到聚焦该配置的时间线(官方点与 ?unofficialrun= 叠加点 均支持),链接携带规范化的图表状态。同步更新文档、路由目录摘要及中英文单元 / 组件 / 端到端测试。 * [First-Token Limits] add measured TTFT-cap cost comparison tab / 新增首 token 延迟约束成本对比页 (#1174) * feat: add First-Token Limits tab comparing cheapest measured configs per TTFT cap New footer-only dashboard tab at /first-token and /zh/first-token. For each time-to-first-token cap it picks the cheapest measured row per chip vendor that also clears an interactivity floor, draws grouped bars with the best-vs-runner-up gap under the axis, and lists every winner with its run link. Unofficial runs render as their own series. TTFT is now carried on GPUDataPoint and kept by the calculator API view. 中文:新增页脚入口的仪表板页面 /first-token 与 /zh/first-token。按每一档首 token 延迟(TTFT)上限,从满足交互性下限的实测数据行中逐厂商选出成本最低的配置,以分组 柱形展示并在坐标轴下方标注最优与次优的成本差距,表格列出各优胜配置及其运行链接。 非官方运行以独立系列呈现。GPUDataPoint 新增 ttft 字段,calculator API 视图保留 TTFT 指标。 * fix: count overlay rows in the first-token empty state Overlay-only rows that carry a TTFT but miss the interactivity floor were described as having no first-token measurement, because the empty state only looked at official rows. Report overlay readable rows separately so the caption stays official-only while the empty state sees both. 中文:仅有非官方运行数据、且带 TTFT 但未达交互性下限时,空状态误报为「未报告首 token 延迟」。现单独统计叠加运行的可读行,标题说明仍只计官方数据,空状态则同时 考虑两者。 * fix: blame the cap ladder when qualifying rows clear no first-token cap The empty state told readers to lower the interactivity floor whenever no bar drew, even when rows cleared the floor and only the TTFT caps were too tight. It now distinguishes three cases: no TTFT reported, nothing reaches the floor, and nothing finishes within the largest cap. Overlay rows count toward the floor check through overlayQualifyingRows; formatCap moves to the shared module so the message prints caps like the axis ticks. 中文:此前只要没有柱形可画,空状态就提示降低交互性下限,即使数据行已达到下限、 只是 TTFT 上限过紧。现在区分三种情况:未报告 TTFT、无配置达到下限、无配置落在 最大上限之内。非官方运行行通过 overlayQualifyingRows 参与下限判断;formatCap 移入共享模块,使提示中的上限格式与坐标轴刻度一致。 * fix: prevent measured power statistic label overflow Size the control grid to its container and preserve statistic label widths. Add English and Chinese desktop/mobile geometry regression coverage. 中文:修复实测功耗统计量标签溢出。根据容器宽度排列控件,并保留按钮标签所需宽度;新增中英文桌面和移动端布局回归测试。 * fix: keep power envelope ties only for provisioned and modelled gauges The tie-keeping added for flat boundary gauges applied to every power curve, so the measured watt axes kept tied points that Optimal Only used to collapse (certified-power-filter rendered 6 of 6 instead of 2 of 6). `upperPowerEnvelope` takes `keepTies`; both charts pass it per series via `isPowerGaugeSeries`, true for a boundary axis or a boundary comparison clone, false for measured series, the measured clone and role clones. 中文:此前为平直的边界量表加入的"保留并列点"逻辑作用到了所有功耗曲线, 实测功率轴上 Optimal Only 本应收起的并列点被保留(certified-power-filter 显示 6/6 而非 2/6)。`upperPowerEnvelope` 新增 `keepTies` 参数,两张图按 序列通过 `isPowerGaugeSeries` 传入:边界轴或边界对比克隆为 true,实测序列、 实测克隆和角色克隆为 false。 * fix: key comparison legend rows by the base series identity `isBase` was "no official point carries this variant", which is wrong when only an `?unofficialrun=` overlay carries the comparison: every row became the base, all swatches drew solid and every toggle hid the base key, so a sibling could never be hidden on its own. Derive the base from `powerCompareBase` for the selected metric instead. A component test covers the overlay-only case. 中文:`isBase` 原来的判断是"没有官方点带这个 variant",当只有 `?unofficialrun=` overlay 带对比序列时所有行都被当成 base,图例全画实线且每次切换都隐藏 base,兄弟序列无法单独隐藏。改为按所选指标通过 `powerCompareBase` 推导 base。新增组件测试覆盖仅 overlay 的情况。 * fix: rename the timeline axis so a "Measured Power" search stays unique The Timeline display adds a fourteenth `y_measured*` option. Its label began with "Measured Power", so the family search matched two options. Lead with "Measured Average Power", as the %TDP display does, and update the selector count test, the timeline specs and the Chinese label. 中文:Timeline 显示模式新增了第 14 个 `y_measured*` 选项,其标签以 "Measured Power" 开头,导致按系列名搜索时匹配到两项。改为与 %TDP 显示 一致地以"Measured Average Power"开头,同步更新选择器数量测试、时间线 测试和中文标签。 * fix: label each power comparison series on its own line Under i_pcompare the line labels were grouped by hardware only, so one label per hardware landed on whichever variant had the most points (usually the TDP boundary) and multi-precision views drew identical labels on every sibling. Group by hardware and variant, keep the plain label on the base measured series, and suffix siblings with the short variant name plus the flat watts of a provisioned boundary (`H200 (SGLang) · TDP 700 W`, `… · Decode GPUs`). Label groups now carry data-series-id / data-power-variant hooks; overlays get the same text. 中文:对比模式(i_pcompare)下曲线标签原先只按硬件分组,每个硬件只有一个 标签且落在点数最多的 variant 线上(通常是 TDP 边界线),多精度视图则给每条 兄弟线画出相同文字。现改为按硬件 + variant 分组:基线保留原标签,兄弟线追加 简短 variant 名及平直边界的瓦数(如 `H200 (SGLang) · TDP 700 W`、 `… · Decode GPUs`);标签组新增 data-series-id / data-power-variant 钩子, 叠加运行同样处理。 * fix: load overlay runs first in the power timeline run cap `?unofficialrun=` telemetry sorted after official runs by run id and was the run dropped by POWER_TIMELINE_MAX_RUNS; overlay runs now take the cap's slots right after the deep-linked run. Adds prioritizeRuns with unit and component coverage and records the order in docs/powerx-permanent-view.md. 中文:功耗时间线每张图最多加载 4 个 run,此前按 run id 排序,`?unofficialrun=` 叠加的 run 通常最新、排在最后而被丢弃;现在叠加 run 紧随深链 run 优先占位。 新增 prioritizeRuns 及单元 / 组件测试,并在 docs/powerx-permanent-view.md 记录顺序。 * fix: route timeline legend clicks through the unified overlay selection With an overlay loaded the timeline reads localOfficialOverride, so the context toggleHwType changed nothing visible. Official rows now solo through setUnifiedOverlaySelection like ScatterGraph; component test added. 中文:加载 `?unofficialrun=` 叠加后,时间线读取 localOfficialOverride,图例点击 只改 activeHwTypes、界面无变化。现在官方硬件行与 ScatterGraph 一样通过 setUnifiedOverlaySelection 独显,并补充组件测试。 * fix: name the hardware instead of the branch on overlay line labels Unofficial-run pills read `✕ <hardware>`, parsed like official pills so the GPU stays bold; a short run tag (` · main`, ` · …<date>-<sha>`) is appended only when several overlay runs draw the same hardware. Branch names run to 70+ characters, so two of them could not share a chart and one pill was hidden. Legend rows keep the full branch. Unit and component coverage added. 中文:叠加曲线标签改为「✕ 硬件名」,与官方标签同样解析、GPU 名加粗;仅当多个 叠加 run 画同一硬件时追加短 run 标签(` · main`、` · …<日期>-<sha>`)。分支名 常超过 70 字符,两条标签放不下会被隐藏。图例行仍显示完整分支。补充单元与 组件测试。 * fix: keep overlay line labels visible when only their overlay row is active The filter-sync opacity pass judged every line label by the official hardware set, so soloing one official hardware hid the pills of every other overlay hardware. Overlay labels now follow activeOverlayHwTypes; unit tests cover the rule and the component test pins the rendered pill visibility. 中文:图例筛选后的标签透明度同步只看官方硬件集合,独显一个官方硬件会把其他 硬件的叠加曲线标签一并隐藏。叠加标签现在跟随 activeOverlayHwTypes;单元测试 覆盖规则,组件测试固定渲染后的可见性。 * fix: merge equal-size pool TDP references and stack coinciding labels In pool mode a prefill pool and a decode pool of the same GPU count drew two references at the same watts, printing their labels over each other. Pools of one hardware that share a size now draw one line labelled `prefill / decode ×16`, and labels of lines at equal watts stack upward. 中文:池模式下 prefill 池与 decode 池 GPU 数相同时,两条参考线重合、标签互相 覆盖。同一硬件、相同 GPU 数的池现在只画一条参考线,标签为 `prefill / decode ×16`;瓦数相同的参考线标签向上错开排列。 --------- Co-authored-by: Alec Ibarra <93070681+adibarra@users.noreply.github.com> Co-authored-by: functionstackx <47992694+functionstackx@users.noreply.github.com> Co-authored-by: Perplexity Computer <computer@users.noreply.github.com> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…licts Brings master (20 commits) into feat/powerx-db-ingest after #1177 landed. Three conflicts, all "both sides inserted at the same anchor": - AGENTS.md: kept master's Pareto cross-repo bullet; our gpu-metrics-point API line auto-merged elsewhere in the same file. - packages/app/cypress/e2e/landing-dashboard-navigation.cy.ts: master's "landing model curation" block is a strict superset of ours (verified: zero lines exist on our side that master lacks), so master's file wins. - packages/app/src/lib/glossary.ts: kept both. Our vera-rubin and extreme-co-design entries plus our KV/CPU/NVMe offload rewrites stay; master's 15 new Engram-article entries are appended. packages/app/src/lib/api-route-catalog.ts auto-merged cleanly this time. Local gate on the merge result: lint, fmt, typecheck, typography, and 6,872 unit tests (app 5,888 + 4 skipped, constants 63, db 921). The one db failure is the known PGlite cold-start timeout in operatorx.test.ts and passes on rerun. 中文:在 #1177 落地后把 master 的 20 个提交合入 feat/powerx-db-ingest。三处冲突 都是两侧在同一锚点插入内容: - AGENTS.md:保留 master 的 Pareto 跨仓库条目;我们新增的 gpu-metrics-point API 说明在同文件其他位置已自动合并。 - landing-dashboard-navigation.cy.ts:master 的 landing model curation 块是 我们版本的严格超集(已核对,我们没有任何一行是 master 缺失的),整体取 master。 - glossary.ts:两侧都保留。我们的 vera-rubin、extreme-co-design 词条以及 KV/CPU/NVMe offload 相关改写保持不变,追加 master 新增的 15 条 Engram 词条。 api-route-catalog.ts 本次自动合并干净。 合并结果的本地检查:lint、fmt、typecheck、typography,以及 6,872 个单元测试 (app 5,888 通过 + 4 跳过,constants 63,db 921)。db 侧唯一失败是已知的 operatorx.test.ts PGlite 冷启动超时,重跑通过。
…证据并保护已发布曲线 (#1152) * feat: enforce required PowerX publication coverage 中文:强制校验 PowerX 必需功耗点的入库与发布完整性。 按实际入库身份匹配并发点,保留过滤后的缺失检查,并核对 AgentX 数据库与 API 结果。 * fix: validate power evidence and preserve published curves 中文:校验必需功耗证据,并在入库前保护已发布曲线。版本化 manifest 绑定来源、物理 GPU、测量窗口和产物哈希;有意替换必须精确声明旧快照及移除点。
Migration 016 declares `on delete cascade` from gpu_metric_series to workflow_runs and from benchmark_result_gpu_metrics to benchmark_results, so a purge already removed telemetry — with no count in the preview and no line in the transcript. Past GitHub's 90-day artifact retention the stored samples are the only copy, which is the reason 016 exists, so that loss should not be invisible at the moment the operator confirms it. The same rows are removed as before; this is accounting, not a change of behaviour. previewPurge now prints the series and sample counts alongside the benchmark and server_log counts, a whole-run purge deletes the series itself before dropping workflow_runs and logs what went, and a point purge drops only the point-to-series links. The series stays with its run in that second case: /api/gpu-metrics?runId= reads series by run rather than through the links, and other points of the same run may still reference it. New lib/telemetry-purge.ts holds the three queries, covered by 11 PGlite tests against the real 001 and 016 migrations so the cascades under test are the ones the migration declares. 中文:迁移 016 声明了 gpu_metric_series 到 workflow_runs、 benchmark_result_gpu_metrics 到 benchmark_results 的 `on delete cascade`,因此 purge 一直在删除遥测数据,但预览里没有计数、日志里没有记录。超过 GitHub 90 天 产物保留期后,库里的采样就是唯一副本,这正是 016 存在的理由,所以操作者确认的 那一刻不应该看不见这笔损失。 删除的行与此前完全相同,这次改的是账目而不是行为。previewPurge 现在会在 benchmark 与 server_log 计数旁一并打印 series 与采样数;整个 run 的 purge 会在 删除 workflow_runs 之前先显式删除 series 并记录;单点 purge 只解除点与 series 的 链接。后一种情况下 series 会随 run 保留,因为 /api/gpu-metrics?runId= 是按 run 读取 series 而不是走这些链接,而且同一个 run 的其他点可能仍在引用它。 新增的 lib/telemetry-purge.ts 收拢这三条查询,由 11 个 PGlite 测试覆盖,直接对 真实的 001 与 016 迁移运行,确保被测的正是迁移声明的那些级联。
This workflow fires on a push to run-overrides.ts, which can land before the next ingest dispatch has applied a pending migration. It ran admin:db:verify with no migrate step of its own, and verify-db counts every table it knows about, so it would fail on a schema the checked-out ref expects but production does not have yet. Added the same Run migrations step ingest-results.yml uses; migrations are idempotent by filename. The worse half was what a red verify did to the two steps after it. Neither carried if: always(), so a verify failure skipped both the production cache invalidation and the warmup while the overrides were already committed to the database. The dashboard then served pre-override data with a red workflow as the only signal. Both now run whenever the overrides step itself succeeded. 中文:该工作流在 run-overrides.ts 的 push 上触发,而这次 push 可能早于下一次 ingest 派发应用待处理的迁移。它自己没有 migrate 步骤就直接跑 admin:db:verify, 而 verify-db 会统计它已知的每一张表,于是会在一个"检出的 ref 期望、但生产还没有" 的 schema 上失败。现已加入与 ingest-results.yml 相同的 Run migrations 步骤;迁移 按文件名幂等。 更严重的一半是 verify 变红对其后两个步骤的影响。两者都没有 if: always(),所以 verify 一失败就会跳过生产缓存失效和预热,而此时 override 已经写进数据库了。仪表板 于是继续提供改写前的数据,唯一的信号只有一个变红的工作流。现在只要 override 步骤 本身成功,这两步就会执行。
A gpu_metrics digest failure was recorded through recordDbError, which feeds tracker.skips.dbError, which the ingest writes into the publication manifest as an ingestError, which verify-power-publication folds into errors and exits non-zero on. That script is the required "Verify PowerX source, database and public API" step with no continue-on-error, so one malformed telemetry CSV would turn a whole production ingest red even though every benchmark row landed correctly. This failure mode does not exist before migration 016. Telemetry failures now have their own counter and their own recordTelemetryError, with its own print budget so telemetry noise cannot suppress real DB errors. The manifest reports them under telemetryWarnings, and the new fatalPublicationErrors makes the fatal set explicit at the one place that decides the exit code. What is lost on a digest failure is one point's PowerX tab, and admin:db:backfill-gpu-metrics --run <id> can re-digest the artifact. Coverage: the two counters are proven independent, and fatalPublicationErrors is proven to ignore telemetry warnings while still failing on real ingest errors and verification mismatches. The catch site itself has no direct test — it sits inside the artifact loop of a script the suite only runs as a subprocess against a purged run, which returns before reaching it. 中文:gpu_metrics 摘要失败此前经 recordDbError 记录,进入 tracker.skips.dbError, 再被 ingest 作为 ingestError 写入发布 manifest,verify-power-publication 把它折进 errors 并以非零码退出。该脚本是必需的 "Verify PowerX source, database and public API" 步骤且没有 continue-on-error,因此一个格式错误的遥测 CSV 就能让整条生产 ingest 变红,哪怕每一行 benchmark 数据都正确落库。这个失败模式在迁移 016 之前 并不存在。 遥测失败现在有独立计数器和独立的 recordTelemetryError,并有自己的打印额度,避免 遥测噪声淹没真正的 DB 错误。manifest 将其归入 telemetryWarnings,新增的 fatalPublicationErrors 在决定退出码的唯一位置显式界定致命集合。摘要失败损失的只是 某一个点的 PowerX 标签页,用 admin:db:backfill-gpu-metrics --run <id> 可以重新摘要 该产物。 覆盖范围:已验证两个计数器互相独立,并验证 fatalPublicationErrors 忽略遥测警告、 同时仍然对真正的 ingest 错误和校验不一致判为失败。catch 处本身没有直接测试——它 位于一个脚本的产物循环内部,而测试套件只以子进程方式对一个已 purge 的 run 运行该 脚本,那条路径在到达此处之前就返回了。
The anchor pass in placeLineLabels tried four fractions along each line and, if all of them collided, emitted visible:false. In the PowerX article's Fig 6 — roles mode with two ?unofficialrun= overlays, so nine lines whose anchors compete in one narrow band — that silently cost the GB300 overall series its pill. Its curve was still drawn, next to two labelled siblings of the same colour, so the reader could only identify it by elimination. The drop was a layering mistake. It decided visibility from a crude nominal box of collisionWidth/2 by 21px, while layoutPills runs right after with the real measured boxes, mirrors, shifts rows and clamps into the plot, and says in its own doc comment that an overlapped label beats a missing one. A rendered pill in this figure measures 223-330px against that 120px model, so the pass that gave up was the one least able to judge. The same function's pinned-anchor branch already emits visible:true unconditionally. A series with no clear slot is now deferred and placed after the clean ones, on whichever of its slots carries the least overlap, and it no longer anchors on points[0] — the axis-hugging index lineCandidates deliberately skips. Deferring matters on its own: a doomed series used to be able to reserve a slot that a later series could have had to itself. keepVisibleOnCollision is gone; it was the single-point special case of the rule this generalises. This ends the promise that line labels never overlap, stated in #132 and restated in #434. It was worth less than it cost: a dropped pill is silent, and with line labels on the PNG export omits the legend, so the series loses its only identifier. Verified at /inference?i_metric=y_measuredAvgPower&i_pcompare=roles&unofficialruns= 35319969159,35319956855 against the branch DB: 17 labels, all rendered, none overlapping, including the pill the figure was missing. The 12 pills still hidden there are the separate GH #470 de-duplication path, which is untouched. Gates: lint, fmt, typecheck, typography, 6,949 unit tests, 521 Cypress component tests. Three of the five new unit tests and the new component test fail on the old code. 中文:placeLineLabels 的锚点阶段沿每条线尝试四个位置,若全部碰撞就发出 visible:false。在 PowerX 文章的 Fig 6 中(roles 模式加两个 ?unofficialrun= 叠加, 九条线的锚点挤在同一窄带里),这让 GB300 整体序列悄悄丢掉了标签。它的曲线仍然 画着,旁边是两条同色且有标签的兄弟线,读者只能靠排除法辨认。 这个丢弃是分层错误。它用 collisionWidth/2 乘 21px 的粗略估算盒来决定可见性,而 紧随其后的 layoutPills 拿着真实测量盒做镜像、错行和边界钳制,其文档注释明确写着 重叠的标签也好过消失的标签。该图里一个实际渲染的标签宽 223 到 330px,而模型只按 120px 估算,所以放弃的恰恰是最没有判断力的那一遍。同一函数的固定锚点分支本来就 无条件发出 visible:true。 现在没有空闲槽位的序列会被推迟到干净标签之后放置,落在重叠代价最小的槽位上,并且 不再锚定到 points[0]——lineCandidates 刻意跳过的贴轴位置。推迟本身也有意义:原先 一个注定重叠的序列可能占掉后面序列本可独享的槽位。keepVisibleOnCollision 已删除, 它只是本规则的单点特例。 这终结了「折线标签永不重叠」的承诺,该承诺由 #132 提出、#434 重申。它的价值抵不上 代价:标签被丢弃是无声的,而开启折线标签后 PNG 导出会省略图例,序列就失去了唯一的 标识。 验证:在指向分支数据库的 /inference?i_metric=y_measuredAvgPower&i_pcompare=roles &unofficialruns=35319969159,35319956855 上,17 个标签全部渲染、互不覆盖,包括该图 原本缺失的那一个。页面上仍隐藏的 12 个标签来自独立的 GH #470 去重路径,未受影响。 闸门:lint、fmt、typecheck、typography、6,949 个单元测试、521 个 Cypress 组件测试。 五个新单元测试中的三个以及新增的组件测试在旧代码上失败。
Remove unused availability and sample-omission paths, identity wrappers, test-only label/source helpers, and write-only telemetry fields. Keep label formatting and nested-source coverage on the production helpers, preserve artifact failure isolation, and refresh the unchanged API contract digest. 中文:删除未使用的 PowerX 遥测查询模式、包装函数与只写不读的字段; 将标签格式及嵌套路径测试迁到实际生产入口,保留逐产物错误隔离, 并同步 API 路由摘要。此次仅落实精简清单 1–8。
Deduplicate each GPU/timestamp before computing metadata and statistics, keeping the first sample for both CSV and per-host power-audit series. Include sidecars and normalized sample counts in the replay check so corrected context is persisted and explicit re-ingest repairs old digests. Add PGlite regressions for duplicate populations, legacy replay, timezone and identity corrections, and JSON key-order no-ops. The focused suite fails on the old implementation and passes all 12 tests with the fix. 中文:统一 PowerX 遥测明细与摘要的采样集合,在统计前按 GPU 和时间戳去重。 重新摄取时同时核对 sidecar 与去重后的样本数,使时区及身份修正能够落库, 并支持显式重跑修复旧摘要;补充真实 PGlite 回归测试及恢复说明。
|
Claude finished @edwingao28's task in 4m 14s —— View job Review complete — 🟢 LGTM, no blocking issues foundI reviewed the full diff against Todo
What I verified
Chinese copyThe diff touches user-visible zh strings in Notes (non-blocking)
This is a carefully structured, well-tested change. No changes requested. |
|
Claude finished @edwingao28's task in 4m 31s —— View job Claude Code Review — finalizing
|

Summary
Adds GB200/GB300 NVL72 planning from measured compute-module power, or measured GPU-board plus Grace-socket power, with modeled rack overhead. Missing CPU measurements remain unavailable.
Base:
master; the diff includes inherited ingestion and chart work from #1167.Dependencies: InferenceX #3296 supplies CPU-side evidence. Rack-source revision
963ead8bstill needs publication confirmation; live pricing requires accepted CPU-enabled runs.Architecture and review map
1. Measurements, ingestion and system power
flowchart TB subgraph Producer["Producer — InferenceX #3296"] GPU["Measured GPU-board power"] CPU["Measured Grace / compute-module power"] ART["Benchmark results and power_audit.cpu<br/>Matching measurement window"] GPU --> ART CPU --> ART end ART --> ETL["benchmark-mapper<br/>Retain validity, metrics and sensor provenance"] ETL --> DB[("Benchmark rows and metrics")] DB --> API["Benchmark API"] API --> GATE{"modelSystemPower<br/>Valid GPU + CPU telemetry?<br/>GPU, socket and worker topology consistent?"} GATE -->|"No"| NONE["Unavailable with reason"] GATE -->|"Yes"| TRAY["Compute trays<br/>4 GPUs + 2 Grace sockets per full tray"] TRAY --> BASIS{"Measured basis"} BASIS -->|"Module sensor"| MODULE["Measured module watts<br/>No duplicate GPU / Grace addition"] BASIS -->|"Grace socket sensor"| SUM["Measured GPU-board + Grace watts<br/>Profile regulator allowance"] MODULE --> RACK SUM --> RACK PROFILE["Pinned GB200 / GB300 rack profiles<br/>Static loads and conversion assumptions"] -.-> RACK RACK["Mean tray load scaled to an 18-tray rack<br/>Modeled rack residual + power-shelf losses"] RACK --> AC["Rack AC power"] AC --> FAC["Apply PUE once<br/>NVL72 default: 1.1"] FAC --> DEP["Allocate rack share to measured deployment<br/>Retain basis, profile and provenance"]2. Planning, comparison and output
flowchart TB PERF["Official benchmark performance<br/>Selected percentile and target"] --> FRONT["Original serving frontier<br/>Exact point or bounding points"] POWER["Validated deployment facility power<br/>From diagram 1"] --> ACCEPT FRONT --> ACCEPT{"Full measured trays?<br/>Target inside measured range?<br/>Same source basis between points?"} ACCEPT -->|"No"| SKIP["Not priced<br/>Missing or unsupported evidence"] ACCEPT -->|"Yes"| POINT["Per-point planning power<br/>Facility kW / GPU x 1.10"] POINT --> SMART["Matched planning kW/GPU<br/>Interpolate only between accepted points"] SPEC["Provisioned kW/GPU"] --> CAP SMART --> CAP["Same facility budget<br/>Compute deployable GPU capacity"] INPUT["Same throughput, token prices,<br/>utilization and unit costs"] --> ECON CAP --> ECON["Revenue, costs and profit"] ECON --> UI["Provisioned / measured + modeled / compare<br/>Chart, tooltips and CSV"] META["Sensor basis, PUE, reserve<br/>Profile revision and source hash"] -.-> UISolid arrows show data flow; dashed arrows supply assumptions or provenance. Planning uses official frontier points, not
?unofficialrun=overlays. Grace/LPDDR power is measured; rack overhead is modeled.AI model disclosure
Validation
Previously reported: typecheck, lint, fmt,
test:unit, typography, API catalog and zh-copy guards pass; 204 Python-generated rack parity cases; Cypress profit-estimator fixtures 53/53. This metadata update did not rerun tests or establish live CPU measurement coverage.中文说明
为 GB200/GB300 新增基于实测计算模块功耗,或实测 GPU 板卡与 Grace socket 功耗的 NVL72 智能预配;机架其余开销由模型估算。缺少 CPU 实测时保持不可用。目标分支改为
master,当前 diff 同时包含从 #1167 继承的入库和图表改动。第一张图说明生产端、入库、有效性检查、tray 拓扑及机架功耗计算:完整 tray 为 4 个 GPU 与 2 个 Grace socket;模块读数不重复累加 GPU/Grace,基于机架平均 tray 负载计算 power-shelf 损耗,PUE 只应用一次。第二张图说明原有性能前沿、目标范围与测量口径检查、10% 规划余量,以及相同吞吐、价格和成本假设下的容量与收益对比。实线表示数据流,虚线表示假设或来源信息;利润估算器使用正式前沿数据,不读取
?unofficialrun=叠加层。Grace/LPDDR 功耗为实测,机架开销为建模。CPU 侧证据依赖 InferenceX #3296;机架模型来源版本
963ead8b仍需确认已发布,实际定价需要有效的 CPU 采集结果。保留既有验证记录:typecheck、lint、fmt、test:unit、排版、API 目录和中文文案检查通过,204 条机架对照用例、Cypress fixtures 53/53。本次只更新 PR 元数据和说明,未改代码或重跑测试,也不代表线上 CPU 测量覆盖已完成。既有实现及本次 Codex 编辑的准确模型名称/版本无法核实。Note
Medium Risk
Changes capacity math for GB200/GB300 in profit and chart power paths and depend on new CPU-side telemetry and a draft rack model; x86 chassis behavior is largely preserved but PUE selection is now hardware-aware.
Overview
GB200/GB300 move from unsupported hardware to pinned NVL72 rack profiles (
rackProfilesat model revision963ead8b).modelSystemPowertreats each measured worker as a compute tray (4 GPUs, 2 Grace sockets), requirescpu_power_valid=1and module or GPU-board + Grace inputs, folds trays into one 18-tray rack at mean load for a single power-shelf evaluation, and applies DLC PUE 1.1 by default (chassis stay at 1.3).The Profit Estimator “measured + modeled” path can now plan on fully measured NVL72 trays (and still on full eight-GPU chassis). It records measured basis, sensor kind, and profile in tooltips, caption notes, and CSV; refuses partial trays for planning and won’t interpolate between frontier knots with different power bases.
Offline export aligns with the dashboard: per-row default PUE unless
--pueoverrides all rows; NVL72 rows carry CPU-side measured fields and rack-specific boundary text. API docs document Grace/module metrics andpower_audit.cpu. Reference generation adds rack parity cases; Cypress covers a priced GB200 tray vs unavailable GB300 without CPU telemetry.Reviewed by Cursor Bugbot for commit 128605b. Bugbot is set up for automated code reviews on this repo. Configure here.