Skip to content

feat(dsv4): add GB300 Dynamo-TensorRT-LLM AgentX recipes / 添加 GB300 Dynamo-TensorRT-LLM AgentX 配置 - #2595

Open
RohitNagraj wants to merge 2 commits into
mainfrom
dsv4-fp4-gb300-dynamo-trt-agentx-recipes
Open

feat(dsv4): add GB300 Dynamo-TensorRT-LLM AgentX recipes / 添加 GB300 Dynamo-TensorRT-LLM AgentX 配置#2595
RohitNagraj wants to merge 2 commits into
mainfrom
dsv4-fp4-gb300-dynamo-trt-agentx-recipes

Conversation

@RohitNagraj

Copy link
Copy Markdown
Collaborator

Description

Add a DeepSeek-V4-Pro FP4 Dynamo–TensorRT-LLM AgentX configuration for GB300.

  • Add ten disaggregated MTP recipe points and three in-repo EPLB placement tables.
  • Add the dsv4-fp4-gb300-dynamo-trt-agentx master configuration with NIXL KV transfer.
  • Integrate the recipes with the GB300 launcher while preserving the existing runner settings and /scratch/models/DeepSeek-V4-Pro.

Validation:

  • Generated the exact configuration matrix.
  • Validated every recipe individually with srt-slurm v1.0.36.
  • Passed matrix-logic, launcher-contract, YAML, topology, and confidentiality checks.

中文说明

为 GB300 添加 DeepSeek-V4-Pro FP4 Dynamo–TensorRT-LLM AgentX 配置。

  • 新增十个解耦式 MTP 配方点及三个仓库内 EPLB 放置表。
  • 新增采用 NIXL KV 传输的 dsv4-fp4-gb300-dynamo-trt-agentx 主配置。
  • 将这些配方接入 GB300 启动脚本,同时保留现有运行器设置和 /scratch/models/DeepSeek-V4-Pro 模型路径。

验证:

  • 已生成并检查精确配置矩阵。
  • 已使用 srt-slurm v1.0.36 逐个验证所有配方。
  • 已通过矩阵逻辑、启动脚本契约、YAML、拓扑及保密性检查。

Related Issue

N/A / 不适用

Type of Change

  • Bug fix
  • New feature
  • Configuration change
  • Documentation update
  • Other (please describe)

Checklist

  • I have tested my changes locally
  • I have updated documentation if necessary
  • For every change that can affect benchmark performance and every recipe addition or modification, I have appended a new entry to the physical end of perf-changelog.yaml and have not edited historical entries
  • Before merging via reuse, an authorized maintainer (OWNER/MEMBER/COLLABORATOR) has commented /reuse-sweep-run on this PR. Do this only once there is a final full sweep that is all green with evals passing, since after this comment the sweep label will no longer automatically kick off new sweeps. Remove and re-add the label to force one.

添加 DeepSeek-V4-Pro GB300 Dynamo TensorRT-LLM AgentX 配置与配套配方。
@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase As a PR reviewer and CODEOWNER, I have reviewed this and have.

For PR verification, add the full-sweep-fail-fast label (strongly recommended) to this PR — the benchmark sweep only runs on labeled PRs. Use full-sweep-enabled only if you need matrix jobs to keep running past a failure.

PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs


感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 As a PR reviewer and CODEOWNER, I have reviewed this and have

如需进行 PR 验证,请为此 PR 添加 full-sweep-fail-fast 标签(强烈推荐)— 基准测试 sweep 仅在带有标签的 PR 上运行。仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled

PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档

1 similar comment
@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase As a PR reviewer and CODEOWNER, I have reviewed this and have.

For PR verification, add the full-sweep-fail-fast label (strongly recommended) to this PR — the benchmark sweep only runs on labeled PRs. Use full-sweep-enabled only if you need matrix jobs to keep running past a failure.

PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs


感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 As a PR reviewer and CODEOWNER, I have reviewed this and have

如需进行 PR 验证,请为此 PR 添加 full-sweep-fail-fast 标签(强烈推荐)— 基准测试 sweep 仅在带有标签的 PR 上运行。仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled

PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档

在性能变更日志中补充 PR 2595 链接。
Comment thread perf-changelog.yaml Outdated
- "Add a DeepSeek-V4-Pro FP4 Dynamo-TensorRT-LLM MTP AgentX configuration for GB300 with ten disaggregated topologies."
- "Use NIXL KV transfer, deterministic MTP acceptance settings, and in-repo EPLB placement tables."
- "Image: nvcr.io/nvidia/tensorrt-llm/release:1.3.0rc24"
pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/XXX

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 The new perf-changelog.yaml entry for dsv4-fp4-gb300-dynamo-trt-agentx ends with a placeholder pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/XXX instead of this PR's real number. Since this is PR #2595, the line should read pull/2595; TBD is the only sanctioned pre-PR placeholder per docs/configuration-procedures.md.

Extended reasoning...

The last entry appended to perf-changelog.yaml (for the new dsv4-fp4-gb300-dynamo-trt-agentx config-key introduced by this PR) ends with:

pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/XXX

XXX is a literal, hardcoded string — it isn't a template variable that gets substituted anywhere, so it resolves to a broken GitHub URL that points at a PR which doesn't exist (there is no PR #"XXX").

docs/configuration-procedures.md documents the expected changelog convention: the pr-link field must follow the .../pull/<number> format, and the only acceptable placeholder before a PR exists is the literal string pr-link: TBD — which must be replaced with the real PR URL immediately once the PR is created. pull/XXX doesn't match either of these: it's neither a valid link nor the sanctioned TBD placeholder, so it violates the documented convention outright.

The code path that "triggers" this is simply: whoever authored the recipe copy-pasted a previous changelog entry as a template and forgot to fill in the actual PR number before generating the diff. Nothing in the repo's tooling currently validates this field at PR-creation time, so a malformed pull/XXX slips through unless a reviewer or CI check catches it — which is presumably why the docs explicitly call out this exact failure mode.

Impact is limited to traceability/documentation: a maintainer or anyone browsing the changelog later to find "which PR introduced this recipe / regression" will follow a dead link instead of landing on the actual PR. It has zero effect on the recipe YAML, the launcher, or benchmark behavior — this is purely a metadata defect in the changelog.

Step-by-step proof:

  1. Open perf-changelog.yaml and look at the newest (physically last) entry, whose config-keys list contains dsv4-fp4-gb300-dynamo-trt-agentx (the recipe this PR adds).
  2. Its final field reads pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/XXX.
  3. Per the PR metadata, this change is submitted as PR feat(dsv4): add GB300 Dynamo-TensorRT-LLM AgentX recipes / 添加 GB300 Dynamo-TensorRT-LLM AgentX 配置 #2595.
  4. Compare against sibling entries in the same file (e.g. pull/2580, pull/2578, pull/2577 for the three most-recently-merged PRs before this one) — every other entry uses a real numeric PR id, confirming pull/XXX is an unintentional leftover rather than an intentional convention.
  5. Fix: change the line to pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/2595.

This is a simple one-line fix with no functional risk, so it doesn't block merging, but it should be corrected since the PR's own checklist explicitly calls out maintaining accurate perf-changelog entries.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

1 participant