[Klaud Cold][agentic experiment][Variant A] feat: add Kimi-K3 MXFP4 B200 aggregated TP8xPP2 Dynamo-vLLM agentic bring-up / 新增 Kimi-K3 MXFP4 B200 聚合式 TP8xPP2 Dynamo-vLLM 智能体编码基准测试(bring-up) - #2355
Conversation
…ecipe Aggregated TP8 x PP2 across 2 B200 nodes (16 GPUs), plain TP (no expert parallelism) for the agentic-coding trace replay. Dedicated bring-up image vllm/vllm-openai:kimi-k3 with VLLM_ENABLE_K3_LATENT_MOE_TAIL_FUSION=1, fastsafetensors load format, kimi_k3 tool-call/reasoning parsers. Model pre-staged at /lustre/fsw/models/Kimi-K3; launch_b200-dgxc.sh gains the kimik3/fp4 model-path mapping, the agentic recipe overlay, and the agentic cache default_mounts used by the GB200/GB300 agentic paths. 中文:新增 Kimi-K3 MXFP4 B200 聚合式 TP8xPP2 Dynamo-vLLM 智能体编码基准测试配方 (2 节点 / 16 GPU,纯 TP,不启用专家并行(EP))。使用专用 bring-up 镜像 vllm/vllm-openai:kimi-k3(VLLM_ENABLE_K3_LATENT_MOE_TAIL_FUSION=1、 fastsafetensors 加载格式、kimi_k3 工具调用/推理解析器)。模型已预置于 /lustre/fsw/models/Kimi-K3;启动器 launch_b200-dgxc.sh 增加 kimik3/fp4 模型路径映射、智能体配方覆盖及智能体缓存挂载。 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase For PR verification, add the PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs 感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 如需进行 PR 验证,请为此 PR 添加 PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档 |
中文:在更新日志条目与 MODELS 表格行中补充 PR #2355 链接。 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30294472010 |
| interval_seconds: 10 | ||
| max_attempts: 1440 |
There was a problem hiding this comment.
🔴 The recipe deliberately sets health_check.max_attempts: 1440 (4h) for the ~1.4TB Kimi-K3 MXFP4 checkpoint, but runners/launch_b200-dgxc.sh:291 unconditionally runs sed -i 's/^ max_attempts: [0-9]*/ max_attempts: 720/' on the copied config right before srtctl apply, silently clobbering it back to 720 attempts (2h) — the same budget sized for DSR1-FP8 at roughly half this checkpoint's weight size. The recipe file will look correctly configured but the wider window never actually takes effect at runtime; either make the sed a floor (only raise, never lower) or special-case kimik3 like the other model-prefix branches nearby.
Extended reasoning...
The bug: runners/launch_b200-dgxc.sh line 291 runs:
sed -i 's/^ max_attempts: [0-9]*/ max_attempts: 720/' "${CONFIG_FILE%%:*}"
right before srtctl apply -f "$CONFIG_FILE" inside the IS_MULTINODE branch. This is an unconditional hard-set, not a max()/floor operation — whatever numeric value follows max_attempts: (two-space indent) in the config file gets overwritten to exactly 720, no matter what it was before.
The interaction this PR introduces: the new recipe benchmarks/multi_node/srt-slurm-recipes/vllm/kimi-k3/agentic/agg-b200-tp8pp2-agentic.yaml sets:
health_check:
interval_seconds: 10
max_attempts: 1440
with a two-space indent that matches the sed's regex exactly. The recipe's own comments make clear this 1440 (= 4h at 10s/attempt) was deliberately sized up from the launcher's usual value, specifically because the native MXFP4 checkpoint is ~1.4TB and has to be pulled over shared Lustre across 2 nodes. But because the sed is unconditional, srtctl apply never sees 1440 — it sees the file after the sed has already rewritten it to 720 (= 2h). The comment directly above the sed ("Bump recipe health-check timeout from 360×10s to 720×10s so large-model loads ... finish in time") shows the sed was written with the mental model of bumping the common 360 default up to 720; it never anticipated a recipe explicitly setting a value above 720, so for this recipe the sed is actually a reduction, not a bump.
Why nothing else catches this: there's no validation step between writing the recipe's health_check block and srtctl apply that would surface the discrepancy — the sed runs silently, will echo the sed invocation but not diff the before/after content of the file, and the recipe file committed to the repo (and reviewed in this PR) legitimately shows 1440. Someone reading the recipe or this diff would have no way to know the value gets overwritten at runtime.
Step-by-step proof:
CONFIG_FILEresolves torecipes/vllm/kimi-k3/agentic/agg-b200-tp8pp2-agentic.yaml(per theconfigs/nvidia-master.yamlentry'sadditional-settings: CONFIG_FILE=...), a copy of this PR's new recipe overlaid into the srt-slurm-nv clone.- That file, as written in this PR, contains the line
max_attempts: 1440. - At runtime, line ~289 first does
sed -i "s/^name:.*/name: \"${RUNNER_NAME}\"/"(unrelated), then line 291 doessed -i 's/^ max_attempts: [0-9]*/ max_attempts: 720/' "${CONFIG_FILE%%:*}". - The regex
^ max_attempts: [0-9]*matchesmax_attempts: 1440(2-space indent, then digits), so the line becomesmax_attempts: 720. srtctl apply -f "$CONFIG_FILE"runs immediately after, reading the now-mutated file — it appliesmax_attempts: 720,interval_seconds: 10→ a 7200s (2h) health-check window, not the 14400s (4h) the recipe declares.
Impact: if 2h genuinely isn't enough to pull ~1.4TB of MXFP4 weights and bring the engine up (workers loading in parallel off shared Lustre, contended by whatever else is running on the same filesystem), the health check gives up and fails the entire multinode job — the exact scenario this recipe's author was trying to prevent by widening the window, on the largest checkpoint in the fleet and its first-ever bring-up. Because the recipe file looks correct, anyone debugging a health-check failure here would have to know to check the launcher script rather than trusting the recipe as documentation of what will actually run — a debugging detour on a bring-up PR that's likely to already have plenty of other failure modes to sort through.
Fix: make the sed a floor instead of a hard-set, e.g. only replace when the existing value is less than 720 (or compare-and-max in the shell), or skip the sed entirely for kimik3/any recipe that already sets a larger value, similar to how other model prefixes get their own branch in this same script.
…oke test The cquil11/srt-slurm-nv cam/sa-submission-q2-2026 fork rejected the recipe (benchmark.aiperf_server_metrics: Unknown field). Switch the b200-dgxc agentic clone to upstream NVIDIA/srt-slurm v1.0.36 (validated in #2302/#2341), drop the aiperf_server_metrics field, pin dynamo wheel/router to 1.2.1 (the combination validated with v1.0.36), and reduce the bring-up to a single conc-8 smoke test. 中文:cquil11/srt-slurm-nv 分支的 srtctl 校验拒绝了配方字段 benchmark.aiperf_server_metrics(Unknown field)。将 b200-dgxc 智能体路径改用 上游 NVIDIA/srt-slurm v1.0.36(已在 #2302/#2341 验证),移除该字段,dynamo wheel/router 固定为 1.2.1,并将 bring-up 缩减为单并发(conc 8)冒烟测试。 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30295407363 |
The dynamo-vllm worker entrypoint rejected --enable-auto-tool-choice --tool-call-parser kimi_k3 (unrecognized arguments; different arg parser than vllm serve). Chat parsing happens at the dynamo frontend — same convention as the DSv4 GB300 agentic recipes. Keep --reasoning-parser kimi_k3 (accepted by the worker). Also drop the explicit max-model-len and let vLLM derive the native 1M window from the model config, mirroring the agentic recipe convention. 中文:dynamo-vllm worker 入口不接受 --enable-auto-tool-choice 与 --tool-call-parser kimi_k3(unrecognized arguments,与 vllm serve 的参数解析器 不同),聊天解析由 dynamo 前端处理,与 DSv4 GB300 智能体配方约定一致;保留 worker 可接受的 --reasoning-parser kimi_k3。同时移除显式 max-model-len, 由 vLLM 从模型配置推导原生 1M 上下文窗口。 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30295861948 |
Third sweep attempt: the engine loaded and served (TP8xPP2 healthy in ~14 min), but dynamo 1.2.1's rust frontend tokenizer rejects Kimi-K3's tiktoken model_type 'kimi_k3' (supported: kimi, kimi_k2, kimi_k25, deepseek_v3), so the model never registered and all chat completions returned 404, aborting the AgentX warmup. Switch to the 1.2.0.dev20260426 wheel used by the DSv4 GB300/B200 Dynamo-vLLM recipes. Upstream published v1.4.0-kimi-k3-dev.1 (2026-07-27) as the day-zero K3 build if this wheel also lacks support. 中文:第三次扫描中引擎已成功加载并提供服务(TP8xPP2 约 14 分钟就绪),但 dynamo 1.2.1 的 rust 前端分词器不支持 Kimi-K3 的 tiktoken model_type 'kimi_k3',模型未能注册,所有请求返回 404,AgentX 预热中止。改用 DSv4 GB300/B200 Dynamo-vLLM 配方所用的 1.2.0.dev20260426 wheel;如仍不支持, 上游已于 2026-07-27 发布 day-zero 构建 v1.4.0-kimi-k3-dev.1。 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Pin dynamo to ba83080ecd31c1ce918559e576d3c5bc9e092ff1 ("feat: Added
support for Kimi-K3", tag v1.4.0-kimi-k3-dev.1) via srt-slurm's
hash-cached source install: it adds the kimi_k3 tiktoken tokenizer to the
rust frontend (dynamo <=1.2.1 404s every request because the model never
registers) and accepts the kimi_k3 tool-call/reasoning parser worker args,
so restore --enable-auto-tool-choice --tool-call-parser kimi_k3
--reasoning-parser kimi_k3.
中文:将 dynamo 固定到 day-zero Kimi-K3 提交 ba83080("feat: Added support
for Kimi-K3",标签 v1.4.0-kimi-k3-dev.1),通过 srt-slurm 的哈希缓存源码
安装:该提交为 rust 前端新增 kimi_k3 tiktoken 分词器(dynamo <=1.2.1 因模型
无法注册而全部返回 404),worker 亦支持 kimi_k3 解析器参数,故恢复
--enable-auto-tool-choice --tool-call-parser kimi_k3 --reasoning-parser kimi_k3。
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30297837112 |
Replace the vLLM OpenAI-frontend spellings (--enable-auto-tool-choice / --tool-call-parser) with dynamo's namespaced worker args: --dyn-tool-call-parser kimi_k3 --reasoning-parser kimi_k3 --dyn-reasoning-parser kimi_k3. 中文:将 vLLM OpenAI 前端风格参数(--enable-auto-tool-choice / --tool-call-parser)替换为 dynamo 命名空间的 worker 参数: --dyn-tool-call-parser kimi_k3 --reasoning-parser kimi_k3 --dyn-reasoning-parser kimi_k3。 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30298269589 |
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30298672988 |
Fifth sweep attempt: the day-zero dynamo registered the kimi_k3 tiktoken tokenizer and the engine served, but all warmup requests got 400 — aiperf's conv-aware routing emits nvext.session_control, a removed POC field this dynamo build rejects (schema moved to router/routing_constraints/ agent_hints). Opt out via AIPERF_USE_DYNAMO_CONV_AWARE_ROUTING=0, matching the GB300 aggregate AgentX recipes; a single aggregate worker has no P/D routing to bind anyway. 中文:第五次扫描中 day-zero dynamo 已成功注册 kimi_k3 tiktoken 分词器并正常 服务,但全部预热请求返回 400——aiperf 的会话感知路由会发送 nvext.session_control(已被移除的 POC 字段,schema 已迁移至 router/routing_constraints/agent_hints)。通过 AIPERF_USE_DYNAMO_CONV_AWARE_ROUTING=0 关闭,与 GB300 聚合式 AgentX 配方 一致;单聚合 worker 本无需 P/D 路由绑定。 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30300683178 |
Sixth sweep attempt (both A and C variants): warmup requests 500 then the model 503s — the image's first decode step crashes in the KDA hybrid-state postprocess (mamba_hybrid.py postprocess_state, IndexError: index_fill_(): Expected dtype int64 for index; torch requires an int64 index but the runner passes the int32 idx_mapping). Ship an in-container patch through srt-slurm's setup_script hook (same pattern as configs/patches/ vllm_numa_bind_hash_fix.py): coerce the index with .long(), idempotent, refuses to run if the image layout changed. 中文:第六次扫描(A、C 两个变体一致):预热请求先 500、随后模型 503——镜像 首个解码步在 KDA 混合状态后处理中崩溃(mamba_hybrid.py postprocess_state, IndexError: index_fill_() 需要 int64 索引,但 runner 传入 int32 idx_mapping)。 通过 srt-slurm 的 setup_script 钩子在容器内打补丁:将索引用 .long() 转换, 幂等,且镜像布局变化时拒绝执行。 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30304440483 |
Seventh sweep attempt (mamba patch confirmed applied, model forward now executes): the first warmup forward OOMs in the flashinfer trtllm MXFP4 MoE kernel, which allocates a ~1.6 GiB runtime workspace outside vLLM's memory pool — at 0.95 a 178 GiB B200 has only ~1.35 GiB free. 0.90 matches the GB200/GB300 agentic recipes. 中文:第七次扫描(mamba 补丁已确认生效,模型前向已可执行):首个预热前向在 flashinfer trtllm MXFP4 MoE 内核中 OOM——该内核在 vLLM 显存池之外分配约 1.6 GiB 运行时工作区,0.95 下 178 GiB B200 仅剩约 1.35 GiB 空闲。改为 0.90, 与 GB200/GB300 智能体配方一致。 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30306482385 |
Eighth sweep attempt at gpu-mem-util 0.90 still OOM'd: a 2.92 GiB MLA long-context prefill transient (kv_b_proj in _compute_prefill_context) failed while 3.39 GiB sat reserved-but-unallocated — allocator fragmentation. Set PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True (what the torch OOM message recommends; the DSv4 recipes set it) and drop NCCL_CUMEM_ENABLE to trim NCCL's share of non-PyTorch device memory. 中文:第八次扫描在 0.90 显存利用率下仍 OOM:MLA 长上下文预填充的 2.92 GiB 瞬时分配(_compute_prefill_context 中的 kv_b_proj)失败,而 3.39 GiB 处于 已保留未分配状态——分配器碎片化。设置 PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True(torch OOM 报错所建议、 DSv4 配方亦采用),并移除 NCCL_CUMEM_ENABLE 以减少 NCCL 占用的非 PyTorch 显存。 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
Closing in favor of the direct-vLLM experiment (Variant D, #2359): serving Kimi-K3 directly with 中文:关闭本 PR,转向直接 vLLM 实验(变体 D,#2359):直接以 vllm serve 提供服务,从根本上消除本 PR 需要逐一解决的 Dynamo 兼容性问题(分词器 404、session_control 400、--dyn-* worker 参数)。调试成果均已由变体 D 继承——mamba_hybrid 索引类型容器补丁、会话感知路由关闭、显存利用率 0.90 及启动器/模型路径接线。进行中的扫描已取消。 |
Inherited from the closed dynamo-frontend variants (#2355/#2358): at gpu-mem-util 0.90 the first long-context MLA prefill OOM'd on a 2.92 GiB transient while 3.39 GiB sat reserved-but-unallocated (fragmentation). Set PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True and drop NCCL_CUMEM_ENABLE. 中文:继承自已关闭的 dynamo 前端变体(#2355/#2358):0.90 显存利用率下首个 长上下文 MLA 预填充因 2.92 GiB 瞬时分配 OOM,而 3.39 GiB 处于已保留未分配 状态(碎片化)。设置 PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True 并 移除 NCCL_CUMEM_ENABLE。 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=30308356213 |
…n't work with Pipeline yet, offloading & TP16 and DEP8PP2 to be done in followup PR) (#2391) * feat: add Kimi-K3 MXFP4 B200 aggregated TP8xPP2 Dynamo-vLLM agentic recipe Aggregated TP8 x PP2 across 2 B200 nodes (16 GPUs), plain TP (no expert parallelism) for the agentic-coding trace replay. Dedicated bring-up image vllm/vllm-openai:kimi-k3 with VLLM_ENABLE_K3_LATENT_MOE_TAIL_FUSION=1, fastsafetensors load format, kimi_k3 tool-call/reasoning parsers. Model pre-staged at /lustre/fsw/models/Kimi-K3; launch_b200-dgxc.sh gains the kimik3/fp4 model-path mapping, the agentic recipe overlay, and the agentic cache default_mounts used by the GB200/GB300 agentic paths. 中文:新增 Kimi-K3 MXFP4 B200 聚合式 TP8xPP2 Dynamo-vLLM 智能体编码基准测试配方 (2 节点 / 16 GPU,纯 TP,不启用专家并行(EP))。使用专用 bring-up 镜像 vllm/vllm-openai:kimi-k3(VLLM_ENABLE_K3_LATENT_MOE_TAIL_FUSION=1、 fastsafetensors 加载格式、kimi_k3 工具调用/推理解析器)。模型已预置于 /lustre/fsw/models/Kimi-K3;启动器 launch_b200-dgxc.sh 增加 kimik3/fp4 模型路径映射、智能体配方覆盖及智能体缓存挂载。 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: link PR #2355 in changelog entry and MODELS rows 中文:在更新日志条目与 MODELS 表格行中补充 PR #2355 链接。 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: pin agentic srt-slurm to NVIDIA v1.0.36, dynamo 1.2.1, conc-8 smoke test The cquil11/srt-slurm-nv cam/sa-submission-q2-2026 fork rejected the recipe (benchmark.aiperf_server_metrics: Unknown field). Switch the b200-dgxc agentic clone to upstream NVIDIA/srt-slurm v1.0.36 (validated in #2302/#2341), drop the aiperf_server_metrics field, pin dynamo wheel/router to 1.2.1 (the combination validated with v1.0.36), and reduce the bring-up to a single conc-8 smoke test. 中文:cquil11/srt-slurm-nv 分支的 srtctl 校验拒绝了配方字段 benchmark.aiperf_server_metrics(Unknown field)。将 b200-dgxc 智能体路径改用 上游 NVIDIA/srt-slurm v1.0.36(已在 #2302/#2341 验证),移除该字段,dynamo wheel/router 固定为 1.2.1,并将 bring-up 缩减为单并发(conc 8)冒烟测试。 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: drop OpenAI-frontend tool-choice flags from dynamo-vllm worker args The dynamo-vllm worker entrypoint rejected --enable-auto-tool-choice --tool-call-parser kimi_k3 (unrecognized arguments; different arg parser than vllm serve). Chat parsing happens at the dynamo frontend — same convention as the DSv4 GB300 agentic recipes. Keep --reasoning-parser kimi_k3 (accepted by the worker). Also drop the explicit max-model-len and let vLLM derive the native 1M window from the model config, mirroring the agentic recipe convention. 中文:dynamo-vllm worker 入口不接受 --enable-auto-tool-choice 与 --tool-call-parser kimi_k3(unrecognized arguments,与 vllm serve 的参数解析器 不同),聊天解析由 dynamo 前端处理,与 DSv4 GB300 智能体配方约定一致;保留 worker 可接受的 --reasoning-parser kimi_k3。同时移除显式 max-model-len, 由 vLLM 从模型配置推导原生 1M 上下文窗口。 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: try dynamo wheel 1.2.0.dev20260426 for Kimi-K3 frontend tokenizer Third sweep attempt: the engine loaded and served (TP8xPP2 healthy in ~14 min), but dynamo 1.2.1's rust frontend tokenizer rejects Kimi-K3's tiktoken model_type 'kimi_k3' (supported: kimi, kimi_k2, kimi_k25, deepseek_v3), so the model never registered and all chat completions returned 404, aborting the AgentX warmup. Switch to the 1.2.0.dev20260426 wheel used by the DSv4 GB300/B200 Dynamo-vLLM recipes. Upstream published v1.4.0-kimi-k3-dev.1 (2026-07-27) as the day-zero K3 build if this wheel also lacks support. 中文:第三次扫描中引擎已成功加载并提供服务(TP8xPP2 约 14 分钟就绪),但 dynamo 1.2.1 的 rust 前端分词器不支持 Kimi-K3 的 tiktoken model_type 'kimi_k3',模型未能注册,所有请求返回 404,AgentX 预热中止。改用 DSv4 GB300/B200 Dynamo-vLLM 配方所用的 1.2.0.dev20260426 wheel;如仍不支持, 上游已于 2026-07-27 发布 day-zero 构建 v1.4.0-kimi-k3-dev.1。 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: pin dynamo to day-zero Kimi-K3 commit, restore kimi_k3 parser flags Pin dynamo to ba83080ecd31c1ce918559e576d3c5bc9e092ff1 ("feat: Added support for Kimi-K3", tag v1.4.0-kimi-k3-dev.1) via srt-slurm's hash-cached source install: it adds the kimi_k3 tiktoken tokenizer to the rust frontend (dynamo <=1.2.1 404s every request because the model never registers) and accepts the kimi_k3 tool-call/reasoning parser worker args, so restore --enable-auto-tool-choice --tool-call-parser kimi_k3 --reasoning-parser kimi_k3. 中文:将 dynamo 固定到 day-zero Kimi-K3 提交 ba83080("feat: Added support for Kimi-K3",标签 v1.4.0-kimi-k3-dev.1),通过 srt-slurm 的哈希缓存源码 安装:该提交为 rust 前端新增 kimi_k3 tiktoken 分词器(dynamo <=1.2.1 因模型 无法注册而全部返回 404),worker 亦支持 kimi_k3 解析器参数,故恢复 --enable-auto-tool-choice --tool-call-parser kimi_k3 --reasoning-parser kimi_k3。 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: use dynamo namespaced --dyn-* kimi_k3 parser args on the worker Replace the vLLM OpenAI-frontend spellings (--enable-auto-tool-choice / --tool-call-parser) with dynamo's namespaced worker args: --dyn-tool-call-parser kimi_k3 --reasoning-parser kimi_k3 --dyn-reasoning-parser kimi_k3. 中文:将 vLLM OpenAI 前端风格参数(--enable-auto-tool-choice / --tool-call-parser)替换为 dynamo 命名空间的 worker 参数: --dyn-tool-call-parser kimi_k3 --reasoning-parser kimi_k3 --dyn-reasoning-parser kimi_k3。 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: disable aiperf conv-aware routing (session_control 400-rejected) Fifth sweep attempt: the day-zero dynamo registered the kimi_k3 tiktoken tokenizer and the engine served, but all warmup requests got 400 — aiperf's conv-aware routing emits nvext.session_control, a removed POC field this dynamo build rejects (schema moved to router/routing_constraints/ agent_hints). Opt out via AIPERF_USE_DYNAMO_CONV_AWARE_ROUTING=0, matching the GB300 aggregate AgentX recipes; a single aggregate worker has no P/D routing to bind anyway. 中文:第五次扫描中 day-zero dynamo 已成功注册 kimi_k3 tiktoken 分词器并正常 服务,但全部预热请求返回 400——aiperf 的会话感知路由会发送 nvext.session_control(已被移除的 POC 字段,schema 已迁移至 router/routing_constraints/agent_hints)。通过 AIPERF_USE_DYNAMO_CONV_AWARE_ROUTING=0 关闭,与 GB300 聚合式 AgentX 配方 一致;单聚合 worker 本无需 P/D 路由绑定。 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: patch kimi-k3 image mamba_hybrid index_fill_ dtype via setup_script Sixth sweep attempt (both A and C variants): warmup requests 500 then the model 503s — the image's first decode step crashes in the KDA hybrid-state postprocess (mamba_hybrid.py postprocess_state, IndexError: index_fill_(): Expected dtype int64 for index; torch requires an int64 index but the runner passes the int32 idx_mapping). Ship an in-container patch through srt-slurm's setup_script hook (same pattern as configs/patches/ vllm_numa_bind_hash_fix.py): coerce the index with .long(), idempotent, refuses to run if the image layout changed. 中文:第六次扫描(A、C 两个变体一致):预热请求先 500、随后模型 503——镜像 首个解码步在 KDA 混合状态后处理中崩溃(mamba_hybrid.py postprocess_state, IndexError: index_fill_() 需要 int64 索引,但 runner 传入 int32 idx_mapping)。 通过 srt-slurm 的 setup_script 钩子在容器内打补丁:将索引用 .long() 转换, 幂等,且镜像布局变化时拒绝执行。 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat: agentic experiment D — direct vllm serve via srt-slurm PR #278 Serve Kimi-K3 directly with vllm serve (srt-slurm PR #278 frontend.type: vllm, branch kylliang/direct-aggregate-vllm): no dynamo frontend/worker/ router, which removes the dynamo tokenizer/schema gaps entirely, and the OpenAI-frontend flags --enable-auto-tool-choice --tool-call-parser kimi_k3 --reasoning-parser kimi_k3 become legitimate. PR #278 validates single-node only, so ship patches/srt-slurm-pr278-direct-vllm-multinode.patch extending it to vLLM-native multi-node serve (--master-addr/--nnodes/--node-rank, headless non-leader ranks) for the 2-node TP8xPP2 topology. Keeps the mamba_hybrid index-dtype container patch (engine bug is frontend-agnostic). 中文:智能体实验变体 D——通过 srt-slurm PR #278(frontend.type: vllm)直接以 vllm serve 提供服务:去除 dynamo 前端/worker/router,从根本上规避 dynamo 的 分词器与 schema 兼容问题,OpenAI 前端参数 --enable-auto-tool-choice --tool-call-parser kimi_k3 --reasoning-parser kimi_k3 因此可用。PR #278 仅 支持单节点,故新增补丁将其扩展为 vLLM 原生多节点 serve(--master-addr/ --nnodes/--node-rank,非主节点 headless),以运行 2 节点 TP8xPP2 拓扑。保留 mamba_hybrid 索引类型容器补丁(引擎缺陷与前端无关)。 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: point changelog and MODELS rows at experiment PR #2359 中文:将更新日志条目与 MODELS 表格行链接指向实验 PR #2359。 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: drop gpu-memory-utilization to 0.90 (flashinfer MoE workspace OOM) Same engine-level OOM as the dynamo-frontend variants: the flashinfer trtllm MXFP4 MoE kernel allocates a ~1.6 GiB runtime workspace outside vLLM's memory pool on the first forward; at 0.95 a 178 GiB B200 has only ~1.35 GiB free. 中文:与 dynamo 前端变体相同的引擎级 OOM:flashinfer trtllm MXFP4 MoE 内核在 首个前向时于 vLLM 显存池外分配约 1.6 GiB 工作区,0.95 下仅剩约 1.35 GiB。 改为 0.90。 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: expandable_segments allocator, drop NCCL_CUMEM_ENABLE (prefill OOM) Inherited from the closed dynamo-frontend variants (#2355/#2358): at gpu-mem-util 0.90 the first long-context MLA prefill OOM'd on a 2.92 GiB transient while 3.39 GiB sat reserved-but-unallocated (fragmentation). Set PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True and drop NCCL_CUMEM_ENABLE. 中文:继承自已关闭的 dynamo 前端变体(#2355/#2358):0.90 显存利用率下首个 长上下文 MLA 预填充因 2.92 GiB 瞬时分配 OOM,而 3.39 GiB 处于已保留未分配 状态(碎片化)。设置 PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True 并 移除 NCCL_CUMEM_ENABLE。 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Revert "fix: expandable_segments allocator, drop NCCL_CUMEM_ENABLE (prefill OOM)" This reverts commit 4370988. The superseded direct-vllm run served the agentic benchmark for 24 minutes on the original env (NCCL_CUMEM_ENABLE=1, no expandable_segments) without any OOM — the allocator change was precautionary carryover from the closed dynamo-frontend variants and was never justified by evidence from this serving path. Restore the env that was demonstrably running. 中文:回滚 4370988。被中断的 direct-vllm 运行在原始环境 (NCCL_CUMEM_ENABLE=1、未设 expandable_segments)下已稳定运行智能体基准测试 24 分钟且无 OOM——该分配器改动只是从已关闭的 dynamo 前端变体沿袭的预防性 措施,并无本服务路径上的证据支持。恢复已被验证可运行的环境。 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat: add VLLM_PREFIX_CACHE_RETENTION_INTERVAL=32768 Keep prefix-cache blocks alive across agentic turn gaps, matching the GB200/GB300 AgentX recipes. 中文:新增 VLLM_PREFIX_CACHE_RETENTION_INTERVAL=32768,使前缀缓存块在智能体 回合间隔内保持留存,与 GB200/GB300 AgentX 配方一致。 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat: widen agentic conc list to 1/8/16/32 中文:将智能体并发列表从单点 8 扩展为 1/8/16/32。 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: remove VLLM_PREFIX_CACHE_RETENTION_INTERVAL (K3 scheduler block) Engine init hard-fails on Kimi-K3 with the GB200/GB300 AgentX value: "VLLM_PREFIX_CACHE_RETENTION_INTERVAL (32768) must be non-negative and a multiple of scheduler_block_size (3145728)" — the KDA hybrid architecture gives K3 a 3.1M-token scheduler block. Default retention served fine in the earlier runs, so drop the override. 中文:移除 VLLM_PREFIX_CACHE_RETENTION_INTERVAL——Kimi-K3 的 KDA 混合架构使 scheduler_block_size 达 3145728,GB200/GB300 AgentX 的 32768 取值导致引擎 初始化直接失败(必须为其整数倍)。此前运行证明默认留存策略可正常服务, 故不再覆盖。 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat: agentic experiment G — variant D + prefix-cache retention 0 Identical GPU-resident direct-vllm config to variant D (#2359) plus VLLM_PREFIX_CACHE_RETENTION_INTERVAL=0. Any positive value must be a multiple of Kimi-K3's KDA-hybrid scheduler_block_size (3145728), so 0 is the only setting below one 3.1M-token scheduler block. 中文:智能体实验变体 G——与变体 D(#2359)完全相同的 GPU 常驻直接 vllm serve 配置,另加 VLLM_PREFIX_CACHE_RETENTION_INTERVAL=0(任何正值都必须是 Kimi-K3 KDA 混合架构 scheduler_block_size 3145728 的整数倍,0 是唯一低于 一个 3.1M token 调度块的取值)。 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: point changelog and MODELS rows at experiment PR #2374 中文:将更新日志条目与 MODELS 表格行链接指向实验 PR #2374。 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat: variant G conc curve 1/2/4/8/16/32 (add 2 and 4) 中文:变体 G 并发曲线扩展为 1/2/4/8/16/32(新增 2 与 4)。 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * probe: variant K — drop the kimi-k3 in-container patch script Variant G (#2374, fully green) minus the kimi-k3-container-deps.sh in-container patch (setup_script reference, script file, and launcher copy), to verify whether the mamba_hybrid index_fill_ dtype patch is still required by the current vllm/vllm-openai:kimi-k3 image. Expected to fail at the first decode step if it is; results will be commented on the PR. 中文:探针变体 K——在全绿的变体 G(#2374)基础上移除 kimi-k3 容器内补丁 脚本(setup_script 引用、脚本文件与启动器复制),验证当前 vllm/vllm-openai:kimi-k3 镜像是否仍需 mamba_hybrid index_fill_ 类型补丁。 如仍需要,预计在首个解码步失败;结果将评论在 PR 中。 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: point changelog and MODELS rows at experiment PR #2391 中文:将更新日志条目与 MODELS 表格行链接指向实验 PR #2391。 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat: clone fork branch with multinode support, drop git-apply patch The srt-slurm PR #278 multi-node extension now lives as commits on functionstackx/srt-slurm-nv branch klaud/direct-vllm-multinode (head df5baa93), so the launcher clones that branch directly instead of applying srt-slurm-pr278-direct-vllm-multinode.patch onto the upstream kylliang/direct-aggregate-vllm branch. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Summary
Bring-up configuration: first Kimi-K3 benchmark config in InferenceX; single-concurrency smoke test, validated via the PR sweep before merge.
Adds an aggregated TP8 × PP2 (plain TP — expert parallelism deliberately OFF) Dynamo-vLLM agentic-coding recipe for Kimi-K3 MXFP4 on B200, following the aggregated tep8pp2 pattern established in #2196 and the srt-slurm v1.0.36 AgentX recipe pattern from #2341/#2302:
tp*pp/gpus_per_node = 8*2/8 = 2). Aggregated mode (prefillnum-worker: 1+ decodenum-worker: 0, RECIPES.md §5): a single worker serves both phases, no P/D KV transfer.vllm/vllm-openai:kimi-k3(verified on Docker Hub, amd64 + arm64), withVLLM_ENABLE_K3_LATENT_MOE_TAIL_FUSION=1and--trust-remote-code --load-format fastsafetensors --moe-backend auto --gpu-memory-utilization 0.95 --no-enable-flashinfer-autotune --reasoning-parser kimi_k3.--enable-auto-tool-choice/--tool-call-parserare not set on the worker: the dynamo-vllm worker entrypoint rejects them (unrecognized arguments — second sweep attempt); chat parsing happens at the dynamo frontend, same convention as the DSv4 GB300 agentic recipes. No explicitmax-model-len(vLLM derives the native 1M window from the model config).cam/sa-submission-q2-2026fork rejected the recipe schema (benchmark.aiperf_server_metrics: Unknown field— first sweep attempt).max-model-lenfor the AgentX trace (K3's KDA layers keep per-token KV small; only the 24 gated-MLA layers hold cache);max-num-batched-tokens: 8192so one long prefill cannot OOM a pipeline stage. Smoke test at concurrency 8 only — the conc curve will be widened once the topology is proven green.runners/launch_b200-dgxc.sh): adds thekimik3/fp4model-path mapping (/lustre/fsw/models/Kimi-K3, pre-staged and verified present on the first sweep attempt), overlays the kimi-k3 agentic recipes onto the srt-slurm clone, and adds the agenticdefault_mounts(/aiperf_mmap_cache,/hf_hub_cache) already used by the GB200/GB300 agentic paths.kimik3*already resolves to the default CC-traces weka loader on main (eb30e57).Files:
benchmarks/multi_node/srt-slurm-recipes/vllm/kimi-k3/agentic/agg-b200-tp8pp2-agentic.yaml(new)configs/nvidia-master.yaml: newkimik3-fp4-b200-dynamo-vllm-agenticentry on thecluster:b200-dgxcpoolrunners/launch_b200-dgxc.sh,perf-changelog.yaml,MODELS.md/MODELS_zh.md(PR link on the existing Kimi-K3 row)Validation:
python -m pytest utils/matrix_logic/ -v: 224 passedgenerate_sweep_configs.py full-sweep --model-prefix kimik3yields exactly one agentic matrix entry (conc 8, TP8 PP2 EP1, decode num-worker 0, router 1.2.1,CONFIG_FILErouted to the new recipe)bash -non the launcher pass中文说明
这是 bring-up 配置:InferenceX 中首个 Kimi-K3 基准测试配置;先做单并发冒烟测试,合并前通过 PR 扫描完成集群验证。
为 B200 上的 Kimi-K3 MXFP4 新增聚合式 TP8 × PP2(纯 TP,特意不启用专家并行(EP))Dynamo-vLLM 智能体编码配方,沿用 #2196 建立的聚合式模式与 #2341/#2302 的 srt-slurm v1.0.36 AgentX 配方模式:
num-worker: 1+ 解码num-worker: 0):单个 worker 同时承担预填充与解码,无 P/D KV 传输。vllm/vllm-openai:kimi-k3(已在 Docker Hub 验证存在),启用VLLM_ENABLE_K3_LATENT_MOE_TAIL_FUSION=1及上述 vLLM 启动参数(保留 worker 可接受的 kimi_k3 推理解析器)。--enable-auto-tool-choice/--tool-call-parser不在 worker 上设置:dynamo-vllm worker 入口不接受这两个参数(第二次扫描报 unrecognized arguments),聊天解析由 dynamo 前端处理,与 DSv4 GB300 智能体配方约定一致;不再显式设置max-model-len,由 vLLM 从模型配置推导原生 1M 窗口。benchmark.aiperf_server_metrics: Unknown field)。max-model-len;max-num-batched-tokens: 8192防止超长预填充撑爆流水线阶段。仅并发 8 冒烟测试——拓扑验证通过后再扩展并发曲线。runners/launch_b200-dgxc.sh):新增kimik3/fp4模型路径映射(模型已预置于/lustre/fsw/models/Kimi-K3,首次扫描已确认存在),将 kimi-k3 智能体配方覆盖到 srt-slurm 克隆中,并补充 GB200/GB300 智能体路径已使用的缓存挂载。已完成验证:
utils/matrix_logic224 项测试通过;扫描生成器输出恰为一个智能体矩阵条目(conc 8);YAML 解析与启动器bash -n通过;首次扫描已端到端验证启动器接线(模型路径解析、配方覆盖生效),仅在旧分支 schema 校验处失败,已通过固定 v1.0.36 修复。Related experiments
Three sibling
[agentic experiment]PRs race the same recipe with different worker parser flags; whichever goes green first merges and the others close:--dyn-tool-call-parser kimi_k3 --reasoning-parser kimi_k3 --dyn-reasoning-parser kimi_k3--tool-call-parser kimi_k3 --reasoning-parser kimi_k3— [Klaud Cold][agentic experiment][Variant B] Kimi-K3 B200 agg TP8xPP2 agentic — vLLM tool-call-parser flags / Kimi-K3 B200 聚合式 TP8xPP2 智能体实验——vLLM tool-call-parser 参数 #2357vllm serve(srt-slurm PR 278 + multinode patch), no Dynamo — [Klaud Cold][agentic experiment][Variant D] Kimi-K3 B200 agg TP8xPP2 agentic — direct vllm serve (srt-slurm PR 278) / Kimi-K3 B200 聚合式 TP8xPP2 智能体实验——直接 vllm serve(srt-slurm PR 278) #2359相关实验
三个兄弟
[agentic experiment]PR 以不同的 worker 解析器参数并行竞跑同一配方,先通过者合并、其余关闭:变体 A(本 PR,dynamo --dyn- 命名空间参数)*;变体 B(vLLM 原生写法)#2357;变体 C(不带解析器参数)#2358;变体 D(直接 vllm serve,srt-slurm PR 278 + 多节点补丁,不经 Dynamo)#2359。🤖 Generated with Claude Code