-
Notifications
You must be signed in to change notification settings - Fork 253
[Klaud Cold][agentic experiment][Variant C] Kimi-K3 B200 agg TP8xPP2 agentic — no parser flags / Kimi-K3 B200 聚合式 TP8xPP2 智能体实验——不带解析器参数 #2358
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Closed
Closed
Changes from all commits
Commits
Show all changes
13 commits
Select commit
Hold shift + click to select a range
b4fd077
feat: add Kimi-K3 MXFP4 B200 aggregated TP8xPP2 Dynamo-vLLM agentic r…
functionstackx 4dbbdc8
docs: link PR #2355 in changelog entry and MODELS rows
functionstackx e8d42a7
fix: pin agentic srt-slurm to NVIDIA v1.0.36, dynamo 1.2.1, conc-8 sm…
functionstackx c1e2a56
fix: drop OpenAI-frontend tool-choice flags from dynamo-vllm worker args
functionstackx ef35fd1
fix: try dynamo wheel 1.2.0.dev20260426 for Kimi-K3 frontend tokenizer
functionstackx be6c56e
fix: pin dynamo to day-zero Kimi-K3 commit, restore kimi_k3 parser flags
functionstackx c6917e6
fix: use dynamo namespaced --dyn-* kimi_k3 parser args on the worker
functionstackx 4586383
feat: parser-flag experiment C — no parser flags on the worker
functionstackx 0ada69d
docs: point changelog and MODELS rows at experiment PR #2358
functionstackx 6ae0aaf
fix: disable aiperf conv-aware routing (session_control 400-rejected)
functionstackx ab8b422
fix: patch kimi-k3 image mamba_hybrid index_fill_ dtype via setup_script
functionstackx 98969cf
fix: drop gpu-memory-utilization to 0.90 (flashinfer MoE workspace OOM)
functionstackx ea62af4
fix: expandable_segments allocator, drop NCCL_CUMEM_ENABLE (prefill OOM)
functionstackx File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
39 changes: 39 additions & 0 deletions
39
benchmarks/multi_node/srt-slurm-recipes/configs/kimi-k3-container-deps.sh
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,39 @@ | ||
| #!/bin/bash | ||
| # Setup script for the Kimi-K3 vLLM bring-up image (vllm/vllm-openai:kimi-k3). | ||
| # srt-slurm runs this in every worker container before dynamo install and | ||
| # worker startup (recipe field: setup_script). | ||
|
|
||
| set -euo pipefail | ||
|
|
||
| # The image's first decode step crashes in the KDA hybrid-state postprocess: | ||
| # vllm/v1/worker/gpu/model_states/mamba_hybrid.py, postprocess_state: | ||
| # IndexError: index_fill_(): Expected dtype int64 for index. | ||
| # torch's index_fill_ requires an int64 index tensor, but the runner passes | ||
| # the int32 idx_mapping (hit by moonshotai/Kimi-K3 agentic bring-up, first | ||
| # decode step, engine v0.1.dev19262+gb6bbf29dd). Coerce the index to int64. | ||
| # Idempotent: exits 0 if the patch is already applied. | ||
| python3 - <<'PY' | ||
| import pathlib | ||
| import re | ||
|
|
||
| import vllm.v1.worker.gpu.model_states.mamba_hybrid as mh | ||
|
|
||
| path = pathlib.Path(mh.__file__) | ||
| src = path.read_text() | ||
| if "idx_mapping.long()" in src: | ||
| print(f"mamba_hybrid index_fill_ patch already applied: {path}") | ||
| raise SystemExit(0) | ||
|
|
||
| new, n = re.subn( | ||
| r"index_fill_\(\s*0,\s*idx_mapping,", | ||
| "index_fill_(0, idx_mapping.long(),", | ||
| src, | ||
| ) | ||
| if n != 1: | ||
| raise SystemExit( | ||
| f"expected exactly one index_fill_(0, idx_mapping, ...) call in " | ||
| f"{path}, found {n} — image layout changed, refusing to patch" | ||
| ) | ||
| path.write_text(new) | ||
| print(f"Patched mamba_hybrid index_fill_ index dtype: {path}") | ||
| PY |
136 changes: 136 additions & 0 deletions
136
benchmarks/multi_node/srt-slurm-recipes/vllm/kimi-k3/agentic/agg-b200-tp8pp2-agentic.yaml
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,136 @@ | ||
| name: "kimik3-vllm-agg-b200-tp8pp2-agentic" | ||
|
|
||
| # Kimi-K3 MXFP4 B200 AGGREGATED TP8 x PP2 agentic recipe (2 nodes / 16 GPUs). | ||
| # The native MXFP4 checkpoint (2.8T total params, ~1.4TB of weights) does not | ||
| # fit one 8xB200 node, so TP8 shards attention/dense (/8) and PP2 splits the | ||
| # 93 layers (/2) across 16 GPUs. Plain TP (NOT TEP): expert parallelism is | ||
| # deliberately off, so the 896 routed experts are TP-sharded inside each | ||
| # pipeline stage. Node allocation = tp*pp/gpus_per_node = 8*2/8 = 2 nodes. | ||
| # Aggregated (single worker, decode num-worker 0) — no P/D split, no NIXL. | ||
| # VLLM_ENABLE_K3_LATENT_MOE_TAIL_FUSION fuses the K3 LatentMoE tail path in | ||
| # the kimi-k3 bring-up image. | ||
| model: | ||
| path: "kimik3" | ||
| container: "vllm/vllm-openai:kimi-k3" | ||
| precision: "fp4" | ||
|
|
||
| identity: | ||
| model: | ||
| repo: "moonshotai/Kimi-K3" | ||
| container: | ||
| image: "vllm/vllm-openai:kimi-k3" | ||
| frameworks: | ||
| dynamo: "1.4.0-kimi-k3-dev.1" | ||
|
|
||
| dynamo: | ||
| install: true | ||
| # Day-zero Kimi-K3 dynamo ("feat: Added support for Kimi-K3", tag | ||
| # v1.4.0-kimi-k3-dev.1 == this commit): adds the kimi_k3 tiktoken tokenizer | ||
| # to the rust frontend (dynamo <=1.2.1 only knows kimi/kimi_k2/kimi_k25/ | ||
| # deepseek_v3, so the model never registers and every request 404s) and the | ||
| # kimi_k3 tool-call/reasoning parser worker args. | ||
| hash: "ba83080ecd31c1ce918559e576d3c5bc9e092ff1" | ||
|
|
||
| # Patches the image's mamba_hybrid postprocess_state: torch index_fill_ | ||
| # requires an int64 index but the runner passes the int32 idx_mapping, | ||
| # crashing the first decode step (IndexError: Expected dtype int64 for index). | ||
| setup_script: kimi-k3-container-deps.sh | ||
|
|
||
| slurm: | ||
| time_limit: "8:00:00" | ||
|
|
||
| health_check: | ||
| interval_seconds: 10 | ||
| max_attempts: 1440 | ||
|
|
||
| resources: | ||
| gpu_type: "b200" | ||
| gpus_per_node: 8 | ||
| agg_nodes: 2 | ||
| agg_workers: 1 | ||
| gpus_per_agg: 16 | ||
|
|
||
| infra: | ||
| etcd_nats_dedicated_node: false | ||
| nats_max_payload_mb: 32 | ||
|
|
||
| frontend: | ||
| type: dynamo | ||
| enable_multiple_frontends: false | ||
|
|
||
| backend: | ||
| type: vllm | ||
| connector: null | ||
| aggregated_environment: | ||
| VLLM_ENABLE_K3_LATENT_MOE_TAIL_FUSION: "1" | ||
| VLLM_SERVER_DEV_MODE: "1" | ||
| # ~1.4TB of MXFP4 weights off shared Lustre: keep the engine-ready window | ||
| # generous, and let one long AgentX request hold a PP stage beyond vLLM's | ||
| # 300-second model-execution default. | ||
| VLLM_ENGINE_READY_TIMEOUT_S: "3600" | ||
| VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS: "1800" | ||
| # expandable_segments: the eighth sweep attempt OOM'd on a 2.92 GiB MLA | ||
| # long-context prefill transient (kv_b_proj in _compute_prefill_context) | ||
| # while 3.39 GiB sat reserved-but-unallocated — fragmentation, exactly | ||
| # the case this allocator mode fixes (and what the torch OOM message | ||
| # recommends; the DSv4 recipes set it too). NCCL_CUMEM_ENABLE dropped at | ||
| # the same time to trim NCCL's share of non-PyTorch device memory. | ||
| PYTORCH_CUDA_ALLOC_CONF: "expandable_segments:True" | ||
| TILELANG_CLEANUP_TEMP_FILES: "1" | ||
| UCX_MEMTYPE_CACHE: "n" | ||
| UCX_MEMTYPE_REG_WHOLE: "n" | ||
| UCX_NET_DEVICES: "mlx5_0:1,mlx5_1:1,mlx5_2:1,mlx5_3:1,mlx5_4:1,mlx5_5:1,mlx5_10:1,mlx5_11:1" | ||
| HF_HUB_CACHE: "/hf_hub_cache" | ||
| HUGGINGFACE_HUB_CACHE: "/hf_hub_cache" | ||
| vllm_config: | ||
| aggregated: | ||
| served-model-name: "moonshotai/Kimi-K3" | ||
| tensor-parallel-size: 8 | ||
| pipeline-parallel-size: 2 | ||
| trust-remote-code: true | ||
| load-format: fastsafetensors | ||
| moe-backend: auto | ||
| # 0.90, not 0.95: the flashinfer trtllm MXFP4 MoE kernel allocates a | ||
| # ~1.6 GiB runtime workspace OUTSIDE vLLM's memory pool on the first | ||
| # forward; at 0.95 a 178 GiB B200 has only ~1.35 GiB free and the first | ||
| # warmup request OOMs (seventh sweep attempt). 0.90 matches the | ||
| # GB200/GB300 agentic recipes. | ||
| gpu-memory-utilization: 0.90 | ||
| no-enable-flashinfer-autotune: true | ||
| # No parser flags on the worker (parser-flag experiment variant C — | ||
| # sibling PRs try dynamo's --dyn-* namespaced args and the plain vLLM | ||
| # spellings). Chat parsing happens at the dynamo frontend, matching the | ||
| # DSv4 GB300 agentic recipes, which run parser-less. | ||
| # No explicit max-model-len: let vLLM derive the native 1M window from | ||
| # the model config (agentic trajectories blow past any small cap, and | ||
| # K3's KDA layers keep per-token KV small — only the 24 gated-MLA | ||
| # layers hold cache). Prefix caching stays on (default) for trajectory | ||
| # reuse. Cap prefill chunks so a single long request cannot OOM a | ||
| # pipeline stage; let vLLM pick max-num-seqs. | ||
| max-num-batched-tokens: 8192 | ||
|
|
||
| sbatch_directives: | ||
| segment: "1" | ||
|
|
||
| srun_options: | ||
| container-remap-root: "" | ||
|
|
||
| benchmark: | ||
| type: custom | ||
| command: bash /infmax-workspace/benchmarks/multi_node/agentic_srt.sh | ||
| env: | ||
| INFMAX_CONTAINER_WORKSPACE: "/infmax-workspace" | ||
| RESULT_DIR: "/logs/agentic" | ||
| PORT: "8000" | ||
| # Keep the aggregate worker in the multinode result schema so ingestion | ||
| # uses the zero decode-worker count instead of duplicating TP into P and D. | ||
| IS_MULTINODE: "true" | ||
| # aiperf's conv-aware routing emits nvext.session_control, a removed POC | ||
| # field this dynamo build 400-rejects at warmup (schema moved to | ||
| # router/routing_constraints/agent_hints). Same opt-out as the GB300 | ||
| # aggregate AgentX recipes — and with a single aggregate worker there is | ||
| # no P/D routing to bind anyway. | ||
| AIPERF_USE_DYNAMO_CONV_AWARE_ROUTING: "0" | ||
| AIPERF_DATASET_MMAP_CACHE_DIR: "/aiperf_mmap_cache" | ||
| HF_HUB_CACHE: "/hf_hub_cache" | ||
| WEKA_LOADER_OVERRIDE: "semianalysis_cc_traces_weka_062126" |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Oops, something went wrong.
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
🟡 Documentation-only: the new header comment for kimik3-fp4-b200-dynamo-vllm-agentic in configs/nvidia-master.yaml (lines 8336-8339) states the setup uses 'the kimi_k3 tool-call/reasoning parsers,' but this is variant C, which deliberately passes NO parser flags to the worker (chat parsing happens at the dynamo frontend instead). The comment appears to be copied from the base variant-A PR (#2355) and contradicts both the recipe yaml's own comment and this PR's own perf-changelog entry, which correctly describe the no-parser-flags behavior.
Extended reasoning...
This PR adds a new header comment block to
configs/nvidia-master.yaml(lines 8325-8339) documenting the newkimik3-fp4-b200-dynamo-vllm-agenticconfig entry. The final sentence of that comment reads: "Dedicated kimi-k3 vLLM bring-up image withVLLM_ENABLE_K3_LATENT_MOE_TAIL_FUSION=1and the kimi_k3 tool-call/reasoning parsers." The "with X and Y" phrasing pairs two things the recipe supposedly configures — andVLLM_ENABLE_K3_LATENT_MOE_TAIL_FUSION=1genuinely is set in the recipe'saggregated_environment, so a reader naturally takes "the kimi_k3 tool-call/reasoning parsers" to also be something this recipe configures on the worker.That is incorrect for this PR. This PR is explicitly variant C of a three-way parser-flag experiment (per its own title, "no parser flags", and its description: "no tool-call/reasoning parser flags are passed to the worker"). The actual recipe yaml this config points to,
benchmarks/multi_node/srt-slurm-recipes/vllm/kimi-k3/agentic/agg-b200-tp8pp2-agentic.yaml(lines 84-87), spells this out unambiguously: "No parser flags on the worker (parser-flag experiment variant C — sibling PRs try dynamo's--dyn-*namespaced args and the plain vLLM spellings). Chat parsing happens at the dynamo frontend." Thevllm_config.aggregatedblock in that same file contains no--tool-call-parser,--reasoning-parser,--dyn-tool-call-parser, or--dyn-reasoning-parserkeys at all.The stale comment almost certainly originates from the base bring-up PR, #2355 (variant A), which does pass
--dyn-tool-call-parser kimi_k3 --reasoning-parser kimi_k3 --dyn-reasoning-parser kimi_k3to the worker — for that PR the "and the kimi_k3 tool-call/reasoning parsers" clause would have been accurate. When the comment block was copied over to build out this variant-C sibling PR, the parser-flag sentence was not updated to match variant C's defining behavior (passing none).Tellingly, two other artifacts touched by this exact same PR were correctly updated to describe the no-parser-flags behavior: the recipe yaml comment quoted above, and the new
perf-changelog.yamlentry for this same config key, which says "...--trust-remote-code, and NO parser flags on the worker (parser-flag experiment variant C; chat parsing happens at the dynamo frontend...)". So within this one PR, two of three descriptive artifacts say "no parser flags" and the third (the nvidia-master.yaml header) says the opposite.Step-by-step proof:
configs/nvidia-master.yaml:8336-8338(added by this PR): "...Dedicated kimi-k3 vLLM bring-up image withVLLM_ENABLE_K3_LATENT_MOE_TAIL_FUSION=1and the kimi_k3 tool-call/reasoning parsers." → implies the worker runs the kimi_k3 parsers.agg-b200-tp8pp2-agentic.yaml:84-87: "No parser flags on the worker (parser-flag experiment variant C ...)." → the worker runs no parsers.vllm_config.aggregatedin that recipe file — there is notool-call-parserorreasoning-parserkey present, confirming point 2 is what actually gets passed to the worker at runtime.perf-changelog.yamlentry added by this same PR for the identical config key — it explicitly says "NO parser flags on the worker (parser-flag experiment variant C...)", matching point 2 and contradicting point 1.Impact: this is comment-only — it has no effect on what actually gets executed (the recipe yaml, not the master-config comment, drives the worker's CLI args), so it does not change behavior or benchmark results. But it is misleading to anyone reading
nvidia-master.yamlto understand what this config does, and it directly contradicts the recipe file and changelog entry sitting right next to it in the same PR.Suggested fix: drop the "and the kimi_k3 tool-call/reasoning parsers" clause from the nvidia-master.yaml comment, or replace it with something like "and no parser flags on the worker (parser-flag experiment variant C; chat parsing happens at the dynamo frontend)" to match the recipe yaml and changelog wording.