-
Notifications
You must be signed in to change notification settings - Fork 253
[Klaud Cold][agentic experiment][Variant G] Kimi-K3 B200 agg TP8xPP2 agentic — direct vllm serve + prefix-cache retention 0 / Kimi-K3 B200 聚合式 TP8xPP2 智能体实验——直接 vllm serve + 前缀缓存留存 0 #2374
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Closed
functionstackx
wants to merge
21
commits into
main
from
klaud/kimik3-b200-agentic-direct-vllm-retention0
Closed
Changes from all commits
Commits
Show all changes
21 commits
Select commit
Hold shift + click to select a range
b4fd077
feat: add Kimi-K3 MXFP4 B200 aggregated TP8xPP2 Dynamo-vLLM agentic r…
functionstackx 4dbbdc8
docs: link PR #2355 in changelog entry and MODELS rows
functionstackx e8d42a7
fix: pin agentic srt-slurm to NVIDIA v1.0.36, dynamo 1.2.1, conc-8 sm…
functionstackx c1e2a56
fix: drop OpenAI-frontend tool-choice flags from dynamo-vllm worker args
functionstackx ef35fd1
fix: try dynamo wheel 1.2.0.dev20260426 for Kimi-K3 frontend tokenizer
functionstackx be6c56e
fix: pin dynamo to day-zero Kimi-K3 commit, restore kimi_k3 parser flags
functionstackx c6917e6
fix: use dynamo namespaced --dyn-* kimi_k3 parser args on the worker
functionstackx a26853a
fix: disable aiperf conv-aware routing (session_control 400-rejected)
functionstackx f61eafb
fix: patch kimi-k3 image mamba_hybrid index_fill_ dtype via setup_script
functionstackx 8fd319c
feat: agentic experiment D — direct vllm serve via srt-slurm PR #278
functionstackx c0ace4d
docs: point changelog and MODELS rows at experiment PR #2359
functionstackx 862024d
fix: drop gpu-memory-utilization to 0.90 (flashinfer MoE workspace OOM)
functionstackx 4370988
fix: expandable_segments allocator, drop NCCL_CUMEM_ENABLE (prefill OOM)
functionstackx 0c5fe11
Revert "fix: expandable_segments allocator, drop NCCL_CUMEM_ENABLE (p…
functionstackx e675dd2
feat: add VLLM_PREFIX_CACHE_RETENTION_INTERVAL=32768
functionstackx 4b0c3a4
feat: widen agentic conc list to 1/8/16/32
functionstackx 479b74b
fix: remove VLLM_PREFIX_CACHE_RETENTION_INTERVAL (K3 scheduler block)
functionstackx 3824d0e
feat: agentic experiment G — variant D + prefix-cache retention 0
functionstackx d002202
docs: point changelog and MODELS rows at experiment PR #2374
functionstackx 58f1652
Merge remote-tracking branch 'origin/main' into klaud/kimik3-b200-age…
functionstackx f14afdc
feat: variant G conc curve 1/2/4/8/16/32 (add 2 and 4)
functionstackx File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
39 changes: 39 additions & 0 deletions
39
benchmarks/multi_node/srt-slurm-recipes/configs/kimi-k3-container-deps.sh
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,39 @@ | ||
| #!/bin/bash | ||
| # Setup script for the Kimi-K3 vLLM bring-up image (vllm/vllm-openai:kimi-k3). | ||
| # srt-slurm runs this in every worker container before dynamo install and | ||
| # worker startup (recipe field: setup_script). | ||
|
|
||
| set -euo pipefail | ||
|
|
||
| # The image's first decode step crashes in the KDA hybrid-state postprocess: | ||
| # vllm/v1/worker/gpu/model_states/mamba_hybrid.py, postprocess_state: | ||
| # IndexError: index_fill_(): Expected dtype int64 for index. | ||
| # torch's index_fill_ requires an int64 index tensor, but the runner passes | ||
| # the int32 idx_mapping (hit by moonshotai/Kimi-K3 agentic bring-up, first | ||
| # decode step, engine v0.1.dev19262+gb6bbf29dd). Coerce the index to int64. | ||
| # Idempotent: exits 0 if the patch is already applied. | ||
| python3 - <<'PY' | ||
| import pathlib | ||
| import re | ||
|
|
||
| import vllm.v1.worker.gpu.model_states.mamba_hybrid as mh | ||
|
|
||
| path = pathlib.Path(mh.__file__) | ||
| src = path.read_text() | ||
| if "idx_mapping.long()" in src: | ||
| print(f"mamba_hybrid index_fill_ patch already applied: {path}") | ||
| raise SystemExit(0) | ||
|
|
||
| new, n = re.subn( | ||
| r"index_fill_\(\s*0,\s*idx_mapping,", | ||
| "index_fill_(0, idx_mapping.long(),", | ||
| src, | ||
| ) | ||
| if n != 1: | ||
| raise SystemExit( | ||
| f"expected exactly one index_fill_(0, idx_mapping, ...) call in " | ||
| f"{path}, found {n} — image layout changed, refusing to patch" | ||
| ) | ||
| path.write_text(new) | ||
| print(f"Patched mamba_hybrid index_fill_ index dtype: {path}") | ||
| PY |
62 changes: 62 additions & 0 deletions
62
benchmarks/multi_node/srt-slurm-recipes/patches/srt-slurm-pr278-direct-vllm-multinode.patch
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,62 @@ | ||
| diff --git a/src/srtctl/backends/vllm.py b/src/srtctl/backends/vllm.py | ||
| index 74f673b..377606a 100644 | ||
| --- a/src/srtctl/backends/vllm.py | ||
| +++ b/src/srtctl/backends/vllm.py | ||
| @@ -716,25 +716,31 @@ class VLLMProtocol: | ||
| if frontend_type == "vllm": | ||
| if mode != "agg": | ||
| raise ValueError("frontend.type: vllm supports aggregate vLLM jobs only") | ||
| - if is_multi_node: | ||
| - raise ValueError("frontend.type: vllm currently supports single-node aggregate jobs only") | ||
|
|
||
| config.pop("host", None) | ||
| config.pop("port", None) | ||
| config.pop("connector", None) | ||
| config.setdefault("served-model-name", served_model_name) | ||
|
|
||
| - cmd.extend( | ||
| - [ | ||
| - "vllm", | ||
| - "serve", | ||
| - model_arg, | ||
| - "--host", | ||
| - "0.0.0.0", | ||
| - "--port", | ||
| - str(runtime.frontend_port), | ||
| - ] | ||
| - ) | ||
| + node_rank = endpoint_nodes.index(process.node) | ||
| + cmd.extend(["vllm", "serve", model_arg]) | ||
| + if node_rank == 0: | ||
| + cmd.extend(["--host", "0.0.0.0", "--port", str(runtime.frontend_port)]) | ||
| + if is_multi_node: | ||
| + # vLLM-native multi-node serve (torchrun-style): the leader owns | ||
| + # the OpenAI server; other node ranks run headless engine workers. | ||
| + cmd.extend( | ||
| + [ | ||
| + "--master-addr", | ||
| + leader_ip, | ||
| + "--nnodes", | ||
| + str(len(endpoint_nodes)), | ||
| + "--node-rank", | ||
| + str(node_rank), | ||
| + ] | ||
| + ) | ||
| + if node_rank > 0: | ||
| + cmd.append("--headless") | ||
| if not self.set_cuda_visible_devices: | ||
| device_ids = ",".join(str(i) for i in sorted(process.gpu_indices)) | ||
| if device_ids: | ||
| diff --git a/src/srtctl/core/schema.py b/src/srtctl/core/schema.py | ||
| index 1263ddc..0ef7ae4 100644 | ||
| --- a/src/srtctl/core/schema.py | ||
| +++ b/src/srtctl/core/schema.py | ||
| @@ -1587,8 +1587,6 @@ class SrtConfig: | ||
| raise ValidationError("frontend.type: vllm supports aggregate jobs only, not disaggregated layouts") | ||
| if self.resources.num_agg < 1: | ||
| raise ValidationError("frontend.type: vllm requires resources.agg_workers >= 1") | ||
| - if (self.resources.agg_nodes or 1) != 1: | ||
| - raise ValidationError("frontend.type: vllm currently supports single-node aggregate jobs only") | ||
|
|
||
| def _validate_het_jobs(self): | ||
| """When ``resources.het_jobs`` is set to True, enforce supported shape. |
144 changes: 144 additions & 0 deletions
144
benchmarks/multi_node/srt-slurm-recipes/vllm/kimi-k3/agentic/agg-b200-tp8pp2-agentic.yaml
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,144 @@ | ||
| name: "kimik3-vllm-agg-b200-tp8pp2-retention0-agentic" | ||
|
|
||
| # Kimi-K3 MXFP4 B200 AGGREGATED TP8 x PP2 agentic recipe (2 nodes / 16 GPUs). | ||
| # The native MXFP4 checkpoint (2.8T total params, ~1.4TB of weights) does not | ||
| # fit one 8xB200 node, so TP8 shards attention/dense (/8) and PP2 splits the | ||
| # 93 layers (/2) across 16 GPUs. Plain TP (NOT TEP): expert parallelism is | ||
| # deliberately off, so the 896 routed experts are TP-sharded inside each | ||
| # pipeline stage. Node allocation = tp*pp/gpus_per_node = 8*2/8 = 2 nodes. | ||
| # Aggregated (single worker, decode num-worker 0) — no P/D split, no NIXL. | ||
| # VLLM_ENABLE_K3_LATENT_MOE_TAIL_FUSION fuses the K3 LatentMoE tail path in | ||
| # the kimi-k3 bring-up image. | ||
| model: | ||
| path: "kimik3" | ||
| container: "vllm/vllm-openai:kimi-k3" | ||
| precision: "fp4" | ||
|
|
||
| identity: | ||
| model: | ||
| repo: "moonshotai/Kimi-K3" | ||
| container: | ||
| image: "vllm/vllm-openai:kimi-k3" | ||
|
|
||
| # Direct vLLM serving (frontend.type: vllm, srt-slurm PR #278 + the | ||
| # InferenceX multinode patch): `vllm serve` owns the OpenAI port itself, so | ||
| # no Dynamo frontend/worker is involved and no dynamo install is needed. | ||
| dynamo: | ||
| install: false | ||
|
|
||
| # Patches the image's mamba_hybrid postprocess_state: torch index_fill_ | ||
| # requires an int64 index but the runner passes the int32 idx_mapping, | ||
| # crashing the first decode step (IndexError: Expected dtype int64 for index). | ||
| setup_script: kimi-k3-container-deps.sh | ||
|
|
||
| slurm: | ||
| time_limit: "8:00:00" | ||
|
|
||
| health_check: | ||
| interval_seconds: 10 | ||
| max_attempts: 1440 | ||
|
|
||
| resources: | ||
| gpu_type: "b200" | ||
| gpus_per_node: 8 | ||
| agg_nodes: 2 | ||
| agg_workers: 1 | ||
| gpus_per_agg: 16 | ||
|
|
||
| infra: | ||
| etcd_nats_dedicated_node: false | ||
| nats_max_payload_mb: 32 | ||
|
|
||
| frontend: | ||
| # Direct vLLM OpenAI server (srt-slurm PR #278): the vllm serve leader owns | ||
| # the public port; rank-1 runs a headless engine worker (vLLM-native | ||
| # multi-node TP8xPP2 via --master-addr/--nnodes/--node-rank, enabled by | ||
| # patches/srt-slurm-pr278-direct-vllm-multinode.patch). | ||
| type: vllm | ||
| enable_multiple_frontends: false | ||
|
|
||
| backend: | ||
| type: vllm | ||
| connector: null | ||
| aggregated_environment: | ||
| VLLM_ENABLE_K3_LATENT_MOE_TAIL_FUSION: "1" | ||
| VLLM_SERVER_DEV_MODE: "1" | ||
| # ~1.4TB of MXFP4 weights off shared Lustre: keep the engine-ready window | ||
| # generous, and let one long AgentX request hold a PP stage beyond vLLM's | ||
| # 300-second model-execution default. | ||
| VLLM_ENGINE_READY_TIMEOUT_S: "3600" | ||
| VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS: "1800" | ||
| # Prefix-cache retention (variant G): 0, on the otherwise-unchanged | ||
| # GPU-resident variant D config. Any positive value must be a multiple of | ||
| # Kimi-K3's KDA-hybrid scheduler_block_size (3145728; the GB recipes' | ||
| # 32768 is hard-rejected at engine init — verified on this PR family), so | ||
| # 0 is the only value below one 3.1M-token scheduler block. | ||
| VLLM_PREFIX_CACHE_RETENTION_INTERVAL: "0" | ||
| # No VLLM_PREFIX_CACHE_RETENTION_INTERVAL: the GB200/GB300 AgentX value | ||
| # (32768) hard-fails engine init on Kimi-K3 — the KDA hybrid gives it a | ||
| # scheduler_block_size of 3145728 and the interval must be a multiple of | ||
| # it ("VLLM_PREFIX_CACHE_RETENTION_INTERVAL (32768) must be non-negative | ||
| # and a multiple of scheduler_block_size (3145728)"). Default retention | ||
| # served fine in earlier runs. | ||
| NCCL_CUMEM_ENABLE: "1" | ||
| TILELANG_CLEANUP_TEMP_FILES: "1" | ||
| UCX_MEMTYPE_CACHE: "n" | ||
| UCX_MEMTYPE_REG_WHOLE: "n" | ||
| UCX_NET_DEVICES: "mlx5_0:1,mlx5_1:1,mlx5_2:1,mlx5_3:1,mlx5_4:1,mlx5_5:1,mlx5_10:1,mlx5_11:1" | ||
| HF_HUB_CACHE: "/hf_hub_cache" | ||
| HUGGINGFACE_HUB_CACHE: "/hf_hub_cache" | ||
| vllm_config: | ||
| aggregated: | ||
| served-model-name: "moonshotai/Kimi-K3" | ||
| tensor-parallel-size: 8 | ||
| pipeline-parallel-size: 2 | ||
| trust-remote-code: true | ||
| load-format: fastsafetensors | ||
| moe-backend: auto | ||
| # 0.90, not 0.95: the flashinfer trtllm MXFP4 MoE kernel allocates a | ||
| # ~1.6 GiB runtime workspace OUTSIDE vLLM's memory pool on the first | ||
| # forward; at 0.95 a 178 GiB B200 has only ~1.35 GiB free and the first | ||
| # warmup request OOMs (seen on the dynamo-frontend variants). 0.90 | ||
| # matches the GB200/GB300 agentic recipes. | ||
| gpu-memory-utilization: 0.90 | ||
| no-enable-flashinfer-autotune: true | ||
| # kimi_k3 parsers via the native vllm serve OpenAI-frontend flags — | ||
| # legitimate here because this recipe serves directly with vllm serve | ||
| # (frontend.type: vllm), not through the dynamo worker entrypoint that | ||
| # rejects them. | ||
| enable-auto-tool-choice: true | ||
| tool-call-parser: kimi_k3 | ||
| reasoning-parser: kimi_k3 | ||
| # No explicit max-model-len: let vLLM derive the native 1M window from | ||
| # the model config (agentic trajectories blow past any small cap, and | ||
| # K3's KDA layers keep per-token KV small — only the 24 gated-MLA | ||
| # layers hold cache). Prefix caching stays on (default) for trajectory | ||
| # reuse. Cap prefill chunks so a single long request cannot OOM a | ||
| # pipeline stage; let vLLM pick max-num-seqs. | ||
| max-num-batched-tokens: 8192 | ||
|
|
||
| sbatch_directives: | ||
| segment: "1" | ||
|
|
||
| srun_options: | ||
| container-remap-root: "" | ||
|
|
||
| benchmark: | ||
| type: custom | ||
| command: bash /infmax-workspace/benchmarks/multi_node/agentic_srt.sh | ||
| env: | ||
| INFMAX_CONTAINER_WORKSPACE: "/infmax-workspace" | ||
| RESULT_DIR: "/logs/agentic" | ||
| PORT: "8000" | ||
| # Keep the aggregate worker in the multinode result schema so ingestion | ||
| # uses the zero decode-worker count instead of duplicating TP into P and D. | ||
| IS_MULTINODE: "true" | ||
| # aiperf's conv-aware routing emits nvext.session_control, a removed POC | ||
| # field this dynamo build 400-rejects at warmup (schema moved to | ||
| # router/routing_constraints/agent_hints). Same opt-out as the GB300 | ||
| # aggregate AgentX recipes — and with a single aggregate worker there is | ||
| # no P/D routing to bind anyway. | ||
| AIPERF_USE_DYNAMO_CONV_AWARE_ROUTING: "0" | ||
| AIPERF_DATASET_MMAP_CACHE_DIR: "/aiperf_mmap_cache" | ||
| HF_HUB_CACHE: "/hf_hub_cache" | ||
| WEKA_LOADER_OVERRIDE: "semianalysis_cc_traces_weka_062126" | ||
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Oops, something went wrong.
Oops, something went wrong.
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
🟡 Lines 77-82 contain a stale comment block copy-pasted from Variant D's config claiming VLLM_PREFIX_CACHE_RETENTION_INTERVAL is intentionally left unset — but line 76 immediately above it sets that exact variable to "0", which is this PR's headline change. The two blocks now directly contradict each other; the stale block should be deleted since the correct rationale is already given in lines 71-75.
Extended reasoning...
What the bug is: In
benchmarks/multi_node/srt-slurm-recipes/vllm/kimi-k3/agentic/agg-b200-tp8pp2-agentic.yaml, theaggregated_environmentblock contains two adjacent, mutually-contradictory comments about the same environment variable:# Prefix-cache retention (variant G): 0, on the otherwise-unchanged GPU-resident variant D config. Any positive value must be a multiple of Kimi-K3's KDA-hybrid scheduler_block_size (3145728; the GB recipes' 32768 is hard-rejected at engine init — verified on this PR family), so 0 is the only value below one 3.1M-token scheduler block.VLLM_PREFIX_CACHE_RETENTION_INTERVAL: "0"# No VLLM_PREFIX_CACHE_RETENTION_INTERVAL: the GB200/GB300 AgentX value (32768) hard-fails engine init on Kimi-K3 ... Default retention served fine in earlier runs.How it manifests / code path: This PR's entire stated purpose (per the description) is "Variant D (#2359) plus
VLLM_PREFIX_CACHE_RETENTION_INTERVAL=0" — isolating this one knob against D's baseline, where the variable was left unset. The diff correctly adds lines 71-76 to set it. But the config was assembled by copying Variant D's file wholesale, and Variant D's original comment explaining why the variable was left unset (lines 77-82) was never deleted. It now sits directly beneath a line that sets the very variable it claims is absent.Why existing code/review doesn't catch it: YAML treats
#-prefixed lines as comments, so nothing in CI, schema validation, or the live config is affected — the env var is emitted exactly once, as"0", with no duplicate key. This is a pure documentation defect, which is why all three independent verifiers confirmed zero functional/runtime impact.Impact: Purely a readability/maintenance hazard, but a meaningful one here: a future engineer reading this recipe (e.g. while interpreting Variant G's benchmark results, or copying this file forward into a future variant) will hit two adjacent comments asserting opposite things about the exact variable that is this PR's sole experimental variable, and will have to consult git blame or the PR history to determine which is current.
Fix: Delete the stale block at lines 77-82 (the "No VLLM_PREFIX_CACHE_RETENTION_INTERVAL: ..." comment inherited from Variant D). The accurate rationale — including the same hard-fail-at-32768 fact — is already fully captured in lines 71-75, so nothing is lost by removing the duplicate/contradictory block.
Step-by-step proof:
VLLM_PREFIX_CACHE_RETENTION_INTERVAL: "0".# No VLLM_PREFIX_CACHE_RETENTION_INTERVAL:— i.e., claims the variable is absent.