Skip to content

WIP: MiniMax-M3 support — GQA + MSA block-sparse attention, o200k tokenizer, converter (follow-up to #418)#601

Open
steve-m wants to merge 53 commits into
JustVugg:devfrom
steve-m:minimax-m3
Open

WIP: MiniMax-M3 support — GQA + MSA block-sparse attention, o200k tokenizer, converter (follow-up to #418)#601
steve-m wants to merge 53 commits into
JustVugg:devfrom
steve-m:minimax-m3

Conversation

@steve-m

@steve-m steve-m commented Jul 24, 2026

Copy link
Copy Markdown
Contributor

What this is

Early, testable support for MiniMax-M3 (428B total / 23B active MoE, the text backbone of the VL checkpoint) — GQA attention, the MSA "Lightning Indexer" block-sparse attention, the o200k-family tokenizer, converter support, and Vulkan offloads. Opening as WIP so people who want to play with M3 can start early; API/knobs may still shift.

Based on #418 (Vulkan backend) — this is a follow-up. #418's commits are re-included here (rebased onto current dev) until that PR merges; once it does, this branch rebases down to the M3 commits only. Of the 53 commits, roughly the first 38 are the #418 stack plus its follow-ups (int3-g64 fmt=5 in the shaders, VRAM pressure-proofing, N-way mirror replicas + striped reads, multi-worker pilot prefetch, optional second-device tier) — please review only the last 15 commits as new material here (bc6adf3 and up).

What works

  • Engine (ARCH_M3, auto-detected from model_type): GQA 64Q/4KV attention with per-head Gemma QK-norm, partial split-half NEOX RoPE (rotary 64, θ=5e6), swigluoai activation, sigmoid+bias router with renormalized top-4 ×2.0, first-3-dense pattern, 64K context.
  • MSA (Lightning Indexer) block-sparse attention: the 4-head scoring branch, causal max-pool into 128-token blocks, top-16 + forced local block, per-KV-group selection. Validated bit-exact against a numpy oracle on a tiny checkpoint (24/24 teacher-forced prefill + 20/20 incremental decode) and with a real-model needle-in-haystack retrieval at 2.5K tokens. Decode attention is capped at ~2048 attended tokens regardless of context. Containers converted without the indexer weights fall back to full causal attention automatically (exact ≤2048 tokens); add them later with convert --indexer.
  • Tokenizer: tok.h now accepts classic space-separated BPE merges (o200k family) alongside the array form; MiniMax chat template (]~b]role … [e~[) in the serve loop and openai_server.py.
  • KV persistence includes the MSA index-key cache, so resumed conversations stay correct past 2048 tokens.
  • Converter (tools/convert_fp8_to_int4.py --arch m3, auto-detected): maps the VL checkpoint onto the container scheme (drops vision), int4-g64 or int3-g64 (--xbits 3) routed experts, indexer kept at int8 (--idx-bits), per-shard resume.
  • Vulkan: expert tier runs M3 (swigluoai as a shader push-constant), generic dense/attention-projection offload (COLI_VK_DENSE=1), fused shared expert, lm_head offload, opt-in GQA attention core (COLI_VK_ATTN=1, throughput-neutral on the reference box — see commit messages).
  • AVX2 kernel for int3-g64 (matmul_i3 was scalar on x86): 3.8 → 24 GB/s effective on a Zen2 12-core, which is what makes int3 containers viable for decode.

Measured (reference box: Ryzen 9 3900X, 64 GB DDR4, RX 9070 16 GB, 2× NVMe striped mirror)

config decode tok/s
CPU only, single drive, int4 0.67
+ VK expert tier 0.91
+ 2-drive mirror ~1.5
+ 512-expert tier ~2.1
+ VK dense/lm_head/paired submits 2.8 (topp 0.7) / 3.5 (0.5)
int3-g64 container + AVX2 kernel 3.4 (topp 0.7) / 3.8 (0.5)

Prefill 516 tokens: 152.7 → 96.8 s over the same span. --topp here is the expert-routing top-p (~228 → ~155 expert loads/token at 0.7).

How to test

Without the weights (validates the engine math end-to-end, minutes):

python3 tools/make_m3tiny.py /tmp/m3tiny
python3 tools/convert_fp8_to_int4.py --indir /tmp/m3tiny --outdir /tmp/m3tiny_i8 --ebits 8 --io-bits 8
python3 tools/oracle_m3.py /tmp/m3tiny_i8          # writes ref_m3.json
SNAP=/tmp/m3tiny_i8 REF=ref_m3.json TF=1 TF_DECODE=1 IDOT=0 ./colibri 8   # expect 24/24 + 20/20

With the real model (BF16 download is ~796 GB; the int4 container is 225 GB, int3 is 177 GB):

# convert (≈35 min from local NVMe; add --xbits 3 for int3-g64: 25% smaller, needs the AVX2 kernel in this branch)
python3 tools/convert_fp8_to_int4.py --indir <MiniMax-M3-BF16> --outdir <container> --group-size 64
make colibri VK=1
COLI_MODEL=<container> DIRECT=1 PIPE=1 COLI_NO_OMP_TUNE=1 \
COLI_VULKAN=1 COLI_VK_SHADERS=$PWD/shaders/qmatmul.spv \
COLI_VK_EXPERTS=512 COLI_VK_RESERVE_GB=7 COLI_VK_DENSE=1 \
./coli chat --ctx 65536 --topp 0.7

A second copy of the container on another drive via COLI_MODEL_MIRROR=<dir> roughly doubles streaming bandwidth. CPU-only works (drop the COLI_V* vars) but is ~5× slower.

Note the model's custom license (minimax-community): convert your own weights; don't redistribute containers.

Known gaps / WIP

  • Cosmetics: banner still prints the GLM identity; PROF verdict mentions DSA while MSA is active
  • No COLI_MSA=0 kill switch for A/B against full attention
  • Tool calls not yet rendered in the MiniMax chat template (server returns 400)
  • Routing-topp quality evaluation is spot-checks only so far (0.5 looks fine, 0.7 is the conservative default)
  • Rebase onto dev once Vulkan backend: expert tier + dense + MLA attention on any Vulkan 1.2 GPU (successor to #84) #418 merges (drops the first ~38 commits)

steve-m added 30 commits July 24, 2026 14:10
… RX 9070

backend_vulkan.c/.h + shaders/qmatmul.comp: opt-in Vulkan int4/int8 quantized
GEMV via RADV, mirroring coli_cuda_matmul. Shader decodes int4 as nibble-8
(offset-binary), numerically consistent with the CPU path.

Validated on the RX 9070 (RADV GFX1201, Mesa 26.1): the built-in VK_TEST harness
passes all int4/int8 cases vs the CPU reference (maxrel ~1e-4). Bypasses ROCm's
dropped Polaris support -> also targets the RX 580.

Head-to-head naive int4 GEMV on the RX 9070 (same shapes, synchronous per-call):
ROCm 0.115ms vs Vulkan 0.306ms (6144->1536, S=1) -- ROCm ~2.7x faster as-is
(naive shader + host-visible memory + heavy submit overhead; not the tuned
coopmat path that wins in llama.cpp). Next: optimize the shader/backend to reach
or beat ROCm. Not yet wired into glm.c (Makefile/hooks pending).
Optimizations toward matching ROCm on the RX 9070:
- backend: cache descriptor set + resubmit the prerecorded command buffer when
  tensor/shape/scratch are unchanged (single int4 GEMV 6144->1536: 0.306->0.214 ms).
- shader: llama.cpp-style mul_mat_vec (MIT techniques, per-row int4/int8): x staged
  once in shared memory, one subgroup per output row + subgroupAdd (no barrier tree),
  grid-stride rows. Needs SPIR-V 1.3 (glslc --target-env=vulkan1.2). Correct (maxrel ~1e-4).
- VK_TEST: batched throughput probe (N dispatches / one submit).

Honest result (noise-controlled, vs colibri's production coli_cuda_expert_group):
production ROCm ~0.179 ms/expert (fused dual gate+up + down, batched); our unfused VK
~0.260 ms/expert -- ~45% slower. VK wins the short-reduction down-proj, loses the
long-reduction gate/up. Reaching ROCm needs fused dual gate+up + expert batching +
a shape-adaptive reduction.
…pert

New qmatmul_gate_up.comp + coli_vk_gate_up: computes hidden=silu(gate(x))*up(x) in
ONE dispatch, reading x once for both projections (VK equivalent of colibri's
grouped_hidden_w4_dual). Second 6-binding compute pipeline; build_pipeline() helper
refactors the pipeline setup for both. Correct vs CPU ref (maxrel ~9e-5).

The fusion is the win: gate+up drops from ~0.13 ms (two separate matmuls) to ~0.080
ms/expert (fair, 8 DISTINCT experts cycled so weights come from VRAM not L2 — added
bench_experts_fair to control for the caching artifact). Full fused expert:
  VK  gate_up ~0.080 + down ~0.050 = ~0.13 ms/expert
  ROCm production coli_cuda_expert_group = 0.179 ms/expert (stable)
-> VK ~25-30% FASTER, holding under cache-controlled measurement (some run variance
0.07-0.10 on gate_up). Reaching this needed the fused dual projection, exactly the
optimization the production HIP path already had and our unfused VK lacked.
…one submit)

The real engine primitive: K experts, fused gate+up+silu then down, hidden staying
on-device, all in one submit. Per-expert descriptor sets (gate_up: 6-binding,
down: 4-binding) sliced into packed x/hidden/y via descriptor offsets; one phase
barrier between gate_up and down. Mirrors coli_cuda_expert_group; correct vs CPU
ref (maxrel ~7e-4 at K<=8).

Honest throughput (RX 9070, K=8, distinct experts):
  GPU-only (resubmit recorded cmd buffer): 0.113 ms/expert -- BEATS ROCm 0.179 (~37%)
  per-call (as-is API):                    1.03  ms/expert -- 5.6x SLOWER than ROCm
The GPU compute is genuinely faster (fused kernel delivers), but the per-call HOST
setup (80 descriptor updates + recording 16 dispatches every call) dominates. The
earlier microbench 0.13 was the GPU-only number; it did not capture per-call cost.

Fix (next): cache descriptor sets + command buffer across calls (hot experts are
reused across tokens), or a bindless/BDA single-dispatch-multi-expert design like
ROCm's grouped kernels. Either drops per-call toward the 0.11 GPU-only floor.
(K=32 maxrel 2e-3, slightly over 1e-3 threshold -- fp accumulation, worth a look.)
…s ROCm

The per-call cost (1.03 ms/expert) was ALL in reading the output back: eg_y/y/h were
allocated in write-combined DEVICE_LOCAL (ReBAR) memory, which the CPU reads at ~40
MB/s. VK_PROF breakdown (K=1): memcpy_x 0.005 | desc 0.027 | record 0.005 | gpu 0.19
| memcpy_y 0.598 ms -- the readback dominated; descriptor updates + recording were
negligible (so no descriptor/command caching needed).

Fix: pick_memtype_cached() (HOST_VISIBLE|HOST_COHERENT|HOST_CACHED) for buffers the CPU
reads back (eg_y, y, h); inputs/hidden stay write-combined (fast CPU writes / GPU-only).

Result (RX 9070, K=8, distinct experts): per-call expert_group 1.03 -> 0.117 ms/expert,
now == the 0.111 GPU-only floor. vs ROCm 0.179 -> VK ~35% FASTER end-to-end, per-call.
The real primitive now beats ROCm; hypothesis confirmed. (K=32 maxrel 2e-3 is fp32
precision on the 6144-elem reduction, fine for greedy argmax.) VK_PROF=1 env-gated.
make glm VK=1 compiles backend_vulkan.o (plain C + vulkan headers), builds the .comp
shaders to .spv via glslc (--target-env=vulkan1.2), links -lvulkan. Independent of
CUDA/HIP (own -DCOLI_VULKAN). Verified on the RX 9070: glm links libvulkan.so.1,
both shaders compiled. glm.c hooks (COLI_VULKAN) come next; default build unchanged.
End-to-end integration of coli_vk_expert_group into the engine:
- QT gains a resident ColiVkTensor *vk (+vk_eligible); qt_vk_reset frees it when a
  slot is reused for another expert (expert_load_impl hook, mirrors qt_cuda_reset),
  so the LRU never computes with stale weights. g_vk_resident tracks the tier size.
- coli_vk_init at startup (COLI_VULKAN=1), shader path COLI_VK_SHADERS.
- moe() VK path (decode S<=4): upload routed int4 experts to VK once (capped by
  COLI_VK_EXPERTS, default 1024), compute the resident ones as one batched
  coli_vk_expert_group (fused gate+up+silu -> down, on-device), CPU-fallback the rest.
- coli_vk_tensor_ensure() backend entry: upload a resident tensor without computing.

Verified on the RX 9070 (make glm VK=1): [VK] expert tier active, greedy decode of
'The capital of France is' -> 'Paris.' CORRECT. Default build (no VK=1) unchanged.

NOTE: a VK-only build has no GPU dense/attention offload (that is the HIP CUDA_DENSE/
COLI_CUDA_ATTN path), so end-to-end tok/s here is CPU-dense-bound (0.06 cold), NOT a
measure of the expert path -- which is the ~35%-faster-than-ROCm primitive. A hybrid
HIP-dense + VK-experts build, or a VK dense/attn port, is the next step for throughput.
…d expert on Vulkan

vk_matmul_qt(t,y,x,S) routes a resident int4/int8 matmul_qt through coli_vk_matmul
(uploads the weight once into t->vk, then reuses it). Wired, with CPU fallback, into
the decode attention projections (q_a, q_b, kv_a, o) and the shared expert
(sh_gate/up/down). Env COLI_VK_DENSE=1.

Together with the expert tier, a Vulkan-only machine now runs experts + dense
projections + shared expert on the GPU; only the MLA attention core (absorb/softmax/
RoPE, small in latent space) and I/O stay on CPU. Verified on the RX 9070
(make glm VK=1, COLI_VULKAN=1 COLI_VK_DENSE=1): greedy 'The capital of France is'
-> 'Paris.' correct. Default build unchanged.

NEXT for a fully-GPU Vulkan path: a dedicated MLA attention-core compute shader
(scores/softmax/weighted-values/RoPE) — the last CPU-bound piece.
…te piece on Vulkan-only machines

New shaders/attention_absorb.comp runs the whole decode absorb core for one
layer in ONE dispatch, one workgroup per (query row, head): absorbed query
(q_nope through the int4/int8 kv_b nope rows), scores over the cache window,
softmax, weighted latent, and the value-row projection. Subgroup-per-token
score dots + subgroup-per-row value projection, q/qabs/clat staged in shared.

The KV latent/rope cache is mirrored in persistent per-layer device buffers
(coli_vk_kv_ensure/_row/_reset), appended ~2.3 KB/token/layer instead of
re-uploading the window each call — same design as the CUDA kv_dev shadow,
with the same invalidation points (row rewrite, kv_bind, kv_alloc resize)
tracked by a vk_kv_valid watermark in glm.c.

Falls back to CPU on DSA top-k selection, ragged KV, the MTP layer, or any
backend failure (mirrors COLI_CUDA_ATTN's guards). Opt-in via COLI_VK_ATTN=1.

Validated on RX 9070 (RADV): 5/5 CPU-ref harness cases maxrel <= 1.7e-4
(GLM decode shape, kv_start window, S=2 causal, int8, T=2000); engine output
byte-identical to the pre-change build under the same env; decode
score-softmax-value 1.75s -> 0.61s per 16 tokens (2.9x) with prefill
untouched. Remaining VK perf work: resident-on-device layer pipeline and a
shape-adaptive long-reduction for the o-projection.
…he parallel CPU loop)

Previously budget 0 still entered the vk_active block, sending every routed
expert through its serial CPU fallback. With the gate, dense+attention can run
on Vulkan while routed experts keep the normal parallel CPU path — measured
the best Vulkan-only config on the RX 9070 (1.31-1.37 tok/s vs 1.46-1.50 HIP;
the GPU expert tier as integrated is slower, 1.11, because uploads and submits
sit serially on the decode critical path).

Bench note: Vulkan-only runs on this box need COLI_NO_OMP_TUNE=1 — the OMP
self-tune (skipped under COLI_CUDA, so HIP never hit it) sets
OMP_WAIT_POLICY=active + GOMP_SPINCOUNT=200000, and 12 pinned spinning threads
starve the PIPE I/O pool and pilot worker (CPU expert rows 28 -> 5 GB/s,
0.52 tok/s). A standalone probe confirmed VK init itself is harmless.
…pert (1 submit each)

CORRECTNESS: qmatmul.comp staged x into shared xsh[6144] unconditionally; the
o-projection's input row is H*vh = 16384, so every VK dense o-proj since the
dense port computed with a truncated/undefined activation tail — deterministic,
so greedy output was plausible and stable, which masked it. Rows with
I > 6144 now skip staging and read x from the storage buffer (uniform branch,
coalesced; harness case fmt=2 I=16384 O=6144 added, maxrel 2e-4). With the fix
the VK build's greedy output matches pure CPU exactly ('Paris.') — the earlier
divergence was this bug, not fp tie-breaks. gate_up/expert_group get D<=6144
host guards (their shader shares the pattern; engine dims are within bounds).

PERF (toward HIP parity):
- coli_vk_attention_absorb_project: absorb + resident o-projection in ONE
  submit, ctx stays on-device (was: absorb submit + ctx readback + o submit).
  Harness: fused 0.51-0.53 ms/call vs 0.73 unfused absorb ALONE, maxrel <= 7e-5.
- Shared expert now runs as coli_vk_expert_group(count=1): fused gate+up+silu
  -> down, hidden on-device, one fence instead of three matmul submits.
- expert_group harness threshold 3e-3 for K>=32 (documented fp32 accumulation
  on the 6144-length double reduction; was a standing false FAIL at 2e-3).
Both projections read the same x row(s): coli_vk_matmul_pair stages x once,
records both dispatches, and waits one fence — replacing two full submit+wait
roundtrips per layer per token. Harness-validated (maxrel 4.4e-5, 0.123 ms/pair
vs ~0.35 for two singles); engine greedy still matches pure CPU ('Paris.').
Also: expert_group harness threshold 3e-3 at any K (the fp32 accumulation
lives in the 6144-length reduction chain, not in the expert count).
…ch overlapped with CPU rows

Replaces the LRU-slot-tied tier (which uploaded on the decode critical path,
churned with evictions, and computed ALL routed experts on the GPU while 12
cores idled — measured slower than CPU-only experts).

- vk_registry_fill(): at startup, upload the top-COLI_VK_EXPERTS experts by
  persistent usage history into a (layer,eid) registry, decoupled from cache
  slots — stable residency like the HIP VRAM tier (RX 9070: 256 experts,
  4.83 GB, 1.0s; pinned RAM slots feed uploads directly, the rest stream
  through one transient slot).
- coli_vk_expert_group_issue/_take: the group gets its own command buffer +
  fence and splits into submit-and-return / join, so moe() ISSUES the GPU
  batch, computes the CPU share concurrently, drains the pipe + loads the
  GPU-side slots (cache invariants unchanged), then takes and accumulates.
  Sync coli_vk_expert_group (shared expert, harness) wraps the same path.
- vk_active now keys on a non-empty registry; empty (no history/budget 0)
  falls back to the normal parallel CPU loop.

Harness: issue/take validated bit-identical to the sync path; full PASS.
Engine: 'Paris.' correct with the tier active.
…etch entirely

The resolve phase now classifies registry-resident experts FIRST (before
pin/LRU): they take no cache slot, dispatch no load, and skip the LRU
recency bump — so they age out of RAM and free capacity for the CPU-served
experts, the same effect CUDA_RELEASE_HOST gives the HIP tier. The pilot
worker, cross-layer lookahead, and next-block readahead treat registry
residency like RAM residency (decode only; prefill still loads normally
since the tier serves S<=4). Counted as a new 'vk' bucket in the hit-rate
split. The group block computes GPU entries straight from the registry
(no ESlot), attributes its CPU share to t_ecpu/rows and only the take-wait
to t_egpu, and on device-loss falls back via a transient ws slot load.
…ath was inverted

Without a RAM pin every candidate loads transiently, and the inverted check
skipped every SUCCESSFUL load (pin-fed fills masked it: the pins covered the
whole top-256). PIN_GB=0 + COLI_VK_EXPERTS=256 now fills from disk in ~1.6s.
Also: bail after 64 consecutive load failures, diagnostics on the first few.
256-384 measured flat within noise (1.73-1.78 tok/s); 320 is the best median
with ~2.3 GB VRAM headroom left for the long-context KV mirror.
Upstream PR JustVugg#399 (KV8 fp8 / TQ4 quantized latent KV) leaves Lc/Rc NULL when a
quantized tier is active — the mirror sync would deref NULL. Guard on the
arrays themselves (env-independent), falling back to the CPU attention path;
an fp8-aware VK mirror (upload Lc8+scale, decode e4m3 in the shader, 4x less
mirror traffic) is the follow-up once JustVugg#399 merges.
Upstream JustVugg#168 (merged as 5e42e70) now carries the engine int3-g64 support this
commit used to port, textually near-identical — including the uring_finalize_load
qt_resolve_fmt fix we carried separately. What remains ours is the Vulkan side:
vk_matmul_qt/pair accept fmt 5, the absorb kv_b/o and shared-expert paths feed
fmt-5 tensors, and vk_registry_fill passes the true fmt through xf/d.fmt instead
of assuming int4.
qmatmul/qmatmul_gate_up/attention_absorb learn the 24B-per-64-group two-plane
layout: lanes stride the 16-value low-plane words, each word's partial dot is
multiplied by its group scale before the subgroup reduction, and the per-row
scale multiply is skipped. The absorb query pass stops pre-folding the kv_b
row scale into q for fmt=5 (scales vary along K) and applies the (row,group)
scale per element instead. upload_tensor accepts fmt=5 (rows are already
word-aligned: ceil(I/64)*24), sizes the scale buffer O*ceil(I/64) via
scale_floats() mirrored in tensor_free's accounting, and the expert-group
prep validates up/down fmts (down may differ per-projection; phase 2 pushes
downs[0]->fmt). Engine gates opened: vk_matmul_qt/pair, absorb kv_b+o, shared
expert, and the pinned tier registry (int4/int3 accepted, fmt passed through).
VK_TEST: fmt=5 cases across dense (incl. tail group + unstaged o_proj shape),
fused gate_up, expert group K=8/32, absorb (+causal window), int3-vs-int4
fair throughput lines. Shaders validated against glslangValidator vulkan1.2.
- sample.h: COLI_LOGIT_DUMP=1 prints top-5 (id:logit) per pick_tok step —
  the tool that separated backend error from fp tie-flips (VK matched CPU
  logits to ~1e-4 at every matched step; run-to-run token divergence turned
  out to be engine-wide threading jitter, present on int4 CPU-only too).
- harness: fmt=5 matmul_pair case (the q_a+kv_a decode path) and count=1
  expert group (the shared-expert shape) — the two production paths the
  first harness round left uncovered.
- tier messages made format-neutral (int4/int3-g64).
… — classify the VK-path call sites

demand=1 on the moe CPU-share miss load (same semantics as the upstream miss
path: routing-driven, FASE A snapshot valid); demand=0 on the device-lost
VK-fallback reload (bookkeeping for vk-served experts diverges from the
snapshot's assumptions — DISK-CLASS leaves it unclassified, per its own doc)
and on the startup tier fill (matches the PIN-load sites).
pipe_ready() (non-blocking peek of the load-done flag; URING reports not-ready
and defers to the wait) lets the block's CPU share compute loaded/resident
experts FIRST, so a still-loading expert's I/O hides behind their matmuls
instead of head-of-line blocking. Ordering is a hint only — every entry still
pipe_waits/loads in its body. Measured neutral on a mirror config (waits
already near their floor) and aimed at wait-heavy setups (single disk: felt
waits 20-37s). t_ecpu now counts kernel time only, like the default path
(it used to swallow the in-loop waits); waits land in t_ewait.
classify / issue / take+accumulate wall plus avg nvk/ncpu per block, printed
under PROF next to t_ecpu/t_ewait/t_egpu — enough to see where the big-tier
regression (420: 1.44 vs 950: 1.20 with LOWER felt wait) actually spends its
extra 9 s: submit overhead growth, take waits, or outside the block entirely
(dense/absorb under VRAM pressure).
…t-capped tier fill

The big-tier regression was never the tier mechanics (VK-BLOCK: classify+issue+
take <=0.3s/run) — it was decode attention degrading with VRAM occupancy (7.8s
at 7.6 GB resident -> 15.1s at 11.6 -> 17.8s at 15.2, identical workload): once
expert weights crowd the heap, the kernel evicts and re-migrates buffers under
the per-token attention submits.

Two runtime-detected extensions fix the two halves:
- VK_EXT_memory_priority: allocations now carry an eviction class — scratches
  and the KV mirror pin at 1.0 (they ride every submit), dense weights 0.75,
  and the engine brackets the expert-tier fill at 0.4. An oversubscribed heap
  sheds cold tier experts instead of thrashing attention.
- VK_EXT_memory_budget: vk_registry_fill checks the device-local budget every
  8 uploads and stops while COLI_VK_RESERVE_GB (default 3) is still free for
  the lazily-allocated dense + KV mirror + staging, logging the stop. Reserve
  0 disables the stop; without the extension the count cap applies as before.

Both are no-ops on drivers without the extensions.
… submit tax

Isolation matrix + harness probes pinned the big-tier regression precisely:
per-submit driver cost grows ~linearly with the number of distinct device-memory
objects the queue actively references (~0.35 ms/submit extra at a 950-expert
tier's ~5.7k allocations), hitting every synchronous fence-waited path — the
dense projections (4-5 submits/layer, accounted in t_attn) most of all. Decode
attention: 7.9s @420 experts -> 17.6s @950; flat when dense+attn run on CPU;
IDLE allocations are free (ballast probe: absorb and dense GEMV unchanged under
5000 dummy allocations); eviction exonerated (2.9 GB free showed the same tax,
scaling continuous in tier size, no budget-edge cliff).

upload_tensor now binds wbuf/sbuf into shared 256 MB arena blocks (alignment-
respecting bump allocator, memory-priority tag carried over) instead of one
vkAllocateMemory per buffer: a 950-expert tier + dense drops from ~7k memory
objects to ~60. Arena slices are not reclaimed per-tensor (registry/dense live
for the process; the rare fill-failure free leaks its slice — tensor mem
handles stay VK_NULL_HANDLE, vkFreeMemory's documented no-op). Arenas are
unmapped/freed at shutdown. Scratches and the KV mirror stay standalone (few,
constant count).
… semantics

The merged JustVugg#298 gave the CUDA kernels per-group scale handling for fmt=4 —
without the same semantics a g64 container on Vulkan would decode with per-row
scales (the exact bug JustVugg#298 fixed). All three shaders gain a fmt==4 branch:
int4 nibble decode with one scale per gs inputs, applied to the packed-word
partial (host gates gs to multiples of 8 so a word never straddles a group;
per-row scaling is skipped like fmt=5). Group size flows as an explicit
parameter through the upload-triggering entries into ColiVkTensor and the
push constants; engine call sites pass QT.gs and the VK gates accept fmt=4
via VK_FMT_OK (word-aligned gs only — anything else stays on the CPU path,
which JustVugg#298 already fixed).

Harness: ref helpers generalized to runtime group size (g_ref_gs); fmt=4
cases at gs=64 across the real shapes (dense, o-proj, batch, expert_group,
matmul_pair, absorb incl. S=2 causal + window) plus a gs=32 sanity case.
The attention section paid three synchronous fence roundtrips per layer per
token (pair, q_b, absorb+o); the middle one existed only because the q-latent
RMS norm ran on the CPU between q_a and q_b. A 40-line rmsnorm.comp (one
workgroup per row; fp32 vs the CPU's f64 accumulate differs ~1e-7 at D=1536)
lets coli_vk_attn_qprep record the whole chain in one command buffer with
compute barriers: [q_a + kv_a] -> norm -> q_b, one fence. q, the kv latent,
and the NORMED q latent return to the host (the DSA indexer consumes the
latter from QR; RoPE + the canonical KV append stay CPU-side). Norm weights
upload once per layer (KV-mirror pattern). The engine tries the chain first
and falls back whenever rmsnorm.spv is absent, formats mismatch, or the call
fails; COLI_VK_QPREP=0 forces the split path, =2 runs both and prints
per-layer chain-vs-split divergence.

Validation: harness chain cases vs a matmul->rmsnorm->matmul CPU reference
(int8/int4, S=1/2/11, real GLM shapes; threshold 1e-2 — two quantized
reductions + the norm through cancellation-heavy random rows compound ~10x a
single GEMV, same reasoning as the expert_group 3e-3). Engine COLI_VK_QPREP=2
on the real container: all 78 layers within 1.8e-4 of the split path, kv
bit-exact; decode-context logits match to 3e-5. Step-1 logits shift O(1) with
the chain on — measured to be the container's sensitivity to ANY fp-order
change, not a chain defect: pure-CPU vs split-VK step-1 logits differ ~4.0 on
the same prompt with the same stable argmax.
steve-m and others added 23 commits July 24, 2026 14:11
Generalize the dual-SSD mirror to N drives: COLI_MODEL_MIRROR takes a
';'/','-separated list of byte-identical model copies, COLI_DISK_WEIGHTS
one positive weight per drive, and expert_route hashes into cumulative
per-replica cuts instead of a single primary/mirror threshold. The
startup bandwidth probe measures every drive. st.h grows per-replica
fd rows (st_mirror_add/st_mirror_reset; st_mirror_init stays as the
single-mirror wrapper) and test_st_mirror covers the multi-replica
paths. Cold decode is disk-bound (~0.8 GB/token): the third NVMe adds
its full bandwidth to the stripe.
The ~8 GB of dense weights (attention projections, absorb kv_b, o-proj,
shared expert) uploaded lazily on the first forwards — after the expert
tier had already filled to its budget. On a 16 GB card the late arrivals
overflowed to GTT (measured 2.1 GB spilled with a 320-expert int4 tier)
and every per-token attention submit paid PCIe latency: absorb core
1.69 -> 2.26 ms/call. Claiming the dense set at startup lets the
budget-capped tier fill self-size to what actually fits: zero spill,
decode attention 13.1 -> 9.4 s per tg64, +7% decode.
Every sync submit's vkWaitForFences pays a scheduler wake (~50-150 us)
on signal; the engine fences ~2 submits per layer per token. Poll
vkGetFenceStatus for a short budget first (default 300 us, the common
decode dispatch completes in 0.5-2 ms), then fall back to the blocking
wait. The spinning thread is stalled on the GPU result anyway.
COLI_VK_SPIN_US=0 restores the pure blocking wait.
…RIPE)

A cold ~19 MB expert read is single-thread latency-bound (~4 GB/s on
one NVMe) and is the felt wait of every demand miss. Split the one
coalesced O_DIRECT pread into disjoint 4K-aligned chunks, one per
replica that holds the shard, read in parallel; chunk 0 stays on the
hash-routed replica so first-chunk load spreads evenly. O_DIRECT only
(no page cache, so no double-caching); any short stripe falls back to
the whole read on the routed replica. Measured on the 3-drive box:
felt wait 22.8 -> 18.6 s per tg128, decode +6% (1.66 -> 1.76 tok/s,
new record) with the drives already near thermal limits.
…guard, rebased on current dev

Deterministic rebase of the multi-worker pilot onto current dev (which carries the
merged JustVugg#474 eviction guard). dev's only colibri.c delta since the branch base was the
guard itself, in the original pilot_realload structure; this branch carries the guard
re-applied into the SPMC multi-worker structure, so the two overlap — resolved by
reparenting the byte-identity-verified tree onto dev.

Contents: SPMC pilot ring + PILOT_WORKERS (default 1 = today's behavior), visible
-(eid+2) reservations on the blocking loader, adversarial-review hardening,
tests/test_pilot_ring.c (+ Makefile rule), and the JustVugg#474 LFRU guard in both
pilot_realload and pilot_uring_batch.

Draft base for the JustVugg#441 good-hardware A/B; not for merge until it shows a win.
Byte-identical output across PILOT off / workers 1|8 / guard on|off (sha256, 9/9),
make check 111/111.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…ide the primary

A self-contained context on a second Vulkan GPU hosts ONLY tier experts
and runs ONLY the async expert-group path; attention, dense, q-prep and
the KV mirror stay on device 0, and both groups fly simultaneously
(slower device issued first, taken last). The dev2 tier fills with the
next heat-ranked experts after dev0's budget stop (COLI_VK_EXPERTS2,
COLI_VK_RESERVE2_GB). COLI_VK_DEV2=auto picks the best real GPU that is
not device 0; an explicit index allows the same physical device (second
logical device — pre-hardware test mode). Tensors carry their device
(coli_vk_tensor_dev) so the moe() partition, free and byte accounting
route correctly. Device-0 hot path untouched: the shared helpers only
gained explicit device/phys parameters. Smoke-tested on RX 9070 +
RX 580 (POLARIS10, chipset x4): tier fill, dual issue/take, correct
output.
The Polaris/x4-slot submit path costs ~0.8ms per dev2-active block
(3.2s/decode serialized, vs 0.15s dev0-only) while the 580's exec
finishes long before take (take-wait 0.1s). Spawn issue2 on a pthread
and join right before take2: the submit cost overlaps dev0's issue and
the CPU expert share. G2 state is private to issue2/take2 and the join
precedes take2, so the one-in-flight invariant is unchanged.
VK_PROF=1 now splits the dev2 issue cost (memcpy_x/desc/record/submit,
[VK_PROF d2iss] every 2048 calls) and VK-BLOCK reports the issue
worker's own elapsed vs the join wait. This localized the dev2 tax to
vkQueueSubmit at 1.33ms/call — the runtime-suspended headless 580
parked at 300MHz/gen1 by DPM during microsecond-burst decode — not
engine code. With the card kept awake: submit 48us, decode 2.07 tok/s.
Maps the minimax_m3_vl VL checkpoint onto the GLM container scheme so the
engine's MoE loader stays single-naming: strips the language_model. prefix,
block_sparse_moe -> mlp (router gate.weight + e_score_correction_bias line
up with the existing loader), Mixtral w1/w3/w2 -> gate/up/down_proj, drops
the vision tower + the MSA index branch (a later pass, like DSA), keeps
q/k/v_proj under the attn bits class, q_norm/k_norm as f32. Writes a
flattened text_config as the container config.json and carries
chat_template.jinja. Validated on a synthetic tiny checkpoint: full name
flow + bit-exact int4-g64 vs quant_int4_grouped.
…igluoai, Gemma norms

model_type minimax_* in config.json switches the engine to the M3 family:
- load_cfg: M3 key set (num_local_experts, head_dim, num_key_value_heads,
  rotary_dim, dense/shared intermediate, swiglu params, moe_layer_freq ->
  first_dense); the GQA KV rows ride the MLA cache aliases (Lc=K, Rc=V,
  kv_lora/qk_rope = n_kv_heads*head_dim) so kv_alloc/bind/persist/mux work
  unchanged; arch-split config validation.
- loader: self_attn.{q,k,v,o}_proj + per-head q_norm/k_norm for ARCH_M3;
  MoE/router/shared names already line up via the converter mapping.
- attention_gqa(): per-head Gemma QK-norm BEFORE partial split-half (NEOX)
  RoPE on the first rotary_dim dims, standard per-token KV rows, grouped
  scores/softmax/values (repeat_kv h -> h/(H/NK)), o_proj; honors the ragged
  kvs/positions contract; full causal attention (MSA block selection exact
  <=2048 ctx, deferred beyond).
- act_glu(): all 11 silu(g)*u sites now dispatch silu vs swigluoai
  (clamps, alpha, and the (up+1) factor); rmsnorm gains the Gemma (1+w)
  variant behind g_gemma_norm. Both default off -> GLM math untouched.
- VK: expert tier + fused shared expert gated off for M3 (shaders hardcode
  silu); the per-matmul dense chain stays available.
- chat: MiniMax template (]~!b[ / ]~b]role / [e~[) in the serve dialog loop,
  coli run, and openai_server (text subset; tools 400 for now); [e~[ EOS
  fallback.
- tools: make_m3tiny.py (tiny random M3 checkpoint) + oracle_m3.py (numpy
  reference on the DEQUANTIZED container -> the engine's REF/TF gate).
TF_LOGITS=1 prints every teacher-forcing logit row (oracle bisection);
TF_DECODE=1 re-runs the same REF continuation through the incremental
S=1 decode path (fresh kv_alloc, prefill np then per-token steps) and
compares argmax against tf_pred — validates KV append across calls,
which the batch check alone cannot see. M3 tiny results: prefill 24/24
(max logit diff 3.2e-6 vs the numpy oracle with IDOT=0 exact kernels;
the idot default's activation quant explains the earlier 23/24),
decode 20/20.
…mily)

tok.h required each tokenizer.json merge to be a JSON array ["left","right"]
(the newer HF tokenizers form GLM ships). MiniMax-M3's tokenizer uses the classic
string form "left right" — so every merge failed and the tokenizer aborted at
entry 0, leaving the prompt un-encodable (the M3 engine loaded the 225 GB int4
container fine, then produced nothing). Accept both: for the string form, split
on the first space, which is unambiguous because byte-level BPE encodes spaces as
U+0120 and no token contains a literal space. GLM's array form is unchanged.
… M3 expert tier

The fused gate+up shader hardcoded silu(gate)*up, so the VK expert tier was gated
off for MiniMax-M3 (swigluoai) and its experts ran on the CPU. Add a push-const-
selected activation: act=1 computes (up+1)*gate*sigmoid(alpha*gate) with gate/up
clamped to +/-limit (alpha 1.702, limit 7.0), matching act_glu and the numpy oracle;
act=0 keeps GLM's silu bit-identical. The activation is model-global, set once via
coli_vk_set_activation from vk_registry_fill (g_act_swigluoai -> the mode), read into
every gate_up dispatch's push constants (both the G and dev2 G2 paths).

Un-gates vk_registry_fill for M3. Validated on the real 225 GB int4-g64 container,
RX 9070: expert tier fills (384 resident, 10.9 GB, vk 40% of lookups), output stays
correct ('Paris' greedy + coherent prose), decode 0.64 -> 0.91 tok/s (+42%) with the
tier active. Attention (GQA) and the shared expert still run on the CPU (P4).
New attention_gqa.comp: one workgroup per (query row, head) computes
scores/softmax/weighted-V for GQA (query head h reads KV group h/(H/NK)),
reading the persistent on-device K/V mirror — K in the L buffer, V in the R
buffer, both NK*hd wide (the mirror already stores two arbitrary-width rows;
the absorb shader's R<=64 cap is its own, not the mirror's). coli_vk_gqa_attn
mirrors the absorb dispatch (5 bindings, one submit+fence+readback per layer).
Wired into attention_gqa behind COLI_VK_ATTN, single-sequence decode only,
CPU fallback on any failure.

Validated on the real 225 GB container, RX 9070: output correct ('Paris'
greedy + coherent prose) with the core active. HONEST PERF NOTE: it is
throughput-NEUTRAL on this box and stays opt-in/default-off, because the core
is only ~0.7s of an ~8.6s decode-attention cost — the projections dominate
(q/k/v proj+norm+rope ~4.8s, o-proj ~2.9s) and are still on the CPU. The per-
layer submit/readback slightly exceeds the tiny core it replaces. Offloading
the projections (fused, GLM-qprep style, to amortize submits) is the actual
lever and the next step. -Wno-missing-field-initializers quiets the shared
struct PC's zero-filled tail.
…evice)

coli_vk_gqa_attn_project runs the GQA core and the o-projection in one command
buffer: the core writes ctx into the att_ctx device scratch (never read back), a
compute barrier, then the o-proj matmul reads it and only [S,hidden] returns to
the host — the absorb_project pattern, for GQA. attention_gqa tries it first
(fills out directly, skips the CPU o-proj), then the core-only path, then CPU.

Validated correct ('Paris' greedy). Profile confirms the offload: decode
output-projection drops to 0.000s on the CPU, and the core+o-proj VK dispatch
(~2.3s/64tok) replaces CPU core+o-proj (~3.9s) — a ~1.6s attention saving.

HONEST e2e NOTE: still throughput-neutral on this box (fused 1.68/1.71 vs CPU
1.70/1.73). M3 decode here is expert-matmul + disk bound (expert-matmul ~16-17s,
disk-wait ~6s, attention ~9s of a ~34s decode; the Zen2 CPU is AVX2-maxed with
no VNNI, and the 16 GB VRAM caps the expert tier at 512). The saving is real but
masked by expert/disk variance. Kept behind COLI_VK_ATTN (default off for M3);
it pays off on a compute-bound or faster-disk box and sets up full projection
fusion.
M3's native attention is block-sparse beyond 2048 tokens: a small 4-head indexer
scores every key, max-pools into 128-token blocks, and the main GQA attention only
reads the top-16 blocks (+ the local block) per KV group. Our port ran full causal
attention (exact <=2048 ctx, divergent + O(ctx) beyond). This adds the real MSA.

Converter: keep + quantize the 228 indexer tensors (index_{q,k}_proj int8,
index_{q,k}_norm f32) instead of dropping them; --indexer mode gains M3 support for a
supplemental add-on pass (no full re-convert). The sparse config already rides the
flattened container config.

Engine: reuses the DSA index-key cache (Ic) + index_hd/index_nh/idx_type slots (M3's
sparse layers == MoE layers). load_cfg reads sparse_attention_config (index dim 128,
4 heads, block 128, top-16, local 1); guards the GLM DSA index read from clobbering
M3's fields. attention_gqa: the indexer (Gemma-norm + same partial-ROT rope as the
main attention, idx_k mirrored to Ic) runs per sparse layer, block-scores in f64, and
picks top-k blocks (local forced in; greedy ties -> lowest index). The main core then
attends only the selected blocks. VK attention is skipped on sparse layers.

Validated on a tiny indexer-bearing checkpoint (block_size 4 so 24 tokens span 6
blocks): engine vs numpy oracle bit-exact — PREFILL 24/24 + incremental DECODE 20/20
(IDOT=0). The 2 near-tie mismatches during bring-up were a config bug (JSON-bool
use_sparse_attention + the unguarded GLM index read zeroing index_hd), now fixed.
…ights absent

A container converted before the indexer was kept still announces sparse_attention_
config, so the loader would try (and fail) to load index_q_proj on every sparse layer.
Probe the first sparse layer's indexer weight; if absent, disable MSA (full causal
attention, exact <=2048 ctx) and point the user at the --indexer pass. Same
auto-detect discipline as the DSA indexer / MTP.
Three fixes so a container converted before the indexer was kept can get it added
without a full 34-min re-convert:
 - drop the '--arch m3 --indexer not implemented' guard (only --mtp is refused now);
 - the shard pre-filter matched GLM's 'indexer' name only — also match M3's
   'self_attn.index_' so the right shards are read;
 - layer_idx() only parsed 'model.layers.N...'; find 'layers.N' anywhere so it also
   handles M3's 'language_model.model.layers.N...' VL prefix.

Extracts the 228 indexer weights (+scales) to out-idx-*.safetensors; drop them beside
the main container (st.h scans any *.safetensors). Applied to the box container +
mirror: MSA now active on the real model (no [MSA] fallback, output correct).
kv_persist.h gated the Ic rows on has_dsa (GLM DSA), which is false for
M3 (index_topk=0), so .coli_kv persisted K/V but not the MSA index keys:
a resumed conversation carried uninitialized Ic for every restored
position, silently corrupting block selection once the conversation
passed 2048 tokens (topk_blocks * block_size). Gate on (has_dsa || msa)
— the same condition kv_alloc uses for Ic — in the header, record size,
append and load. Old M3 .coli_kv files now mismatch the header
(h[3] 0 -> index_hd) and take the existing 'different model or version'
reject path (start over, no misparse).

Validated on the real 225 GB container:
- record size 245764 -> 274948 B/token (= +57 sparse layers x 128 f32)
- old-format file: clean reject + start-over
- end-to-end: a 2363-token needle document (needle in block 15 of 19)
  saved by one process; a FRESH process resumed 2365 tokens (0.1 s, no
  re-prefill) and answered the needle question exactly. Without the fix
  the zero-key tie-break selects blocks {0..14, local} and drops block
  15, so this recall only works through the restored index keys.
M3's GQA has no MLA low-rank bottleneck: q_proj and o_proj are 50 MB of
int8 each across 60 layers, ~6 GB of weight reads EVERY token that rode
the same ~28 GB/s DDR4 the routed experts stream over — the reason a
428B/23B-active model barely beat the 744B GLM (whose attention/dense
lives in VRAM via CUDA_DENSE). matmul_qt_ex had Metal and CUDA branches
but no Vulkan one; COLI_VK_DENSE only wired GLM's MLA call sites.

Add a generic VK branch to matmul_qt_ex behind COLI_VK_DENSE=1, gated on
a per-tensor vk_gemm flag set at load for M3's resident weights (q/k/v/o,
indexer, dense MLP; shared experts already had wired sites) — routed
expert QTs are slab-transient and never marked. At decode only tensors
>= COLI_VK_GEMM_MB (default 8) are worth the per-call submit; a batched
forward (S>=8) offloads every marked tensor. Weights upload lazily at
the dense priority class (0.75 > tier 0.4), so the launcher reserve
grows to 8 GB and the expert tier self-sizes into the rest.

Measured (RX 9070, 512->~170-200-expert tier, 2-drive mirror, ngen 64):
  decode topp 0.7: 2.12-2.20 -> 2.6-2.84 tok/s (+25-30%)
  decode topp 0.5: 2.49 -> 3.0-3.10 tok/s
  prefill 516 tok: 152.7 -> 102.3 s (-33%); attention projections -67%
Output verified (greedy Paris + coherent lighthouse prose). Launchers
updated to COLI_VK_DENSE=1 RESERVE=8 and the stale MSA caveat rewritten.
Four levers on top of the VK dense matmul (95a565d):
 - lm_head vk_gemm: the 1.2 GB int8 logit matmul rides COLI_VK_DENSE.
 - q+k pair: k_p's 3 MB joins q_p's submit via vk_matmul_pair_qt (a lone
   k submit costs more than its CPU read; paired it is free).
 - fused shared expert un-gated for M3: the gate_up shader takes the
   activation as a push constant since 8dd7fae, and fmt=4/g64 rides the
   same grouped path the expert tier uses — the !g_act_swigluoai and
   fmt gate predated both. 3 submits/layer -> 1 across 57 layers. Also
   set the activation BEFORE vk_registry_fill's early returns, so a
   tier-less run (COLI_VK_EXPERTS=0) can not dispatch silu on an
   swigluoai model.
 - MSA selection: t-outer loop (each cached key row read once for all 4
   index heads), AVX2 4-lane f64 FMA scoring dot (the 64K-context decode
   hot spot), OMP over prefill rows (was single-threaded on 12 cores).

Oracle gate after the selection rewrite: PREFILL 24/24 (exercises the
OMP path), incremental DECODE 20/20 (IDOT=0, tiny checkpoint).

Measured (RX 9070, R8 tier 168 + dense, 2-drive mirror, ngen 64):
  decode topp 0.5: 3.02-3.10 -> 3.49-3.52 tok/s
  decode topp 0.7: 2.78-2.84 (unchanged — 154 experts x 28.3 MB =
    4.4 GB/token sits ON the dual-SSD read floor; compute offloads only
    pay below it, which is the standing argument for topp<=0.55 or int3)
  prefill 516 tok: 102.3 -> 96.8 s (152.7 before VK dense)
RAM_GB=58 probe: no gain (hit 84.8->85.9, wash) — RAM budget is spent.
Quality: Paris + coherent lighthouse prose at 0.5 and 0.7.
…ain pass

Since the indexer weights are kept in the main pass (MSA), they fell into the
generic resident class and silently followed --ebits. Block selection is
discrete: quantization noise flips top-k choices rather than blurring them,
and the validated real-model MSA config used int8 (supplemental pass). New
'idx' class + --idx-bits (default 8) pins them regardless of ebits/xbits;
the tensors are tiny (~215 MB) so the cost is nil.

Flag-path smoke on the tiny checkpoint: --ebits 4 --xbits 3 --group-size 64
converts (routed experts fmt=5 int3-g64, [MIXED] idx=8bit) and the engine
loads and runs the result.
matmul_i3 shipped with a NEON path only; on x86 every 3-bit weight went
through the scalar bit-extraction tail. Measured on the Zen2 12-core box
(MiniMax-M3 int3-g64 routed experts): routed-CPU effective bandwidth
3.7-3.9 GB/s vs 27 GB/s for the int4 IDOT path — expert-matmul at 88%
of decode time, 0.95 tok/s where int4 does 2.8.

Low plane reuses matmul_i2's 2-bit unpack (shift/mask + unpacklo into
16 ordered bytes); high plane broadcasts the two mask bytes with
pshufb, bit-tests against per-lane masks, and adds 4 where set; the
[0,7] byte then drops to [-4,3] at the epi32 widen, two FMAs per 16
weights, one hsum per 64-input group. Ragged tails stay scalar.

Verified against a double-precision scalar reference over S=1/S>1,
multiple-of-64 and ragged shapes: max rel err 5.4e-6.
@JustVugg

Copy link
Copy Markdown
Owner

Really promising direction (GQA + MSA block-sparse + o200k), and thanks for pushing it. Current status so it's clear where it stands:

  • It's marked WIP and is currently CONFLICTING with dev — needs a rebase before it can be evaluated.
  • It touches the shared core (colibri.c, openai_server.py, sample.h, quant.h, kv_persist.h) and brings a Vulkan backend + shaders in the same PR (+5,050 lines). As one unit that's too much surface to merge safely against the GLM/Inkling path we're actively tuning.

To move it forward:

  1. Rebase onto current dev and resolve the conflicts.
  2. Split it the way we're asking Add Qwen3.6-35B-A3B engine: Vulkan MoE backend, resident expert pinning, serve GPU support + fixes #602 to split: the MiniMax engine/attention/tokenizer as one reviewable piece, and the Vulkan backend as a separate PR — and please coordinate the Vulkan work with Vulkan backend: expert tier + dense + MLA attention on any Vulkan 1.2 GPU (successor to #84) #418 (steve-m) and Add Qwen3.6-35B-A3B engine: Vulkan MoE backend, resident expert pinning, serve GPU support + fixes #602 so the project gets one Vulkan backend, not three overlapping ones.
  3. Keep the core-file changes minimal and called out explicitly (what in colibri.c/sample.h/quant.h MiniMax actually needs), so we can see the blast radius.

Flag it ready-for-review (un-WIP) once it's rebased and split, and we'll go through it. No rush — better landed cleanly than fast.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants