Skip to content

Latest commit

 

History

History
235 lines (210 loc) · 34.6 KB

File metadata and controls

235 lines (210 loc) · 34.6 KB

Model Architecture Cards

English | 中文

This folder is the canonical, per-model reference for every architecture that TensorSharp can run. Each card is a self-contained brief: it walks an engineer or researcher from "I have never heard of this model" all the way to "I can explain the forward graph and reproduce the inference path in TensorSharp." If you only need a top-level pointer, use the table below; otherwise jump into the individual cards.

For BERT/XLM-R sentence encoders (Snowflake Arctic Embed L v2.0 and MiniLM), see the embedding model guide for downloads, the forward graph, and APIs.

What every card contains

Each card follows the same shape so you can diff architectures cleanly:

  1. Origin and intent — who designed the model, what the GGUF arch keys are, and which capabilities (modalities, thinking, tools) it exposes.
  2. Model architecture — the high-level block diagram, layer counts, and any per-layer heterogeneity.
  3. Forward graph — the exact ordered list of ops a single token (decode), a multi-token sequence (prefill), or a diffusion denoising step flows through, including residuals and normalizations.
  4. Components — every sub-block (attention, FFN/SSM, routing, normalization, RoPE flavor, vision/audio encoder) explained in detail with the math that governs it.
  5. Parameters and settings — the GGUF metadata keys, weight tensor naming convention, and dtype expectations.
  6. TensorSharp implementation — pointers to the C# source files, the instantiation order, the cache layout, and the way the model plugs into ModelBase / Ops / native GGML kernels.
  7. Prefill optimization — chunking, fused per-layer kernels, parallelization, cross-layer caches.
  8. Decode optimization — fused single-call kernels, pre-resolved weight pointers, batched MoE, in-place kernels, cache reuse.
  9. Memory and KV cache strategy — circular vs. linear caches, mmap-backed weights, pre-allocated decode buffers.
  10. Multimodal pipeline — how images / audio / video are processed, encoded, and injected into the language model.
  11. Output / chat template — protocol parser, stop tokens, thinking / tool formats.
  12. Optimization opportunities — work that has not been done yet but that we know would unlock more performance or capability.

Verified start lane

The verified native GGML family/path tier is Gemma 4 E4B Q8_0; the recommended public artifact is ggml-org/gemma-4-E4B-it-GGUF. Run it on ggml_cuda, ggml_metal, or ggml_vulkan; this lane exercises fused native kernels. See the Gemma 4 card. Its matching mmproj is optional for text and required for image, video, or audio input.

For a continuous learning path through that example—from tensor foundations to a complete multimodal inference engine—use Zhongkai Fu's From Tensors to Tokens book guide, or view the paperback on Amazon.

Implementation matrix

Architecture Card Verified download (HF) Source class GGUF keys Modalities Reasoning Tools Batched / paged forward Notable acceleration
DeepSeek V4.1 Flash deepseek41.md vcruz305/DeepSeek-V4.1-Flash-GGUF, seven Q2_K shards with embedded Engram; validation status DeepSeek4Model with a dedicated native V4.1 graph, plus the pure-C# DeepSeek4CpuExecutor, which implements that same V4.1 graph for --backend cpu deepseek41 Text; image and video through the separately prepared vision companion (--mmproj; it is native, so it needs a ggml backend — LoadVisionEncoder throws on --backend cpu), audio rejected with HTTP 400 Reference-format renderer Spaced DSML; model-level validation tracked separately Per-sequence slots; per-slot decode fallback ggml_cuda, layer placement, CPU MoE, and experimental routed-MoE TP; attention/distributed TP remain unimplemented, and V4.1 DSpark is experimental (--draft-model loads a deepseek41-dspark drafter on ggml_cuda / ggml_cpu only; initial text/image HTTP probes with trained weights passed using two-GPU layer split on ggml_cuda; broad quality and throughput remain unqualified). --backend ggml_cpu (the same native graph on scalar ggml kernels) and --backend cpu (the pure-C# executor, with no ggml, no native library and no GPU) are correctness and portability paths, not serving paths
DeepSeek V4 Flash deepseek4.md unsloth/DeepSeek-V4-Flash-0731-GGUF (multi-shard per quant directory; point --model at the -00001-of- shard). DSpark drafters: MODEL_DOWNLOADS.md DeepSeek4Model (+ DeepSeek4CudaExecutor, DeepSeek4CpuExecutor) deepseek4 Text Yes Yes (DSML markup) Native per-sequence slots (DeepSeek4Model.PerSeqCache.cs) rather than IBatchedPagedModel — servable with continuous batching through the same engine Three whole-model executors (direct CUDA, native ggml, pure C#), whole-layer placement across N local GPUs with --layer-split N (default: one device), on-device compressed KV state (SWA ring + CSA/HCA + lightning indexer), shape-signature graph cache replaying a captured CUDA graph, fused decode index-gather over [ring | top-512] K, and DSpark block speculative decoding (1.3–1.4× decode)
Qwen 3.8 Flash Next qwen38-flash-next.md unsloth/Qwen3.8-Flash-Next-GGUF (multi-shard; point --model at the -00001-of- shard) Qwen4ExpModel (whole-token fused graph on GGML) qwen4exp Text + image + video (video_url) Yes Yes Per-sequence state holders (SupportsPerSequenceFusedForward): per-request KV + GDN + PLE state, round-robin fused decode One captured graph per token incl. in-graph PLE and fused LM head; IMRoPE vision; KV reuse across turns (extend-only); per-token MTP speculative decoding with a separate shared MTP head GGUF (--draft-model, GGML backends only); multi-GPU layer split across N selected local GPUs (--layer-split N: contiguous whole layers per GPU — a capacity feature, not tensor parallelism, since qwen4exp shards no weights; byte-identical output, TS_Q4E_LAYER_SPLIT overrides the balance)
GLM 5.x glm.md unsloth/GLM-5.2-GGUF, unsloth/GLM-5.3-GGUF (UD-Q2_K_XL is seven shards, 236.4 GiB), unsloth/GLM-5.3-Flash-GGUF (multi-shard per quant directory; point --model at the -00001-of- shard) GlmDsaModel (+ the native ggml_ops_glm_dsa.cpp whole-model executor); GLM-5.3 (the same 79-block glm-dsa shape as 5.2, 256 experts, text only) loads on the GLM-5.2 path with no new code and no new flag — see the card glm-dsa (GLM-5.2 and GLM-5.3), glm5next (5.3-Flash) — one architecture descriptor serves all three, Flash vs non-Flash being a single boolean on general.architecture Text (5.2 and 5.3 — the 5.3 repo publishes no mmproj at any quant, and LoadVisionEncoder warns and ignores an --mmproj on glm-dsa instead of failing the run); text + image (5.3-Flash via mmproj) Yes Yes (XML tool calls) Native per-sequence slots (TSGgml_GlmSlotAlloc) rather than IBatchedPagedModel — servable with continuous batching through the same engine Native whole-model ggml executor and a pure-C# per-op reference (GLM 5.2, GLM-5.3 and GLM-5.3-Flash all run on --backend cpu, but it is a reference implementation to A/B against rather than bit-parity: on GLM-5.3-Flash the prefill-logit cosine is 0.9567 against ggml_cpu and the greedy token differs — see the card), automatic whole-layer placement with --layer-split N or Megatron tensor parallelism (--tp N: column/row-parallel heads, every routed expert split row-wise), --cpu-moe host-resident experts served straight from the GGUF mapping, MLA weight absorption with a 576-wide cache row, DSA lightning indexer with a selection reused across 57 of 78 layers, and a shape-keyed graph cache replaying a captured CUDA graph; on GLM-5.3 the blk.78 NextN block drafts with --spec with single-device or explicit --layer-split N placement (no active tensor parallelism), because it ships no LM head of its own and borrows the trunk LM head, which --tp splits column-wise
Gemma 4 gemma4.md E4B Q8_0 is the verified native-GGML family/path tier; ggml-org/gemma-4-E4B-it-GGUF is the recommended public artifact Gemma4Model gemma4 (gemma4-assistant / gemma4_assistant load only as the MTP draft) Text, image, video, audio Yes Yes Default (toggle off with TS_GEMMA4_BATCHED=0) Single-graph fused decode (all layers in one GGML dispatch), fused whole-model prefill/verify with in-kernel PLE + shared-KV handling, chunked prefill, circular SWA cache, and MoE variants. Batched path matches legacy logits within FP noise (Gemma4BatchedForwardTests); reaches ~1.5× legacy at batch=8 and ~1.6× at 4×800-token prompts. Concurrent decode runs the token-batched fused kernel (PLE + KV-donor + SWA wrap, so E2B/E4B qualify): short-horizon streams token-for-token equal to round-robin (long greedy runs can flip a low-margin token, as with any batched GEMM), 2.2× aggregate decode at 4 concurrent on E4B Q8_0 / A40.
DiffusionGemma diffusiongemma.md unsloth/diffusiongemma-26B-A4B-it-GGUF; vision tower: model-00011-of-00011.safetensors (2.84 GB) from google/diffusiongemma-26B-A4B-it DiffusionGemmaModel + DiffusionGemmaSampler diffusion-gemma, diffusion_gemma Text + image (Gemma 4's gemma4v vision tower, from an mmproj GGUF or the raw HF model-00011-of-00011.safetensors shard); ordinary chat refuses audio and video_url (Web UI video arrives as plain image frames); Jev adds documents, sampled video and configured ASR transcripts No (not prompted; a thought block the model writes anyway is returned only with think: true) No (refused with HTTP 400) Separate Web UI DiffusionBatchScheduler; not an autoregressive IBatchedPagedModel path EntropyBound block denoising over [prompt | canvas], prompt-KV caching on the GPU backends and the pure-C# cpu backend (host K/V, batched MoE), self-conditioning, fused GGML whole-model diffusion decode and fused lm-head tail; the same model serves the Jev decision API (POST /v1/systemone)
Qwen-Image-2.1 qwenimage21.md Abiray/Qwen-Image-2.1-GGUF; a metadata-free Unsloth Q8_0 GGUF is recognised from its tensor layout (its mmproj-BF16.gguf needs --qwen-image-mmproj) QwenImageModel qwen_image, qwen-image Text-to-image and image edit No No Serialized diffusion requests Qwen3-VL-8B conditioning, dedicated 2.1 VAE, FlowMatch-Euler, unmerged LoRA plug-ins (--lora), a default-on prefix KV cache for the text and reference-image tokens (TS_QWEN21_PREFIX_CACHE=0 disables), and diffusion-transformer tensor parallelism (--tp on ggml_cuda / ggml_vulkan; the encoders and VAE stay on GPU 0). Runs on the GGML backends and on the pure-C# cpu backend, where the DiT, text encoder, vision encoder and VAE are managed code (single process; --tp refused)
Qwen 3.5 / 3.6 family qwen35.md unsloth/Qwen3.5-9B-GGUF; NextN MTP: unsloth/Qwen3.6-35B-A3B-MTP-GGUF (base-repo Qwen3.6 GGUFs strip the NextN block and silently fall back to standard decode) Qwen35Model qwen35, qwen35moe, qwen3next Text, image Yes Yes Default (toggle off with TS_QWEN35_BATCHED=0 or --no-continuous-batching). Per-slot recurrent-state pool + optional native GatedDeltaNet kernel (TS_QWEN35_BATCHED_GDN_NATIVE=1) Hybrid FullAttention + GatedDeltaNet recurrent, fused attention layer decode, fused prefill attention, fused output-projection + FFN, fused output-projection + norm + router, batched MoE (routed + shared + residual in a single kernel), fused vision encoder blocks
Bonsai2 27B bonsai2.md Local Ternary-Bonsai-2-27B-PQ2_0.gguf / PTQ1_0.gguf; exact hashes in the card Qwen35Model qwen35 plus prism.hadamard.* Text, image with companion projector Yes Yes Qwen 3.5 per-sequence KV/recurrent state; single-device GGML Lossless publisher-format repacking and TensorSharp-owned signed Hadamard transforms; on ggml_metal the 27B dense hybrid geometry it shares with Qwen3.8-27B prefills in 512-token chunks by default (TS_PREFILL_CHUNK overrides); see validation scope in the card
GPT OSS gptoss.md ggml-org/gpt-oss-20b-GGUF GptOssModel gptoss, gpt-oss Text Yes (always) Yes Default (toggle off with TS_GPTOSS_BATCHED=0). Per-head attention sinks via TSGgml_PagedAttentionForwardWithSinks (or TS_GPTOSS_PAGED_ATTN_MANAGED=1 for the C# fallback). 100% greedy match vs legacy in GptOssBatchedCorrectnessTests. Stacked MoE prefill kernel (mul_mat_id + add_id + swiglu_oai), attention sinks, MXFP4 expert weights
Nemotron-H nemotron.md bartowski/nvidia_Nemotron-H-8B-Reasoning-128K-GGUF; Omni: unsloth/NVIDIA-Nemotron-3-Nano-Omni-30B-A3B-Reasoning-GGUF (+ mmproj-BF16.gguf for image); Nemotron 3.5 Lightning: unsloth/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-GGUF NemotronModel nemotron_h, nemotron_h_moe, nemotron_h_omni Text, image (Omni-class); audio only with a companion GGUF carrying the Parakeet tower Yes Yes Default (toggle off with TS_NEMOTRON_BATCHED=0). Per-slot Mamba2 conv + SSM state pool; optional native batched Mamba2 step (TS_NEMOTRON_MAMBA2_BATCHED_NATIVE=1). 100% greedy match vs legacy; up to 3.95× tps at batch=3 on Apple M4 Pro. Mamba2 + attention + MoE FFN hybrid stack, batched GPU MoE, RADIO/v2_vl image encoder, Parakeet/FastConformer audio tower (it runs only when a companion GGUF carrying it is loaded through --mmproj or TS_NEMOTRON_AUDIO_MMPROJ; the public GGUF distributions ship none)
Mistral 3 mistral3.md bartowski/mistralai_Mistral-Small-3.1-24B-Instruct-2503-GGUF Mistral3Model mistral3 (plus llama-labelled Mistral Small 3.x files) Text, image No No Default — reference IBatchedPagedModel implementation. End-to-end validated on Ministral-3-14B; native paged-attention kernel is ~21% faster than the legacy per-seq path on long context. YaRN-corrected RoPE with position-dependent Q scaling, fused QKV / gate_up, Pixtral vision encoder
Hunyuan Dense hunyuan-dense.md Tencent dense Hunyuan GGUFs declaring hunyuan-dense, e.g. the Hy-MT2 releases (tencent/Hy-MT2-1.8B supplies the reference chat template) HunyuanDenseModel hunyuan-dense Text No No (the protocol renders neither tool declarations nor tool results) Default (IBatchedPagedModel over the generic per-op kernels; a q8_0 / q4_0 KV cache or TS_HUNYUAN_BATCHED=0 selects the K/V-snapshot swap path) None yet — QKV and gate/up fusion at load, single device, no fused whole-model graph, no TP or layer split
Muse-Glimmer muse-glimmer.md unsloth/Muse-Glimmer-30B-GGUF (Muse-Glimmer-30B-*.gguf + mmproj-Muse-Glimmer-30B-*.gguf; DFlash drafter dflash-kquant.gguf in the same repo) MuseGlimmerModel muse-glimmer, muse_glimmer Text, image Yes Yes No (legacy per-seq) Interleaved SWA with NoPE full layers, attention output gate, 4 RMSNorms/layer (post-norms at eps 1e-8), logit scale + tanh softcap, sparse-window 2D-RoPE ViT with 2x2 pixel shuffle, optional DFlash block drafter (--draft-model, lossless), tensor parallelism (--tp 2 on GGML CUDA/Vulkan — 2 KV heads cap the degree at 2)
MiniMax-H3 minimax-h3.md Denoisers (separate checkpoints, not settings): minimax_h3_fl2va_pruned-Q4_K.gguf (text + keyframes) and minimax_h3_ref2va_pruned-Q4_K.gguf (text + references), plus the shared Qwen3-VL-32B text encoder qwen3vl_32b_minimax_h3-Q4_K_M.gguf, all from unsloth/MiniMax-H3-GGUF; minimax_h3_video_vae_fp16.safetensors and minimax_h3_audio_vae_fp32.safetensors (omit for silent video) from Comfy-Org/MiniMax-H3. The text-encoder GGUF ships no tokenizer — put vocab.json and merges.txt from MiniMaxAI/MiniMax-H3 beside it, or point TS_VIDEO_TOKENIZER at them MiniMaxH3Model (+ MiniMaxH3Pipeline) minimax-h3, minimax_h3 — but neither published GGUF carries any metadata at all, so ModelBase.Create() recognises H3 from its tensor table (MiniMaxH3Architecture.DetectFromTensors) instead of an architecture string Video + native 32 kHz stereo audio out, generated together in one packed latent (text → video, image → video, first/last frame, reference → video) No No None — Forward() throws; generation runs through GenerateVideo() Native whole-network ggml graphs, one graph per network, weights bound resident straight from the GGUF/safetensors mmap; CFG-distilled, so --cfg 1.0 at 4–8 steps is the operating point (M5 Pro / Metal, 22 frames, 8 steps: 2.4× faster than stable-diffusion.cpp at 256×256 and 1.7× at 640×384); learned AdaLN curve table instead of a timestep MLP; 3-axis continuous-float RoPE putting video and audio on one timeline; video VAE decode chunked at 5 latent frames and tiled at 256 px as a correctness requirement rather than an optimization; FP16-safe h3_attend that pre-scales V by a power of two derived from the key count, which is what keeps a 107-frame clip finite
Wan video wan.md Base: QuantStack/Wan2.2-TI2V-5B-GGUF, QuantStack/Wan2.2-I2V-A14B-GGUF, city96/Wan2.1-T2V-14B-gguf. Step-distilled (25× less denoising work, same flags): hum-ma/Wan2.2-TI2V-5B-Turbo-GGUF, jayn7/WAN2.2-I2V_A14B-DISTILL-LIGHTX2V-4STEP-GGUF. (+ UMT5-XXL encoder and video VAE, see the card) WanVideoModel (+ WanVideoPipeline) wan, wan2.1, wan2.2 Video out (text -> video, image -> video) No No None - Forward() throws; generation runs through GenerateVideo() and is serialized Step-distilled checkpoints auto-detected from the DiT file name (100 DiT passes -> 4; M5 Pro 1088x832x121f: 3 h 30 m -> 17 m 30 s), one resident-weight ggml graph per denoise step (CUDA-graph-captured, flash attention over F16 keys/values -- 2.02x at 27 k tokens, per-token-timestep modulation for TI2V i2v), --cfg-cache-stride guidance reuse (1.30x / 1.43x on base checkpoints), causal 3D video VAE encode and decode each as a single graph with the convs on MPSGraph on Metal (VAE decode 1.99x), A14B's two 14B experts hot-swapped at the timestep boundary, stagewise VRAM handoff (TE -> DiT -> VAE), memory-sized im2col budget and 720p decode tiling

Backend notes

Model code is intentionally backend-agnostic. ModelBase selects tensor storage through BackendType and the registered execution plan, then delegates the actual ops to the backend that owns those allocators:

Backend type Package Notes
Cpu TensorSharp.Core Pure managed tensors with SIMD/managed quantized fast paths (RMSNorm, RoPE, softmax, fused activations, GEMM, dequant). The quantized matmuls run on a persistent worker pool (TensorSharp.Models/CpuWorkerPool.cs) rather than a per-matmul Parallel.For, which is what lets decode scale past a handful of cores; it is sized at half the usable CPUs on purpose, and TS_CPU_* tunes it. Quantized weights are bound zero-copy from the GGUF mapping — the same file-backed binding the GGML backends already used — instead of being copied into fresh anonymous memory at load (TS_DIRECT_QUANT_WEIGHTS=0 restores the old expand-to-F32 behaviour for an A/B). ManagedQuantizedOps also covers IQ2_XS and IQ4_XS (managed dequantizers plus entry into the CPU quantized-storage matrix, so they stay quantized instead of expanding to F32 at load), and has direct IQ2_XS x Q8_K and IQ3_XXS x Q8_K dot kernels with AVX2 paths. Families whose generation path bypasses the ggml graph (Wan, MiniMax-H3) share the direct primitives in TensorSharp.Models/Direct/.
Cuda TensorSharp.Backends.Cuda Direct CUDA Driver-API allocator and storage, cuBLAS GEMM, PTX kernels for hot ops (RMSNorm, softmax, RoPE/RoPEEx, SDPA, GQA prefill/decode, causal mask, gather/concat, activation fusions), native quantized matmul / get_rows for supported quant types, CPU fallback for ops that are not yet implemented.
Mlx TensorSharp.Backends.MLX Apple Silicon mlx-c bridge with quantized / fused / compiled kernels, async worker dispatch, MoE expert offload, and a CPU fallback layer. Requires libmlxc.
GgmlCpu / GgmlMetal / GgmlCuda / GgmlVulkan TensorSharp.Backends.GGML + TensorSharp.GGML.Native Native ggml bridge with quantized graph dispatch and platform backends. mmap-backed quantized weights are bound zero-copy through host-pointer buffers. Includes the paged-attention kernel (TSGgml_PagedAttentionForward, plus the GPT OSS sinks variant) that powers the batched / paged execution path.

When a card mentions a fused GGML kernel (for example Qwen35AttentionLayerDecode, Gemma4LayerPrefill, or MoEExpertsSwiGLUResidual), the kernel is compiled from TensorSharp.GGML.Native/ggml_ops_*.cpp and exposed through TensorSharp.Backends.GGML/GgmlBasicOps.cs. The native bridge is the place to look when a fused path engages on GGML CPU / Metal / CUDA but not on the pure managed CPU or direct CUDA backends.

Continuous batching & paged KV cache

All autoregressive architectures listed above run through the shared InferenceEngine + ContinuousBatchScheduler + BatchExecutor stack documented in docs/PAGED_ATTENTION_AND_CONTINUOUS_BATCHING.md. Models that implement IBatchedPagedModel.ForwardBatch execute one batched forward per scheduler step (with slotMapping-based K/V scatter into a shared paged buffer and per-sequence attention via the native paged kernel); the others run through the per-sequence KV-swap fallback inside the same engine. Prompt-prefix reuse on every autoregressive family in the matrix goes through the engine's Radix prefix cache, which is on by default (TS_PREFIX_CACHE_MODE=tree; legacy selects the older block-hash sharing, and TS_SCHED_PREFIX_CACHE=0 or --no-prefix-cache disables reuse). DiffusionGemma does not support autoregressive Forward(), so it uses DiffusionGemmaSampler and the server-side DiffusionBatchScheduler instead. Its Jev decision API reads typed answer probabilities directly from a seeded canvas with one denoising step and a selected-label output projection, over a state that may include images, uploaded documents, sampled video frames and configured ASR transcripts as well as text. Qwen-Image-2.1 is likewise not autoregressive: Forward() throws, generation runs through QwenImageModel.GenerateImage() and editing through QwenImageModel.EditImage() over a FlowMatch-Euler diffusion loop, and concurrent requests are serialized (the diffusion nets are not thread-safe). The opt-in env vars are summarised in the matrix above and in the project root README.

Solo (non-concurrent) sequences on architectures that ship a multi-token-prediction draft head — Qwen 3.6 and Qwen 3.8-27B (embedded NextN block), GLM 5.2 and GLM-5.3 (embedded NextN block, single-device or layer-split placement only), Gemma 4 (separate gemma4-assistant draft GGUF) and Qwen 3.8 Flash Next (separate shared MTP head GGUF, GGML backends only) — can additionally run lossless MTP speculative decoding through the same engine (--spec opts in the embedded NextN blocks; for Gemma 4 and Qwen 3.8 Flash Next, naming the draft GGUF on --draft-model enables speculation by itself. Both flags are accepted on both hosts, since TensorSharp.Cli and TensorSharp.Server.Host share one flag parser (SpeculativeCliFlags); the TS_SPEC_* / legacy TS_MTP_* env vars work too). The shared draft / verify / rollback core is SpeculativeExecution; per-architecture mechanics are in the Qwen 3.5/3.6 (§12) and Gemma 4 (§12) cards.

DeepSeek V4 plugs a block drafter into that same core: its DSpark support module ships as a separate GGUF loaded with --draft-model (on both the CLI and the server) and proposes a whole block of tokens per step instead of one at a time. Because the drafter's weights must be counted by the layer split, it is passed to ModelBase.Create() at load time rather than attached afterwards. See the DeepSeek V4 card. DeepSeek V4.1 accepts only a deepseek41-dspark drafter, on ggml_cuda or ggml_cpu; that path is experimental; initial text/image HTTP probes with trained weights passed using two-GPU layer split on ggml_cuda; broad quality and throughput remain unqualified. Muse-Glimmer (DFlash / DFlash2) and Qwen 3.8-27B (DFlash2) also take block drafters through --draft-model.

A weight-free n-gram (prompt-lookup) speculator, selected with --spec-type ngram, needs no drafter and runs on the Qwen 3.5 family, Gemma 4, GLM 5.x and Qwen 3.8 Flash Next, and on DeepSeek V4 / V4.1 and Muse-Glimmer only while their drafter is loaded. It does not run on GPT OSS, Mistral 3 or Hunyuan Dense, and Nemotron-H refuses every speculator; see speculative decoding.

Architecture comparison

Feature DeepSeek V4 Gemma 4 DiffusionGemma Qwen 3.5 / 3.6 family GPT OSS Nemotron-H Mistral 3 Muse-Glimmer
Layer type MoE (256 routed experts, top-6 + 1 shared) Dense / MoE Gemma-4-derived MoE encoder/decoder Hybrid (Attn + Recurrent) ± MoE MoE Hybrid (Mamba2 + Attn + FFN, dense or MoE) Dense Dense (52 layers, 32 Q / 2 KV heads)
Attention Raw SWA-128 + compressed CSA 4:1 / HCA 128:1 (lightning-indexer top-512 on CSA layers) SWA + Global Region-aware prompt/canvas attention Full GQA + Sigmoid Gate Full + Sinks Full GQA (no RoPE) Full GQA Interleaved SWA-2048 + full NoPE layers (39 + 13), sigmoid attention output gate
FFN activation SwiGLU with a per-layer clamp GeGLU Dense GeGLU + top-8 MoE SwiGLU SiLUAlphaLimit (clamped GLU) ReLU² SwiGLU SwiGLU
RoPE variant Interleaved-pair + YaRN; separate raw and compress bases, inverted after attention NeoX + proportional / partial NeoX, local/global bases NeoX / MRoPE NeoX + YaRN None GPT-J + YaRN ggml NORM (interleaved pairs) on the SWA layers only; the full layers are NoPE
QK-norm Q only (per-head RMS) Yes Yes Yes No No No Yes (per-head; the Q norm carries the folded qk_scale_factor)
V-norm No Yes (unweighted) Yes (unweighted) No No No No No
Bias in projections No (router selection bias only) No No No Yes (all linear) No No No
Per-layer scaling No (per-layer swiglu clamp and compress ratio instead) Yes Encoder / decoder scalars No No No No No (logit scale 0.19612 + tanh softcap 20.0 on the output instead)
Per-Layer Embedding (PLE) No Yes No No No No No No
KV sharing Yes (one shared 512-dim K=V head for all queries) Yes (tail layers) Prompt-KV cache across denoising steps No No No No No
Attention sinks Yes No No No Yes No No No
Circular KV cache Yes (raw SWA-128 ring) Yes (SWA layers) No autoregressive KV No No No No Yes (SWA ring on the GPU backends; TS_MUSE_GLIMMER_SWA_RING=0 disables)
SSM / recurrent layers No (4-stream hyper-connections replace the plain residual) No No Yes (GatedDeltaNet) No Yes (Mamba2) No No
Shared experts Yes No No Yes (qwen35moe / qwen3next) No Yes (optional) No No (dense FFN)
Latent bottleneck FFN No (LoRA-factored Q / output projections instead) No No No No Yes (optional) No No
Position-dependent Q scaling No No No No No No Yes (with YaRN) No
Vision No Yes Yes (Gemma 4's gemma4v tower, from an mmproj GGUF or the HF safetensors shard) Yes No Yes (Omni) Yes (Pixtral) Yes (sparse-window 2D-RoPE ViT with 2×2 pixel shuffle)
Audio No Yes No (refused) No No Only with a companion GGUF carrying the Parakeet tower (--mmproj or TS_NEMOTRON_AUDIO_MMPROJ); the public Omni GGUFs ship none No No
Video No Yes No No No No No No
Thinking Yes Yes No (not prompted) Yes Yes (always) Yes No Yes (assistant to=self channel)
Tool calling Yes (DSML markup) Yes No (refused) Yes Yes Yes No Yes (ATEM XML markup)
MTP / NextN speculative decoding DSpark block drafter (separate GGUF via --draft-model) Yes (separate gemma4-assistant draft GGUF) No Yes: embedded NextN on Qwen 3.6 and Qwen 3.8-27B (--spec); a DFlash2 block drafter via --draft-model on Qwen 3.8-27B No No No DFlash block drafter (separate 5-layer GGUF via --draft-model, lossless)
Fused QKV n/a (LoRA-factored Q, single shared K=V head) Yes Yes Mixed (full attention layers split, recurrent layers fuse a 5-way pack) Yes Yes Yes No
Fused single-graph decode Yes (whole-model executor, one graph per ubatch, CUDA-graph replayed) Yes (Gemma4ModelDecode) Yes (DiffusionModelDecode + lm-head tail) Per-layer fused (Qwen35AttentionLayerDecode, FusedOutProjFFN, FusedOutProjNormRouter) Per-layer Per-layer / batched MoE No Yes (persistent whole-model decode graph on GGML CUDA / Vulkan / Metal / CPU)
Fused single-graph prefill Yes (same whole-model executor, chunked ubatches) Yes (whole-model NativeGemma4ModelVerify + per-layer Gemma4LayerPrefill fallback) Prompt-KV prefill cache Yes (FusedPrefillAttention, FusedOutProjFFN, MoE prefill) Yes (MoE prefill via mul_mat_id) No No Yes (same fused kernel, chunked with on-device causal+SWA band masks)
Batched GPU MoE Yes (grouped expert kernels) Yes for all-MoE variants (fused whole-model MoE decode/verify); mixed dense+MoE pending Fused per-canvas MoE; concurrent requests batched by diffusion scheduler Yes (routed + shared + residual fused) Yes (stacked weight slabs) Yes n/a n/a (dense FFN)
Fused vision encoder n/a Standard Standard (Gemma 4's) Yes (FusedVisionAttention + FusedVisionMLP) n/a Standard (RADIO ViT) Standard (Pixtral) Yes (fused vision block + flash attention on CUDA)
Output parser DeepSeek4OutputParser (always required) Gemma4OutputParser (always required) Gemma4OutputParser (always required) Qwen35OutputParser HarmonyOutputParser (always required) ChatMlOutputParser PassthroughOutputParser MuseGlimmerOutputParser (always required)

Adding a new architecture

When you add a new model:

  1. Create TensorSharp.Models/Models/<Name>/<Name>Model.cs inheriting ModelBase.
  2. In the constructor: read GGUF metadata via _gguf.GetXxx(), call ParseBaseConfig() and ParseTokenizer(), call LoadWeights(), fuse weights, then initialize caches.
  3. Implement Forward(int[] tokens) → float[] for autoregressive models: embedding → optional multimodal injection → transformer blocks → final norm → LM head → logit copy. For diffusion models, document the alternate sampler entry point and make unsupported autoregressive paths explicit.
  4. Implement ResetKVCache() and Dispose(). Implement TruncateKVCache() when KV-cache reuse is supported.
  5. Declare the architecture plug-in in TensorSharp.Models/Models/<Name>/<Name>Architecture.cs -- a ModelArchitectureDescriptor with the GGUF general.architecture aliases, the factory, and anything non-default (multi-GPU mode and why, mmproj companion file hints, a tensor-based detector for metadata-free GGUFs, process-wide native tunables) -- then add ONE line for it to TensorSharp.Models/Architecture/BuiltInArchitectures.cs. There is no switch to extend: ModelBase.Create() resolves through ModelArchitectureRegistry.
  6. If the model is multimodal, implement the capability interfaces on it: IVisionCapableModel (load the tower, receive an embedding span) and IMultimodalPromptExpander (expand your own placeholders), plus IAudioCapableModel / IAudioEncoderLoader / IMRoPEPositionSink as applicable. ModelMultimodalInjector owns all the generic bookkeeping and names no model types, so nothing there needs editing.
  7. If the model has its own chat format, add ONE ChatProtocol entry to TensorSharp.Runtime/ChatProtocolRegistry.cs. That single record carries the renderer, whether to bypass the GGUF's Jinja template, the media placeholder tokens, the IOutputParser (implemented in TensorSharp.Runtime/OutputParser.cs), whether that parser is mandatory, where a structured-output grammar may arm, the KV-cache generation suffix that keeps multi-turn prefix reuse working, and video-frame capping.
  8. Add a card under docs/models/<name>.md (and <name>_zh-cn.md if you want bilingual coverage), update this README's matrix, and link the card from the project root README.
  9. Update the capability gates (_detect_capabilities in TensorSharp.Server.Host/testdata/test_multiturn.py) if the model exposes new modalities, thinking, or tool capabilities.