Part of the TensorSharp documentation.
- Sentence embeddings — GGUF
bert/ XLM-R encoders: Snowflake Arctic Embed L v2.0 (1024 dimensions, CLS pooling) and all-MiniLM-L6-v2 (384 dimensions, mean pooling). 100% pure C# CPU plus native GGML CPU/Metal/CUDA, with bidirectional attention; managed execution uses a model-owned worker pool and ARM Q8 SIMD tiles; native execution retains quantized weights, packs projections, isolates attention by sequence, and reuses graphs; OpenAI/Ollama single and batched inputs, normalized vectors, base64, and dimension reduction. See the embedding guide and its validation report. - Multi-architecture support -- DeepSeek V4 Flash, DeepSeek V4.1 Flash (
deepseek41, served onggml_cuda;cpuis the 100% pure C#DeepSeek4CpuExecutorwith no ggml and no native dependency, and it,cudaandggml_cpuare correctness and portability paths, not serving paths), GLM 5.x (GLM-5.2 and GLM-5.3 both onglm-dsa, GLM-5.3-Flash onglm5next), Gemma 4, DiffusionGemma, Qwen 3.5/3.6-family, Qwen 3.8 Flash Next (qwen4exp), Bonsai2 27B (qwen35), GPT OSS, Nemotron-H, Mistral 3, Hunyuan Dense (hunyuan-dense), Muse-Glimmer, Qwen-Image-2.1 (text-to-image and image editing), MiniMax-H3 (video with native 32 kHz stereo audio), and Wan 2.1/2.2 (video only) - Multimodal inference -- image, video, and audio inputs (Gemma 4); image and
video_urlvideo for Qwen 3.8 Flash Next (ordered frame pairs with Qwen-VL temporal M-RoPE coordinates) and DeepSeek V4.1; images for Qwen 3.5/3.6-family / GLM-5.3-Flash / Mistral 3 / Muse-Glimmer / Nemotron-H Omni, each through its ownmmprojtower (GLM-5.2 and GLM-5.3, bothglm-dsa, are text-only); and images for DiffusionGemma, through Gemma 4's vision tower loaded from an mmproj GGUF or straight from the upstreammodel-00011-of-00011.safetensorsshard, since no published DiffusionGemma GGUF carries one (it has no video path:video_urlis refused, and a Web UI video upload reaches it only as separate frame images). Audio input is served by Gemma 4, and by Nemotron-H / Nemotron 3 Nano Omni only when an audio companion GGUF carrying its Parakeet tower is loaded (no public GGUF ships one); audio sent to a model without a loaded audio tower (DeepSeek V4.1 and DiffusionGemma always, Nemotron-H with only the public files) is refused with HTTP 400 through one table (AudioInputSupport.UnsupportedReasonFor), never decoded and silently dropped.--pdfis architecture-agnostic: a born-digital PDF's text layer is inlined into the prompt for any model, and only scanned PDFs fall back to page images (which then need a vision model). Generated media is its own axis: Qwen-Image-2.1 emits an image, Wan 2.1/2.2 emit an H.264 MP4, and MiniMax-H3 is the one family whose output is audio as well as video — a 32 kHz stereo track denoised jointly with the picture and written as a sidecar.wavbeside the MP4 - Thinking / reasoning mode -- structured chain-of-thought output with
<think>/<|channel>thought/<|channel>analysis/to=selftags (Qwen 3.5/3.6-family, Qwen 3.8 Flash Next, Gemma 4, GPT OSS, Nemotron-H, Muse-Glimmer, DeepSeek V4, DeepSeek V4.1, GLM 5.x) - Tool calling / function calling -- models can invoke user-defined tools; multi-turn tool-call conversations supported across all three API styles
- Agent Skills -- folders of model-facing instructions (
SKILL.md+ scripts / references / assets) that load only when a task needs them. Selected per request with"skills": ["pdf"]on every chat API or--skillon the CLI; the model pulls the rest through built-inskills_list/skills_readtools that TensorSharp answers in process, so an ordinary OpenAI client still just receives a finished completion. Without--skills-dirorTS_SKILLS_DIR, every.agents/skillsfolder from the working directory up to the Git root is searched, plusskillsbeside the binary (on the server that folder is also the upload directory and is always scanned first, even with--skills-dir). The Playwright skill inTensorAgent/skills/playwright(not copied beside the desktop binaries; point--skills-diratTensorAgent/skills) drives a real browser from skill scripts on desktop hosts (browser skills). → Agent Skills - Sub-agents (automatic delegation) -- on the server chat paths (
/v1/chat/completions,/v1/responses,/api/chat/ollama, the Web UI's/api/chat) and in TensorAgent, a model that can call tools may hand independent parts of a request to bounded child agents (spawn_agent,wait_agent,send_input,close_agent,list_agents) and combine their reports into one answer. On by default, and the model decides whether to delegate; children are read-only unless the operator optsworkerchildren into mutable tools, and--no-multi-agent, a request's"multi_agent": falseor TensorAgent's "Sub-agents" switch turns it off. The CLI has no sub-agents. No latency or quality gain is published for it. → Sub-agents - Code execution -- with
--code-execthe model drives a real shell in a sandboxed workspace, typing a command line and reading back the exit code and everything it printed. Web/CLI chats retain one workspace for their session; each OpenAI/Ollama request retains a private workspace across its internal agent rounds and deletes it afterward, so separate HTTP requests remain stateless while create → run → diagnose → edit → verify can use the same file. When code execution is enabled, skill scripts share that workspace; script-only execution retains per-call scratch. File work uses dedicated tools:read_fileshows bounded current bytes with line numbers,write_filecreates new files, andapply_patchmodifies one or multiple existing files atomically, including single-line edits.apply_patchalso supports creating, renaming, and deleting files. Models are instructed to read relevant file contents first and useapply_patchfor every existing-file modification. When a diagnostic identifies a bounded regular workspace source file, a code failure returns that source region and directs the model to patch the smallest faulty span before rerunning the check. Anchored patches preserve unrelated contents without re-emitting the whole file. Off by default, and it needs a real OS sandbox (sandbox-execon macOS,bwrap0.12.0+ on Linux) — without one the tool refuses rather than running unconfined. Generated commands also start with Internet/IP access denied;--code-exec-allow-network/TS_CODE_EXEC_ALLOW_NETWORKexplicitly grants unrestricted host IP-network access without removing the macOS/Linux workspace-write or home-read boundaries (subject to the documented macOS shared-temp exception). That access includes LAN/loopback services and IP listening sockets and can expose other host-readable data. Linux additionally bounds descendants with a PID namespace. On macOS, children inherit Seatbelt and ordinary process groups are stopped, but a deliberately detached child can outlive the request; every result reports that gap. macOS denies common/private/tmp/com.apple.launchd*pathname sockets while permitting runtime-required Mach lookup and the exact mDNSResponder pathname socket required for DNS, and Linux hides common/runendpoints, but local Unix IPC is not a complete isolation boundary: macOS retains shared-temporary-directory Unix IPC for compatibility, and Linux's host network namespace may expose abstract sockets and pathname sockets outside/run. The switch is independent of--skills-allow-network. Host-installer package/domain allow-lists cannot constrain direct downloads in this unrestricted mode. Windows still needs the explicit--code-exec-unconfinedescape hatch. - Quantized model support -- loads GGUF files with Q4_K_M, Q8_0, F16, MXFP4,
Q1_0(GGML tensor type 41: one F16 scale plus 128 one-bit signs per block, 1.125 bits/weight), the Bonsai2 publisher typesPQ2_0(142) andPTQ1_0(143), which are transcoded losslessly to GGMLQ2_0at load and need a single-device GGML backend (see the Bonsai2 card), and other quantization formats; performs native quantized matmul without dequantizing to FP32, including memory-efficient pure C# CPU loading for large GGUFs - GPU-accelerated -- GGML Metal on macOS, GGML CUDA on Windows/Linux with NVIDIA GPUs, GGML Vulkan on Windows/Linux with AMD/Intel/NVIDIA GPUs, a direct CUDA/cuBLAS backend with PTX kernels, and an MLX backend for Apple Silicon (mlx-c / Metal), all with CPU fallbacks for unsupported ops
- Optimized pure C# CPU backend -- managed GEMM fast paths plus fused SIMD kernels for RMSNorm, RoPE, softmax, fused activations, and other inference hot paths, with the managed matmuls dispatched through a persistent spin-then-park worker pool instead of a
Parallel.Forper matmul -- worth ~+15% prefill and ~2.8x decode on a 122-CPU host. → Pure C# CPU Backend - Continuous batching & paged KV cache -- vLLM-style block-paged KV pool, a Radix-tree prefix cache that reuses prompt prefixes across requests, iteration-level scheduler that admits / preempts sequences mid-batch, and a native fused paged-attention kernel (
TSGgml_PagedAttentionForward) that drivesggml_flash_attn_exton Metal/CUDA/Vulkan. The engine runs by default inTensorSharp.Server, inTensorSharp.Cligeneration and in TensorAgent;--no-continuous-batchingturns off only the batched forward, so the engine still schedules requests but every sequence takes the per-sequence KV-swap path. The Radix cache is the default prefix-reuse mode (TS_PREFIX_CACHE_MODE=tree;legacyselects the older block-hash sharing for diagnosis) on Qwen 3.5/3.6-family, Gemma 4, GLM 5.x, Qwen 3.8 Flash Next, DeepSeek V4/V4.1, GPT OSS, Mistral 3, Hunyuan Dense, Muse-Glimmer and Nemotron-H — not on DiffusionGemma or the image/video models;--no-prefix-cache(orTS_SCHED_PREFIX_CACHE=0) disables all prefix reuse, the hosts' startup warm-up and the server's checkpoint persistence, and--specno longer turns it off. (The RAM / SSD / Redis tiers of the standalonePagedKvCacheManagerare exercised only by the CLI's--paged-bench; they are not on the serving path.) A cached prefix only helps a request that arrives while it is still resident, so it is joined by shared-prefix checkpoints: on Gemma 4 and Qwen 3.5/3.6 over the GGML backends, and on Qwen 3.8 Flash Next, the engine checkpoints the complete model state at the end of the prompt every conversation begins with — system prompt, tool declarations, selected skills — and starts each new chat from a clone of it, so a new chat re-prefills only its own message (TS_PREFIX_CHECKPOINTS, on by default;TS_PREFIX_CHECKPOINTS_MAXbounds how many distinct prefixes stay resident). A host that attaches anIPrefixCheckpointStoremakes one outlive the process — the executor reads it at admission and writes it when it takes a checkpoint, with the model family owning and validating the byte format — which is how TensorAgent turns the first message of a launch into a restore instead of a full prefill. The server does the same: it warms the shared prefix at startup and keeps up to two checkpoint files per model underprefix-cache/<model>/beside the binary (TENSORSHARP_PREFIX_CACHE_DIRmoves the root); the CLI keeps its checkpoints in memory only. Reuse is isolated per conversation: state of another conversation (a retained holder, the live cache, pooled blocks) is shared only up to the public system-prompt-and-tools prefix, a stateless OpenAI/Ollama request continues a conversation only when its history reproduces a turn this server generated, and an assistant message a client wrote or edited is rendered from its own text, never from another client's generated tokens. Media is identified by content (SHA-256): a resent base64 image is stored once and encoded once, and the text before an image stays reusable; the turns after an image continue the cache too, including on Qwen 3.5/3.6, whose holders and checkpoints carry the M-RoPE position delta. See docs/PAGED_ATTENTION_AND_CONTINUOUS_BATCHING.md. Note what this does and does not buy: the paged KV cache is host-resident, so it delivers admission control and cross-request prefix reuse rather than throughput that scales with concurrency: theBatchedPagedroute saturates at roughly 69 tok/s aggregate no matter how many sequences are in flight, because the pool it scatters into lives in host memory. (That ceiling is measured on the PAGED route. GLM 5.x has a separate, default-on, non-paged batched fused decode — see the Batched / parallel inference bullet below — which is reported at 1.81x aggregate decode at 4 concurrent; that is a different mechanism and is not evidence that the paged path scales.) (A device-resident paged KV pool exists in the tree —TensorSharp.GGML.Native/ggml_ops_paged_kv_pool.cppandTensorSharp.Models/Paged/DevicePagedKvCache.cs— but it is not wired into any model and is not a shipped feature.) GLM 5.x is the exception: MLA with weight absorption (one 576-wide cache row per token per layer) and the DSA lightning indexer have no paged layout, so concurrency is served by native per-sequence slots instead — each request owns its MLA and indexer caches and its ownn_past, and binding a request switches the active slot without moving KV bytes. Qwen 3.8 Flash Next (qwen4exp) is the same shape for the same reason — its GatedDeltaNet, PLE and QSA-indexer state have no paged layout either — and is served through per-sequence state holders: each in-flight request owns its attention KV and indexer caches, its GDN conv + delta-net state and its PLE history, the native kernel keys its device-resident recurrent state by the holder, so switching requests is a reference swap, and the engine round-robins sequences through their own captured fused decode graphs. Its finished conversations and shared-prefix checkpoints stay resident as retained holders (TS_Q4E_RETAINED_CACHE, on by default;TS_Q4E_RETAINED_CACHE_MB, default 4096). DeepSeek V4.1 retains a finished conversation's native state for reuse only whenTS_DSV41_RETAINED_CACHE=1(budgetTS_DSV41_RETAINED_CACHE_MB, default 2048). - Speculative decoding -- a pluggable algorithm layer (
--spec-type:auto/draft-head/block/ngram) over a shared draft-verify-rollback runtime; the weight-freengramspeculator needs no drafter but does need a family that can speculate (GPT OSS, Mistral 3, and Hunyuan Dense always decode plainly, and Nemotron-H refuses every speculator), and learned draft heads accelerate solo (non-concurrent) decode. Qwen 3.6, Qwen 3.8 27B, GLM 5.2 and GLM-5.3 ship their NextN block fused into the trunk GGUF; Gemma 4 loads a separate EAGLE-stylegemma4-assistantdraft GGUF via--draft-modelwhose draft layers attend the target's own KV cache, and Qwen 3.8 Flash Next loads its shared MTP head the same way. The draft proposes up to--spec-drafttokens per step (kept while draft confidence ≥--spec-pmin) and the trunk verifies them in a single batched forward; the request's own sampler — penalties included — drives both drafting and verification, so output is identical to standard decode. Opt in with--specon either host (off by default). OnTensorSharp.Cliit engages on every single-sequence path —--input,--multi-turn-jsonland--interactive. On ggml backends fused multi-token-verify / draft-step kernels make it a clear win; the directcudabackend runs a fully GPU-resident per-op verify/draft and is also a win. On the CPU and MLX backends it depends on the family: Gemma 4 speculates on every GGML backend includingggml_cpubut not on pure-C#cpuor MLX, while Qwen 3.5/3.6 is not gated by backend and leaves the runtime cost governor to park drafting wherever it measures slower (its default 8-token window measured 1.21x on Qwen3.6-35B-A3Bggml_cpu). Env:TS_SPEC_*(shared; the legacyTS_MTP_*spellings still work) andTS_GMTP_*(Gemma 4 tuning). - Batched / parallel inference --
IBatchedPagedModel.ForwardBatchimplementations for Mistral 3, Hunyuan Dense, Gemma 4, GPT OSS, Qwen 3.5/3.6-family, and Nemotron-H all run by default and pack N sequences into a single forward pass with paged K/V scatter and per-sequence attention via the native kernel. Gemma 4, Qwen 3.5/3.6, GPT OSS, and Nemotron-H expose a per-familyTS_<FAMILY>_BATCHED=0escape hatch (TS_GEMMA4_BATCHED=0,TS_QWEN35_BATCHED=0,TS_GPTOSS_BATCHED=0,TS_NEMOTRON_BATCHED=0) to fall back to the per-sequence KV-swap path for A/B comparison or regression isolation; Hunyuan Dense hasTS_HUNYUAN_BATCHED=0(its per-sequence path swaps K/V snapshots). Mistral 3 has no per-family switch — use the globalTS_SCHED_DISABLE_BATCHED=1. A repeated prompt longer than one KV block adopts the earlier request's blocks where they actually live: in the model's paged arrays when a batched step wrote them, from pool snapshots when the per-sequence path captured them (before this, a batched model without linear-to-paged migration restored never-captured bytes and answered a repeated long prompt with fluent garbage). GLM 5.x has no pagedForwardBatch; instead its default-on batched fused decode runs one graph with one token per sequence so the weights are read once — 1.81x aggregate decode at 4 concurrent requests. SetTS_BATCHED_FUSED_DECODE=0to use serial fused decode for A/B or regression isolation; batching changes GEMM shapes, and a 2-bit MoE can amplify that into different expert picks. Dense Gemma 4 has the same kind of token-batched fused decode for concurrent requests (TSGgml_Gemma4ModelDecodeBatchedEx2), and it covers per-layer embeddings, KV-donor layers, a wrapped SWA ring and sequences whose caches have different capacities (such as a retained-prefix clone beside a fresh request), so the E2B/E4B checkpoints take it instead of declining to round-robin: on gemma-4-E4B-it-Q8_0 (A40, ggml_cuda) the AgentTurnBenchconcstreams equal the round-robin ones token for token and aggregate decode goes from 66 to 149 tok/s at 4 concurrent requests (2.2x; 2 concurrent 65 -> 99 tok/s, solo unchanged at 68 tok/s); over hundreds of greedy tokens a low-margin token can flip, exactly as with the v1 kernel, since batching changes GEMM shapes.TS_GEMMA4_BATCHED_CAPS=0forces the old gates for A/B, andTS_GEMMA4_BATCHED_CAPS=7keeps everything but the per-sequence capacities. - Tensor parallelism, layer split & distributed inference —
--tp N/TENSORSHARP_TP_DEGREE=Nselects tensor parallelism only: it shards weights inside layers across N local GPUs.--layer-split N/TENSORSHARP_LAYER_SPLIT_DEGREE=Nselects whole-layer placement on Qwen 3.8 Flash Next, DeepSeek V4 / V4.1, and GLM 5.x. The options are mutually exclusive; unsupported requests fail at startup instead of changing modes or using one GPU. Whole-layer placement is local to one node and mainly increases capacity. Supported tensor-parallel architectures can use--tp-node-id/--tp-peersfor multi-node inference; GLM and Qwen-Image TP remain local-only. The default is one device; explicit legacy native GPU-count variables can still request automatic placement. DeepSeek V4.1 experimental routed-MoE TP additionally requiresTS_DSV41_TP=Nwith the matching--tp N; it does not implement attention TP or multi-node execution. Migrate old layer-placement commands from--tp Nto--layer-split N. → Multi-GPU modes - Ollama & OpenAI API compatibility -- drop-in replacement endpoints for existing tooling
- TensorAgent (iPhone and iPad) -- a native iOS/iPadOS app (
net10.0-ios, iOS 17 or later) that links the TensorSharp engine and answers its own phone-shaped page from an in-process loopback server with the same chat pipeline the Web UI uses: chat with photo, camera, video and file attachments, Agent Skills, sandboxed code tools (an in-process shell with embedded CPython and JavaScriptCore, since iOS allows no child processes) and sub-agents, all on the device throughggml_metal(the simulator runs on the CPU). Its catalog lists five hash-pinned downloads; an entry that needs a larger memory tier than the device has is shown greyed as "Too big" and cannot be loaded: the experimental Bonsai 2 27B needs the 16 GB tier (iPads and Macs), the other four the 12 GB tier. Speculative decoding is on by default there, and switching it applies to the running engine from the next reply. Sub-agents are on by default too and have their own Settings switch. A turn survives leaving the chat screen, but iOS does not let an app that is not frontmost submit GPU work, so generation pauses while the app is in the background and resumes when it returns. The "Ask TensorAgent" share extension takes text, links, web pages, images, movies and documents from Safari, Photos, Mail, Files and other apps, offers Summarize / Explain / Key points / Action items / Translate presets, and hands the share to the app through a durable App Group inbox as an unsent draft. → TensorAgent - Configurable sampling -- temperature, top-k, top-p, min-p, repetition/presence/frequency penalties, seed, stop sequences
- Structured outputs -- the OpenAI
response_formatJSON schema is compiled to a grammar and enforced by grammar-constrained decoding: any token that would break the schema is removed from the distribution before sampling, so the response is structurally valid by construction rather than repaired afterwards. Supported:type,enum,const,properties,required,additionalProperties,items,prefixItems,min/maxItems,anyOf,oneOf,allOf,$ref/$defs(recursive included),min/maxLength,pattern, the date/time/date-time/uuid formats, and integerminimum/maximum. Keywords a CFG cannot express (not,if/then/else,dependentSchemas,dependentRequired,multipleOf,patternProperties) are refused up front.TS_JSON_GRAMMAR=0falls back to prompt-and-repair. - Chat templates -- auto-loaded from GGUF metadata (Jinja2), with hardcoded fallbacks per architecture
- Inference engine -- the new
InferenceEngine(worker-thread scheduler + paged block pool) replaces the legacy single-request FIFO queue insideTensorSharp.Server. The old queue object is now a compatibility shim for status/event shapes; the engine itself handles concurrency. - Batch processing -- JSONL input support in the console application, plus a built-in inference benchmark for prefill/decode throughput
- Streaming -- token-by-token output via SSE (web) or stdout (console), with abort/stop support for in-flight generations
- Text-diffusion generation -- DiffusionGemma uses an iterative EntropyBound denoising sampler instead of autoregressive
Forward(). The CLI exposes--diffusion-steps,--diffusion-seed, and--diffusion-blocks; the Web UI streams whole-messagereplaceevents for live denoising previews and batches concurrent diffusion requests throughDiffusionBatchScheduler. Image attachments (and the CLI's--image) reach it through the vision tower described under multimodal inference. The server also answers Jev typed-decision requests atPOST /v1/systemonefrom the same checkpoint: the requiredstate(a string, JSON object or array) can come with up to 8 inlineimages, each typed question (noulyes/no,choice,score) gets one answer-canvas slot, and the label probabilities are read after one denoising step rather than generated as text;jev-latest/jev-previeware aliases for the loaded model, andTS_JEV_MAX_BODY_MB,TS_JEV_MAX_CANVASandTS_JEV_MAX_PENDINGbound the request body, the canvas and the queue. → Jev decision inference - Image generation and editing (Qwen-Image-2.1) -- a prompt produces an image, and a prompt plus one or more reference images produces an edited image. The loaded
qwen_imageGGUF is the Qwen-Image-2.1 diffusion transformer — recognized from its tensors, so metadata-free files such as Unsloth's Q8_0 load too, while an earlier Qwen-Image or Qwen-Image-Edit (e.g. 2511) checkpoint is refused at load with exit code 2; TensorSharp resolves its companions alongside it — the dedicated 2.1 VAE, the Qwen3-VL-8B text encoder and, for editing, itsmmprojvision encoder — or takes them from--qwen-image-vae/--qwen-image-vl/--qwen-image-mmproj. Native RGBA input and PNG output preserve transparency. Omitted settings select 2048×2048 output (for edits, about that area at the first reference's aspect ratio), 40 FlowMatch-Euler steps on the official Qwen exponential dynamic-shift schedule, and CFG 1.0, which runs one transformer prediction per step;--width 1024 --height 1024or--diffusion-steps 25are faster draft settings. The diffusion transformer runs as one complete GGML graph with resident quantized weights, and CUDA and Metal reuse that graph across predictions (TS_QWEN21_GRAPH_REUSE); on CUDA that retained graph replays as a CUDA graph. A prefix KV cache (on by default,TS_QWEN21_PREFIX_CACHE) stores the text and reference-image keys and values at the first step, so later steps compute only the target image: editing is roughly twice as fast per step with one reference and about three times with two, and on Metal the output PNG is byte-identical to running without the cache.--tp Nonggml_cuda/ggml_vulkanshards the transformer's heads and MLP columns across the GPUs of one machine, while the text encoder, vision encoder and VAE stay on the first GPU (1.34–1.57x per step on two A40s under CUDA; slower than one GPU on Vulkan, where every reduction goes through host memory). LoRA plug-ins (--lora, repeatable, with--lora-scale/--lora-config) are applied unmerged on top of the quantized weights on every GGML backend, with the prefix KV cache and under--tp; they accept diffusers/PEFT, ComfyUI, kohya, DiffSynth, DoRA and VideoX-Fun PDD files, and twelve ready-made plug-ins inconfig/lora/include step-distillation adapters whose sampling recipes cut a run to 4–8 transformer passes. Driven from C# viaQwenImageModel.GenerateImage/QwenImageModel.EditImage, from the CLI (--prompt,--image,--cfg,--diffusion-steps,--diffusion-seed,--width/--height), from the Web UI with live denoising previews, and over HTTP at/api/image-generateand/api/image-edit(each with a/streamvariant; there is no OpenAI/v1/images/*route). On the server,--widthand--heighttogether set the default size for requests that name none. → Qwen-Image-2.1 card - Video generation with audio (MiniMax-H3) -- a prompt produces an H.264 MP4 and a native 32 kHz stereo soundtrack, generated together: one diffusion transformer denoises a packed video+audio latent in a single token sequence, so the track is part of the model output rather than something added afterwards. Up to 15 s at 24 fps. TensorSharp runs it as native whole-network ggml graphs — one graph per network, weights bound resident straight from the GGUF/safetensors mmap: the Qwen3-VL-32B text encoder (
TSGgml_MiniMaxH3TextEncode), its vision tower (TSGgml_MiniMaxH3VisionEncode), the DiT step (TSGgml_MiniMaxH3DitForward), and video/audio VAE encode and decode; on--backend cputhose same networks run instead as managedMiniMaxH3Direct*implementations (see Pure C# CPU Backend). The DiT is 50 blocks and ~19.3 B parameters, single-stream with no cross-attention — text, conditioning frames, target audio and target video are ONE sequence under full bidirectional attention (hidden 5376, 56 heads × 128 = 7168 inner, patch (t, h, w) = (1, 2, 2)); AdaLN interpolates a learned[8, 1025]curve table instead of running a timestep MLP, and 3-axis RoPE over continuous float positions puts both streams on one timeline measured in audio-latent units (1/40 s). Two checkpoints, not two settings:minimax_h3_fl2va_pruned-Q4_K.gguftakes text and keyframes (--video-mode t2v,i2v— the image IS the first frame and gets animated — orfl2v, first and last frame),minimax_h3_ref2va_pruned-Q4_K.gguftakes identity/appearance references for a new scene (--video-mode ref: up to nine--ref-image, plus--ref-video/--ref-video-audio/--ref-audio, addressed positionally in the prompt as<Picture 1>,<Video 1>,<Audio 1>), and asking one checkpoint for the other's mode fails with a message naming the file to load instead of quietly dropping the input. It ships CFG-distilled, so--cfg 1.0is required (TensorSharp refuses anything higher, and--negative-promptconsequently does nothing — no unconditional pass runs) and 4-8--diffusion-stepsis the fast operating point against a 20-step default.--video-framessnaps up onto a17k+5grid (5, 22, 39, 56, 73, 90 …),--width/--heightround up to a multiple of 32, and fps is pinned to 24 whatever is asked for. The video VAE decodes 5 latent frames at a time with a 2-frame look-ahead and cross-faded seams, and tiles at 256 px — both correctness requirements rather than optimizations, because its decoder is a pure 36-layer transformer whose RoPE coordinates are length-normalized over whatever extent it is handed. Long clips needed one more fix: ggml's flash-attention kernels keep the softmax numerator in FP16 with three bits of headroom, and H3 attends bidirectionally over the whole clip (2364 packed tokens at 22 frames, 8646 at 107), soh3_attendpre-scales V by a power of two derived from the key count and undoes it on the output — exact, because attention is linear in V — which is what turned a 107-frame clip that came back pure black into a correct one; the sampler also refuses to save a diverged sample, failing the request at the step where the latent went non-finite rather than writing a black file (TS_H3_TRACE=1prints the per-step magnitudes). Measured on an M5 Pro (ggml_metal, 22 frames, 8 steps, identical seed) against stable-diffusion.cpp at its best-performing configuration: 20.9 s vs 49.3 s at 256×256 (2.4× faster) and 63.1 s vs 108.5 s at 640×384 (1.7×); on a 16 GB RTX 3080 Laptop (ggml_cuda, same workload, stable-diffusion.cpp at--auto-fit --stream-layers --diffusion-fa --rng cpubecause its default--offload-to-cpupath cannot run this model there) the end-to-end result reverses — 43.6 s vs 37.8 s at 256×256 and 63.7 s vs 59.8 s at 640×384, stable-diffusion.cpp 1.15× / 1.07× — while the per denoise step cost stays TensorSharp's (3.325 s vs 3.338 s by the 8-vs-16-step slope): the difference is fixed setup on a machine with 16 GB of VRAM and 31.7 GB of RAM against a ~35.5 GB model set, and roughly 3 s of the 3.9 s at 640×384 is H.264 encoding against sd.cpp's MJPEG+PCM AVI plus .NET process startup against a native binary, not inference. On a card that size the pipeline manages residency rather than assuming the weights fit: the finished denoiser's device copy is handed back before the video VAE loads when the two would not fit together (peak VRAM during decode 16 041 → ~5 600 MiB, worth 22 s at 640×384 — on Windows/WDDM the oversized allocation does not fail, it is silently backed by host memory and the decode runs at PCIe speed), and the denoiser GGUF is sequentially prefaulted as soon as the text trunk produces its hidden states, pipelined with its own upload rather than joined before it, because weights bound as pointers into the mmap otherwise fault every page in from inside the host-to-device copy — 0.91 GB/s, against 5.97 GB/s once the pages are resident. Output is byte-identical with the prefault on or off; the first denoise step goes 14.87 s → ~10.2 s and the whole run 89.0 s → 63.7 s at 640×384 (67.2 → 43.6 s at 256×256) on that card.TS_H3_PREFAULTpicks the mode (0off,1serial,2overlapped with text conditioning,3pipelined with the upload — the default; mode2loses because the encoder streams its own 17 GB through the same page cache and evicts what was just placed) andTS_H3_PREFAULT_THREADSthe reader count (default1; 4 and 16 streams both measured slower, since this read competes with the teardown and the upload it is warming).TS_H3_PHASE=1prints the per-stage breakdown — encoder open / trunk / teardown, prefault, every denoise step, VAE open / decode — andTS_H3_TE_GROUP=<n>runs the 50-layer text-encoder trunk in groups ofnlayers, releasing each group's device copy; it is off by default because it removes the encoder's own spill (peak 16 041 → 12 981 MiB) bit-identically and is still 3 s slower, a one-shot prefill over a ten-token prompt reading each weight exactly once, so the overflowed ~1.3 GB costs a single PCIe crossing while grouping still moves all 17 GB. Every network is checked against the reference implementation rather than against itself — text encoder cos 0.999999, DiT step video cos 0.9983 / audio cos 0.9998, video VAE encode and decode cos 1.000000, audio VAE decode cos 0.999995. Driven from C# viaMiniMaxH3Model.GenerateVideo(prompt, VideoGenerationParams), from the CLI (--prompt,--image,--end-image,--ref-image,--video-mode,--width/--height,--video-frames,--diffusion-steps,--cfg,--audio-vae,--no-audio), and from the server API —/api/video-generate[/stream]and/v1/videos/generationsacceptvideoMode,endImage,referenceImages/referenceVideos/referenceAudios/referenceVideoAudiosandgenerateAudio, and hand backaudioUrl/audio_urlbeside the MP4, whileGET /api/modelsreports what the loaded checkpoint accepts (video.family=minimax-h3,supportsAudio,supportsEndImageConditioning,supportsReferenceConditioning,maxReferenceImages) so the Web UI offers exactly the attachment controls that will work. The soundtrack is written as a sidecar.wavnext to the MP4 rather than muxed in, because muxing needs an encoder that cannot be assumed present —ffmpeg -i fox.mp4 -i fox.wav -c:v copy -c:a aac fox_with_audio.mp4. → MiniMax-H3 card - Video generation, video-only (Wan 2.1 text-to-video, Wan 2.2 text/image-to-video) -- a prompt (plus an optional first-frame image on the Wan 2.2 models) produces an H.264 MP4 with no audio track. The loaded
wanGGUF is the Wan DiT — Wan 2.1 T2V, Wan 2.2 TI2V-5B (48-channel 16×16×4 latent, 24 fps) and Wan 2.2 A14B (two 14B experts switched at a timestep boundary, second GGUF auto-resolved) are auto-detected; TensorSharp resolves the companions alongside it — the UMT5-XXL text encoder GGUF (prompt → 512×4096 conditioning, exact unigram-Viterbi SentencePiece tokenization) and the matching causal 3D video VAE (wan_2.1_vae.safetensors/Wan2.2_VAE.safetensors). The FlowMatch CFG denoise (UniPC or Euler) runs the whole DiT (self-attention with 3D RoPE + flash attention over F16 keys/values, cross-attention, AdaLN time modulation — per-token-timestep for TI2V image-to-video) as ONE resident-weight ggml graph per step, CUDA-graph-captured per shape (TSGgml_WanDitForward); the video VAE decodes all temporal chunks in one graph with the causal feature cache carried in-graph (TSGgml_WanVaeDecode) -- convs go through MPSGraph on Metal (a 736x544x81f decode: 159 s -> 80 s, 1.99x, numerics unchanged at 93.9 dB PSNR;TS_WAN_VAE_MPS_CONV=0restores ggml's im2col+GEMM lowering) and through a banded im2col+GEMM path elsewhere, with the im2col budget and the tiling threshold now sized from free device memory instead of a fixed 16 GB card's budget, so large-memory devices decode a 720p plane whole (565 s vs 655 s banded, peak RSS 4.85 vs 5.37 GB) while small cards still tile; and image-to-video conditioning encodes the first frame through the causal VAE encoder in one graph (TSGgml_WanVaeEncode). Each stage releases its VRAM before the next, so TI2V-5B 81-frame 480p image-to-video and both A14B Q4_K_M experts fit a 16 GB GPU. Generation runs on every backend except MLX: the GGML paths (ggml_cuda,ggml_metal,ggml_vulkan,ggml_cpu) share the whole-graph kernels, while--backend cudaand--backend cpurun a ggml-independent direct implementation (WanDirect*: resident-quantized linears on TensorSharp's MMQ/dp4a/cuBLAS routing with streaming online-softmax attention kernels on CUDA, parallel SIMD GEMM/attention on CPU, and a channels-last banded-im2col causal video VAE shared by both). Step-distilled checkpoints are auto-detected from the DiT file name (Turbo,distill,Lightning,lightx2v,FastWan,-dmd, or an explicit…-4steps-…) and are by far the biggest speed lever: the official 50-step x CFG recipe costs 100 DiT passes, a 4-step distilled checkpoint costs 4, and the pipeline switches to that step count with guidance off automatically (--diffusion-steps/--cfgoverride). Measured on an M5 Pro at 1088x832x121f = 27 404 tokens,ggml_metal, Wan2.2-TI2V-5B Q8_0: the base checkpoint runs 100 passes at 120.2 s for ~3 h 30 m end to end, and the identical request on a Turbo checkpoint runs 4 passes for 17 m 30 s -- only the--modelpath differs. On base checkpoints--cfg-cache-stride 2/3reuses the guidance direction between steps for a further 1.30x / 1.43x. Numerics verified against diffusers (DiT cosine > 0.995, VAE encoders > 0.999, decoders 59.9 dB / >35 dB PSNR) and across backends (final-latent cosine ≥ 0.999 on identical seeds); the F16 attention keys/values that make one 27 k-token self-attention 2.02x faster (~1.7x per DiT pass together with the VAE work) score the same 0.999964 DiT cosine as F32. Driven from C# viaWanVideoModel.GenerateVideo(prompt, WanVideoParams), the CLI (--prompt,--image,--video-frames,--fps,--flow-shift,--negative-prompt), the server API (/v1/videos/generationswith base64image,/api/video-generate[/stream]withimagePath), and the Web UI chat (type a prompt — with an attached image for image-to-video — and get the video with live progress: per-pass timings, a running ETA and a 30 s heartbeat, since one pass over a 5-second 720p latent is minutes of GPU work). → Wan card - Hybrid SSM-Transformer -- Nemotron-H mixes Mamba2 SSM layers, attention-only layers, and MoE FFN layers in a single model. The Mamba2 step has both a per-sequence native kernel and a batched native kernel (
TSGgml_NemotronMamba2BatchedStepF32, NEON SIMD + GCD parallelism) used by the batched path. On GGML backends the attention layers decode through the device-side flash-attention kernel against the resident KV cache (TS_NEMOTRON_FLASH_DECODE=0restores the host path), so decode no longer degrades with context length. - Hybrid Attention-Recurrent -- Qwen 3.5/3.6-family models mix full-attention layers with GatedDeltaNet recurrent layers; the batched path keeps recurrent running state in a per-slot recurrent-state pool
- Mixture of Experts -- Gemma 4 MoE variants (e.g. gemma-4-26B-A4B), GPT OSS MoE (e.g. gpt-oss-20b), Qwen 3.5/3.6-family MoE (
qwen35moe/qwen3nextvariants such as Qwen3.5-35B-A3B), Nemotron-H MoE FFN layers, and GLM 5.2 (744B-A40B: 256 routed experts at top-8 plus one shared expert, sigmoid-gated routing with a selection-only bias and a x2.5 routed scale, after 3 leading dense SwiGLU layers), GLM-5.3 (the same 79-blockglm-dsashape as 5.2 -- 78 trunk layers and one NextN block, 256 routed experts at top-8 plus one shared expert, the same sigmoid gating and x2.5 routed scale), GLM-5.3-Flash (320B: 288 routed experts at top-8 plus one shared expert and the same x2.5 routed scale, with a SwiGLU clamp limit of 10 on every FFN), and Qwen 3.8 Flash Next (512 experts, 10 used per token, interleaved with GatedDeltaNet recurrent layers) - MoE CPU offload --
--n-cpu-moe N/--cpu-moe(llama.cpp's-ncmoe/-cmoeequivalent) keeps the routed expert weights of the first N layers in system RAM and multiplies them on the host, leaving attention, the norms, the router and the always-active shared expert on the accelerator. The offloaded layers stay inside the fused whole-model graph on every architecture that has one (Qwen 3.5/3.6, Gemma 4 MoE, GPT OSS, DiffusionGemma) — the accelerator pauses after each offloaded layer's router, the host multiplies the selected experts straight out of the GGUF mmap, and the result is handed back before the next segment — so only ~8 KB of activation crosses the bus per layer at decode. Gemma 4 MoE and Qwen 3.5/3.6 segment their prefill graphs the same way, where the host side becomes a real GEMM over the whole prompt chunk. It also composes with tensor parallelism: under--tp Nthe seams merge into the ranks' own AllReduce segment schedule, so the fused multi-rank graph is kept and each offloaded layer is evaluated once on the host over the unsharded expert stack (Qwen3.5-35B-A3B--tp 2: 17.4 GB of resident weights across 2 GPUs falls to 3.2 GB; gemma-4-26B-A4B: 12.9 GB falls to 2.4 GB, byte-identical output). Measured on a 16 GB RTX 3080 Laptop: Qwen3.6-35B-A3B 13.4 -> 4.6 GB at--cpu-moe; gemma-4-26B-A4B 16.1 -> 4.8 GB (decode 39.7 -> 17.7 tok/s, and 38.6 tok/s at--n-cpu-moe 8for 3 GB back); gpt-oss-20b 16.2 -> 2.9 GB, which takes it off the WDDM spill cliff and turns 0.3 tok/s into 25.4 at--n-cpu-moe 12. That is what makes these models fit beside a long-context KV cache on a 12-16 GB GPU. DeepSeek V4 Flash uses the same seam on the GPU backends: it is 91% routed-expert bytes, and the loader sizes its layer split against each device's actual free VRAM. Offload stays opt-in there too -- a checkpoint that does not fit is refused at load with the exact--n-cpu-moe Nthat would make it fit, rather than silently trading away decode throughput -- and with that flag 3x48 GB RTX A6000 hosts the UD-Q8_K_XL checkpoint at 10 tok/s decode / 126 tok/s prefill. GLM 5.x offloads the same way (92% of that checkpoint is routed-expert bytes) and its host-resident experts are served straight from the GGUF mapping rather than copied. Note that offload is for fitting, not speed: on 3x RTX PRO 6000 where GLM-5.2 already fits,--n-cpu-moe 30costs pp2048 915.9 -> 94.7 and tg64 43.9 -> 16.4 tok/s. On GLM 5.2 the host-resident experts are multiplied straight out of the GGUF mapping with no private copy, and offload composes with--tp N— host-resident layers keep their experts whole and rank 0 evaluates them. It also nearly doubles the context the loader can size there, 342,272 -> 646,400 tokens. -> MoE CPU offload - Batched GPU MoE -- a single fused GGML graph dispatch handles all selected experts (plus the optional shared expert and residual add) for Qwen 3.5/3.6-family and Nemotron-H decode, eliminating per-expert round-trips
- Whole-model fused decode graphs -- Gemma 4 (dense and MoE), Qwen 3.5/3.6 and GPT OSS run an entire decode token — every layer, the MoE router and experts, the final norm and the LM head — as ONE GGML graph dispatch instead of one submission per layer, so the GPU is never left waiting on the host between layers. On CUDA/Vulkan the graph is built once with stable tensor addresses and replayed (
ggml_set_rowsKV write with the row as an I64 input, a stride-padded attention window with an F16 mask input), which is what lets ggml-cuda capture it as a CUDA graph. GPT OSS decode: 24 → 154 tok/s on an A40, and flat in context length (133 tok/s at 16K) where the per-layer path collapsed to 2.3. Disable per model withTS_GPTOSS_MODEL_DECODE=0/TS_GEMMA4_FD_PERSIST=0/TS_QWEN35_FD_PERSIST=0. - KV cache codecs -- pluggable codec interface (
IKvBlockCodec) with a built-in TurboQuant (2-bit affine / Q4 / Q8) compressed codec for the blocks of the standalonePagedKvCacheManager, which only the CLI's--paged-benchruns. Both hosts accept--paged-kv-quant-bits 0|2|4|8(TS_KV_PAGED_QUANT_BITS), but on the server the value is only stored inTS_KV_PAGED_QUANT_BITSand has no effect, because no serving path builds that manager. The 2-bit tier reaches ~10x compression on fp32 blocks for very long contexts. - KV cache precision --
--kv-cache-dtype <f32|f16|q8_0|q4_0>(CLI and server, envKV_CACHE_DTYPE; default auto — the backend/model pick) trades a small numerical drift for memory.q4_0(~0.56 bytes/element, ~1/7 of f32) is the most aggressive tier and is aimed at the very long (128K–256K) contexts where the KV cache dominates memory; the block-quantized tiers (q8_0/q4_0) require the native GGML flash path. - Message editing -- edit or delete previous messages in the web chat UI and regenerate from that point
- Text/Image/Audio/Video/PDF uploads -- the web UI accepts file uploads up to 500 MB and preserves text content in full. Born-digital PDFs have their complete text layer extracted and inlined into the prompt (cap pages explicitly with
TS_PDF_MAX_PAGES); scanned PDFs are rendered to page images for vision-capable models. The final prompt is checked against the model's actual context window instead of an arbitrary upload budget. The CLI accepts a PDF in one-shot mode via--pdf <file> - Per-turn observability -- structured logs capture the full user input and the full raw assistant output (both
<think>reasoning and the final result) plus the KV cache hit ratio. The same cache-hit stats are surfaced through every API:prompt_cache_hit_tokens/prompt_cache_hit_ratio(Ollama),usage.prompt_tokens_details.cached_tokens(OpenAI), andpromptTokens/kvReusedTokens/kvReusePercentin the Web UI SSEdoneevent
Models that support thinking mode (Qwen 3.5/3.6-family, Qwen 3.8 Flash Next, Gemma 4, GPT OSS, Nemotron-H, Muse-Glimmer, DeepSeek V4, DeepSeek V4.1, GLM 5.x) can produce structured chain-of-thought reasoning before generating the final answer. The thinking content is separated from the main response and can be displayed or hidden by the client.
- Qwen 3.5/3.6-family / Qwen 3.8 Flash Next / Nemotron-H: uses
<think>...</think>tags. On Nemotron-H (nemotron_h,nemotron_h_moe,nemotron_h_omni)response_formatcombines with"think": true: the JSON grammar arms after</think>. Where the vocabulary has</think>as one token (Nemotron 3.5, Nemotron Omni),TS_THINKING_BUDGETcloses the block with it and the answer continues insidemax_tokens. Whitespace between</think>and the JSON is allowed: the Nemotron-H Reasoning-128K vocabularies merge>with the line break that follows it, and masking that token made the model write</think}, so the grammar never armed - Gemma 4: uses
<|channel>thought\n...<channel|>tags.response_formatcombines with"think": true(the JSON grammar arms after<channel|>). The model opens that channel itself, so it is budgeted from<|channel>: with thinking on,TS_THINKING_BUDGET(75% of an allowance of at least 512 tokens) closes it with<channel|>and the answer continues insidemax_tokens. With thinking off, a channel the model opens anyway (E4B does after tool results) is hidden and closed at the first line break past a quarter of the allowance (at most 64 tokens), and after a tool result a<channel|>with no open channel is masked, so the answer is neither empty nor streamed twice.TS_THINKING_BUDGET=0disables both caps - Muse-Glimmer: reasons in an
assistant to=selfmessage and answers inassistant to=user.response_formatcombines with"think": true: the JSON grammar arms afterto=user<|message|>; with thinking off it enforces from the first token and the parser reads the headerless object as the answer - GPT OSS: uses Harmony format with
<|channel|>analysisfor thinking and<|channel|>finalfor the response. Reasoning cannot be switched off, only shortened: the OpenAIreasoning_effortfield (low/medium/high, defaultmedium; other values are HTTP 400) sets the HarmonyReasoning:line, and an explicit"think": falsewithout an effort renders atlow.response_formatcombines with"think": true(the grammar arms at the final channel either way) - DiffusionGemma: thinking is not prompted — the prompt always renders with thinking off — but the model sometimes writes Gemma 4's channel syntax anyway; that thought block is parsed out of every preview and the final text and returned as reasoning only on
"think": true, and a reply with nothing outside the thought block (the checkpoint often leaves it unclosed) is returned as the answer instead of an empty message. Tools /tool_choiceare refused with HTTP 400 - DeepSeek V4: uses
<think>...</think>tags; the chat template closes the block for you unless--thinkis passed, so reasoning is opt-in - DeepSeek V4.1: also
<think>...</think>, with a trained closing transition — whenTS_THINKING_BUDGETis reached the model emits</think>and continues the final answer inside the originalmax_tokens. The budget defaults to 75% of the output allowance when that allowance is at least 512 tokens; smaller allowances get no automatic budget, and0disables it. The repetition guard asks for the same closing transition when reasoning enters a loop - GLM-5.2 / GLM-5.3 (
glm-dsa): uses<think>...</think>tags, opt-in like the rest ----thinkadds theReasoning Effort: Maxsystem line and leaves the generation prompt's block open for the model to close; without it the prompt emits an empty<think></think>so the model answers directly. Past turns' reasoning is always dropped from the prompt, matching the template'sclear_thinkingdefault - GLM-5.3-Flash (
glm5next): also<think>...</think>, but it always reasons: the system line mapsreasoning_effort: "low"/"high"toLow/High, withMaxfor omitted or"medium"effort, the generation prompt always opens<think>, and past turns keep their reasoning (clear_thinkingdefaults to false)."think": falsecannot turn reasoning off; the reply is still parsed as reasoning up to</think>, and only what follows is the answer
Enable via --think (console), "think": true (Ollama API), or the thinking toggle in the web UI.
DeepSeek V4 ships DSpark ("Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation") as a support module in the checkpoint: three DSV4 blocks that read the trunk's hidden states and propose a whole BLOCK of tokens per step instead of one, a Markov head that conditions each block position on the token before it, and a confidence head that predicts each position's acceptance probability. TensorSharp loads it as a separate drafter GGUF (--draft-model, built with eng/dsv4-dspark-to-gguf.py) and runs it on both GPU engines (--backend cuda and --backend ggml_cuda) for greedy single-sequence generation — on ggml the drafter is three extra graph layers whose key ring the trunk graph commits itself, so speculation costs no host round trips; the trunk verifies each block in one batched forward and keeps only the prefix it would have produced anyway. Measured 1.3-1.4x decode on 4xA40 with the drafter's cumulative-confidence gate at its default; the gate matters because each extra verify row pulls a fresh set of MoE experts through VRAM. See the DeepSeek V4 card.
DeepSeek V4.1 (deepseek41) DSpark is experimental: the loader accepts a deepseek41-dspark drafter through --draft-model on ggml_cuda and ggml_cpu only; initial text/image HTTP probes with trained weights passed using two-GPU layer split on ggml_cuda; broad quality and throughput remain unqualified.
Muse-Glimmer and Qwen 3.8 both have a block drafter, DFlash: a separate 5-layer GGUF (general.architecture = dflash) that proposes the whole speculative window in one forward. It is architecture-agnostic on the TensorSharp side - a target model gets it by tapping the per-layer residuals the drafter's encoder reads, and nothing else. It borrows the target's token embedding and LM head, keeps its own sliding-window KV ring, and runs three passes per step — encode the trunk's per-layer input residuals at dflash.target_layers into one wide row, inject that row as the K/V of every draft layer, then draft [anchor, MASK x (block-1)] through the five blocks and score it with the target's LM head. The trunk verifies the block in one batched forward and keeps only the prefix it would have produced anyway, so the emitted token stream is the plain-greedy stream.
DFlash2 is the same backbone plus two additions, both keyed off the GGUF so one code path serves either generation: a grouped dynamic depthwise convolution around every attention and every FFN sublayer (which gives a block-diffusion draft a local left-to-right signal without a second forward), and a candidate selector that scores the top-K candidates of adjacent positions pairwise through two low-rank [vocab, r] codebooks and reads the block off as a walk through that lattice - so position i+1 is no longer chosen without knowing what i chose. Attach either with --draft-model; the file says which it is. See speculative_decoding.md.
Both halves are fused native graphs that CUDA-graph-capture and replay, and the draft block finishes with an on-device argmax (or, for DFlash2, a ~7 KB lattice), so the 202048-wide probability block never crosses PCIe. A runtime cost governor measures speculation against plain decoding and parks the drafter while it is measurably slower — speculation can therefore only help, but it needs a few hundred generated tokens to settle. Load it with --draft-model (CLI) or TS_MUSE_GLIMMER_DFLASH/TS_QWEN35_DFLASH. Sampling composes — verification draws each token from a trunk row with the run's own sampler — but note that a block drafter proposes its whole block in one pass, so penalties are not applied to the proposal and acceptance falls as a penalized history grows. Measured against llama.cpp's own DFlash on one RTX PRO 6000 Blackwell (Q8_0, greedy, 60-token prompt): 50.9 tok/s vs llama.cpp's 45.5 and 35.0 plain. See the Muse-Glimmer card.
Nemotron 3.5 Lightning refuses speculative decoding, its DSpark drafter included. The official nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4-DSpark module (6 SWA-1024 layers, per-head attention sinks, rank-512 Markov head; eng/nemotron-dspark-to-gguf.py rebuilds the GGUF) is recognized, but --draft-model does not attach it and --spec / --spec-type ngram serve plain decoding with a one-time warning naming the reason. Speculation's contract is that the output equals plain greedy decoding, and this hybrid trunk cannot keep it: its multi-token verify and its single-token decode run different attention and MoE kernels whose logits for the same tokens differ by 0.2-1.1 on ggml_cuda, which flipped greedy picks on every solo request of the 2026-09-16 campaign; a verify that runs those layers row by row is bit-exact but costs 4.1 plain steps for 4 rows (116 ms against a 28 ms decode step, plus 70 ms per DSpark block), so no mode is both exact and faster. The Mamba-2 snapshot/rollback itself is exact. Separately, the shared DFlash attention-sink path added the sink to the post-softmax weights instead of treating it as an extra softmax logit, which made this drafter propose noise (0 of 47 teacher-forced next tokens; 17 of 47 after the fix). See speculative_decoding.md.
A drafter proposes several future tokens cheaply, the trunk verifies all of them in one batched forward, and accepted tokens are committed in a single step. Because the request's own sampler — temperature, top-k/p, and repetition/presence/frequency penalties — drives both the draft and the verify, the output is identical to standard decode; speculation only changes how many forward passes it takes to produce it. Engages for solo (non-concurrent) sequences.
Multi-turn chat. Speculation used to arm only on a turn that started from an empty KV cache, so a Web UI conversation sped up on its first turn and then silently never again — which is what made DFlash2 look useless from turn 2 onward. Whether an algorithm can start drafting on top of a reused KV prefix is now the algorithm's own call (ISpeculator.CanArmAfterPrefixReuse): the block-draft and n-gram speculators opt in, and so does a draft head that can resume after a gap (Gemma 4's assistant head, and the NextN head of Qwen 3.5/3.6 and Qwen 3.8 27B; see Prefix reuse below). Measured 1.02x → 1.85x in the server chat path.
The design is three independent layers — model architecture ≠ speculation algorithm ≠ speculator weights — so a new model that implements ISpeculativeTarget reuses every algorithm, and a new algorithm reuses every such model. A model implements ISpeculativeTarget (a multi-row forward with per-row logits, plus the KV rollback trio); an algorithm implements ISpeculator (Propose / Commit) and is registered by name; a checkpoint's trained drafter sits behind IDraftHead, a thin model-specific adapter, because those weights are bound to one target model and do not transfer. Adding EAGLE, Medusa, PARD or a tree-drafting variant is a new class plus one SpeculatorRegistry.Register call — no model, executor or scheduler code changes. See Speculative Decoding in TensorSharp.
Four algorithms ship today, selected with --spec-type:
--spec-type |
Drafts by | Needs trained weights? |
|---|---|---|
auto (default) |
whatever drafter the checkpoint carries | — |
draft-head |
one token per pass through a NextN/MTP head, chaining its own hidden state (Qwen 3.6, Qwen 3.8 27B, GLM 5.2, GLM-5.3, Gemma 4's separate assistant GGUF, Qwen 3.8 Flash Next's separate shared MTP GGUF) | yes, per target model |
block |
a whole block per pass with a confidence head (DeepSeek V4 DSpark, DFlash / DFlash2 on Muse-Glimmer and Qwen 3.8) | yes, per target model |
ngram |
suffix match over the sequence's own tokens — where did these last few tokens occur before, and what followed? | no |
ngram is the weight-free one: it needs no trained weights, so it works on checkpoints that ship no speculator at all, and is strong wherever the answer quotes its input — summarizing, editing, translating or answering about a document, repetitive structured output, code with repeated identifiers, agentic tool loops. Measured on Qwen3.5-9B (Q8_0, ggml_metal, M5 Pro), which ships no draft head: 45.2 tok/s vs 31.4 plain (1.44x) on a reproduce-this-config prompt, with byte-identical output. On free-form prose it finds nothing, every step degrades to a plain decode, and the runtime cost governor keeps that cheap. It still needs an architecture that can verify a draft: Qwen 3.5/3.6/3.8, Gemma 4, GLM 5.x and Qwen 3.8 Flash Next take it; DeepSeek V4 and Muse-Glimmer speculate only while their own drafter is loaded; GPT OSS, Mistral 3 and Hunyuan Dense have no speculative path and decode plainly; and Nemotron-H refuses every speculator, ngram included (see above).
Speculative decoding is off by default. --spec (env TS_SPEC=1), on the server or on TensorSharp.Cli, is the explicit opt-in for drafters embedded in the trunk checkpoint (the NextN blocks of Qwen 3.6, Qwen 3.8 27B, GLM 5.2 and GLM-5.3), because loading them pages extra weights into VRAM; a drafter that ships as its own GGUF is enabled by --draft-model alone, with an explicit --no-spec as the veto. Env vars are still published under both TS_SPEC_* and TS_MTP_* — the glm-dsa native loader reads TS_MTP_SPEC / TS_MTP_DRAFT from C++ at load time, so those names are a cross-language contract:
# Qwen 3.6 — use the -MTP- repository GGUF so the embedded NextN block is retained
dotnet TensorSharp.Server.Host/bin/TensorSharp.Server.Host.dll --model models/Qwen3.6-35B-A3B-UD-Q4_K_M.gguf --backend ggml_cuda \
--spec --spec-draft 8 --spec-pmin 0.75
# Gemma 4 — load the separate gemma4-assistant draft GGUF that matches the target
dotnet TensorSharp.Server.Host/bin/TensorSharp.Server.Host.dll --model models/gemma-4-E4B-it-Q8_0.gguf --backend ggml_cuda \
--draft-model models/gemma-4-E4B-it-assistant.Q8_0.ggufFour draft-head shapes:
- Qwen 3.6 / Qwen 3.8 27B (embedded NextN) — the GGUF carries one extra decoder block past the main stack (
{arch}.nextn_predict_layers) plus the NextN projection/norm tensors. No separate file is required;--draft-modelis ignored. The recurrent trunk state (GatedDeltaNet) is snapshotted so a partially-rejected verify batch can be rolled back. - GLM 5.2 and GLM-5.3 (embedded NextN) — same shape, and the stock unsloth/GLM-5.2-GGUF already carries it (
blk.78.nextn.*plus a full MLA + 256-expert decoder block). Nothing extra to download:--specis the whole configuration, on the CLI (--input,--multi-turn-jsonl,--interactive) as well as the server. The block is only paged in when that flag is set, because it is a whole extra decoder layer (~3 GiB at IQ2_XXS) competing with the KV cache for the same VRAM the loader sizes the context against. glm-dsa has no recurrent state, so a partially-rejected verify keeps the accepted prefix's KV and only rewinds the position — no re-forward. unsloth/GLM-5.3-GGUF carries the sameblk.78block on the same terms — nothing extra to download, and--specon the command line before load, because that flag is what pages it in. Neither block ships anextn.shared_head_head.weight(GLM-5.2's has nonextn.embed_tokenseither), so both borrow the trunk LM head, and under--tp N > 1that head is column-parallel and the loader refuses to draft from one rank's strip of the vocabulary — decided before any weight byte is placed, with a stderr line and standard decode for the rest of the run (no test covers that refusal); the CLI also warns when--specand--tpare combined. Speculation therefore engages on single-device or explicit--layer-split Nplacement, with no active tensor parallelism. See the GLM card and the GLM-5.3 section. GLM-5.3-Flash (glm5next) is the family's exception: its NextN block is not built, so--spec --spec-type ngramis what speculates there, and because its KDA layers carry recurrent state the trunk snapshots that state before every verify and restores it on a partial rejection instead of rewinding a position — see speculative decoding on GLM-5.3-Flash. - Gemma 4 (separate
gemma4-assistantGGUF) — an EAGLE-style recurrent drafter loaded with--draft-model, which enables speculation by itself. It holds no K/V of its own: every draft layer queries the target model's existing per-layer KV cache (last local + last global layer), so the drafter is stateless given(token, hidden). The draft's hidden size must match the target — pair the 12B target with its 12B draft, not the 26B-A4B draft. A mismatched, missing, or incomplete draft GGUF fails fast at startup with a remediation hint instead of silently disabling speculation. - Qwen 3.8 Flash Next (separate shared MTP GGUF) —
--draft-model mtp-Qwen3.8-Flash-Next-shared-Q8_0.ggufattaches a per-token MTP block on the GGML backends. It keeps its own K/V, so it drafts only for a solo request that prefilled from position 0; turns that continue a retained holder or a shared-prefix clone decode plainly. On a 192-token code-copy stream it measured about 1.7x (83.2 against 49.1 tok/s plain) with the exact verify-row kernels added on 2026-09-17 (1.75-1.96x before that change), with output identical to plain greedy, while prose ran slower than plain decode — see the Qwen 3.8 Flash Next card.
Prefix reuse. --spec does not disable the Radix prefix cache (an explicit --no-prefix-cache still does). The NextN head of Qwen 3.5/3.6 and Qwen 3.8 27B (qwen35) restarts only its private attention cache after a reused prefix — startup-warmed, disk-restored or retained — so prefill takes the normal fast path with speculation on; this covers the default fused-verify route for a solo request. Gemma 4's assistant head holds no state and drafts from any position, and the block and n-gram speculators arm after a reused prefix too. The GLM-5.2 / GLM-5.3 NextN head does not resume after a gap, so, like a Qwen 3.8 Flash Next turn that continues a retained holder or clone, a GLM turn that adopts a cached prefix decodes plainly.
Where it's profitable (engaged automatically; otherwise the engine serves standard decode). The GLM column describes the shared glm-dsa path, so it covers GLM-5.3 as well as GLM-5.2, but the runs behind those verdicts were measured on GLM-5.2 — no GLM-5.3 speculation run exists yet:
| Backend | Qwen 3.6 | GLM 5.2 / GLM-5.3 (glm-dsa) |
Gemma 4 |
|---|---|---|---|
| GGML CUDA / GGML Metal | ✅ fused multi-token-verify + draft-step kernels | ✅ one graph per verify window in the native whole-model executor | ✅ fused multi-token-verify + draft-step kernels |
Direct CUDA (cuda, driver-API/cuBLAS) |
✅ GPU-resident per-op verify/draft | — (GLM takes the managed per-op path on cuda as well as cpu, and on a GGML backend with TS_GLM_NATIVE=0; it is a reference path, not a fast one) |
✅ GPU-resident per-op verify/draft |
GGML CPU (ggml_cpu) |
speculates: nothing gates it by backend, and the cost governor parks drafting while it measures slower (the default 8-token window measured 1.21x on Qwen3.6-35B-A3B) | the same native whole-model executor as the GGML GPU rows | speculates (every GGML backend qualifies) |
Pure C# CPU (cpu) / MLX |
speculates, left to the cost governor as on ggml_cpu |
per-op reference path (correct, not fast) | standard decode |
Tuning: --spec-draft (default 8) bounds tokens drafted per step; --spec-pmin is the confidence gate (0 = never gate), and drafting stops at the first token below it. What that number means is the algorithm's business, so each brings its own default rather than sharing one: 0.15 for a per-token head (top-1 probability over its top-10 logits), 0.35 cumulative for a block drafter, 0 for n-gram (where it scales the required match length instead). The two knobs interact — a wide window occasionally forms a long chain that is mostly rejected, and those verify rows are paid for either way — so they are worth sweeping together on a new model/host pair. On GLM 5.2, --spec-draft 4 --spec-pmin 0.55 was best or tied-best in every measured run, by ~4% over the defaults. Gemma 4 draft-path A/B switches are the TS_GMTP_* env vars (see the MTP / speculative-decoding tunables table under Web Application). Per-architecture mechanics are in the Qwen 3.5/3.6 card, the GLM card and the Gemma 4 card.
Greedy output and floating point. Verification draws every emitted token from a trunk row, so speculation cannot change which distribution a token comes from. It does change the arithmetic: a K+1-row verify runs the trunk's matmuls at a different batch size than a 1-row decode, which selects different kernels and reduction orders. On a dense model that is usually invisible but not always: on ggml_cuda a half-precision (BF16/F16) matmul over one row and over several rows differs by up to 0.2%, and Gemma 4 E4B's per-layer-embedding projection is BF16, so a verify row sits 0.4-0.65 logits from the one-row decode of the same token (Metal: 0.003-0.014, ggml_cpu: bit-identical); across 512 greedy prose tokens that flipped 4-5 near-ties, every one with a top-two margin under 0.25 (measured). On GLM-5.2 — 2-bit weights, 256 experts at top-8 — a last-bit difference in a router logit changes which experts run, and 78 layers amplify it. Measured over 140 verify rows against per-token decode: the top token differs on 2.9% of rows, so a long greedy run eventually takes a different (equally valid) branch. Runs with drafting suppressed reproduce greedy exactly, which is what locates the effect in the batch size rather than in the speculation.
--backend cpu is TensorSharp's 100% pure C# path. Its managed matmuls now run on
a persistent spin-then-park worker pool (TensorSharp.Models/CpuWorkerPool.cs)
instead of a Parallel.For per matmul. Two problems were fixed together: the
work-item count used to scale with the thread count -- 1024 tiny tasks per matmul
at 122 threads, which stopped scaling past 8 -- and every matmul paid a ThreadPool
fork/join.
Measured on gemma-4-E4B-it-Q8_0 with a 122-CPU allocation, --backend cpu, with the
pool-off baseline A/B-ed inside one binary via TS_CPU_POOL and the widths set by
TS_CPU_THREADS (each cell is the two interleaved runs, tok/s):
| pool width | prefill | decode |
|---|---|---|
| off (before) | 21.7 / 21.0 | 2.0 / 2.4 |
| 32 | 24.9 / 24.1 | 4.9 / 5.0 |
| 48 | 25.6 / 28.5 | 5.4 / 6.0 |
| 61 (default = cores/2) | 24.2 / 24.9 | 6.3 / 5.9 |
| 122 (every core) | 13.5 | 4.8 |
So prefill gains ~15% and decode ~2.8x. The default width is deliberately half
the usable cores rather than all of them: pool workers spin between jobs while the
rest of the CPU path still uses the ThreadPool, so a pool that owns every core
starves the work it is waiting on. At 122 threads that shows up as a PREFILL
regression (13.5 against 21.7 / 21.0 with the pool off) while decode still beats
the pool-off baseline (4.8 against 2.0 / 2.4). The 61-wide default is simply the
best row overall: it beats the 122-wide one on both axes, and gives up a little
prefill against the 48-wide one for equal-or-better decode. Tune with TS_CPU_THREADS, TS_CPU_POOL,
TS_CPU_SPIN, TS_CPU_TASK_BYTES and TS_CPU_TASKS_PER_WORKER
(env var matrix).
GEMM and SIMD kernels. Quantized weights (Q4_K, Q5_K, Q6_K, Q4_0, Q5_0, Q8_0)
now go through a multi-row int8 GEMM that decodes each pair of weight rows once for
every activation row (AVX-512BW 8x2 and AVX2 4x1 register tiles), and F16 / BF16 /
F32 or dequantize-only types through a float-panel GEMM. F32 matmuls behind
Ops.Addmm use a packed, cache-blocked SGEMM (AVX-512 8x32, AVX2 6x16, portable
fallback), and elementwise, norm, softmax and RoPE ops use Vector512 / Vector256
kernels; TensorSharp.Models binds Core's parallel loops to the same worker pool.
DiffusionGemma gained a host prompt-KV cache and batched MoE on this backend, and
Qwen-Image-2.1 a managed pipeline (below). Measured on an i7-11800H (8 cores / 16
threads, AVX-512, 32 GB):
| Workload | cpu before |
cpu now |
ggml_cpu |
|---|---|---|---|
| DiffusionGemma-26B-A4B Jev read, new 54-token prompt, width 16 | 9.5 s | 0.92 s | 1.57 s |
| same, further reads of the same prompt | 9.5 s | 0.22–0.25 s | 1.57 s |
| Qwen-Image-2.1 256×256, 2 steps | refused | 24.6 s | 36.1 s |
| Qwen-Image-2.1 512×512, Pruna 5-step LoRA | refused | 169.6 s | 295.7 s |
Jev label decisions matched ggml_cpu on all 36 labels of the probe's quality set.
The Qwen-Image images are 42.0 dB (256×256) and 31.4 dB (512×512) PSNR from the
ggml_cpu ones, which quantize activations to 8 bits where the managed transformer
does not. Every kernel family has a switch that restores the previous code in the
same binary (TS_CPU_QGEMM=0, TS_CPU_FGEMM=0, TS_CPU_SGEMM=0,
TS_CPU_SIMD_ELEMENTWISE=0, DIFFUSION_CPU_LEGACY=1). TS_CPU_DISABLE_AVX512=1
runs the AVX2 kernels on an AVX-512 host; to emulate an AVX2-only host use
DOTNET_EnableAVX512=0 (.NET 10 ignores DOTNET_EnableAVX512F=0). The AVX2 kernels
were not run on AVX2-only hardware, and the portable paths that ARM64 and x64
without AVX2 take were not run on such hosts either.
Zero-copy quantized weights in the ModelBase loader. BackendType.Cpu was the only backend missing
from CanUseFileMappedQuantizedWeights, so it alone copied every quantized
tensor into fresh anonymous memory at load instead of binding it straight from
the GGUF mapping, as every GGML backend already did. It now binds them zero-copy
-- ManagedQuantizedOps reads a weight through a raw pointer and never writes to
it -- and the loader reports the split, e.g. for GLM-5.3-Flash UD-Q2_K_XL:
Quantized: 103255 MB (103255 MB file-backed), F32: 983 MB
Any model whose weights used to be copied benefits and the effect is largest on big quantized checkpoints: that one went from a load that never completed (resident set 412 GB and still climbing) to ~48 s, most of which is the page-cache prefault. The embedding encoder has a separate loader: it owns compact quantized arrays and repacks supported projections; see embedding execution.
IQ2_XS / IQ4_XS, and direct i-quant dots. ManagedQuantizedOps gained
managed dequantizers for IQ2_XS and IQ4_XS (verified against ggml's own
dequantize_row_*) plus entries in the CPU quantized-storage matrix, so weights
of those types are kept quantized instead of expanded to F32 at load -- an
expansion that came to 765 GB for GLM-5.3-Flash UD-Q2_K_XL, which is why that
load grew rather than failing. It also gained direct IQ2_XS x Q8_K and
IQ3_XXS x Q8_K dot kernels with AVX2 paths (VecDotIq2XsQ8KAvx2,
VecDotIq3XxsQ8KAvx2); previously both types fell to the generic
dequantize-the-row-into-scratch path. This is backend-wide, not GLM-specific: it
applies to any model on --backend cpu.
One trap worth recording, since it does not announce itself: ggml folds a
constant into the result of some i-quant dots -- 0.125 for IQ2_XS, 0.25 for
IQ3_XXS, 1.0 for IQ3_S -- rather than into each per-block scale. Omitting it
is an 8x error that produces fluent-looking garbage rather than a crash.
GLM-5.3-Flash on --backend cpu. glm5next first could not load and then
could not run on the pure C# path; four fixes, three of them silent rather than a
clean error, make it work for text. NoPE MLA (GLM-5.3-Flash sets
rope.dimension_count = 0, so there is no rope half anywhere and the compressed
latent IS the whole cache row, where the GLM-5.2 path narrowed a zero-width
slice; plain GLM-5.3 keeps rope — base 8e6, n_rot 64 — and stays on the
GLM-5.2 path), an MLA-absorption check that demanded attn_k_b / attn_v_b on
all 45 trunk layers when glm5next carries them on only the 12 full-attention ones
starting at layer 3, and the two loader problems above. Measured on the 122-CPU
box, 22-token prompt, 16 tokens out (tok/s):
| prefill | decode | |
|---|---|---|
cpu, scalar i-quant dots |
0.9 | 0.4 |
cpu, AVX2 i-quant dots |
3.1 | 1.6 |
ggml_cpu, same file and box |
17.7 | 3.9 |
So ~5.7x off ggml_cpu on prefill and ~2.4x on decode. About 89% of the time is
the MoE expert path -- 8 experts x 3 matrices x 45 layers of small matmuls per
token -- and it sits OUTSIDE the Linear timing bucket, so the built-in
breakdown reports it as "Other"; that is not unaccounted overhead.
This route is not claimed as parity. The prefill-logit cosine against
ggml_cpu is 0.9567 over the 154880-wide vocabulary (compared with
TS_DUMP_LOGITS, which writes the first real forward's logits and skips the
warm-up ones). The greedy text differs because native's top two logits are 0.11
apart and the managed path ranks them the other way, putting native's pick at
managed rank 2. That is consistent with 2-bit expert-pick sensitivity -- the same
effect that makes TS_BATCHED_FUSED_DECODE=0 useful for an exact serial-path A/B -- but it is not
proven to be only that: 0.96 is lower than the ~0.999 a higher-precision
checkpoint would be expected to give, and no higher-precision GLM-5.3-Flash GGUF
was available as a control. Treat the managed path as a reference implementation
to A/B against, not as bit-parity. Every number in this subsection is glm5next:
no managed-path run of the non-Flash GLM-5.3 exists anywhere.
→ GLM card
DeepSeek V4.1 Flash on --backend cpu. deepseek41 runs on the pure C#
DeepSeek4CpuExecutor — no ggml, no native library and no GPU, so the
architecture runs anywhere .NET runs. The managed executor implements the whole
V4.1 graph: the ratio-1 and ratio-2 block compressors, the shared compressed and
indexer caches with the lightning indexer's top-k, candidate block pruning, the
host-mapped Engram tables gathered one 256-value row at a time, the delayed
hyper-connection gates, the shared expert, and the checkpoint's trained cache
quantization (FP8 E4M3 raw rows, MXFP4 indexer, NVFP4 compressed). It is held to
the independent PyTorch oracle eng/dsv41-reference.py at atol=rtol=2e-5 plus
exact greedy-argmax agreement by InferenceWeb.Tests.Dsv41CpuExecutorTests, over
one-shot prefill, chunked prefill at 1/3/5/8 checked at every position, and
reset. Two caveats travel with that gate: it is fixture-scale — a five-layer,
256-hidden, 16-token F32 synthetic model, so it proves architectural agreement
with the oracle, not CPU/CUDA parity on the real 246 GiB Q2_K weights — and the
tests return silently unless TS_DSV41_FIXTURE_DIR is set. The direct-CUDA
engine's layer-0 trace treats this managed executor as its pinned baseline, and
blk00_raw_k is bit-identical between the two.
Like --backend ggml_cpu, this is a correctness and portability path, not a
serving path. No throughput, load time or resident footprint has ever been
measured for a full V4.1 checkpoint here, so the pool-width table above —
gemma-4-E4B on the generic CPU worker pool, swept with TS_CPU_THREADS /
TS_CPU_SPIN — says nothing about V4.1, and there are no V4.1 numbers to put in
its place. It keeps no per-sequence slots and no multi-turn KV prefix reuse,
so every diverging turn re-prefills and concurrent requests serialize. Engram configuration
is read directly from the GGUF. TS_DSV4_THREADS defaults to
ProcessorCount on this backend rather than min(cores, 32), and
TS_DSV4_CPU_TRACE_DIR writes the same per-tensor files
eng/dsv41-reference.py --output writes, so the two directories diff tensor by
tensor; TS_DSV41_ENGRAM_WARM / _THREADS / _RANDOM,
TS_DSV41_SPARSE_FA and TS_DSV41_COMPACT_RAW_GATHER are native-loader knobs and
are inert here. Image and video do not work on this backend — the vision
companion is a native ggml component and LoadVisionEncoder throws, so the card's
"the vision companion follows the text model onto the same backend" is about
--backend ggml_cpu, not this path — and distributed TP groups, any draft model
or TS_DSV4_DSPARK, TS_DSV41_TP != 0 and TS_DSV41_ENGRAM_DEVICE with any
value but 0 are refused before a weight is read. The advertised 1M context is
not reachable here either: MAX_CONTEXT caps to 65,536 by default and every
compressed-cache row lives in host memory. --backend ggml_cpu and
--backend cpu are different paths — the first is the native graph on scalar
ggml CPU kernels, the second is this managed executor with no ggml at all. →
DeepSeek V4.1 card
Shared direct primitives. The non-ggml execution primitives behind
BackendType.Cuda and BackendType.Cpu were family-neutral, so
WanDirect{Context,Linear,Ops} moved to TensorSharp.Models/Direct/DirectOps.cs
as Direct{Context,Linear,Ops} and are now shared by the Wan video networks and
MiniMax-H3. DirectOps' row loops go through the same worker pool, so the Wan CPU
paths benefit from it too. DirectLinear on CPU also stopped
expanding a quantized weight to F32 at load: it keeps the GGUF storage type and
calls ManagedQuantizedOps.AddmmQuantizedToFloat32. On Wan (256x160x5f, 1 step,
--backend cpu) that is 80.9 s vs 121.4 s at 4x less weight memory, and
against the native ggml_cpu render it scores 43.51 dB where the old path scored
43.39 dB -- marginally closer to native, not a quality regression. F16/BF16/F32
weights keep the plain GEMM; TS_DIRECT_QUANT_WEIGHTS=0 restores the old
behaviour.
MiniMax-H3 on --backend cpu. MiniMax-H3 was the only video-generation model without
a pure C# path -- every stage was an unconditional whole-model native ggml call. It
has one now: MiniMaxH3Direct{DiT, TextEncoder, VideoVae, AudioVae, VisionEncoder, VideoVaeEncoder3D, AudioVaeEncoder}, selected by one predicate
(MiniMaxH3Model.UsesDirectBackend). t2v, i2v, fl2v and reference conditioning
(images, clips, soundtracks) all work there. Wan and DiffusionGemma
already had pure C# CPU paths and did not change in coverage. Qwen-Image-2.1 has one
too now: --backend cpu runs its transformer, text encoder, vision encoder, VAE,
LoRA plug-ins and prefix KV cache in managed code
(card); on the GGML
backends it runs whole GGML graphs.
Parity against the GGML path on identical inputs with a fixed --diffusion-seed
(256x160, 5 frames, 1 step). The control column is GGML measured against
itself -- its own flash kernel against its explicit-softmax fallback
(TS_H3_NO_FLASH=1) -- which is the yardstick for how far two correct
implementations may legitimately drift:
| route | managed vs GGML (cosine) | control (cosine) | render PSNR |
|---|---|---|---|
| text encoder alone (64 layers, 32B Q4_K_M) | 0.99999899 | -- | -- |
| t2v | 0.998740 | 0.997032 | 31.95 dB (control 28.17 dB) |
| i2v (3-D VAE encode) | 0.999897 | -- | 34.87 dB |
| ref-audio (audio VAE encode) | 0.999410 | -- | 35.27 dB |
| ref-image (vision tower) | 0.998554 | 0.999275 | 28.80 dB (control 30.80 dB) |
| vision tower output alone | 0.999919 | 0.999952 | -- |
t2v, i2v and ref-audio agree with GGML more closely than GGML's own two attention kernels agree with each other. The vision tower is the exception: its residual is ~1.4x the control rather than below it. At cosine 0.9999 over 737k elements it is structurally correct, but the residual is unexplained -- disabling flash made agreement slightly worse, so the F16 K/V cast is not the cause. That route is not claimed as parity.
The DiT is not bit-identical to GGML and cannot be: truncating the trunk on both
paths with TS_H3_DIT_LAYERS shows 1-cosine ~1e-5 through 25 layers and then
non-monotonic amplification (1.15e-3 at depth 40, 1.5e-4 at 44, 1.26e-3 at 50) --
and non-monotonic is what rules out a bug.
Speed is mixed rather than uniformly slower, --backend cpu against ggml_cpu on
the same clip: t2v 69 s vs 14 s, i2v 70 s vs 176 s, ref-audio 63 s vs 71 s,
ref-image 112 s vs 20 s. The vision tower is the expensive managed stage.
Diagnostics: TS_H3_DUMP_TE, TS_H3_DUMP_VEL_V, TS_H3_DUMP_VEL_A,
TS_H3_DUMP_VIS and TS_H3_DIT_LAYERS, plus the pre-existing TS_H3_NO_FLASH.
Tensor parallelism (TP) splits a single model across multiple GPUs using the Megatron-LM column/row-parallel pattern. Each transformer block runs column-parallel projections (QKV, gate/up) that split output heads or intermediate dimensions across GPUs, independent per-GPU attention or activation computation, and row-parallel projections (output, down) followed by an AllReduce that reconverges the hidden state. Norms, embeddings, and the LM head are replicated.
Two multi-GPU modes, and they are not the same thing. Tensor parallelism shards the weights inside each layer and pays a collective (one or two AllReduces per layer) to reconverge, so every rank does part of every layer's work and the mode can buy latency as well as capacity. A layer split instead gives each GPU a contiguous run of whole layers: nothing is sharded, no collective is issued, and the GPUs take a token in turn rather than working on it together. A layer split is therefore a capacity feature — it is how a model that does not fit one card runs at all — and should not be expected to raise throughput. Which mode each architecture uses:
| Architecture | Multi-GPU mode |
|---|---|
Mistral 3, Gemma 4, Qwen 3.5/3.6-family, GPT OSS, Nemotron-H, Muse-Glimmer (--tp 2 max) |
Tensor parallelism, opt in with --tp N |
Bonsai2 27B (qwen35 with PRISM Hadamard transforms) |
Single device — the loader rejects --tp |
Qwen-Image-2.1 (qwen_image) |
Tensor parallelism of the diffusion transformer only, opt in with --tp N on ggml_cuda / ggml_vulkan (N must divide its 32 heads); the text encoder, vision encoder and VAE stay on the first GPU, and multi-node groups are refused |
GLM-5.2 (glm-dsa) |
Local whole-layer placement with --layer-split N; --tp N switches it to native local/single-process tensor parallelism |
GLM-5.3 (glm-dsa) |
Local whole-layer placement with --layer-split N; --tp N switches it to native local/single-process tensor parallelism on GGML GPU backends, with the KV and indexer caches replicated per rank and the cross-node --tp-node-id / --tp-peers hard-refused before the model is built. No --tp N > 1 configuration has been measured — the only recorded arithmetic is a non-fit, 41.7 GiB per rank at --tp 8 against 46 GB cards |
GLM-5.3-Flash (glm5next) |
Local whole-layer placement with --layer-split N; --tp N selects native local tensor parallelism on GGML GPU backends |
| DeepSeek V4 Flash | Local whole-layer placement with --layer-split N; --layer-split N only caps how many GPUs the split uses (as TS_DSV4_NGPU does) |
| DeepSeek V4.1 Flash | --layer-split N selects local whole-layer placement. Experimental routed-MoE TP requires --tp N plus matching TS_DSV41_TP=N on GGML CUDA; attention TP and distributed groups are unsupported. The two modes cannot be combined. |
Hunyuan Dense (hunyuan-dense) |
Single device — neither mode; startup says so on stderr |
| DiffusionGemma, Wan 2.1/2.2, MiniMax-H3 | Single device; --tp N > 1, --layer-split N > 1, and distributed groups are refused at startup. |
Qwen 3.8 Flash Next (qwen4exp) |
Layer split, opt in with --layer-split N (GGML CUDA / Vulkan) |
Startup prints which mode actually ran, so it never has to be inferred from
nvidia-smi.
Local TP runs within a single process. On the direct cuda backend one
thread issues commands to all GPUs and CUDA streams provide the parallelism; on
the GGML backends a rank worker pool drives the GPUs concurrently, because a
GGML op submits and synchronizes in one call. Enable with --tp N on either
TensorSharp.Cli or TensorSharp.Server.Host (or TENSORSHARP_TP_DEGREE=N);
TENSORSHARP_TP_DEVICES=0,2 picks which physical GPUs the ranks map to.
Distributed TP extends across machines via a peer-to-peer TCP mesh. Each node
runs its own process with its own local GPUs; AllReduce is hierarchical — local
P2P within each node, TCP across node representatives, then broadcast back — so
only 1/tp_local of the data crosses the network. Enable with --tp-node-id and
--tp-peers (or TENSORSHARP_TP_NODE_ID / TENSORSHARP_TP_PEERS). The server
can join such a cluster as node 0 — the driver that owns sampling and serves
HTTP — with every other node running a TensorSharp.Cli worker.
Architecture-specific strategies handle heterogeneous layers:
| Architecture | Strategy |
|---|---|
| Dense transformer (Mistral 3) | Standard column/row-parallel QKV + FFN |
| MoE (GPT OSS, Nemotron-H) | Expert slicing — each GPU holds 1/tp of every expert's weights; router is replicated |
| MoE on GGML (Qwen 3.5/3.6) | Expert parallelism — whole experts partition across GPUs (128 of 256 per rank), so each rank keeps the single batched ggml_mul_mat_id dispatch per projection; the shared expert stays Megatron-split |
| MoE on GGML (Gemma 4) | Megatron split inside each expert (gate/up column-parallel, down row-parallel) so the fused whole-model MoE trunk kernel keeps working with global expert ids; the expert sum becomes a third row-parallel AllReduce per layer. TS_GEMMA4_TP_FUSED_MOE=0 falls back to the whole-expert per-op path |
| GatedDeltaNet SSM (Qwen 3.5/3.6) | Block-cyclic V-head assignment — each rank runs its own packed GDN kernel on its V-head subset with independent delta/conv state, resident on its GPU; no cross-rank communication for the recurrent path |
| Mamba2 SSM (Nemotron-H) | Replicated on rank 0, result broadcast to all ranks |
| MLA / KDA + sparse-attention MoE on GGML (GLM 5.x; native TP is local/single-process) | GLM-5.2 and GLM-5.3 (both glm-dsa) shard MLA heads and the hidden rows inside every routed expert, and leave router, norms, indexer, shared expert, the dense layers and the embedding replicated; no --tp N > 1 run of GLM-5.3 has been measured. GLM-5.3-Flash additionally head-shards KDA and its per-rank recurrent state, while MLA heads and routed-expert hidden rows are sharded the same way. Attention partials reduce before the nonlinear Sinkhorn hyper-connections. On GLM-5.3-Flash's eligible segmented fast path, routed-MoE partials reduce first, then every rank computes and adds the replicated shared expert locally; hyper-connections, pooled indexer, router, norms, dense layers and embedding also remain unsharded and execute per rank, while output norm / LM head stay on rank 0. CPU MoE, tracing, partial TS_GLM_TP_SHARD, oversubscription, or missing native hyper-connection kernels select the combined scheduler fallback, where the shared expert runs once on rank 0; TS_GLM_TP_FUSED=0 forces that diagnostic fallback. The segmented fast path and TS_GLM_TP_FUSED are gated on the Flash boolean, so neither applies to plain GLM-5.3 |
TP runs on the cuda backend and on the GGML CUDA / Vulkan backends
(ggml_cuda, ggml_vulkan); MLX is single-device. On the GGML backends each
rank owns a ggml backend on its own GPU with its own weight shards and KV
cache, and cross-GPU AllReduce goes through ggml-cuda's collective (NCCL when
available) or a host reduction for small payloads. CUDA graph capture stays on
under TP — a tensor-parallel token is dozens of small per-rank submissions, and
replaying them is worth ~45% of decode throughput (4×A40: Qwen 3.5-9B --tp 4
88 → 128.5 tok/s, Qwen 3.5-35B-A3B --tp 2 71.3 → 104.1, the latter being the
difference between TP losing and winning against a single GPU). Disable with
TS_GGML_TP_CUDA_GRAPHS=0. The collective is chosen by
measurement, not by capability flags: at startup the group verifies that peer
copies between the advertised device pairs actually deliver their bytes and that
a real NCCL AllReduce completes, and it picks the fastest transport that passes.
Hosts that advertise peer access which never arrives (common on virtualized
cloud instances) keep the NCCL collective with peer transport disabled rather
than losing it — which matters past two GPUs, where the pinned-host pipeline
does not apply and the alternative is reducing through host RAM at every layer
boundary (measured on 4×A40: 53.5 → 75.1 tok/s decode on Qwen 3.5-9B Q8_0). GGML TP delivers both
capacity and latency: fused per-rank block graphs (attention, dense FFN, MoE
trunk, GatedDeltaNet) replaced the op-at-a-time forward, so on 2× RTX 2000 Ada
--tp 2 decodes 1.39× a single GPU on Gemma 4 E4B Q8_0 (51.7 vs 37.3 tok/s)
and 1.06× on Qwen 3.5-9B Q8_0, with Gemma 4 output byte-identical to the
single-GPU run. Models that do not fit one card run only under TP:
Qwen 3.5-35B-A3B IQ4_XS (16.6 GB) splits across two 16 GB cards at 184 tok/s
prefill / 18 tok/s decode. Full measurements: TENSOR_PARALLELISM_PLAN.md
(Stages 1b and 1c).
How far that generalizes depends on the interconnect and on how much of a layer
can be split at all. On hosts without NVLink the two AllReduces per layer
dominate: GLM-5.2 UD-IQ2_XXS on 3× RTX PRO 6000 (PCIe) runs pp2048 505.6 /
tg64 17.6 tok/s under --tp 3 against 915.9 / 43.9 on the plain single-node
layer split, and every rank holds a full-length cache, which drops the fitted
context from 342,272 to 91,136 tokens. TP there is a capacity feature
rather than a latency one — and because
it changes the reduction order, a 2-bit MoE reproduces 3 of the 6 recorded
llama.cpp goldens where the layer split reproduces 5 of 6. (Against llama.cpp
running on the same backend, the layer split is 6/6.)
Layer split on Qwen 3.8 Flash Next. --layer-split N on qwen4exp runs a layer
split, not tensor parallelism: none of
its weights are sharded, its decode is one persisted single-device GGML graph
per token, and its GDN/PLE recurrent state lives in device buffers owned by a
single backend. It is also the same (and only) multi-GPU mode llama.cpp offers
this architecture — -sm row refuses to load it.
Measured on 2x A100-80GB, Qwen3.8-Flash-Next-UD-Q2_K_XL (73.4 GiB):
- greedy output is byte-identical between the 1-GPU and the 2-GPU run (same SHA-256);
- VRAM 24.2 GB + 26.2 GB — roughly half the model on each card instead of all of it on one;
- throughput unchanged: prefill ~1520-1550 t/s and decode ~56 t/s either way.
For reference, llama.cpp on the same box: 1 GPU pp1536 1094 / tg128 61.2;
2 GPUs
-sm layer1200 / 61.5 — so llama.cpp also gains ~10% prefill and ~0 decode from the second card.
TS_Q4E_LAYER_SPLIT=20,28 overrides the automatic balance with explicit layer
counts per GPU (llama.cpp's --tensor-split in spirit) and throws rather than
silently ignoring a value it cannot honour — useful because the automatic
balance prices weights and cannot see the vision tower, which loads later and
lands on GPU 0. Details: Qwen 3.8 Flash Next card.
Hunyuan Dense, DiffusionGemma, Wan 2.1/2.2, and MiniMax-H3 are single-device
architectures. Unsupported multi-GPU requests fail at startup; neither
--tp nor --layer-split silently falls back to a single GPU.
Batched/continuous-batching forward under TP is implemented for Mistral 3; the other families, MoE models included, fall back to per-sequence forward under TP.
Local collectives prefer CUDA peer-to-peer DMA, but the group self-tests every
peer-capable device pair at startup and permanently demotes any pair whose
round-trip comes back corrupt (seen on some L4 PCIe topologies), so hosts
without working P2P — A16 vGPU profiles, most consumer cards — fall back to host
staging automatically. Diagnostic overrides: TENSORSHARP_TP_DISABLE_P2P=1
(stage every cross-GPU copy through host memory) and
TENSORSHARP_TP_HOST_ALLREDUCE=1 (run the local AllReduce on the CPU).
Multi-node connect and receive windows are tuned with
TENSORSHARP_TP_CONNECT_TIMEOUT_SECONDS (default 120 s) and
TENSORSHARP_TP_RECV_TIMEOUT_SECONDS (default 300 s).
The server also supports optional Redis-backed shared state: a Redis-backed
Responses API store (--redis-url or TS_RESPONSES_STORE_REDIS_URL) for
durable response storage. --redis-url also sets the KV-cache tier variable
TS_KV_CACHE_REDIS_URL, and --paged-kv-redis-url / --paged-kv-redis-ttl
are accepted (they set the TS_KV_CACHE_REDIS_* variables), but nothing on the
serving path reads them: that Redis KV tier belongs to the
standalone PagedKvCacheManager behind the CLI's --paged-bench, and no
serving path reuses KV through Redis.
Full configuration reference and examples: Usage → Tensor Parallelism & Distributed Inference.
Models can invoke user-defined tools and participate in multi-turn tool-call conversations. Define tools as JSON and pass them via --tools (console) or the tools parameter in the API.
Each architecture uses its own wire format for tool calls:
- Nemotron-H:
<tool_call>{"name": "...", "arguments": {...}}</tool_call> - Qwen 3.5/3.6-family and Qwen 3.8 Flash Next (
qwen4exp): the same<tool_call>block, but with an XML body —<function=NAME><parameter=key>value</parameter></function>(the JSON form is still accepted) - Gemma 4:
<|tool_call>call:function_name{args}<tool_call|> - Muse-Glimmer: ATEM XML —
<atem:function_calls><atem:invoke name="NAME"><atem:parameter name="key">value</atem:parameter></atem:invoke></atem:function_calls>, routed on the assistant'sto=recipient header - GPT OSS (Harmony): tools are declared as a TypeScript namespace in the developer message, and calls are emitted on the commentary channel as
<|channel|>commentary to=functions.NAME <|constrain|>json<|message|>{args}<|call|> - DeepSeek V4: DSML markup — the system prompt teaches the syntax and carries one JSON schema per function, and the model answers with
<|DSML|tool_calls><|DSML|invoke name="NAME"><|DSML|parameter name="key" string="true|false">value</|DSML|parameter></|DSML|invoke></|DSML|tool_calls>.string="false"marks a JSON-typed argument - DeepSeek V4.1: spaced DSML —
<|DSML| calls>and<|DSML| invoke name="tool">, which V4's unspaced form is not compatible with. On/v1/chat/completionsthe declared tools are enforced by a request-local grammar:tool_choiceauto/required/none/ a named function andparallel_tool_calls: falseare all honoured at the token level, strings can be sent losslessly as JSON withstring="false", and unsupported schema assertions are refused with HTTP 400 before generation rather than silently dropped - GLM 5.x: XML with per-argument tags —
<tool_call>NAME<arg_key>k</arg_key><arg_value>v</arg_value>...</tool_call>; the function name follows the opening tag directly as bare text, and every argument is its own<arg_key>/<arg_value>pair (values the template rendered withtojsonare parsed back into numbers, arrays and objects)
The output parser (OutputParser.cs) automatically extracts tool calls from the model's raw output regardless of architecture. Mistral 3, Hunyuan Dense and DiffusionGemma render no tool declarations; DiffusionGemma refuses a request that carries tools with HTTP 400.
A skill is a folder holding a SKILL.md — YAML frontmatter plus Markdown instructions written for the model — together with the scripts, reference documents and assets those instructions refer to. TensorSharp scans one or more skill directories (--skills-dir, repeatable, or TS_SKILLS_DIR; without either, every existing .agents/skills folder from the working directory up to the nearest Git root — nearest first, or only the working directory outside a repository — plus a skills folder next to the binary, which is created if missing. On the CLI that folder comes last; on the server it is also the upload directory, is always scanned first (so it wins a name clash) and is kept even with an explicit --skills-dir / TS_SKILLS_DIR, which replace only the .agents/skills defaults. Personal or global skill folders are never imported), advertises each skill's one-line description to the model, and loads the rest only when the model asks for it.
That on-demand part is served by built-in tools that TensorSharp executes itself, in process:
skills_list()— every skill reachable in this conversation, with its description and its bundled file pathsskills_read(skill, path, offset)— one page of one file;path="SKILL.md"is the skill's own instructionsskills_run(skill, path, args)— run a bundled script. Off by default;--skills-allow-execenables it, and it then runs sandboxed or not at all (--skills-sandbox requiredis the default)
Answering those calls inside the engine is what makes the feature work for clients that know nothing about skills: an ordinary OpenAI client sends "skills": ["pdf"] and gets back a finished completion, never a tool call it has no implementation for. The caller's own tools are never executed — they are returned to the caller as usual.
Progressive disclosure. On families that can render tool declarations and parse tool output, the prompt contains metadata only — name and description — even for explicitly selected skills. Selection scopes reach and preference; the model activates a skill by calling skills_read(skill, "SKILL.md"), at which point the file index is returned too. The metadata budget is about 2% of context, clamped to approximately 1,024–10,000 tokens. Bundled files are never inlined; contents come from skills_read in 48 KB pages. The older body budget (one quarter of context, clamped to approximately 1,024–48,000 tokens) is used only by the fallback for a family that cannot complete a tool round trip.
Prompt shape. The block is merged into the leading system/developer message rather than appended as a second one, which is the only injection point every chat template in the repository handles. Its bytes are a pure function of the sorted skill selection — no timestamps, paths or counters — so a conversation re-hashes identically turn to turn and the KV prefix cache keeps matching from block 0.
Confinement. Every path the model names is resolved through SkillPathGuard, which closes lexical (.., absolute, ~, UNC, drive-qualified), canonical and symlink escapes, and confines each skill to its own directory. ZIP installs run every entry through the same guard (zip-slip), enforce size on the decompressed stream, and cap per-entry (64 MB), per-archive (256 MB), entry-count (4096) and compression ratio (200×).
Script sandbox. When script execution is enabled, the child runs under sandbox-exec on macOS and bwrap 0.12.0+ on Linux — network denied by default, the user's home directory unreadable, writes confined to the working directory (plus, on macOS, the shared /private/tmp and per-user temp directories — the same exception the code-execution bullet documents) — plus, on every platform, an interpreter allow-list, no shell, a scrubbed environment that withholds host credentials, a time limit and an output cap. That directory is the shared request/chat workspace with --code-exec, otherwise per-call scratch; the skill itself stays read-only. --skills-allow-network is the separate opt-in for these bundled scripts; it neither enables nor is enabled by --code-exec-allow-network. Windows bounds the process tree through a job object but cannot confine the filesystem or network, and says so: every result names what was not enforced. --skills-sandbox required (the default) refuses to run scripts on a host that cannot confine them; Windows therefore needs an explicit --skills-sandbox preferred opt-in for its weaker isolation.
Model families. Mistral 3, Hunyuan Dense and DiffusionGemma render no tool declarations, and a family with no registered output parser cannot complete a structured tool round trip. TensorSharp therefore withholds skill/code tools, inlines selected skill bodies, and drops the discovery catalog for these families. Qwen 3.8 Flash Next (qwen4exp) parses Qwen's XML/JSON tool calls and takes part in the full skills, code-tool and sub-agent loop. Mistral 3 and Hunyuan Dense also drop role: "tool" messages, so any loop result is fed back as a user turn. Tool support requires both a declaration renderer and an output parser; model reasoning support alone is not enough.
Structured output. A request using JSON mode or a JSON schema suppresses the built-in tool loop and inlines selected skill instructions so that schema-constrained output is not interrupted by an internal tool call.
Browser automation. TensorAgent/skills/playwright is an ordinary skill bundle, not a runtime feature: its script runs a pinned @playwright/cli through npx, and the model drives a real browser — navigation, forms, snapshots, screenshots, data extraction, and a handoff to the user for a login or CAPTCHA — entirely through skills_run. There is no screen or mouse tool. It needs a skills root that contains it, --skills-allow-exec and --skills-allow-network, real Node.js / npm / npx on the host, and in practice --code-exec with --code-exec-allow-install and --code-exec-allow-network; on macOS the model writes a workspace .playwright/cli.config.json that turns off Chromium's inner sandbox, because a process already confined by Seatbelt cannot start a second Seatbelt sandbox (the outer one stays on). It has been validated on macOS arm64 only. The skill is left out of TensorAgent's iOS app bundle, since iOS starts no child processes; it stays in TensorAgent/skills for the desktop hosts. → Running browser and native-runtime skills
Selecting skills:
# CLI
dotnet TensorSharp.Cli/bin/TensorSharp.Cli.dll --model models/gemma-4-E4B-it-Q8_0.gguf \
--backend ggml_metal --skills-dir ~/skills --skill pdf --input prompt.txt
# Any chat API — /v1/chat/completions, /v1/responses, /api/chat/ollama (Ollama), /api/chat (Web UI)
curl -X POST http://localhost:5000/v1/chat/completions -H "Content-Type: application/json" \
-d '{"model": "gemma-4-E4B-it-Q8_0.gguf",
"messages": [{"role": "user", "content": "Pull the totals table out of this statement."}],
"skills": ["pdf"], "skills_discovery": false}'The server also exposes the registry itself — GET /v1/skills, GET /api/skills, POST /api/skills (upload a .zip), DELETE /api/skills/{name} — and /api/models reports a skills block (enabled, installable, allowScripts, count) so a UI knows whether to offer the control and whether script execution is available.
Full reference, including the frontmatter fields, the budget, the security model and the C# SkillsChatClient API: Agent Skills in TensorSharp. Open-source skills to start from: https://github.com/anthropics/skills.
A model that can call tools may split a request across bounded child agents. The host adds five coordination tools — spawn_agent(task_name, task, agent_type?), wait_agent(agent_id?, timeout_ms?), send_input(agent_id, message), close_agent(agent_id) and list_agents() — and a short coordination block merged into the leading system message. The model decides whether a substantial, independent piece of work is worth delegating; nothing forces a spawn. Every child runs on the same loaded model with the parent's generation settings, starts from the parent's system/developer instructions plus a self-contained task (never the parent's transcript or attachments), and returns a bounded report. The host collects any outstanding reports before the parent gives its final answer.
- Where it runs. On by default on
/v1/chat/completions,/v1/responses,/api/chat/ollamaand the Web UI's/api/chat, and in TensorAgent, where a Settings switch ("Sub-agents", on by default) turns it off from the next message and the page shows no agent panel. C#SkillsChatClientlocal delivery delegates too (SkillsChatRequest.MultiAgent = falseopts a request out). The CLI and Ollama/api/generatehave no sub-agents. Delegation does not depend on skills or--code-exec. - Which models. Any family that renders tool declarations and has a tool-call parser; there is no model-size gate. Mistral 3, Hunyuan Dense and DiffusionGemma are excluded, and a JSON-mode or JSON-schema request runs without the coordination tools.
- Permissions.
explorer(the default role) andreviewerchildren are always read-only:skills_list,skills_read, andread_filewhen the parent offers it. Aworkeris read-only too unless the server starts with--agents-allow-worker-tools; it then shares the parent's workspace and sandbox, with no separate worktree and no merging. Client-owned tools are never passed to children, and host tool calls within one request tree run one at a time while their generations overlap. - Limits.
--agents-max-concurrent(default 3),--agents-max-count(8),--agents-max-depth(2),--agents-max-rounds(8),--agents-max-generations(48),--agents-timeout(180 s) and--agents-max-result-chars(8000).--no-multi-agentorTS_NO_MULTI_AGENTturns delegation off on the server. A request can opt out with"multi_agent": false, but it cannot turn delegation on, raise a limit or enable worker tools. - KV reuse. Each child owns its conversation state and copies KV rather than sharing pages with the parent. Before the parent prefills, the host renders every child role/tool profile and checkpoints the prefix they share with the parent, so a child restores that prefix instead of prefilling it; this matters for the Qwen 3.5-family recurrent models, where only a checkpoint at the exact boundary can be reused. The scheduler holds a cold sibling until the producer has published its checkpoint.
- Visibility. Only the parent's text is streamed. While
wait_agentruns, the Web UI shows a live card per sub-agent, fed byagentsin the SSEtool_progressframes (skill_stepframes carryagent_id), and usage totals include the children.
No latency, quality or delegation-rate numbers are published for this feature, and enabling it does not guarantee a faster or better answer. Full reference: Multiple agents.
Gemma 4 models support image, video, and audio inputs. For the E4B example above, pass the repository's mmproj-gemma-4-E4B-it-Q8_0.gguf explicitly with --mmproj (use the projector matching other target sizes).
- Images: PNG, JPEG, HEIC/HEIF
- Video: MP4 (time-based extraction at 1 fps using OpenCV; tune with
VIDEO_SAMPLE_FPS/VIDEO_MAX_FRAMES) - Audio: WAV (16kHz mono), MP3, OGG Vorbis
DiffusionGemma accepts images through the same Gemma 4 vision tower (gemma4v).
No published DiffusionGemma GGUF carries it, so --mmproj takes either an
mmproj GGUF or the raw upstream shard model-00011-of-00011.safetensors from
google/diffusiongemma-26B-A4B-it,
which TensorSharp reads directly. Because the prompt is re-run for every block
(and, on backends without prompt-KV caching, for every denoising step), the
image embeddings are kept for the whole turn rather than consumed once, and
concurrent image requests keep their own. Images work in Web UI and API chat,
with the CLI's --image, and alongside the Jev state at /v1/systemone. The
checkpoint has no audio tower, so audio is refused with HTTP 400. There is no
video path: an OpenAI video_url part is refused, and a video uploaded in the
Web UI reaches the model only as its extracted frames, each rendered as a plain
image.
- Images: PNG, JPEG, HEIC/HEIF
All Qwen 3.5/3.6-family variants (qwen35, qwen35moe, and qwen3next) load through the same Qwen35Model implementation. Image inputs are supported via the dynamic-resolution Qwen35VisionEncoder; pass the selected repository's projector explicitly (for the 9B and Qwen 3.6 examples, mmproj-F16.gguf; Bonsai2 27B has its own Ternary-Bonsai-2-27B-mmproj-BF16.gguf / -Q8_0.gguf). The MoE variants (e.g. Qwen3.5-35B-A3B and Qwen3.6-35B-A3B GGUFs that report the same architecture keys) additionally enable a fused MoEExpertsSwiGLUResidual GGML kernel during decode that runs all selected experts, the optional shared expert, and the residual add in a single GPU graph dispatch.
Qwen3.8-Flash-Next (qwen4exp) supports image input through the Qwen3.5-VL
vision tower with (T, H, W) IMRoPE positions; put the repository's
mmproj-BF16.gguf beside the model to enable it. Multi-image prompts and
multi-turn image sessions both work, with KV reuse across turns — extend-only,
because the GatedDeltaNet recurrence cannot rewind, so a cached prefix is
reused only when the new prompt extends it exactly.
Video input arrives as an OpenAI video_url content part (base64 MP4 / WebM /
MOV data URI, fps and max_frames in the part, VIDEO_SAMPLE_FPS /
VIDEO_MAX_FRAMES as the defaults). The clip is sampled into ordered, timed
frames that the prompt renders in the Qwen3-VL video layout — one
<|video_pad|> block per pair of consecutive frames, each labelled
<t seconds> with the pair's mean source time, wrapped once per clip in the
template's <|vision_start|> … <|vision_end|> — and the tower encodes each
pair through its two temporal patch-embedding slices (v.patch_embd.weight /
.weight.1), which is the temporal merging the Qwen-VL processor does. Every
pair is positioned like a still image at the running position of that pair, so
consecutive pairs carry increasing temporal M-RoPE ids and the QSA position
history, MTP catch-up and the post-clip cache gap all see the same coordinates.
Frames are still sampled rather than decoded by a temporal encoder, and an odd
clip repeats its last frame to complete the final pair.
- Images: PNG, JPEG, HEIC/HEIF
- Video: MP4, WebM, MOV through
video_url(time-based sampling; the Qwen3-VL per-clip pixel budget scales long clips down as a whole)
GLM-5.3-Flash (glm5next) supports image input through the GLM-OCR ViT in the
repository's mmproj-BF16.gguf — RMS norms, fused QKV, per-head q/k RMS norms,
2D vision RoPE, a SwiGLU-clamp MLP and a 2x2 conv merger, with all 24 blocks
running as one device-resident GGML graph. The projected embeddings override
the <|image|> placeholder rows inside the native executor; the text tower is
NoPE, so image tokens need no MRoPE bookkeeping. --image, multi-image prompts
and multi-turn image sessions are all supported. GLM-5.2 and GLM-5.3 (both
glm-dsa) are text-only, and GLM-5.3 is text-only twice over: the published
unsloth/GLM-5.3-GGUF ships no
mmproj at any quant, and LoadVisionEncoder warns and ignores an --mmproj
on glm-dsa rather than failing, so passing one yields a text-only run instead
of an error.
- Images: PNG, JPEG, HEIC/HEIF
Mistral 3 supports image inputs via the Pixtral vision encoder. The example repository uses mmproj-mistralai_Mistral-Small-3.1-24B-Instruct-2503-f16.gguf; pass it explicitly with --mmproj.
- Images: PNG, JPEG, HEIC/HEIF
Muse-Glimmer-30B supports image inputs through a 50-layer sparse-window ViT with
2D RoPE and a 2x2 pixel shuffle. Pass the companion projector explicitly with
--mmproj (e.g. mmproj-Muse-Glimmer-30B-Q8_0.gguf). The chat template renders
an image content part as a single <|patch|>, which the multimodal injector
expands to <|image_start|> + N merged-patch rows + <|image_end|> — up to 4096
merged tokens per image, chosen by an aspect-preserving stretch (no tiling, no
padding).
- Images: PNG, JPEG, HEIC/HEIF
The Nemotron Omni distribution adds a RADIO / v2_vl ViT image encoder. Pass the matching multimodal projector with --mmproj (e.g. nvidia_Nemotron-H-Omni-mmproj.gguf); the language-model GGUF stays the same. Image tokens are inserted at <image> placeholders and expanded into <img> + N tile tokens + </img> automatically by the multimodal injector. Audio is refused (400 / CLI error) unless an audio companion GGUF is loaded: the public Omni mmproj carries only the vision tower, so there is nothing to fill the <so_embedding> placeholder with (see nemotron.md §4.6-4.7).
- Images: PNG, JPEG, HEIC/HEIF
- Audio: with only the public files, refused at every entry point (HTTP 400 from
/v1/chat/completions,/v1/responsesand the Web UI;--audio//audiodeclined by the CLI) withNemotronModel.AudioInputUnsupportedMessage: the public GGUFs ship no Parakeet/FastConformer audio tower to fill the<so_embedding>placeholder.NemotronAudioEncoderruns that tower (managed CPU) from a companion GGUF with NVIDIA'ssound_encoder.*/sound_projection.*tensors, loaded through--mmprojorTS_NEMOTRON_AUDIO_MMPROJ; only a companion whose tensors validate lifts the refusal. Trained BF16-compute parity, GPU execution and spoken-audio quality are not yet qualified (see nemotron.md §4.7).