Repository navigation
Integrate unified memory budgets, KV paging and Qwen/Gemma weight streaming - #262
Draft
zhongkaifu wants to merge 42 commits into
Draft
zhongkaifu wants to merge 42 commits into
zhongkaifu wants to merge 42 commits into
Conversation
Engine comparison — TensorSharp vs llama.cpp (PR smoke)No report artifact was produced — the benchmark failed before generating results (see the workflow logs). |
Prefer qualified resident execution with hardware-aware admission, account owned graph buffers, and overlap bounded weight reads. Preserve token suppression across generation paths and protect checkpoint publication during retained-state eviction. Invalidate file-read visibility when history is compacted. Validate unchanged ggml ffa4e8b with CPU regressions, real CUDA lifecycle and streaming checks, and alternating E4B/Qwen comparisons. Keep Gemma IQ2 repetition, Flash numerical differences, absolute Qwen decode speed and unavailable multi-GPU scenarios explicitly unresolved.
Add bounded expert file staging, adaptive read/upload prefetch, decaying frequency eviction and shared host staging credit without modifying ggml. Handle published GLM architecture keys and explicit native offload policy; omit unused Qwen NextN payload when the speculation policy excludes it. Avoid redundant single-token Qwen CUDA copies and materialized broadcasts. Account owned Qwen graph arenas, recurrent state and device snapshots in shared graph budgets, including refusal, retry and physical-release handling. Add reusable diagnostics, benchmark verifiers and lifecycle regression tests. Record real-model and Windows/Linux evidence, performance gaps and incomplete resource coverage in the unified-memory design and model documentation. Generated evidence stays ignored under artifacts; upstream ggml is unchanged. Validation on unchanged ggml ffa4e8b8: - Windows/Linux managed CUDA: 18 unique / 21 passes; CPU policy: 150 each. - Native: Windows 85 pass + 1 skip; Linux 84 pass + 6 skip + 2 failures. Preserve upstream attention-oracle failure and default NCCL gather timeout; gather passes separately with NCCL_P2P_DISABLE=1 on this VM. - Changed Python tools: 85 pass + 1 platform skip. Historical full discovery remains non-green (missing evidence fixtures, CRLF pins, platform assumptions). - Real Qwen IQ1_M: 34 full-vocabulary rows bit-identical in four checked arms. Interleaved application decode: 55.47 -> 56.65 tokens/s; simultaneous llama.cpp 59.95. Both retain the same 3/15 strict raw-JSON format failures; TS tool round trips and schema checks pass.
… shared budgets Charge actual GGML host pool and weight-buffer allocations until physical release, derive dense request growth/workspace peaks, and consume admitted envelopes in the qualified serial execution lane. Record exact ledger high-water marks and preserve failed cleanup ownership. Keep paged-attention CUDA cache payload out of Windows thread destructors, connect its graph buffers to budget allocation/free, and release retired caches during quiescent trim. All native changes are TensorSharp-owned; ggml remains unchanged. Add fresh-process long-context and delegated-cgroup validation tools. Document 8K/32K/64K parity, measured performance variation, accounting gaps and unavailable physical hard-cap enforcement. Windows/Linux related regressions: 398 passed and 5 skipped each, plus final allocation/worker-exit checks; Linux memory harness: 54/54. No all-model or default-on readiness claim.
Qualify omitted-draft Qwen trunks, forecast retained decode graphs, and expose bounded allocation-refusal diagnostics. Move host observations into the memory layer, add Apple providers, and correct MLX wired headroom. Record local CUDA/CPU model validation and the remaining platform and physical-limit gaps.
Integrate main 0ec3277 while preserving budgeted KV snapshot leases, physical accounting, model ownership, and reply-budget fallback. Keep recurrent short payloads distinct from tiered page capacity and preserve exact media measurement before relaxing the reply reserve. Retain failed model cleanup for retry, trim borrowed GGML pools, and fix the admission lost-wakeup race when capacity shrinks during reclamation. Use sparse Windows catalog fixtures and capability-gated symbolic-link coverage. Validation: Release solution build; InferenceWeb 8893 passed / 410 skipped; TensorAgent 1278 passed / 181 skipped plus 225 passed / 2 skipped on latest Chat dependencies; portable memory 58 passed / 1 skipped; 4 native CUDA and 4 CPU copy-fault tests. Qwen 0.8B 32K dual-request and Gemma E4B shared-budget outputs match ordinary mode and pre-merge tokens; image follow-up and real Gemma disposal/reload pass. Unmodified ggml ffa4e8b80930029a35991f94e7c8a93cd67730ab. Changed Python modules: 37 passed / 1 skipped. Full Python suite retains pre-existing evidence, source-pin, and Windows environment failures; attribution and generated validation evidence remain ignored under artifacts/merge-main-20261010. Apple and multi-GPU hardware were not available.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
TensorSharp's request admission, weight caches and KV snapshots previously used separate limits, while forced streaming made models that fit in VRAM unnecessarily slow. This draft connects allocation ownership and hardware/request-aware dense-model placement, repairs cache and agent-history lifetime failures, and adds measured Q8 decode/prefill experiments. Broad quality and independent-baseline performance targets remain unfinished.
Implemented behavior:
AdaptiveModelSessionprefers resident dense Gemma4/Qwen3.5 CUDA execution, then a smaller prefill chunk, then a supported bounded file adapter. Physical availability, context/state, workspace and operator ceilings inform admission; boundary refresh preserves live owners.TS_GGML_Q8_PARALLEL_VECTOR=1,TS_GGML_Q8_PARALLEL_SMALL_BATCH=1); wider-model numerical failures prevent default promotion.ffa4e8b80930029a35991f94e7c8a93cd67730ab.Actual local evidence on RTX 3080 Laptop 16 GiB; artifact identities and execution conditions remain separate:
e0bfdff5ARM64 found one missing iOS Release export-retention entry (7426 passed, 82 skipped); fixed in5d38902e, local project/export checks 21/21. Linux x64/ARM64 both passed at093b2fac, including host-cache and ordered-vector changes (run). At65ec9e04, ARM64 passed; x64 had one reproducible completion/publication race (8094 passed, 82 skipped, 1 failed), fixed incc92d66e.f0baf26aPR Unit Tests passed (run), including the publication fix. New automatic-policy CI at5bce5840is pending. GPU CI remains unexecuted4ae03489…a76d9b93…: 12/12 tests, including exact minimum/one-byte refusal, injected physical-availability limits, rollback, LRU order and queued-copy retirement. Ten final-version model processes; 14 matched-history comparisons, 366 complete vocabulary rows byte-identical. Two synthetic layouts plus local trained Flash9be5eabf…--checkpasses3c6f0fb7…4ebdf2c7…, same checkpoint, 258 prompt IDs, context512, F16 KV, all 66 layers on CUDA, 16 fixed-history full vocabularies. Old/new common-offset relative L2 medians 0.007010/0.004672, maxima 0.026925/0.017443. Argmax 16/16 for both, but both fail 0.001; some individual rows worsen. Preserving stored weights is not full-model numerical qualificationeb3781a9…, unchanged native9be5eabf…: 134 related CPU, 25 CUDA, 2 comparison-tool tests pass without skips. Qwen 32 full vocabularies pass original numerical gates (max relative L2 0.0008168431); Gemma 8 rows are byte-identical to resident. Actual cache hits, read-ahead failure draining, trim/reload, pressure/reset and final zero-owner checks pass. New kernels, language quality and multimodal scheduling are not claimed74c1b4f8…c2404c9e…: 13 native and 129 related managed checks pass, no skips; CUDA memcheck 0 errors. Qwen 32 full vocabularies pass original gates (max relative L2 0.0008168431); Gemma E4B 8 rows are byte-identical to resident. Actual device retention, pressure/reset and zero final owners verified. Combined device payload peaks 135,736,576 / 251,354,112 B within 256 MiB ceilingse1162225…, unchanged native74c1b4f8…: 74 related tests pass with no skips, including changed inputs/weights, incompatible shapes, shared-owner pressure and failed create/upload/free recovery. Qwen 32 full vocabulary rows pass original gates; Gemma E4B 8 rows are byte-exact to resident. Both pass pressure/reset and zero-owner checks. Combined device payload peaks 138,409,472 / 265,132,800 B under 256 MiB ceilings55529501…, unchanged native74c1b4f8…: 78 related managed/CUDA tests pass, no skips, including budget competition, request aging and failed-release retry. 27 relevant comparator tests pass. Qwen 32 full-vocabulary rows pass original gates (max relative L2 0.0008168431); Gemma E4B 8 rows remain byte-exact to resident. Pressure/reset and final zero owners verifiedWarmups are excluded where specified; hashing warms the file cache. Prefill remains substantially below the independent baseline. A separate bounded Q8-to-F32/pedantic-SGEMM research target found original absolute-gate failures in both candidate tiles and the old serial control. Failures are retained without gate relaxation or production promotion.
The follow-up chunked-K research uses short pedantic F32 GEMMs and budgeted FP64 accumulation tiles. Chunk 128/64 clears the original 527-sample gate on one large synthetic geometry but is slower; chunk 256 still has failing tiles. This GEMM experiment is not in production dispatch. The small suite has 30 shape/admission cases and 201 admitted tiles; CUDA memcheck reports zero errors/leaked bytes. CLI/reporting revisions retain distinct executable identities, and a benchmark with zero admitted candidates now fails instead of counting as a pass. A subsequent Tensor Core component expansion also remained slower; chunk256 failed the absolute gate. Rejected code/binaries and failed outcomes remain ignored artifacts. These are exploratory projection results, not model-prefill qualification.
Semantic validation does not treat existing engines as ground truth:
Remaining limits: hooks do not cap total RSS/VRAM or every backend/driver allocation. Wider-model and Gemma multi-token device-weight retention, async H2D/D2H, global expert-cache fairness/regrowth and automatic pressure scheduling, all-model/media integration, production concurrency SLOs and streamed multi-GPU coverage remain incomplete. Slot reduction is geometric, not globally optimal; explicit cache ceilings and the native physical headroom floor remain. Current SSH to the supplied VM is refused. Historical A40/P2P results are separate; unavailable/skipped scenarios never count as passes. This PR remains a draft.
Generated evidence stays ignored under
artifacts/anddocs/validation/; reusable tools/fixtures are checked in. Design and evidence · Adaptive comparison · Batch probe