Skip to content

Integrate unified memory budgets, KV paging and Qwen/Gemma weight streaming - #262

Draft
zhongkaifu wants to merge 42 commits into
mainfrom
codex/unified-memory-scheduler
Draft

zhongkaifu wants to merge 42 commits into
mainfrom
codex/unified-memory-scheduler

Conversation

@zhongkaifu

@zhongkaifu zhongkaifu commented Oct 8, 2026 •

Copy link
Copy Markdown
Owner

TensorSharp's request admission, weight caches and KV snapshots previously used separate limits, while forced streaming made models that fit in VRAM unnecessarily slow. This draft connects allocation ownership and hardware/request-aware dense-model placement, repairs cache and agent-history lifetime failures, and adds measured Q8 decode/prefill experiments. Broad quality and independent-baseline performance targets remain unfinished.

Implemented behavior:

  • Atomic multi-pool budgets, versioned leases, bounded transfers, verified file spill, request peak admission and bounded engine KV snapshots, including Qwen recurrent state.
  • AdaptiveModelSession prefers resident dense Gemma4/Qwen3.5 CUDA execution, then a smaller prefill chunk, then a supported bounded file adapter. Physical availability, context/state, workspace and operator ceilings inform admission; boundary refresh preserves live owners.
  • Native cache/selected graph ownership, retryable cleanup, bounded staging/read-ahead, Gemma cache/refill corrections and model-defined suppression. Empty vocabulary-contract rebinding preserves pending device tokens.
  • Radix promotion retains node lifetime through retries. File deduplication requires visible content after compaction. Reply reservation adapts to protected instructions/task evidence while preserving the hard context limit.
  • Streamed/deferred MoE weight leaves are marked as external lifetime inputs/outputs by TensorSharp, preventing allocator reuse while a fused consumer still needs them.
  • Selected-expert admission tries smaller measured slot tables when shared credit or current physical availability cannot fit the initial table. Warm weights retain their reservation. A quiescent, process-wide LRU trim drains queued copies and physically frees graphs before refunding shared credit; model/KV state and counters remain intact.
  • Q8/F32 prefill automatically chooses 32/64/128 columns from device SM count, actual kernel occupancy, matrix geometry and padding. It preserves K-increasing FMA order and adds no global workspace. Explicit tile overrides remain. Parallel-K decode and small-batch reductions remain opt-in (TS_GGML_Q8_PARALLEL_VECTOR=1, TS_GGML_Q8_PARALLEL_SMALL_BATCH=1); wider-model numerical failures prevent default promotion.
  • Mixed Gate/Up formats stay in their original bytes on single-rank GGML CPU/CUDA families with a split FFN. Matching formats still fuse. The Gemma/Qwen single-CUDA budget no longer forecasts lossy requantization copies/scratch. Tensor-parallel and other backend behavior is unchanged.
  • Optional immutable file-weight ranges now reuse shared-budget RAM across tokens. The adaptive loader derives capacity from physical availability and the execution-phase peak, releases loading-only holdout, and trims before shrinking capacity would starve the next workspace. Scan misses do not churn the retained working set; raw bytes are copied into existing staging before device use. Explicit streaming keeps its default cache disabled. No graph pointer or weight arithmetic changes.
  • Default single-token Q8/F32 projection now assigns one ordered output row per lane and uses aligned two-byte packed reads. It removes shared transposes/barriers with no extra global workspace, preserving the complete increasing-K FMA sequence. The new strict native comparison checks completed, isolated ABBA processes and every warmup/measured logit hash.
  • Dense file adapters now optionally retain complete immutable CUDA weight/input/output/scratch arenas under the same shared budget. Hardware/request slack and user ceilings bound admission; temporary workspaces evict LRU weights before shrinking tiles. Qwen FullPrecision reuses prefill/decode arenas; Gemma retains N=1 decode only. Normal KV reset preserves valid weights, failed release preserves quota, and FullPrecision promotion removes duplicate RAM ranges. Explicit streaming defaults to device cache off.
  • Normal request completion/token-channel closure now follows finished-state retention/release handoff. This fixes a reproducible checkpoint-publication race under retained-state pressure; failed physical releases still keep ownership/quota for retry.
  • The adaptive entry now automatically reuses idle row-tiled CUDA sessions while replacing inputs and weights; explicit zero disables it. Rank/type/K and capacity must match; ResidentCuda also matches logical N/M. Admission uses at most one quarter of discretionary device slack after other owners and execution headroom. Weight admission leaves that share available without charging idle workspace twice. Up to 16 healthy sessions remain; request reset physically retires unused entries and shrinks to current slack. Failed cleanup preserves credit and blocks reuse until reset succeeds. Complete-matrix multi-token Gemma is excluded. Low-level manual streaming options retain their zero-cache defaults.
  • All native changes are TensorSharp-owned. Upstream ggml is unchanged at ffa4e8b80930029a35991f94e7c8a93cd67730ab.

Actual local evidence on RTX 3080 Laptop 16 GiB; artifact identities and execution conditions remain separate:

Scope Result and boundary
CPU / export regression Full Windows lane: 8084 passed, 76 skipped, 1 unchanged symlink-fixture privilege failure. e0bfdff5 ARM64 found one missing iOS Release export-retention entry (7426 passed, 82 skipped); fixed in 5d38902e, local project/export checks 21/21. Linux x64/ARM64 both passed at 093b2fac, including host-cache and ordered-vector changes (run). At 65ec9e04, ARM64 passed; x64 had one reproducible completion/publication race (8094 passed, 82 skipped, 1 failed), fixed in cc92d66e. f0baf26a PR Unit Tests passed (run), including the publication fix. New automatic-policy CI at 5bce5840 is pending. GPU CI remains unexecuted
Bug regressions Pending-device-token/sampling/context 90/90. Old radix coordinator fails all 5 negative controls, repaired one passes. Real reply-budget preparation old 2/2 fail, repaired 2/2 pass
Native 4ae03489… Five CUDA Q8 suites pass; CPU 96 cases twice; two FullPrecision streaming entry tests pass. Small-batch suite covers 184 cases twice, independent FP64, strides, guards, graph reuse and underflow
Opt-in Qwen parallel vector Fixed-history full-vocabulary numerical gate passes on 0.8B. Separate ABBA decode 20.05→186.01 tokens/s; prefill 2210.58→2106.53, regression retained. Independent llama: 183.98 decode / 8792.20 prefill. Wider-model failures keep this opt-in
Qwen small batch Context 8192, prompts 1024/2048/4096/6000, widths 2/3/4, 32 steps: matched-state logits identical, greedy continuations match. Separate four-process ABBA: batch 2 3686.17→658.82 ms (5.60x), batch 4 3853.44→1228.56 ms (3.14x), six samples/arm/width. Decode microbenchmark, not an end-to-end or independent-engine batch result
Qwen prefill tiles Six balanced fresh processes, six samples/tile: prefill 1988.08/2346.67/2549.06 tokens/s for 32/64/128, complete logits byte-identical. Decode 178.61/175.78/171.85 with overlapping ranges. No concurrent inference/build; no clock lock and prior image workload heated the device
Budgeted file refill Original logits gates pass; 64 MiB host/device ceilings, staging peaks 18,350,080/8,519,680 bytes, 256 actual read-ahead operations, pressure/reset recovery and zero final owners. Prior qualified managed binaries with new native, explicitly recorded
MoE repair Original Flash expert geometry, N=1/8/9/38: streamed/resident outputs byte-identical after repair; old native negative control fails. Full-model Flash CPU/GPU N=38 still fails separately
Expert pressure/reclamation Native a76d9b93…: 12/12 tests, including exact minimum/one-byte refusal, injected physical-availability limits, rollback, LRU order and queued-copy retirement. Ten final-version model processes; 14 matched-history comparisons, 366 complete vocabulary rows byte-identical. Two synthetic layouts plus local trained Flash
Original-native negative controls Same managed binaries and budgets: synthetic cache coverage 756/864→864/864 at 26,000,000 bytes; trained Flash 1692/1728→1728/1728 at 5,630,000,000 bytes. Flash's last layer shrinks 24→18 slots, observed shared peak 5,629,208,012 bytes. Reclaim/refill of 2,101,714,944 expert-cache bytes also passes with zero final owners. Captured short teacher forcing, not throughput or language-quality qualification
Native 9be5eabf… Q8 CPU/CUDA tests 7/7 and resident/streaming tests 4/4 pass, no skips. Automatic policy covers grid saturation and padding; complete matrices match the previous dispatch bytewise and retain independent FP64 checks. Rebuilt standalone prefill research --check passes
Default ordered Q8 3c6f0fb7… 12/12 native tests, no skips; production CUDA memcheck reports zero errors. Independent production FP64 gates unchanged. Synthetic ordered research: 60 shapes/stride layouts × 10 routes byte-exact, zero memcheck errors; its separate oracle uses a sequential-FMA forward-error bound, not the stricter production absolute-error gate
Default Qwen0.8B ABBA Six measurements/arm, complete warmup/measured logits and history byte-identical: decode 19.513→84.985 tokens/s (4.355x), prefill 2783.714→2804.476 with overlapping ranges. Still below independent llama reference 183.98/8792.20. No additional global kernel workspace
Default Qwen27B ABBA Local UD-IQ4_XS, prompt258/context512/16 fixed-history predictions: full capture byte-identical. Separate capture-disabled ABBA: decode 7.503→14.749 tokens/s (1.966x), prefill 397.602→391.744 with overlapping ranges. Six measurements/arm, all full logits/history identical. Existing independent-engine numerical/semantic gaps remain
Final streamed Qwen New native plus identical final host-cache managed assemblies: 32 full-vocabulary rows pass original gates, maximum relative L2 0.0008168431. 128 MiB host/64 MiB device/64 MiB cache ceilings, 67,108,160 B actual retention, pressure/reset and zero final owners pass. Correctness/budget run, not throughput
Automatic Q8 prefill Two independent four-process ABBA comparisons, six measurements/arm, Qwen0.8B Q8 / 643 prompt tokens / context2048 / 64 predictions. Default decode arithmetic: prefill 2106.63→2791.37 tokens/s (+32.50%), decode 19.519→19.442. With parallel decode explicitly enabled in both arms: prefill 1994.47→2747.54 (+37.76%), decode 177.943→181.950. Complete logits byte-identical within each comparison; unlocked clocks and overlapping decode ranges
Mixed Gate/Up preservation Old loader fails 1 byte-preservation test and 2 forecast cases; repaired CPU checks 38/38 and CUDA fixture 1/1 pass. Local Qwen3.8 27B preserves 28 mixed pairs, still fuses 37 matching pairs. Single warm-cache observations: load 85.14→7.36 seconds, post-load working set 16.398→14.484 GB. Not repeated/cold-load or quiet throughput evidence
27B independent comparison Clean llama 4ebdf2c7…, same checkpoint, 258 prompt IDs, context512, F16 KV, all 66 layers on CUDA, 16 fixed-history full vocabularies. Old/new common-offset relative L2 medians 0.007010/0.004672, maxima 0.026925/0.017443. Argmax 16/16 for both, but both fail 0.001; some individual rows worsen. Preserving stored weights is not full-model numerical qualification
Wider parallel-decode failure The repaired 27B model's parallel-K vs serial fixed-history comparison fails: maximum relative L2 0.011107. This remains experimental despite earlier small-model speedups
Adaptive host-weight reuse Final Models eb3781a9…, unchanged native 9be5eabf…: 134 related CPU, 25 CUDA, 2 comparison-tool tests pass without skips. Qwen 32 full vocabularies pass original numerical gates (max relative L2 0.0008168431); Gemma 8 rows are byte-identical to resident. Actual cache hits, read-ahead failure draining, trim/reload, pressure/reset and final zero-owner checks pass. New kernels, language quality and multimodal scheduling are not claimed
File-mode cache ABBA Qwen0.8B, context512, 67 prompt tokens, 16 predictions, fixed serial arithmetic/tile32, 512 MiB host/device ceilings, 256 MiB cache ceiling: 249,109,568 bytes actually retained. Six measured requests/arm give prefill 67.224→77.612 tokens/s (+15.45%), decode 2.416→2.772 (+14.74%), logical reads 13,839,729,920→9,356,312,576 bytes/request (−32.40%). Complete logits/history are identical, processes do not overlap and cleanup/shutdown succeeds. Warm file cache, no clock lock, no concurrent build/download/inference; still far below resident baseline
Gemma E4B regression Fixed32/unset configurations give byte-identical full outputs for 16 predictions. Its existing resident arithmetic does not invoke the automatic Q8/F32 tiler, so this is not new-kernel coverage or a speedup claim
Device-weight reuse 74c1b4f8… Final Models c2404c9e…: 13 native and 129 related managed checks pass, no skips; CUDA memcheck 0 errors. Qwen 32 full vocabularies pass original gates (max relative L2 0.0008168431); Gemma E4B 8 rows are byte-identical to resident. Actual device retention, pressure/reset and zero final owners verified. Combined device payload peaks 135,736,576 / 251,354,112 B within 256 MiB ceilings
Device-cache ABBA Final four fresh Qwen0.8B file-mode processes, six measured requests/arm, 512 MiB host/device and fixed 256 MiB host-cache ceilings. Device cache off/on retains 184,048,640 B. Prefill 75.096→86.556 tok/s (+15.26%), decode 3.035→3.851 (+26.90%), weight H2D −20.53%, logical source reads −30.33%. Identical complete logits/history, successful exits/shutdown and zero owners. Warm file cache, unlocked clocks, no concurrent inference/build/download; still far below full-resident baseline
Finished-state boundary Deterministic worker-pausing negative control fails with old Runtime. Repaired related prefix/engine/budget checks: 422 passed, 33 skipped. Skipped model fixtures remain unvalidated
Idle CUDA workspace reuse Models e1162225…, unchanged native 74c1b4f8…: 74 related tests pass with no skips, including changed inputs/weights, incompatible shapes, shared-owner pressure and failed create/upload/free recovery. Qwen 32 full vocabulary rows pass original gates; Gemma E4B 8 rows are byte-exact to resident. Both pass pressure/reset and zero-owner checks. Combined device payload peaks 138,409,472 / 265,132,800 B under 256 MiB ceilings
Workspace ABBA Same Qwen file-mode geometry, 512 MiB host/device, 256 MiB RAM-cache, fixed 128 MiB device-weight cache, workspace ceiling 0/64/64/0 MiB. Six measurements/arm: prefill 82.297→89.414 tok/s (+8.65%), decode 3.574→4.577 (+28.08%). 31,555,584 B idle workspaces replace 2,224 allocations/request with zero new sessions after warmup. Full logits/history, file reads and weight uploads are unchanged; zero final owners. Quiet GPU, warm file cache, unlocked clocks; still below resident parity. Comparator regressions 6/6 pass
Automatic workspace policy Models 55529501…, unchanged native 74c1b4f8…: 78 related managed/CUDA tests pass, no skips, including budget competition, request aging and failed-release retry. 27 relevant comparator tests pass. Qwen 32 full-vocabulary rows pass original gates (max relative L2 0.0008168431); Gemma E4B 8 rows remain byte-exact to resident. Pressure/reset and final zero owners verified
Default-policy ABBA Four separate off/auto/auto/off series, 16 completed processes, automatic RAM/device weight caching in both arms. Auto omits the workspace setting. Qwen 512 MiB host/device: prefill 84.753→89.950 tok/s (+6.13%), decode 3.554→4.550 (+28.02%), weight uploads +6.48%. Qwen 384 MiB: 66.774→74.779 (+11.99%), 2.702→3.436 (+27.17%), uploads +1.46%. Gemma E4B 1 GiB: 10.977→11.457 (+4.37%), 0.574→0.650 (+13.13%), uploads +2.96%. Six measured requests/arm for each fixed workload; complete logits/history match and owners clear. Quiet GPU, warm file cache, unlocked clocks
Mixed requests and limits Qwen 512 MiB alternates 69/258-token prompts: short-shape decode 2.897→3.276 (2 samples/arm); long 3.295→3.284, prefill 123.844→122.702 (4 samples/arm, overlapping ranges). First long shape includes initial setup and can retire the idle pool after actual reuse. Comparator checks per-request reuse and byte conservation, reports cache/transfer tradeoffs, and retains strict isolated-comparison rules. Quarter share is a heuristic; no global optimum, latency-feedback controller, full-resident parity or new semantic-quality pass is claimed
Validation-tool boundary Cache comparison regressions 4/4 pass. Broader Python suite ran 582 tests: 5 failures, 49 errors, 2 skipped, with missing historical artifacts/pins and Windows Bash/path/encoding issues. Not counted as passing; explicit UTF-8 HTTP rerun 15/15 passes

Warmups are excluded where specified; hashing warms the file cache. Prefill remains substantially below the independent baseline. A separate bounded Q8-to-F32/pedantic-SGEMM research target found original absolute-gate failures in both candidate tiles and the old serial control. Failures are retained without gate relaxation or production promotion.

The follow-up chunked-K research uses short pedantic F32 GEMMs and budgeted FP64 accumulation tiles. Chunk 128/64 clears the original 527-sample gate on one large synthetic geometry but is slower; chunk 256 still has failing tiles. This GEMM experiment is not in production dispatch. The small suite has 30 shape/admission cases and 201 admitted tiles; CUDA memcheck reports zero errors/leaked bytes. CLI/reporting revisions retain distinct executable identities, and a benchmark with zero admitted candidates now fails instead of counting as a pass. A subsequent Tensor Core component expansion also remained slower; chunk256 failed the absolute gate. Rejected code/binaries and failed outcomes remain ignored artifacts. These are exploratory projection results, not model-prefill qualification.

Semantic validation does not treat existing engines as ground truth:

  • Qwen arithmetic/extraction/squares: serial, parallel and independent llama match templates/tokens, but all pass only 1/3 semantic checks.
  • Verified Unsloth Gemma 12B UD-IQ2_M still repeats/fails the FF7 request in both engines. Suppression fixes a real sampler defect, not this quality failure. Same-source Q4 finishes naturally at 1152/1177 tokens but retains factual/translation errors. Three original dense low-bit projections receive independent sampled decoding/FP64 diagnosis; no full-model correctness claim follows. An embedding-only Q3_K→Q4_K diagnostic leaves 666 other tensors byte-identical but also fails: original and variant hit 3072 tokens without EOS, with repetitive tails and factual errors. It is not presented as a repaired original checkpoint.
  • Flash Next short math and banner OCR pass narrow checks. Original agent suite was 2/6; repaired code-generation/edit workflows both pass at original 8k context, with real shell execution, downloaded artifacts and independent extra inputs. Roughly 275/261-second runtimes are not qualified performance.
  • Qwen Image 2.1 has one narrow text-to-image pass. Banner edits at 512² and1152×768 still corrupt subtitle text in TensorSharp and independently built clean stable-diffusion.cpp using unchanged ggml. The larger reference run recovers from VAE OOM with spatial tiling; recovery is not semantic success. Raw images/failures are preserved.

Remaining limits: hooks do not cap total RSS/VRAM or every backend/driver allocation. Wider-model and Gemma multi-token device-weight retention, async H2D/D2H, global expert-cache fairness/regrowth and automatic pressure scheduling, all-model/media integration, production concurrency SLOs and streamed multi-GPU coverage remain incomplete. Slot reduction is geometric, not globally optimal; explicit cache ceilings and the native physical headroom floor remain. Current SSH to the supplied VM is refused. Historical A40/P2P results are separate; unavailable/skipped scenarios never count as passes. This PR remains a draft.

Generated evidence stays ignored under artifacts/ and docs/validation/; reusable tools/fixtures are checked in. Design and evidence · Adaptive comparison · Batch probe

@github-actions

github-actions Bot commented Oct 8, 2026

Copy link
Copy Markdown

Engine comparison — TensorSharp vs llama.cpp (PR smoke)

No report artifact was produced — the benchmark failed before generating results (see the workflow logs).

@zhongkaifu zhongkaifu changed the title Add unified memory foundation and generalize Qwen placement Integrate unified memory residency, tiered KV snapshots and request admission Oct 8, 2026
@zhongkaifu zhongkaifu changed the title Integrate unified memory residency, tiered KV snapshots and request admission Integrate unified memory budgets, KV paging and Qwen weight streaming Oct 8, 2026
@zhongkaifu zhongkaifu changed the title Integrate unified memory budgets, KV paging and Qwen weight streaming Integrate unified memory budgets, KV paging and Qwen/Gemma weight streaming Oct 8, 2026
Prefer qualified resident execution with hardware-aware admission, account owned graph buffers, and overlap bounded weight reads. Preserve token suppression across generation paths and protect checkpoint publication during retained-state eviction. Invalidate file-read visibility when history is compacted.

Validate unchanged ggml ffa4e8b with CPU regressions, real CUDA lifecycle and streaming checks, and alternating E4B/Qwen comparisons. Keep Gemma IQ2 repetition, Flash numerical differences, absolute Qwen decode speed and unavailable multi-GPU scenarios explicitly unresolved.
Add bounded expert file staging, adaptive read/upload prefetch, decaying
frequency eviction and shared host staging credit without modifying ggml.
Handle published GLM architecture keys and explicit native offload policy;
omit unused Qwen NextN payload when the speculation policy excludes it.

Avoid redundant single-token Qwen CUDA copies and materialized broadcasts.
Account owned Qwen graph arenas, recurrent state and device snapshots in
shared graph budgets, including refusal, retry and physical-release handling.
Add reusable diagnostics, benchmark verifiers and lifecycle regression tests.

Record real-model and Windows/Linux evidence, performance gaps and incomplete
resource coverage in the unified-memory design and model documentation.
Generated evidence stays ignored under artifacts; upstream ggml is unchanged.

Validation on unchanged ggml ffa4e8b8:
- Windows/Linux managed CUDA: 18 unique / 21 passes; CPU policy: 150 each.
- Native: Windows 85 pass + 1 skip; Linux 84 pass + 6 skip + 2 failures.
  Preserve upstream attention-oracle failure and default NCCL gather timeout;
  gather passes separately with NCCL_P2P_DISABLE=1 on this VM.
- Changed Python tools: 85 pass + 1 platform skip. Historical full discovery
  remains non-green (missing evidence fixtures, CRLF pins, platform assumptions).
- Real Qwen IQ1_M: 34 full-vocabulary rows bit-identical in four checked arms.
  Interleaved application decode: 55.47 -> 56.65 tokens/s;
  simultaneous llama.cpp 59.95. Both retain the same 3/15 strict
  raw-JSON format failures; TS tool round trips and schema checks pass.
… shared budgets

Charge actual GGML host pool and weight-buffer allocations until physical release,
derive dense request growth/workspace peaks, and consume admitted envelopes in the
qualified serial execution lane. Record exact ledger high-water marks and preserve
failed cleanup ownership.

Keep paged-attention CUDA cache payload out of Windows thread destructors, connect
its graph buffers to budget allocation/free, and release retired caches during
quiescent trim. All native changes are TensorSharp-owned; ggml remains unchanged.

Add fresh-process long-context and delegated-cgroup validation tools. Document
8K/32K/64K parity, measured performance variation, accounting gaps and unavailable
physical hard-cap enforcement. Windows/Linux related regressions: 398 passed and
5 skipped each, plus final allocation/worker-exit checks; Linux memory harness:
54/54. No all-model or default-on readiness claim.
Qualify omitted-draft Qwen trunks, forecast retained decode graphs, and expose bounded allocation-refusal diagnostics. Move host observations into the memory layer, add Apple providers, and correct MLX wired headroom. Record local CUDA/CPU model validation and the remaining platform and physical-limit gaps.
Integrate main 0ec3277 while preserving budgeted KV snapshot leases, physical accounting, model ownership, and reply-budget fallback. Keep recurrent short payloads distinct from tiered page capacity and preserve exact media measurement before relaxing the reply reserve.

Retain failed model cleanup for retry, trim borrowed GGML pools, and fix the admission lost-wakeup race when capacity shrinks during reclamation. Use sparse Windows catalog fixtures and capability-gated symbolic-link coverage.

Validation: Release solution build; InferenceWeb 8893 passed / 410 skipped; TensorAgent 1278 passed / 181 skipped plus 225 passed / 2 skipped on latest Chat dependencies; portable memory 58 passed / 1 skipped; 4 native CUDA and 4 CPU copy-fault tests. Qwen 0.8B 32K dual-request and Gemma E4B shared-budget outputs match ordinary mode and pre-merge tokens; image follow-up and real Gemma disposal/reload pass. Unmodified ggml ffa4e8b80930029a35991f94e7c8a93cd67730ab.

Changed Python modules: 37 passed / 1 skipped. Full Python suite retains pre-existing evidence, source-pin, and Windows environment failures; attribution and generated validation evidence remain ignored under artifacts/merge-main-20261010. Apple and multi-GPU hardware were not available.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant