From 63595a3ec7420040d38a2c5e77a5a335658567f3 Mon Sep 17 00:00:00 2001 From: Tin Dang Date: Mon, 7 Sep 2026 13:40:43 +0700 Subject: [PATCH] refactor(shard): delete CoalescedReadBatch, and let the #416 harness toggle the fast path MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit `CoalescedReadBatch` (src/shard/dispatch.rs) was the xshard-read-fastpath C3 type: N single-key foreign reads from different connections batched into one SPSC message to one owner shard. It was defined with the C1 types in PR #177, deferred at C3, and never wired: no `ShardMessage` arm, no `spsc_handler` arm, no config, no metric. Its only references were its own definition and a type-existence test. In the type index it read as a shipped cross-shard optimisation; it was not one. It is deleted rather than built because the measurement no longer supports building it. Its premise — "a foreign lock-free read is storage-impossible in place, so SPSC is the only door" — ended when L4 made `Database` `Send + Sync` and `try_foreign_db_read` began serving foreign reads on the calling thread (#777, default `auto` since #785). Re-measured at HEAD on moon#773's own fixture (GET p=1 c200, `--shards 8`, keyspace populated to `DBSIZE` 100,000, `keyspace_misses` 0) from a dedicated load generator: 99.0% of foreign reads served in place, `cross_spsc` 0.90% of commands, 0.0089 parks/cmd, against 87.5% / 0.866 parks/cmd with the flag off. Coalescing reduces `msgs/cmd`, the term the cost model fits at ~0, and would leave every remaining park in place (dead end #1); the wake-batching variant D1 was retired by measurement in #778. The re-take also found two things the docs now record (cost model §6, §8.3): the residual is the owner's exclusive guard on its own read paths (the inline GET at blocking.rs:2515 and every SPSC execute arm take `databases.write()`), and a same-host `redis-benchmark --threads 4` inflates that residual 15x (13.5-15.1% vs 0.90%) through lock-holder preemption on the oversubscribed box. §8's "100% in place / 0.0003 parks/cmd" is the c50 light-client number and is now stated as such. `scripts/gcloud-xshard-absolute.sh` (moon#416) passed a fixed server-argument list, so its `s4-c1-GET` cell silently stopped measuring the SPSC hop it is named for once the default flipped to `auto`. `start_moon`/`cell` now take extra server args (`--extra-args` / `MOON_EXTRA_ARGS`), read cells populate the keyspace and report `DBSIZE`, the default cell set measures the hop (`--cross-shard-fast-path off`) and the fast path as two cells of one A/B plus the write hop (`s4-c1-SET`), and `--self-test` fails closed if the passthrough stops reaching the server argv (mutation-tested: dropping `"$@"` fails gate 12). No behaviour changes: nothing constructed, sent or matched on the deleted type. Both runtimes checked with `--all-targets`; clippy clean. Refs: moon#773, moon#416, moon#777, moon#785, moon#778 author: Tin Dang --- .add/tasks/xshard-read-fastpath/TASK.md | 14 +++ CHANGELOG.md | 45 ++++++++ docs/internal/cross-shard-cost-model.md | 136 ++++++++++++++++++++++++ scripts/gcloud-xshard-absolute.sh | 103 +++++++++++++++--- src/shard/dispatch.rs | 20 ---- tests/xshard_fastpath_api.rs | 16 +-- 6 files changed, 290 insertions(+), 44 deletions(-) diff --git a/.add/tasks/xshard-read-fastpath/TASK.md b/.add/tasks/xshard-read-fastpath/TASK.md index 5068b8ae1..84986696a 100644 --- a/.add/tasks/xshard-read-fastpath/TASK.md +++ b/.add/tasks/xshard-read-fastpath/TASK.md @@ -92,6 +92,20 @@ Must: into fewer SPSC messages; the s4 P1 GET cell recovers measurably toward parity vs the M0 baseline. Coalescing preserves per-connection command order and read-your-writes exactly as the lock-free path does today. + [RETIRED 2026-09-07 — not built, and will not be. Superseded by the L4 S4 + foreign-read fast path (#777, default-`auto` since #785), which removes the + hop entirely rather than making it cheaper: on a saturated keyspace 99.0% + of foreign reads served in place at #773's c200 fixture from a dedicated + load generator (0.0089 parks/cmd vs 0.866 with the flag off, cost model + §8.3; §8's "100%" is the c50 light-client number), the remainder declining + on the owner's exclusive guard, not on anything a batch could fix. The + wake-batching variant + (D1) was separately retired by measurement in #778. Coalescing optimises + `msgs/cmd`, whose fitted coefficient is ~0 — this is dead end #1. The + `CoalescedReadBatch` type and its xrf2 type-pin test were deleted; the + reasoning and the ordering invariant are preserved in + `docs/internal/cross-shard-cost-model.md` §9. History above is left as + written.] - M3 (cleanup, hard-remove): `--cross-shard-fast-path` (the `cross_shard_fast_path` field + `CrossShardFastPath` enum + `cross_shard_fast_path_enabled`), the `moon_cross_shard_lock_contention_total` metric, the diff --git a/CHANGELOG.md b/CHANGELOG.md index 7009dd0b9..7fbd84c63 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -201,6 +201,51 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 A `record_dispatch_cross_read_fast_batch` counterpart was added for the L4 shared-guard read path, which had no batched variant. +### Removed + +- **`shard`: `CoalescedReadBatch`, the cross-connection read-coalescing type + that never had a producer or a consumer (moon#773).** Defined in PR #177 with + the `xshard-read-fastpath` C1 types and deferred at C3, it sat in + `src/shard/dispatch.rs` for three months with no `ShardMessage` arm, no + `spsc_handler` arm, no config and no metric — its only references were its own + definition and a type-existence test. In the type index and in code review it + read as a shipped cross-shard optimisation. It was not one. + + It is deleted rather than wired up because all three arguments for building it + have since failed. Its stated premise — "a foreign lock-free read is + storage-impossible in place, so SPSC is the only door" — stopped being true + when L4 made `Database` `Send + Sync` and `try_foreign_db_read` began serving + foreign reads on the calling thread with one CAS (#777, default `auto` since + #785); coalescing would now be optimising the fallback. It also attacks the + wrong term: it reduces `msgs/cmd`, whose fitted coefficient in the cost model + is ~0, while every connection still parks on its own `ResponseSlot` — that is + dead end #1, predicted 0.99x. And the variant that batches the *wake* instead + was pre-flighted and retired in #778, because monoio's `EventWaker` already + coalesces 87% of cross-thread wakes. + + The reasoning, the load-bearing ordering invariant the type's doc comment + asserted, and the pointer to where the remaining cross-shard work actually is + (writes at 0.875 parks/cmd, the tokio handler, the read-side declines) are + preserved in `docs/internal/cross-shard-cost-model.md` §9. No behaviour + changes: nothing constructed, sent or matched on this type. + +### Fixed + +- **`bench`: `scripts/gcloud-xshard-absolute.sh` could not toggle the + cross-shard fast path, so its `s4-c1-GET` cell silently stopped measuring the + thing it is named for (moon#416).** `start_moon` passed a fixed server-argument + list. Since #785 flipped `--cross-shard-fast-path` to `auto` by default, that + cell on a populated keyspace is served in place by the L4 fast path — it was + reporting the fast path while the harness, the CSV and the frozen + XSHARD-READ-01 contract row all still called it the SPSC hop. + + `start_moon` and `cell` now take an optional extra-server-args string, and + `--extra-args` / `MOON_EXTRA_ARGS` plumb it from the command line, so the hop + and the fast path can be measured as two cells of one A/B instead of one + ambiguous number. The default cell set now runs `s4-c1-GET-hop` + (`--cross-shard-fast-path off`, the actual hop) alongside `s4-c1-GET` + (default `auto`), and a `--self-test` gate fails closed if the passthrough + ever stops reaching the server process. ## [0.8.9] — 2026-09-04 diff --git a/docs/internal/cross-shard-cost-model.md b/docs/internal/cross-shard-cost-model.md index 130e09ca5..bc6fd772d 100644 --- a/docs/internal/cross-shard-cost-model.md +++ b/docs/internal/cross-shard-cost-model.md @@ -222,6 +222,15 @@ Each of these silently produced a plausible, wrong result at least once: - **`moon_dispatch_path_total` is not a remote-fraction metric.** The inline path records `path="local_inline"`, and the counter was absent on the prototype binary — it read a false 100%. +- **A same-host `redis-benchmark --threads 4` inflates the read fast path's + SPSC residual 15x.** Same binary, same c200 p=1 saturated-keyspace fixture: + `cross_spsc` 13.5-15.1% of commands with the client on the server's 8 vCPUs, + **0.90%** with the client on `moon-bench-client`. The `off` control reads + 87.5% either way. Twelve runnable threads on eight vCPUs preempt a shard + thread while it holds the exclusive guard, and every foreign `try_read` that + lands in that quantum declines to SPSC (§8.3). Score the fast path only from a + dedicated load generator, and treat any same-host residual above ~1% as + oversubscription, not as the mechanism. - **redis-benchmark 8.x emits `\r`** for progress lines: `tr '\r' '\n'` before grepping, and match the RPS by position — `awk '{print $2}'` yields `summary:`. @@ -331,6 +340,133 @@ read side by removing exactly this batch boundary. **Do not score a cross-shard path with these counters without first checking that the path increments them.** +### 8.3 Re-take at moon#773's fixture (2026-09-07): the residual is the owner's exclusive guard, and a same-host client inflates it 15x + +The §8 table is c50, n=200k. moon#773's 87.54% was c200. Re-measured at HEAD +(`6251429f`, `release-fast`, monoio, `--shards 8 --appendonly no --admin-port +9413`, 2M `SET`s over `-r 100000` so `DBSIZE` = 100,000 and `keyspace_misses` += 0 in every leg, GET p=1 n=2M, counters as deltas because `CONFIG RESETSTAT` +does not reset them). Two load-generator placements, same binary: + +| client | c | `local_inline` | `cross_read_fast` | `cross_spsc` | parks/cmd | rps | +|---|---:|---:|---:|---:|---:|---:| +| same host, `--threads 4` | 200 `auto` | 12.5% | 72.3-74.0% | **13.5-15.1%** | 0.135-0.151 | 250-286K | +| same host, `--threads 4` | 200 `off` | 12.5% | 0 | 87.5% | 0.872 | 228K | +| same host, `--threads 4` | 25 / 50 / 100 / 400 `auto` | 12.5% | — | 2.4 / 4.2 / 7.8 / 20.1% | = `cross_spsc` | 250-296K | +| same host, single thread | 50 `auto` | 12.5% | 86.7% | **0.71-0.91%** | 0.007-0.009 | 100K | +| **`moon-bench-client`, `--threads 4`** | **200 `auto`** | 12.5% | **86.63%** | **0.90%** | **0.0089** | 275K | +| `moon-bench-client`, `--threads 4` | 200 `off` | 12.5% | 0 | 87.52% | 0.866 | 266K | +| `moon-bench-client`, `--threads 4` | 50 `auto` | 12.5% | 86.89% | 0.62% | 0.0062 | 235K | + +Three reps at c200 same-host, two per point on the sweep, one per two-host +cell; spreads are the ranges shown. `total_remote_awaits_parked` equals the +`cross_spsc` delta to within 0.1% in every `auto` leg: **every read that +declines the fast path still parks** — the residual is pure park, never a +cheaper message. + +What the residual is. `keyspace_misses` = 0 rules out the cold-key decline +(a); p=1 with `pending_mask` reset per batch rules out (b); a single-key `GET` +rules out (c); the binary is monoio, ruling out (e). What is left is (d), +`try_foreign_db_read` returning `None` because the owner holds the exclusive +guard — and the owner takes it for its *reads*: the inline `GET` at +`server/conn/blocking.rs:2515` goes through `with_shard_db` (`set.write()`) +because `promote_inflight_if_present` needs `&mut`, and every SPSC execute arm +takes `databases.write()` (§3 dead end 8). On a dedicated 8-vCPU server that +hold is ~100 ns and collides 0.6-0.9% of the time. With four client threads +sharing those vCPUs, a holder is preempted mid-hold and the window becomes a +scheduler quantum, which is why the same-host residual tracks client CPU +(`--threads 4` vs one thread: 4.2% vs 0.75% at c50) and connection count (the +sweep), not n (200k vs 2M: identical). §3 dead end 8 tested the SPSC arms at +63% keyspace coverage, where cold keys masked everything; it never tested the +inline path, and it never tested at saturation. + +**Corrected claim.** "100% in place / 0.0003 parks/cmd" holds at c50 with a +light client. The citable HEAD number for the #773 fixture is **99.0% of +foreign reads in place at c200 (0.9% residual, 0.0089 parks/cmd) from a +dedicated load generator**, against 0.866 parks/cmd with the flag off. The +same-host 13.5% is a measurement artifact (§6) and must not be quoted as the +mechanism's rate. + +Raw outputs: `tmp/perf-campaign/G2-DELETE.md` (this repo's campaign +directory), copied from `retake773*.out` on `moon-bench-x86`. + +## 9. Cross-connection read coalescing (C3) — retired unbuilt + +`src/shard/dispatch.rs` carried a `CoalescedReadBatch` type from 2026-06-13 +(ADD task `xshard-read-fastpath`, C1) until 2026-09-07. It had **no producer and +no consumer** for its entire life: no `ShardMessage` arm, no `spsc_handler` arm, +no config, no metric. Its only references were its own definition and a +type-existence test. PR #177 shipped C1 (types) and C2 (idle-gated reply spin) +and deferred C3; the follow-up task was never created. This section records why +the line is closed, so the idea is not re-invented from the type's absence. + +**The design.** Accumulate N independent single-key foreign reads from DIFFERENT +connections on the origin shard into one message to a single owner shard, and +route each result back to its own connection's `ResponseSlot`. + +**The ordering invariant it asserted** (worth preserving — any future +cross-connection batching owes the same proof, and §5.4 of the G2 analysis +states it as a formal obligation): coalescing groups reads ACROSS connections +only. Within one connection, submission order and read-your-writes hold exactly +as on the un-batched path — a read is never reordered before that connection's +own acked write. The oracle for this is the consistency suite +(`scripts/test-consistency.sh` at 1/4/12 shards), not a unit test. + +**Why it was retired.** Three independent lines, none of them opinion: + +1. **It attacks the term measured at zero.** Coalescing reduces `msgs/cmd`. + Each of the N connections still parks on its own `ResponseSlot` and is still + woken individually by `slot.fill`, so `parks/cmd` is unchanged. Under the §1 + fit (`cost = 0.413 - 0.046*msgs/cmd + 2.488*parks/cmd`) the best case is + `0.046 x 0.875 = 0.04` CPU%/kops out of 2.586 — about 1.5%. This is + **dead end #1 restated**, and #1 was already predicted at 0.99x. +2. **Batching the wake instead does not help either.** That variant is D1, and it + was pre-flighted and retired in #778: monoio's `EventWaker` already coalesces + 87% of cross-thread wakes (0.111 `syscw`/cmd against 0.872 parks/cmd at c200 + p=1 s8). The park is a fixed 13.29 us (bootstrap, 90% CI [12.22, 14.28], 30 + rows in `.add/tasks/xshard-read-fastpath/d1_preflight.csv`); forcing that term + to zero collapses the fit's R^2 from 0.9937 to 0.69. The signal is not the cost. +3. **Its premise stopped being true.** C3 was justified by + "a foreign lock-free read is storage-impossible in place, so SPSC is the only + door and we make IT cheaper" (`TASK.md:32-36`). L4 made `Database` + `Send + Sync` behind a per-`(shard, db)` `RwLock` (`src/shard/db_plane.rs`, + static assertion at `:498-506`), and `try_foreign_db_read` + (`src/shard/slice.rs:581-589`) now serves the read on the calling thread with + one CAS. C3 would have been optimising the fallback. §8 measures the + replacement at 100% served in place, 0.0003 parks/cmd on a saturated + keyspace — **at c50 with a light client**. The 2026-09-07 re-take at #773's + own fixture (c200, n=2M, `--threads 4`, `DBSIZE` 100,000, §8.3) measures + **99.0% of foreign reads in place, 0.90% of commands still on SPSC, 0.0089 + parks/cmd** from a dedicated load generator, against 87.5% / 0.866 with the + flag off. (A same-host client reads 13.5-15.1% instead — lock-holder + preemption on an oversubscribed box, §6/§8.3, not the mechanism.) The + residual is the owner's exclusive guard (decline (d)) and each of its reads + still parks on its own slot, so it is not something coalescing could touch + either: batching the residual's messages would leave every one of its parks + in place — dead end #1 again, on under 1% of traffic instead of 87.5%. + +**The hazard it also removes.** After #768, reply assembly for spanning +multi-key reads goes through `ReplySink::Part` + the `fanout_state` fold, with +`multikey_placement` as the ONE placement decision function. A cross-connection +C3 producer would have had to understand that fold too, giving reply assembly +two answers to the same question — the exact shape #708 and #768 built +`multikey_placement` to close. + +**Where the energy goes instead.** Every remaining lever removes a *park* (the +2.488 coefficient) or a *batch cut*, never a message: cross-shard **writes**, +still at 0.875 parks/cmd because the fast path is gated on `!is_write` (§8 +"What this does NOT touch"; D3, `docs/internal/d3-concurrent-keyspace.md`); the +**owner's exclusive guard on its own read paths** (§8.3 — the inline GET at +`server/conn/blocking.rs:2515` and every SPSC execute arm take +`databases.write()`, and each such hold turns a concurrent foreign `try_read` +into a parked SPSC hop — 0.6-0.9% of commands on a dedicated host, and the +whole of the fast path's remaining park budget); the **tokio/Windows** +handler, which has no fast-path twin and still routes every foreign read +through SPSC; narrowing `pending_mask` to writes only; and serving #768's +spanning-read parts through `try_foreign_db_read`. + +--- + ## See also - [`env-knobs.md`](env-knobs.md) — spin governor, THP soak, io_uring dead ends. diff --git a/scripts/gcloud-xshard-absolute.sh b/scripts/gcloud-xshard-absolute.sh index a44104c06..b75ce016b 100755 --- a/scripts/gcloud-xshard-absolute.sh +++ b/scripts/gcloud-xshard-absolute.sh @@ -34,6 +34,18 @@ STEAL_MAX="${STEAL_MAX:-1}" # vCPU steal-% gate (~0 on dedicated cores) NOISE_PCT="${NOISE_PCT:-8}" # s1-LOCAL "flat" tolerance + guard-cell regression tolerance SERVER_CORES="${SERVER_CORES:-0-3}" # moon shards: 4 cores CLIENT_CORES="${CLIENT_CORES:-4-7}" # redis-benchmark: disjoint 4 cores +MOON_EXTRA_ARGS="${MOON_EXTRA_ARGS:-}" # extra server args appended to every moon start (per-cell override) +# moon#416: populate before a read cell. The L4 cross-shard fast path DECLINES a +# key that is not resident (the moon#610 class: `dispatch_read` cannot consult the +# cold tier), so on the EMPTY keyspace this harness used to benchmark, every GET +# missed, every read declined to the SPSC hop, and `--cross-shard-fast-path` +# made no difference at all. Cost model §8: the in-place rate tracks the key HIT +# rate to within 0.3 points, so an unpopulated read cell scores a working +# mechanism as a broken one. POPULATE_FACTOR is N/keyspace for the SET preload: +# coverage is 1-exp(-factor), so 8 => 99.97%. Set POPULATE=0 to reproduce the +# pre-2026-09 (empty-keyspace) cells — they are NOT comparable to populated ones. +POPULATE="${POPULATE:-1}" +POPULATE_FACTOR="${POPULATE_FACTOR:-8}" # cores 8-15 on a 16-vCPU instance left IDLE as a steal-time buffer (C2 sizing rationale) # GCloud (C2 instrument — AMENDED 2026-06-23): c2-standard-16 blocked by C2_CPUS quota=8; @@ -96,6 +108,30 @@ self_test() { [[ "$(us_per_op 25000)" == "40.00" ]] && _ok "us_per_op 25000 -> 40.00us" || _bad "us_per_op 25000 should be 40.00 (got $(us_per_op 25000))" grep -q 's4-c1-GET' "$0" && grep -q 's1-LOCAL' "$0" && _ok "harness declares the frozen cells" || _bad "harness must declare the cells" + # ---- moon#416: the extra-server-args passthrough ------------------------ + # These gates exist because the harness previously ACCEPTED no flag at all and + # its s4-c1-GET cell was labelled "the SPSC hop" while measuring whatever the + # build's default happened to be. Each gate below fails closed if the argv + # builder stops carrying the flag through to the server process. + local argv + argv="$(moon_argv /bin/moon 4 /tmp/d --cross-shard-fast-path off | tr '\n' ' ')" + [[ "$argv" == *"--cross-shard-fast-path off"* ]] \ + && _ok "extra server args reach moon's argv (the moon#416 passthrough)" \ + || _bad "extra server args DROPPED from argv (moon#416 regression): $argv" + [[ "$argv" == *"--admin-port 0"* && "$argv" == *"--shards 4"* && "$argv" == *"--appendonly no"* ]] \ + && _ok "argv keeps the fixed cell args alongside the extras" \ + || _bad "argv lost a fixed cell arg: $argv" + argv="$(moon_argv /bin/moon 4 /tmp/d | tr '\n' ' ')" + [[ "$argv" == *"--cross-shard-fast-path"* ]] \ + && _bad "argv invented a flag with no extras passed: $argv" \ + || _ok "no extras => argv carries no fast-path flag (negative control)" + grep -q 's4-c1-GET-hop' "$0" && grep -q -- '--cross-shard-fast-path off' "$0" \ + && _ok "harness declares the hop cell with the fast path OFF (XSHARD-READ-01)" \ + || _bad "harness must declare an s4-c1-GET-hop cell pinning --cross-shard-fast-path off" + grep -q 'populate_keyspace' "$0" && [[ "$POPULATE_FACTOR" -ge 3 ]] \ + && _ok "read cells populate the keyspace (cost model §8: in-place rate == hit rate)" \ + || _bad "read cells must populate to saturation; POPULATE_FACTOR=$POPULATE_FACTOR too low" + if [[ "$fails" -eq 0 ]]; then log "=== self-test PASS (all gates fail-closed correctly) ==="; return 0; fi log "=== self-test FAIL ($fails) ==="; return 1 } @@ -137,16 +173,30 @@ build() { # build -> echoes binary path (read-only checkout mkdir -p "$BINS"; cp "$WORK/target/release/moon" "$out"; echo "$out" } -start_moon() { # start_moon - local bin="$1" shards="$2"; cleanup +# Pure argv builder, so --self-test can PROVE the extra-args passthrough reaches +# the server process. moon#416's whole failure mode was a flag that the harness +# accepted and silently dropped, leaving a cell labelled "the SPSC hop" while it +# measured something else. A guard that cannot observe the argv cannot catch that. +moon_argv() { # moon_argv [extra...] -> one arg per line + local bin="$1" shards="$2" dir="$3"; shift 3 + printf '%s\n' "$bin" --port "$PORT" --shards "$shards" --dir "$dir" \ + --appendonly no --admin-port 0 "$@" +} + +start_moon() { # start_moon [extra_server_args] + local bin="$1" shards="$2" extra="${3-$MOON_EXTRA_ARGS}"; cleanup local dir; dir="$(mktemp -d /tmp/moon-abs.XXXXXX)" - taskset -c "$SERVER_CORES" "$bin" --port "$PORT" --shards "$shards" --dir "$dir" \ - --appendonly no --admin-port 0 >/dev/null 2>&1 & + local xa=(); [[ -n "$extra" ]] && read -ra xa <<< "$extra" + local argv=() a + while IFS= read -r a; do argv+=("$a"); done < <(moon_argv "$bin" "$shards" "$dir" "${xa[@]+"${xa[@]}"}") + taskset -c "$SERVER_CORES" "${argv[@]}" >/dev/null 2>&1 & MOON_PID=$! local deadline=$((SECONDS+10)) until redis-cli -p "$PORT" ping >/dev/null 2>&1; do if ! kill -0 "$MOON_PID" 2>/dev/null; then - taskset -c "$SERVER_CORES" "$bin" --port "$PORT" --shards "$shards" --dir "$dir" --appendonly no >/dev/null 2>&1 & + # Same argv on retry. It used to drop --admin-port 0, so a cell that hit + # the retry path silently paid the metrics-endpoint cost (~11%, §6). + taskset -c "$SERVER_CORES" "${argv[@]}" >/dev/null 2>&1 & MOON_PID=$! fi [[ $SECONDS -ge $deadline ]] && { log " ERROR moon did not start"; return 1; } @@ -154,25 +204,39 @@ start_moon() { # start_moon done } +# Preload the keyspace with the SAME key distribution the read cell will query +# (redis-benchmark -r expands __rand_int__ identically for -t set and -t get, so +# DEBUG POPULATE would NOT match). Echoes the resulting DBSIZE, which must be +# reported with any read cell: cost model §8 requires it. +populate_keyspace() { + [[ "$POPULATE" != "1" ]] && { redis-cli -p "$PORT" DBSIZE 2>/dev/null | tr -d '\r'; return 0; } + local n=$((KEYSPACE*POPULATE_FACTOR)) + taskset -c "$CLIENT_CORES" redis-benchmark -p "$PORT" -t set -n "$n" -r "$KEYSPACE" \ + -c 50 -P 64 -q >/dev/null 2>&1 + redis-cli -p "$PORT" DBSIZE 2>/dev/null | tr -d '\r' +} + bench_rps() { # bench_rps

taskset -c "$CLIENT_CORES" redis-benchmark -p "$PORT" -c "$1" -P "$2" -n "$REQUESTS" \ -r "$KEYSPACE" -t "$3" --csv 2>/dev/null \ | tr '\r' '\n' | grep "\"${3}\"" | awk -F',' '{gsub(/"/,"",$2); printf "%.0f\n",$2}' | tail -1 } -cell() { # cell

-> "name|best|reps_csv|n" - local bin="$1" shards="$2" c="$3" p="$4" cmd="$5" name="$6" vals=() n=0 attempts=0 r +cell() { # cell

[extra_server_args] -> "name|best|reps_csv|n|dbsize" + local bin="$1" shards="$2" c="$3" p="$4" cmd="$5" name="$6" extra="${7-$MOON_EXTRA_ARGS}" + local vals=() n=0 attempts=0 r db="" local max_attempts=$((BEST_OF_N*3)) # bounded retries so a flaky rep can't loop forever while [[ "$n" -lt "$BEST_OF_N" && "$attempts" -lt "$max_attempts" ]]; do attempts=$((attempts+1)) wait_quiesced # BLOCK until load decays (not skip) — the c2d-VOID fix - start_moon "$bin" "$shards" || { sleep 2; continue; } + start_moon "$bin" "$shards" "$extra" || { sleep 2; continue; } + db="$(populate_keyspace)" # moon#416: read cells need a resident keyspace r="$(bench_rps "$c" "$p" "$cmd" || echo 0)"; cleanup [[ -z "$r" || "$r" -le 0 ]] 2>/dev/null && { log " retry $name (bad rep: '$r')"; continue; } vals+=("$r"); n=$((n+1)) done local best=0; [[ "$n" -gt 0 ]] && best=$(printf '%s\n' "${vals[@]}" | sort -n | tail -1) - printf '%s|%s|%s|%s\n' "$name" "$best" "$(IFS=,;echo "${vals[*]:-}")" "$n" + printf '%s|%s|%s|%s|%s\n' "$name" "$best" "$(IFS=,;echo "${vals[*]:-}")" "$n" "${db:-0}" } measure() { @@ -193,21 +257,28 @@ measure() { log "=== gates OK (steal=$steal%, clean) — measuring on $(uname -srm) ===" echo "# absolute cells (raw)" - echo "# machine|runtime|commit|cell|best_rps|us_per_op|reps|n" + echo "# machine|runtime|commit|cell|best_rps|us_per_op|reps|n|dbsize" declare -gA BEST local rt commit for rt in $RUNTIMES; do for commit in $COMMITS; do local bin; bin=$(build "$commit" "$rt") || die "build $commit/$rt" [[ "$SETTLE_AFTER_BUILD" == "1" ]] && { log " settling compile heat before cells..."; wait_quiesced; } - while IFS='|' read -r name best reps n; do + while IFS='|' read -r name best reps n db; do BEST["$rt|$commit|$name"]="$best" - printf '%s|%s|%s|%s|%s|%s|%s|%s\n' "${GCE_MACHINE:-local}" "$rt" "$commit" "$name" "$best" "$(us_per_op "$best")" "$reps" "$n" + printf '%s|%s|%s|%s|%s|%s|%s|%s|%s\n' "${GCE_MACHINE:-local}" "$rt" "$commit" "$name" "$best" "$(us_per_op "$best")" "$reps" "$n" "$db" done < <( cell "$bin" 1 1 1 GET "s1-LOCAL" - cell "$bin" 4 1 1 GET "s4-c1-GET" + # moon#416 / G2 §6: the read hop and the L4 fast path are now TWO cells of + # one A/B, not one ambiguous number. XSHARD-READ-01 is the -hop cell. + cell "$bin" 4 1 1 GET "s4-c1-GET-hop" "--cross-shard-fast-path off" + cell "$bin" 4 1 1 GET "s4-c1-GET" "" cell "$bin" 4 1 16 GET "s4-P16" cell "$bin" 4 100 1 GET "s4-c100-GET" + # The write hop, which the fast path does NOT touch (gate is !is_write) and + # which D3 targets. G2 §6 proposes this as XSHARD-READ-01's replacement in + # docs/PRODUCTION-CONTRACT.md once it has a measured value. + cell "$bin" 4 1 1 SET "s4-c1-SET" ) done done @@ -283,6 +354,12 @@ gcloud_sweep() { # run gcloud_run_one for each machine in MACHINES, each in its } # ============================================================================= dispatch +# --extra-args "" sets MOON_EXTRA_ARGS for every cell that does not pin its +# own (moon#416). Consumed before the subcommand so either order works. +if [[ "${1:-}" == "--extra-args" ]]; then + MOON_EXTRA_ARGS="${2:-}"; shift 2 +fi + case "${1:---gcloud}" in --self-test) self_test ;; --measure) measure ;; diff --git a/src/shard/dispatch.rs b/src/shard/dispatch.rs index aae644636..b68f59097 100644 --- a/src/shard/dispatch.rs +++ b/src/shard/dispatch.rs @@ -27,26 +27,6 @@ impl std::fmt::Debug for ResponseSlotPtr { } } -/// One batched cross-shard READ message carrying N independent single-key foreign -/// reads from DIFFERENT connections on the ORIGIN shard to ONE owner shard. Each -/// result is routed back to its own connection's response slot. -/// -/// xshard-read-fastpath C3 (cross-connection coalescing). The TYPE is defined with -/// C1; the producer (origin-shard accumulation) and consumer (owner-shard drain + -/// per-slot reply) are wired in the C3 build stage. INVARIANT (load-bearing, proven -/// by the consistency suite, not a unit test): coalescing groups reads ACROSS -/// connections only — within a connection, submission order and read-your-writes are -/// preserved exactly as the current lock-free path (a read is never reordered before -/// that connection's own acked write). -pub struct CoalescedReadBatch { - /// Logical DB index the batched reads target (all reads in a batch share it). - pub db_index: usize, - /// Each foreign read: its dispatch frame + the response slot to route the - /// reply back to. `SmallVec` inline-stores up to 8 (the common fan-in) with no - /// heap allocation; the buffer is reused per event-loop turn, never keyspace-scaled. - pub reads: smallvec::SmallVec<[(std::sync::Arc, ResponseSlotPtr); 8]>, -} - /// Lock-free response slot for accumulating cross-shard PUBLISH subscriber counts. /// /// Instead of N-1 oneshot channels (one per target shard), a single `Arc` diff --git a/tests/xshard_fastpath_api.rs b/tests/xshard_fastpath_api.rs index c20667297..3347c00c2 100644 --- a/tests/xshard_fastpath_api.rs +++ b/tests/xshard_fastpath_api.rs @@ -3,7 +3,11 @@ //! //! These are RED until §5 BUILD defines the symbols: //! - `moon::shard::slice::xshard_may_spin` / `XshardWaitGuard` / `XSHARD_SPIN_GATE` -//! - `moon::shard::dispatch::CoalescedReadBatch` +//! +//! The C3 pin (`moon::shard::dispatch::CoalescedReadBatch`) was REMOVED with the +//! type itself — cross-connection read coalescing was retired unbuilt, superseded +//! by the L4 foreign-read fast path. See `docs/internal/cross-shard-cost-model.md` +//! §9 for the reasoning and the measurement that closed the line. //! //! Before build the imports below are UNRESOLVED → this test crate fails to compile. //! That compile failure IS the red signal for this file (the other test crates are @@ -153,13 +157,3 @@ fn batch_gate_and_inflight_gate_compose() { "singleton read on a now-idle shard spins again once waiters drain" ); } - -/// xrf2 — the cross-connection coalescing message type exists and is referenceable. -/// Routing CORRECTNESS (read-your-writes, per-connection submission order) is proven -/// by the consistency suite (`scripts/test-consistency.sh` 197/197 @1/4/12), NOT here: -/// the frozen contract names the 197-suite + xrf-ryw as the oracle for C3. -#[test] -fn coalesced_read_batch_type_exists() { - fn _assert_type_exists() {} - _assert_type_exists::(); -}