Skip to content

Replay: split the noise headline into re-derived vs as-recorded; re-derive behavior, hot-dir, process ancestry and the S7 payload case - #64

Open
opencdlee-dotcom wants to merge 4 commits into
agent/precision-s0b/integrationfrom
agent/precision-p2a/harness
Open

opencdlee-dotcom wants to merge 4 commits into
agent/precision-s0b/integrationfrom
agent/precision-p2a/harness

Conversation

@opencdlee-dotcom

Copy link
Copy Markdown
Owner

Precision plan, pass 2, step P2a: the replay harness. Base: agent/precision-s0b/integration (PR #59).

This branch also contains S7. origin/agent/precision-s7/integration (PR #61) is merged in at the coordinator's request, because with S7's payload case the harness could no longer re-derive payload-only persistence changes. The last commit fixes that.

Why

The done-condition is "0 judged-noise incidents re-alert on replay". noise re-opened: N of M mixed two groups. One is incidents the current code still alerts on after re-deriving their evidence. The other is incidents replayed as recorded because the harness could not ask their evidence again: the subject is gone from disk, the sensor is not modelled, or the record lacks a field. No code change can move the second group, so the total could not serve as a target.

What changed (harness only: _backtest_replay, _reobserve*, _replay_*, cmd_backtest_replay)

  1. Headline split. Each re-opened noise incident is classified by the recorded findings that re-open it. If a finding interrupts, that finding counts. Otherwise every finding held by the open scratch case it joined counts. An incident is re-derived when all of those findings were re-observed with current code. It is as recorded when any of them was replayed as recorded, and it carries the first such finding's reason. The old total stays first, so earlier numbers remain comparable. Rule 17: R + A == N, and the two halves cover the re-opened set. Both are asserted, and a failure prints SELF-CHECK FAILED at the top. Without --reobserve the headline says "not split".
  2. behavior. A complete command_preview is re-scored by the current _argv_signals and re-keyed by _argv_case_identity. No signals means dropped. Four kinds of preview are counted and replayed as recorded instead:
    • elided (…)
    • clipped at the 240-character budget. Before 2026-09-19 the preview was the first 240 characters of argv with no elision mark. For a harness line that is only the wrapper prologue, so reading it as the whole command would wrongly drop every such finding.
    • redacted ([REDACTED]). redact_sensitive can swallow a token a rule reads.
    • absent
  3. process ancestry. Recorded ancestry is exe paths and was already passed as parents; this branch adds a test for it. The sensor records ancestry only while a vouch exists. When a record has no ancestry and a vouched program could have earned the binary the supervised rung, the finding is counted field missing: ancestry and replayed as recorded. It is no longer graded as if it had no parent. This follows the persistence rebuild's existing rule for an unrecorded field that decides the grade.
  4. hot-dir. Re-derived by check_hot_dirs itself over the item's own directory, while the executable is still the recorded bytes. The freshness window is opened to the epoch, because the record already answered freshness. If the bytes moved on or the item is gone, the finding is counted. If the directory is no longer watched, the finding is dropped.
  5. S7 payload case. _persistence_rederives now also accepts the payload-update finding. It must name the recorded plist in referrer_paths and print the recorded payload and bytes; its severity and custody are then taken. A recorded payload case is rebuilt on every job it names.

Live replay on the reference Mac (backtest replay --days 30 --reobserve)

  • noise re-opened: 74 of 214 — re-derived 17 (the target), as recorded 57. Before, on this base: 76 of 214, unsplit.
  • re-derived: #344 #352 #382 #406 #413 #414 #429 #432 #433 #434 #435 #449 #452 #453 #495 #514 #521
  • "the rebuilt change does not reproduce the record": 1413 with the payload matcher disabled, 0 with it. #383 #392-#397 re-derive to MEDIUM/local-commit and do not re-open.
  • The self-check passes. Assay recall is unchanged at 9/21 interrupt (1 not run, 11 predicate).

Not done / found

  • #289 #290 #384 and #534's first finding carry no ancestry. They were recorded before the sensor captured it, so there are no exe paths to recover. They now count as field missing: ancestry.
  • #352 (beacon, Runner.Worker) is real residue. The beacon sensor never passes parents to _grade_binary, so the supervised rung cannot answer for a beacon under current code, and the beacon records carry no ancestry either. This is S5's "wire _grade_binary into the emitters" item.
  • A behavior record's ancestry is only {pid, name} enrichment that the sensor never grades on. It is left alone.

Tests

tests/test_backtest_replay_split.py has 18 new tests. tests/test_backtest_replay_persistence.py now records the payload test in both shapes.

  • Replay and S7 files (test_backtest_replay*.py, test_persistence_payload_case.py): 69 passed.
  • Full suite at 6a1aa76, python3 -m pytest tests/ -q -p no:cacheprovider: 2212 passed, 6 skipped, 30 xfailed, 41 subtests passed in 1043.27s. The tree sha was a1fa7f24 both before and after the run, and ~/.aegis/custody.jsonl was byte-identical before and after (S7's conftest guard is in the tree; it refused 48 writes from pre-existing tests).

🤖 Generated with Claude Code

opencdlee-dotcom and others added 4 commits September 23, 2026 15:07
…hat explained it was lost to every git timeout

Precision plan step S7 (Phase 3): the operator's own LaunchAgents.

Gap 1: one payload edit, seven cases. #383 and #392-#397 were one edit to
~/Ai/Universe/tools/aikit/schedule/run.py (4a1646366adf -> 25ee9a4593c2),
one incident per com.aikit.* plist that runs it. check_persistence now
collects a change confined to a payload's bytes (same program, argv, env and
payload path; never an attack-defined job) and emits ONE finding per
(payload, new sha): case persistence:payload-update:<payload>, the referring
labels in the detail and their plist paths in referrer_paths. Same shape as
the OS-program case. _accept_into_baseline promotes every job a verdicted
payload case names, so one benign-positive ends the re-assertion for all of
them. Store migration persistence_payload_case_20260923 folds open per-plist
incidents whose newest evidence is payload-only into that case (evidence
matched, created_at < now, adjudicated rows untouched).

Gap 2: why custody read null. The rung did reach the payload.
_custody(run.py, sha) answered local-commit (MEDIUM) on 193 of the 217 scans
that recorded the change. The 24 scans that recorded custody null and HIGH
were the scans where git did not answer: 11-190 s per plist there, against
0-1 s when it answered (probe timeouts are 10-15 s). Seven plists asked git
about the same file seven times a scan. Each non-answer flipped the case
back to HIGH and wrote a new event. The payload is now graded once per scan,
by _custody_payload. It asks the same intent-then-git questions. When git
gives no answer at all, it asks whether the bytes were already explained:
by the custody ledger (a rung this payload earned on an earlier scan, now
remembered) or by an intent receipt for the same sha at another path. Either
one carries as copy-of-graded, the weakest rung. On the recorded corpus that
takes run.py from 24 HIGH scans to 1 (the first sighting, before any answer
existed).

#319 (~/.local/bin/improver bytes -> a588f96d11b4) was never asked. An
extension-less payload fails _intent_worthy, and its receipt binds the source
path the agent wrote (.../VSCode Projects/improver/improver.py), not the
installed copy. _custody_payload drops that hook-mode prefilter, and
_intent_receipt finds the receipt by sha. Result: copy-of-graded, MEDIUM.

Also touched in the same branch: the per-job payload custody lookup now
honours attack_defined, as its comment already promised. Before this, a
DYLD-injected job whose payload was committed recorded
custody=self-committed on an attack-defined finding. The severity was
untouched, but the chain and routing tiers read that field.

Gap 3: producer fields. On current code the NEW finding already resolves a
producer class through its subject. The seven cited incidents lack one only
in their pre-2026-08-29 evidence, which predates subjects and the runner
subcommand fix (script_target None) and cannot be recovered. The NEW finding
now also carries program_sha and target_sha flat, and target_sha on the
subject, so recorded evidence names both halves of what the job executes.

Also touched: tests/conftest.py. The new ledger writer turned an existing
test leak into a write. FamiliesAreTheAdjudicationSurface re-diffs the REAL
LaunchAgents through _accept_into_baseline, and during this branch's suite
run it appended a local-commit row for run.py to the operator's
~/.aegis/custody.jsonl. That row has been removed by hand. The suite now
refuses and reports any _custody_remember aimed at the real ledger, the same
way it handles the notary anchor. The guard also stops a pre-existing leak:
TestWritEnforcementIsActuallyWired reached the real ledger 45 times through
_grade_binary.

Tests: tests/test_persistence_payload_case.py (28).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…not ask again, so it was not a target

`noise re-opened: N of M` counted every re-opened noise incident alike,
whether the current code still alerts on it after re-deriving its evidence,
or the harness replayed that evidence as recorded (a subject gone from disk,
a sensor it does not model, a field the record never carried). The second
group cannot reach 0 by any code change.

backtest replay --reobserve now classifies each re-opened incident by the
recorded findings that re-open it: the finding itself when it interrupts,
else everything the open scratch case it joined holds. Re-derived when all
of them were re-observed with current code; as recorded, with the first
such finding's reason, otherwise. The headline keeps its old total first
and adds `— re-derived R (the target), as recorded A`, the re-derived ids,
and the as-recorded ids grouped by reason. Rule 17: R + A == N and the two
halves cover the re-opened set, asserted before printing (SELF-CHECK FAILED
at the top otherwise). Without --reobserve nothing is split, since every
finding is replayed as recorded and a split would read as a met target.

_reobserve and its helpers now return (finding, as-recorded reason) so the
reason travels with each finding instead of being inferred from counters.

Three more sensors are re-derived:

- behavior: a complete command_preview is re-scored by the current
  _argv_signals and re-keyed by _argv_case_identity; no signals means
  dropped. A preview that is elided, clipped at its 240-character budget
  (the pre-2026-09-19 head-of-argv preview carried no elision mark, and
  read as the whole command it would drop every harness finding),
  redacted (redact_sensitive can swallow a token a rule reads), or absent
  is counted with its reason and replayed as recorded.
- process: recorded ancestry (exe paths) was already passed to
  _grade_binary as parents. A record with no ancestry, where a vouched
  program could have earned the binary the `supervised` rung, is now
  counted `field missing: ancestry` and replayed as recorded rather than
  graded as if it had no parent — the rule the persistence rebuild already
  follows for an unrecorded field that decides the grade.
- hot-dir: re-derived by check_hot_dirs itself over the item's own
  directory, freshness window opened to the epoch (the record already
  answered freshness), while the executable is still the recorded bytes;
  moved-on bytes and a gone item are counted, a directory no longer
  watched is dropped.

Tests: tests/test_backtest_replay_split.py (18).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… it replaced, so 1413 findings fell back to as-recorded HIGH

With S7 merged, check_persistence answers a change confined to a payload's
bytes with ONE finding per (payload, new sha) — persistence:payload-update,
naming every job in referrer_paths — instead of one CHANGED finding per
plist. The persistence rebuild still required the current sensor to print
the recorded per-plist line, so every recorded payload-only change (the
live corpus's shape) fell to "the rebuilt change does not reproduce the
record" and replayed as recorded, and #383 #392-#397 (com.aikit.*, run.py
4a1646366adf -> 25ee9a4593c2) re-opened.

_persistence_rederives accepts, besides the recorded line, the payload
case when it names the recorded plist among its referrer_paths and prints
the payload and bytes the recorded line printed; that finding IS the
record re-derived, and its severity and custody are taken (the record keeps
its identity, as for every same-titled re-grade). A recorded payload case
(what the sensor writes from now on) is rebuilt on every job it names by
_reobserve_payload_case_answer, each accepted only while it still runs the
payload at the recorded new bytes.

On the live store: "the rebuilt change does not reproduce the record" goes
from 1413 to 0, and every in-window finding of the seven re-derives
HIGH/- -> MEDIUM/local-commit or stays MEDIUM/local-commit.

tests/test_backtest_replay_persistence.py: the payload test recorded its
finding with the current sensor, which now writes the payload case; it
records both shapes (the per-job one by holding _payload_update off, as the
sensor was before S7) and checks each re-derives, and each is refused once
the payload moved on.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant