Repository navigation
Replay: split the noise headline into re-derived vs as-recorded; re-derive behavior, hot-dir, process ancestry and the S7 payload case - #64
Open
opencdlee-dotcom wants to merge 4 commits into
Conversation
…hat explained it was lost to every git timeout Precision plan step S7 (Phase 3): the operator's own LaunchAgents. Gap 1: one payload edit, seven cases. #383 and #392-#397 were one edit to ~/Ai/Universe/tools/aikit/schedule/run.py (4a1646366adf -> 25ee9a4593c2), one incident per com.aikit.* plist that runs it. check_persistence now collects a change confined to a payload's bytes (same program, argv, env and payload path; never an attack-defined job) and emits ONE finding per (payload, new sha): case persistence:payload-update:<payload>, the referring labels in the detail and their plist paths in referrer_paths. Same shape as the OS-program case. _accept_into_baseline promotes every job a verdicted payload case names, so one benign-positive ends the re-assertion for all of them. Store migration persistence_payload_case_20260923 folds open per-plist incidents whose newest evidence is payload-only into that case (evidence matched, created_at < now, adjudicated rows untouched). Gap 2: why custody read null. The rung did reach the payload. _custody(run.py, sha) answered local-commit (MEDIUM) on 193 of the 217 scans that recorded the change. The 24 scans that recorded custody null and HIGH were the scans where git did not answer: 11-190 s per plist there, against 0-1 s when it answered (probe timeouts are 10-15 s). Seven plists asked git about the same file seven times a scan. Each non-answer flipped the case back to HIGH and wrote a new event. The payload is now graded once per scan, by _custody_payload. It asks the same intent-then-git questions. When git gives no answer at all, it asks whether the bytes were already explained: by the custody ledger (a rung this payload earned on an earlier scan, now remembered) or by an intent receipt for the same sha at another path. Either one carries as copy-of-graded, the weakest rung. On the recorded corpus that takes run.py from 24 HIGH scans to 1 (the first sighting, before any answer existed). #319 (~/.local/bin/improver bytes -> a588f96d11b4) was never asked. An extension-less payload fails _intent_worthy, and its receipt binds the source path the agent wrote (.../VSCode Projects/improver/improver.py), not the installed copy. _custody_payload drops that hook-mode prefilter, and _intent_receipt finds the receipt by sha. Result: copy-of-graded, MEDIUM. Also touched in the same branch: the per-job payload custody lookup now honours attack_defined, as its comment already promised. Before this, a DYLD-injected job whose payload was committed recorded custody=self-committed on an attack-defined finding. The severity was untouched, but the chain and routing tiers read that field. Gap 3: producer fields. On current code the NEW finding already resolves a producer class through its subject. The seven cited incidents lack one only in their pre-2026-08-29 evidence, which predates subjects and the runner subcommand fix (script_target None) and cannot be recovered. The NEW finding now also carries program_sha and target_sha flat, and target_sha on the subject, so recorded evidence names both halves of what the job executes. Also touched: tests/conftest.py. The new ledger writer turned an existing test leak into a write. FamiliesAreTheAdjudicationSurface re-diffs the REAL LaunchAgents through _accept_into_baseline, and during this branch's suite run it appended a local-commit row for run.py to the operator's ~/.aegis/custody.jsonl. That row has been removed by hand. The suite now refuses and reports any _custody_remember aimed at the real ledger, the same way it handles the notary anchor. The guard also stops a pre-existing leak: TestWritEnforcementIsActuallyWired reached the real ledger 45 times through _grade_binary. Tests: tests/test_persistence_payload_case.py (28). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…not ask again, so it was not a target `noise re-opened: N of M` counted every re-opened noise incident alike, whether the current code still alerts on it after re-deriving its evidence, or the harness replayed that evidence as recorded (a subject gone from disk, a sensor it does not model, a field the record never carried). The second group cannot reach 0 by any code change. backtest replay --reobserve now classifies each re-opened incident by the recorded findings that re-open it: the finding itself when it interrupts, else everything the open scratch case it joined holds. Re-derived when all of them were re-observed with current code; as recorded, with the first such finding's reason, otherwise. The headline keeps its old total first and adds `— re-derived R (the target), as recorded A`, the re-derived ids, and the as-recorded ids grouped by reason. Rule 17: R + A == N and the two halves cover the re-opened set, asserted before printing (SELF-CHECK FAILED at the top otherwise). Without --reobserve nothing is split, since every finding is replayed as recorded and a split would read as a met target. _reobserve and its helpers now return (finding, as-recorded reason) so the reason travels with each finding instead of being inferred from counters. Three more sensors are re-derived: - behavior: a complete command_preview is re-scored by the current _argv_signals and re-keyed by _argv_case_identity; no signals means dropped. A preview that is elided, clipped at its 240-character budget (the pre-2026-09-19 head-of-argv preview carried no elision mark, and read as the whole command it would drop every harness finding), redacted (redact_sensitive can swallow a token a rule reads), or absent is counted with its reason and replayed as recorded. - process: recorded ancestry (exe paths) was already passed to _grade_binary as parents. A record with no ancestry, where a vouched program could have earned the binary the `supervised` rung, is now counted `field missing: ancestry` and replayed as recorded rather than graded as if it had no parent — the rule the persistence rebuild already follows for an unrecorded field that decides the grade. - hot-dir: re-derived by check_hot_dirs itself over the item's own directory, freshness window opened to the epoch (the record already answered freshness), while the executable is still the recorded bytes; moved-on bytes and a gone item are counted, a directory no longer watched is dropped. Tests: tests/test_backtest_replay_split.py (18). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…into agent/precision-p2a/harness
… it replaced, so 1413 findings fell back to as-recorded HIGH With S7 merged, check_persistence answers a change confined to a payload's bytes with ONE finding per (payload, new sha) — persistence:payload-update, naming every job in referrer_paths — instead of one CHANGED finding per plist. The persistence rebuild still required the current sensor to print the recorded per-plist line, so every recorded payload-only change (the live corpus's shape) fell to "the rebuilt change does not reproduce the record" and replayed as recorded, and #383 #392-#397 (com.aikit.*, run.py 4a1646366adf -> 25ee9a4593c2) re-opened. _persistence_rederives accepts, besides the recorded line, the payload case when it names the recorded plist among its referrer_paths and prints the payload and bytes the recorded line printed; that finding IS the record re-derived, and its severity and custody are taken (the record keeps its identity, as for every same-titled re-grade). A recorded payload case (what the sensor writes from now on) is rebuilt on every job it names by _reobserve_payload_case_answer, each accepted only while it still runs the payload at the recorded new bytes. On the live store: "the rebuilt change does not reproduce the record" goes from 1413 to 0, and every in-window finding of the seven re-derives HIGH/- -> MEDIUM/local-commit or stays MEDIUM/local-commit. tests/test_backtest_replay_persistence.py: the payload test recorded its finding with the current sensor, which now writes the payload case; it records both shapes (the per-job one by holding _payload_update off, as the sensor was before S7) and checks each re-derives, and each is refused once the payload moved on. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Precision plan, pass 2, step P2a: the replay harness. Base:
agent/precision-s0b/integration(PR #59).This branch also contains S7.
origin/agent/precision-s7/integration(PR #61) is merged in at the coordinator's request, because with S7's payload case the harness could no longer re-derive payload-only persistence changes. The last commit fixes that.Why
The done-condition is "0 judged-noise incidents re-alert on replay".
noise re-opened: N of Mmixed two groups. One is incidents the current code still alerts on after re-deriving their evidence. The other is incidents replayed as recorded because the harness could not ask their evidence again: the subject is gone from disk, the sensor is not modelled, or the record lacks a field. No code change can move the second group, so the total could not serve as a target.What changed (harness only:
_backtest_replay,_reobserve*,_replay_*,cmd_backtest_replay)--reobservethe headline says "not split".command_previewis re-scored by the current_argv_signalsand re-keyed by_argv_case_identity. No signals means dropped. Four kinds of preview are counted and replayed as recorded instead:…)[REDACTED]).redact_sensitivecan swallow a token a rule reads.parents; this branch adds a test for it. The sensor records ancestry only while a vouch exists. When a record has no ancestry and a vouched program could have earned the binary thesupervisedrung, the finding is countedfield missing: ancestryand replayed as recorded. It is no longer graded as if it had no parent. This follows the persistence rebuild's existing rule for an unrecorded field that decides the grade.check_hot_dirsitself over the item's own directory, while the executable is still the recorded bytes. The freshness window is opened to the epoch, because the record already answered freshness. If the bytes moved on or the item is gone, the finding is counted. If the directory is no longer watched, the finding is dropped._persistence_rederivesnow also accepts the payload-update finding. It must name the recorded plist inreferrer_pathsand print the recorded payload and bytes; its severity and custody are then taken. A recorded payload case is rebuilt on every job it names.Live replay on the reference Mac (
backtest replay --days 30 --reobserve)Not done / found
field missing: ancestry.parentsto_grade_binary, so thesupervisedrung cannot answer for a beacon under current code, and the beacon records carry no ancestry either. This is S5's "wire_grade_binaryinto the emitters" item.{pid, name}enrichment that the sensor never grades on. It is left alone.Tests
tests/test_backtest_replay_split.pyhas 18 new tests.tests/test_backtest_replay_persistence.pynow records the payload test in both shapes.test_backtest_replay*.py,test_persistence_payload_case.py): 69 passed.python3 -m pytest tests/ -q -p no:cacheprovider: 2212 passed, 6 skipped, 30 xfailed, 41 subtests passed in 1043.27s. The tree sha was a1fa7f24 both before and after the run, and~/.aegis/custody.jsonlwas byte-identical before and after (S7's conftest guard is in the tree; it refused 48 writes from pre-existing tests).🤖 Generated with Claude Code