Skip to content

Incidents opened by fixed code stayed open for a week, because nothing asked the current code again - #69

Merged
opencdlee-dotcom merged 1 commit into
agent/fable-precision/assemblyfrom
agent/precision-s9/healer
Sep 24, 2026
Merged

opencdlee-dotcom merged 1 commit into
agent/fable-precision/assemblyfrom
agent/precision-s9/healer

Conversation

@opencdlee-dotcom

@opencdlee-dotcom opencdlee-dotcom commented Sep 24, 2026 •

Copy link
Copy Markdown
Owner

Precision plan, step S9: current code re-judges what it already opened. Base: agent/fable-precision/assembly.

Why

Every machine exit waits for the sensor to say something new:

  • re-grade reads the newest evidence;
  • re-verify asks about one signature again;
  • cleared-state and removed-file need the sensor to look again.

An incident opened by code that has since been fixed hears none of that unless the same subject is observed again, so it waits out the 7-day age-out. Two live examples:

  • #527: a Spotify risk case built from listener findings recorded unsigned while codesign was not answering. Spotify is Developer ID.
  • #537: a staging plugin-container that the current code grades MEDIUM on its build-output rung.

_rejudge_open_incidents(db, now)

It is registered in record_security_state, after the removed-file exit and before age-out.

What it looks at: each OPEN or ACK incident of kind signal, risk or correlation that is not CRITICAL and was created before this scan.

How it re-judges: every observation finding the incident holds is re-derived through _reobserve, the same path backtest replay --reobserve uses. It uses the live suppression memory and dismissal weights, with the learning period off and the seen-ledger empty. The custody ledger is not written while the ladder is asked.

The incident stays open if:

  • any of its evidence is CRITICAL, attack-defined, or has a never-tolerate fingerprint;
  • any finding was replayed as recorded: its subject is gone from disk, its sensor is not modelled, or the record lacks a field.

Otherwise, by kind:

kind closes only if
signal every finding is dropped, or route_findings routes it below the interrupt tier
risk re-scored by _risk_buckets / _risk_score over the findings inside one RISK_WINDOW, anchored at each moment one was observed (as _accumulate_risk scores it), and no window still crosses the threshold
correlation _unjoinable_chain_resolution says the current join rules would not form the chain; lineage chains are left alone

The risk case needs windowing. Summed whole, #527's evidence is 68 distinct listener signals (one new port per Spotify launch) scoring 11.9. Scored per window, as the accumulator scores it, it peaks at 0.9.

How it closes: the status becomes FALSE_POSITIVE with the resolution re-judged by current code (logic <version>:<code sha12>): <one reason per distinct change> — reopens on new evidence. No dismissals row is written and nothing is added to actions.jsonl, following the other machine exits.

When it runs: at most once an hour, and immediately when _REJUDGE_LOGIC_VERSION or the running aegis.py changes. Each run is capped at 25 incidents or 30 seconds. A capped run logs how many it left, and the next run resumes after the last incident it examined.

Also touched: _accumulate_risk. Its per-entity summing and scoring are extracted into _risk_buckets and _risk_score with no behaviour change, so a pile is re-scored by the same code that opened it. _reobserve_stats() is now shared with _backtest_replay.

Sandboxed real scan of 3ff8ea2 against a copy of the live state

OPEN before the scan (6):
  #526  HIGH     -> OPEN            Suspicious running process
  #527  HIGH     -> FALSE_POSITIVE  Accumulated risk on /Applications/Spotify.app/Contents/MacOS/Spotify (
  #537  HIGH     -> FALSE_POSITIVE  Suspicious running process
  #538  MEDIUM   -> FALSE_POSITIVE  Persistence followed by execution   (chain-leg migration, not S9)
  #539  HIGH     -> OPEN            Suspicious running process
  #540  HIGH     -> OPEN            Suspicious running process
#527: re-judged by current code (logic 1:b9d0f4942208): net-listener Spotify: now classifies developer-id, custody publisher-signed, MEDIUM -> LOW; net-beacon Spotify: no longer emitted; the pile now peaks at 0.9 from 5 signal(s) in any 30-minute window, under the risk threshold (4.0 from 3) — reopens on new evidence
#537: re-judged by current code (logic 1:b9d0f4942208): process plugin-container: custody build-output, HIGH -> MEDIUM (digest: below-floor) — reopens on new evidence
run.log: re-judged 5 open incident(s) with current code (logic 1:b9d0f4942208), closed 2
real custody ledger lines: 151 before, 151 after; dismissals rows 201 = the live store's 201

Why the others stay open:

  • #526: the subject is gone from disk (App Translocation), so its finding is replayed as recorded.
  • #539 and #540: rustup's cargo and rustc still re-derive to HIGH with no rung. They close once the parallel receipt work lands.

This run also opened #543, a CRITICAL "tamper-evidence chain does not verify". It is a race in the sandbox copy, not caused by S9:

  1. The driver copied notary.jsonl at 21:14:57, ending at seq 2569.
  2. The live agent appended seq 2570 at 21:17:05 and anchored it in the OS log.
  3. The sandbox scan read that live anchor against the older copy.

The two earlier runs of the driver did not hit this race.

Tests

tests/test_rejudge_open_incidents.py has 16 tests:

  • A signal incident closes when every finding re-derives below the interrupt tier, and stays open while one still interrupts.
  • It stays open on any finding replayed as recorded, on CRITICAL, attack-defined or never-tolerate evidence, and when opened by this scan.
  • A risk incident closes below the threshold and stays open above it.
  • A chain closes once every leg is explained and stays open while one is not.
  • No dismissals row and no custody ledger row are written.
  • A new interrupting finding opens the case again.
  • It runs at most hourly and immediately on a new logic.
  • The per-run cap resumes where it stopped.
  • record_security_state runs it.

Full suite on this commit's tree (tree sha checked before and after, live custody ledger unchanged): 2 failed, the rest passed. Both failures are TestHostileArgsSeverity::test_benign_interpreter_agent_stays_low and TestExpandedHostileArgs::test_benign_args_stay_low ('INFO' != 'LOW'), an interaction between S5 and S7 that is fixed on the assembly branch by ccc01b0 (demotion never goes below LOW); not caused by this change.

🤖 Generated with Claude Code

…g asked the current code again

Every machine exit waits for the sensor to say something new: re-grade reads
the newest evidence, re-verify re-asks one signature, cleared-state and
removed-file need the sensor to look again. An incident opened by code that
has since been fixed hears none of that unless the same subject is
re-observed, so it waits out the 7-day age-out. Live: #527, a Spotify risk
case built from listener findings recorded `unsigned` while codesign was not
answering (Spotify is Developer ID), and #537, a staging plugin-container the
current code grades MEDIUM on its build-output rung.

_rejudge_open_incidents(db, now) runs in record_security_state after the
removed-file exit. For each OPEN/ACK signal, risk or chain incident created
before this scan, every observation finding it holds is re-derived through
_reobserve, the path `backtest replay --reobserve` scores with, with the live
suppression memory and dismissal weights, the learning period off and the
seen-ledger empty. It stands if any evidence is CRITICAL, attack-defined or
never-tolerate, or if any finding is replayed as recorded (gone from disk,
unmodelled sensor, missing field). Then, by kind:

- signal: every finding dropped, or routed below the interrupt tier by
  route_findings;
- risk: re-scored by _risk_buckets / _risk_score over the findings inside one
  RISK_WINDOW at each moment one was observed, as _accumulate_risk scores
  it, and no window may still cross. Summed whole, a program that picks a
  new port per launch is 68 distinct signals (#527); per window it peaks at
  0.9;
- correlation: a chain the current join rules would not form
  (_unjoinable_chain_resolution); lineage chains are left alone.

A closed incident goes to FALSE_POSITIVE with `re-judged by current code
(logic <version>:<code sha12>): <one reason per distinct change> — reopens
on new evidence`. It writes no dismissals row and no custody ledger row (the
ladder is asked with _custody_remember off). It runs at most hourly, at once
when _REJUDGE_LOGIC_VERSION or the running aegis.py changes, and is capped by
_REJUDGE_MAX_INCIDENTS (25) and _REJUDGE_BUDGET (30 s) per run. A capped run
logs how many it left, and the next resumes after the last incident it
examined.

Also touched: _accumulate_risk. Its per-entity summing and scoring are
extracted as _risk_buckets / _risk_score, with no behaviour change, so a pile
is re-scored by the code that opened it. _backtest_replay's counter setup is
now _reobserve_stats(), shared with the healer.

Sandboxed real scan of this code against a copy of the live state: #527 and
#537 closed as re-judged. #526 (subject gone), #539 and #540 (rustup, still
HIGH with no rung) stay. No incident was opened, no outward effect was taken,
and the real custody ledger read 151 lines before and after.

Tests: tests/test_rejudge_open_incidents.py (16).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant