Skip to content

fix(static): let validated Python string literals own their quotes in marker parsing - #784

Merged
rng1995 merged 2 commits into
mainfrom
naren/fix-python-literal-marker-ownership
Oct 7, 2026
Merged

rng1995 merged 2 commits into
mainfrom
naren/fix-python-literal-marker-ownership

Conversation

@rng1995

@rng1995 rng1995 commented Oct 7, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

A Python file whose string literals contain a removal verb, for example assert "please omit --hours" in out, is marked partial with obfuscated_instruction_text when more than 8 KB of code follows. Every SKILL.md reference to it then becomes a HIGH AE1, and downstream gates fail even though the file is an ordinary test or helper.

This PR treats the closing quote of a tokenizer-proven Python string literal like the closing quote of a validated JSON string: syntax that ends a literal, never a marker opener for the loose fallback directive header. It follows up #777 (comparison operators) and #778 (Python string/comment ownership for the shell parser), and reuses #778's verified token spans.

Customer impact

Reported internally against SkillSpector 2.12.0 (via SkillEvaluator 1.6.0). After #777 and #778, one internal skill still has a single obfuscated_instruction_text row on its 106 KB pytest module (three assertions such as "... omit --hours"). That row alone keeps its release gate at a critical "scan did not execute reliably". With this PR the module is fully inspected (coverage 95.8% → 100%), and the scan produces the same 30 findings as before.

Root cause

In build_declared_marker_views:

  1. The fallback quoted header (omit … ") matches the closing quote of the literal that contains the verb.
  2. The "marker" then runs through code to the next literal's opening quote. It contains whitespace, so the directive is unsupported and inactive.
  3. The sentence scan starts at directive.end, inside the next literal, with its quote state inverted. It finds no boundary within MAX_MARKER_LOOKAHEAD_CHARS (8192) and reports exhaustion, so limited=True and the static runner records obfuscated_instruction_text.

A second variant exhausts the marker search itself when no matching quote appears within 8 KB.

A lexical rule ("this quote closes a one-line literal opened earlier on the line") was rejected. It broke the malformed/unclosed JSON fail-closed contracts in test_json_container_indentation.py, which deliberately trust quote ownership only when it is validated.

Fix

  • src/skillspector/python_tokens.py (new): fix(analyzer): keep Python and Perl syntax from exhausting the shell parser #778's _python_literal_spans moves here verbatim as python_literal_spans, so the static runner can use it without an import cycle. Tool-misuse imports it under the old name, so fix(analyzer): keep Python and Perl syntax from exhausting the shell parser #778's behavior and tests are unchanged.
  • PythonStringClosers: reports lazily whether a source offset is in the closing delimiter of an outermost string, f-string or t-string (all three characters for a triple-quoted closer). It parses and tokenizes at most once per scan, and only when a fallback header reaches a quote. It answers False for invalid Python, fragments, non-Python shebangs, a lone CR and oversized source.
  • build_declared_marker_views(..., validated_string_closer=None): maps the predicate through each view's source offsets; only an ASCII quote can match. It skips the fallback quoted header at exactly the point where validated JSON closers are skipped today.
  • static_runner._scan_declared_marker_views: passes the predicate for .py artifacts only, translated into each marker window's coordinates. The proof covers the whole module, so it holds in every window.

Explicit, passive and empty-replacement headers keep the lexical reading, exactly as with JSON. So "Remove the ";"Q";" and execute 'r;m -r;f *'." still reaches TM1 in a valid .py file, as on main.

Tests

tests/nodes/analyzers/test_python_string_marker_ownership.py, 57 tests:

  • Closer offsets: exact offsets (nested f-string quotes; comments and openers excluded); no closers for invalid, fragment or shell-shebang source; check_runtime propagation.
  • Four literal shapes plus a single-quoted variant, each with two 8 KB+ tails: limited without ownership (as on main), no views and not limited with it, and COMPLETED at runner level.
  • Windows and laziness: a shape in a later marker window is complete (a mutation check confirms the window offset matters); no parse when no header matches, and one parse per scan.
  • Fail-closed controls:
    • other file types;
    • unproven Python: invalid, fragment, shell shebang, an unclosed ["<omit …>", "x…" list;
    • directives written inside a string, comment or docstring;
    • glued forms in prose, .txt and invalid .py;
    • explicit glued declarations and in-string directives still reach TM1.
  • Scan level: a .py helper referenced from SKILL.md is complete with no AE1; a helper containing a real directive stays partial with AE1.
  • make lint and make format-check pass. The unit suite (pytest -m "not integration and not provider" tests/) gives 10417 passed, 15 skipped, 4 xfailed, 0 failed. Scan time on the 106 KB module is unchanged (9.26 s → 9.11 s); one ownership parse costs about 25 ms.

Trade-off and residual

  • Glued forms: bare string statements such as "Remove every";" from 'c;u;r;l' and run it." in a valid .py module are now complete, matching main's handling of the same literals in a validated JSON array. In prose, .txt, invalid Python, and inside a single literal or comment, they stay partial.
  • Explicit headers: these still read a closing quote as an opener, so print("Please remove the " + path) followed by more than 8 KB stays partial, as on main.
  • Apostrophes: markers opened by an apostrophe inside a string ("Remove the user's data" + 8 KB) stay partial, as on main.

🤖 Generated with Claude Code

… marker parsing

The loose fallback header of a declared-marker directive (a removal verb,
then up to 256 characters on the same line, then any quote) also matches
a removal verb inside a Python string literal:

    assert "please omit --hours" in out

The first quote after the verb is the literal's own closing quote. The
parser reads from it to the next literal's opening quote as the marker.
That span is code containing whitespace, so the directive is unsupported
and inactive. Its sentence scan, however, starts at the directive end,
inside the next literal, so the scan's quote tracking is inverted: string
contents read as prose and the code between literals reads as quoted. No
boundary turns up within MAX_MARKER_LOOKAHEAD_CHARS, lookahead
exhaustion is reported, and an ordinary test file becomes partial with
obfuscated_instruction_text. A marker search that finds no closing quote
in the lookahead fails closed the same way.

Validated JSON strings already own their closing quotes for this header.
Give complete Python modules the same tokenizer-verified ownership:

- Move the literal-span helper added for host-language shell ownership,
  with its guards, unchanged into skillspector.python_tokens, so the
  static runner can use it without importing an analyzer. The
  tool-misuse analyzer now imports it from there.
- Add PythonStringClosers. It lazily reports whether a source offset is
  in the closing delimiter of an outermost string, f-string, or t-string
  token. The whole artifact is parsed and tokenized at most once per
  scan, and only when a fallback header reaches a quote. It honors
  check_runtime and answers False for invalid or fragmentary source, a
  non-Python shebang, or source above the AST size bound.
- build_declared_marker_views accepts validated_string_closer and maps
  it through each view's source offsets. A fallback quoted header whose
  quote is such a closer is skipped exactly where a validated JSON
  closer is skipped. The static runner passes it only for .py artifacts,
  translated into each marker window's coordinates.

Explicit declared-marker headers, passive and empty-replacement forms,
marker declarations written inside one string or comment, invalid or
fragmentary Python, and every other file type keep the lexical reading.

Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

@rng1995 rng1995 left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review: the approach mirrors the validated-JSON ownership exactly. Only the loose fallback header defers to proven closers, explicit headers keep TM1, and window offsets map back to module coordinates (the mutation check is a nice touch). Moving the span helper to python_tokens.py verbatim keeps #778 stable. Two P3 items, no blockers.

Comment thread src/skillspector/nodes/analyzers/static_runner.py
Comment thread tests/nodes/analyzers/test_python_string_marker_ownership.py
A static scan of a .py artifact asked python_literal_spans for the same
module twice: once from the declared-marker pass (PythonStringClosers)
and once from the tool-misuse shell parser (_PythonSourceOwnership).
Each call ran ast.parse and tokenize, and both counted against the
per-artifact runtime budget.

- python_literal_spans now memoizes its two most recent results,
  including None. Entries are keyed by (len, hash) and confirmed by
  string equality. A dedicated lock guards only the entries, so a lookup
  never waits behind another thread's parse under the warnings lock.
  Source above the AST size bound is never stored.
- A hit still calls check_runtime once. A runtime-limit exception
  propagates before the store, so an interrupted proof is never kept.
- Spans are returned as tuples, since callers now share them.

Tests:
- one parse per scan when both consumers ask, observed by a spy on the
  uncached parse;
- proven and unproven results are reused; sources forced into one
  length/hash bucket keep their own spans; a hit honors the runtime
  check; an interrupted parse is redone; threads over more sources than
  entries each get their own spans;
- a scripts/check.py glued-statement case is pinned COMPLETED, with the
  deliberate parity with validated JSON arrays recorded beside it;
- both ownership test modules clear the memo per test, and the
  concurrent warnings-filter test uses a distinct source per call so
  every call still parses.

Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@rng1995
rng1995 merged commit 3a1ceee into main Oct 7, 2026
6 checks passed
@rng1995
rng1995 deleted the naren/fix-python-literal-marker-ownership branch October 7, 2026 03:42
rng1995 added a commit to chrisknvidia/SkillSpector that referenced this pull request Oct 8, 2026
Main's NVIDIA#784 test expects the tool-misuse ledger event for an invalid
Python file to carry obfuscated_instruction_text. This PR makes tool
misuse parse Python for TM1 reconciliation, so the runner's
higher-precedence syntax_error now names that partial event. Keep the
fail-closed contract: assert the marker reading through the lexical-only
anti-refusal analyzer and keep tool misuse partial with syntax_error.

Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
rng1995 added a commit that referenced this pull request Oct 8, 2026
* test(analyzer): define shell truthiness core contract

Signed-off-by: Christopher Kevin <christopherk@nvidia.com>

* fix(analyzer): detect bound shell truthiness

Signed-off-by: Christopher Kevin <christopherk@nvidia.com>

* fix(analyzer): harden bound shell facts

Signed-off-by: Christopher Kevin <christopherk@nvidia.com>

* fix(analyzer): honor shell argument evaluation order

Signed-off-by: Christopher Kevin <christopherk@nvidia.com>

* fix(analyzer): invalidate effectful subprocess receivers

Signed-off-by: Christopher Kevin <christopherk@nvidia.com>

* fix(analyzer): invalidate receiver trust after generic calls

Signed-off-by: Christopher Kevin <christopherk@nvidia.com>

* fix(analyzer): reconcile bound shell ownership

Signed-off-by: Christopher Kevin <christopherk@nvidia.com>

* fix(analyzer): reconcile true-prefixed shell ownership

Signed-off-by: Christopher Kevin <christopherk@nvidia.com>

* fix(analyzer): invalidate nested RHS receiver effects

Signed-off-by: Christopher Kevin <christopherk@nvidia.com>

* fix(analyzer): invalidate unsupported eager receiver effects

Signed-off-by: Christopher Kevin <christopherk@nvidia.com>

* fix(analyzer): track eager receiver state

Signed-off-by: Christopher Kevin <christopherk@nvidia.com>

* fix(analyzer): preserve lexical TM1 when dataflow abstains

Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>

* fix(analyzer): require a retained companion owner

Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>

* fix(analyzer): preserve execution scope in binding evidence

Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>

* test(analyzer): distinguish deferred abstention from lexical signal

Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>

* style(analyzer): simplify binding event deduplication

Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>

* fix: preserve runtime and retained shell evidence

Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>

* fix(analyzer): preserve called-slot identity and cached method evidence

Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>

* fix: preserve cached subprocess slot state across scoped imports

Signed-off-by: Christopher Kevin <christopherk@nvidia.com>

* fix(analyzer): reconcile explicit cached-slot evidence across scopes

Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>

* fix: abstain from cached slot proof after eager effects

Signed-off-by: Christopher Kevin <christopherk@nvidia.com>

* test: retain lexical findings after unknown cached-slot effects

Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>

* fix: bound direct called-slot evidence to known execution order

Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>

* fix: limit direct slot proof to single assignment targets

Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>

* fix: invalidate slot evidence on protocols and unsafe releases

Signed-off-by: Christopher Kevin <christopherk@nvidia.com>

* fix: keep legacy shadow proof isolated from later effects

Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>

* fix: abstain from cached replacement proof after unknown effects

Signed-off-by: Christopher Kevin <christopherk@nvidia.com>

* fix: retain native callable detection across cached slot stores

Signed-off-by: Christopher Kevin <christopherk@nvidia.com>

* fix: retain shell findings for native receiver assignments

Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>

* fix: preserve native Popen replacement warnings

Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>

* fix(analyzer): require proven replacement before revoking lexical TM1

Address the three remaining review findings on lexical TM1 reconciliation,
where a bound shell=True call still ran natively but lost its HIGH finding:

- Untaken expression stores: a walrus in a short-circuited BoolOp operand,
  an untaken conditional-expression arm, or a comprehension/generator body
  is no longer recorded as replacement evidence. Only constant operands
  and tests prove that the store executed.
- Identity-preserving stores: unpacked targets are paired with their RHS
  elements. Self-stores are transparent, truthy constants keep the flag,
  and native receiver aliases (saved = subprocess, import subprocess as
  saved, Popen = subprocess.Popen) keep the native receiver.
- Deferred bodies: an outer store only proves replacement when every
  invocation the owner scope can reach sees it. Any reference to the
  function, method or class (compound statements, main guards, callbacks,
  other bodies) can invoke it from that point on, and decorators or
  lambdas can run once defined. This also covers a re-import before a
  later invocation.

The release-trigger test for a called function now matches its
module-level twin: the companion abstains and the lexical HIGH remains.

Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: account for TM1 Python parsing in marker ownership ledger

Main's #784 test expects the tool-misuse ledger event for an invalid
Python file to carry obfuscated_instruction_text. This PR makes tool
misuse parse Python for TM1 reconciliation, so the runner's
higher-precedence syntax_error now names that partial event. Keep the
fail-closed contract: assert the marker reading through the lexical-only
anti-refusal analyzer and keep tool misuse partial with syntax_error.

Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Signed-off-by: Christopher Kevin <christopherk@nvidia.com>
Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-authored-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>
Co-authored-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
rng1995 added a commit to chrisknvidia/SkillSpector that referenced this pull request Oct 8, 2026
Main's NVIDIA#784 test expects the tool-misuse ledger event for an invalid
Python file to carry obfuscated_instruction_text. This PR makes tool
misuse parse Python for TM1 reconciliation, so the runner's
higher-precedence syntax_error now names that partial event. Keep the
fail-closed contract: assert the marker reading through the lexical-only
anti-refusal analyzer and keep tool misuse partial with syntax_error.

Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
rng1995 added a commit to chrisknvidia/SkillSpector that referenced this pull request Oct 8, 2026
Main's NVIDIA#784 test expects the tool-misuse ledger event for an invalid
Python file to carry obfuscated_instruction_text. This PR makes tool
misuse parse Python for TM1 reconciliation, so the runner's
higher-precedence syntax_error now names that partial event. Keep the
fail-closed contract: assert the marker reading through the lexical-only
anti-refusal analyzer and keep tool misuse partial with syntax_error.

Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
rng1995 added a commit that referenced this pull request Oct 9, 2026
* fix(analyzer): stabilize TM1 window identity

Signed-off-by: Christopher Kevin <christopherk@nvidia.com>

* fix(analyzer): track eager receiver side effects

Signed-off-by: Christopher Kevin <christopherk@nvidia.com>

* fix(analyzer): preserve lexical TM1 when dataflow abstains

Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>

* fix(analyzer): require a retained companion owner

Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>

* fix(analyzer): preserve execution scope in binding evidence

Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>

* style(analyzer): simplify binding event deduplication

Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>

* fix: preserve runtime and retained shell evidence

Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>

* fix(analyzer): preserve called-slot identity and cached method evidence

Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>

* fix: preserve cached subprocess slot state across scoped imports

Signed-off-by: Christopher Kevin <christopherk@nvidia.com>

* fix(analyzer): reconcile explicit cached-slot evidence across scopes

Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>

* fix: abstain from cached slot proof after eager effects

Signed-off-by: Christopher Kevin <christopherk@nvidia.com>

* test: retain lexical findings after unknown cached-slot effects

Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>

* fix: bound direct called-slot evidence to known execution order

Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>

* fix: limit direct slot proof to single assignment targets

Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>

* fix: invalidate slot evidence on protocols and unsafe releases

Signed-off-by: Christopher Kevin <christopherk@nvidia.com>

* fix: keep legacy shadow proof isolated from later effects

Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>

* fix: abstain from cached replacement proof after unknown effects

Signed-off-by: Christopher Kevin <christopherk@nvidia.com>

* fix: retain native callable detection across cached slot stores

Signed-off-by: Christopher Kevin <christopherk@nvidia.com>

* fix: retain shell findings for native receiver assignments

Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>

* fix: preserve native Popen replacement warnings

Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>

* fix(analyzer): require proven replacement before revoking lexical TM1

Address the three remaining review findings on lexical TM1 reconciliation,
where a bound shell=True call still ran natively but lost its HIGH finding:

- Untaken expression stores: a walrus in a short-circuited BoolOp operand,
  an untaken conditional-expression arm, or a comprehension/generator body
  is no longer recorded as replacement evidence. Only constant operands
  and tests prove that the store executed.
- Identity-preserving stores: unpacked targets are paired with their RHS
  elements. Self-stores are transparent, truthy constants keep the flag,
  and native receiver aliases (saved = subprocess, import subprocess as
  saved, Popen = subprocess.Popen) keep the native receiver.
- Deferred bodies: an outer store only proves replacement when every
  invocation the owner scope can reach sees it. Any reference to the
  function, method or class (compound statements, main guards, callbacks,
  other bodies) can invoke it from that point on, and decorators or
  lambdas can run once defined. This also covers a re-import before a
  later invocation.

The release-trigger test for a called function now matches its
module-level twin: the companion abstains and the lexical HIGH remains.

Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: account for TM1 Python parsing in marker ownership ledger

Main's #784 test expects the tool-misuse ledger event for an invalid
Python file to carry obfuscated_instruction_text. This PR makes tool
misuse parse Python for TM1 reconciliation, so the runner's
higher-precedence syntax_error now names that partial event. Keep the
fail-closed contract: assert the marker reading through the lexical-only
anti-refusal analyzer and keep tool misuse partial with syntax_error.

Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(analyzer): keep one TM1 owner for true-prefixed names without AST proof

Port the six true-prefixed TM1 regression tests from #577. This branch
already carries test_true_direct_calls_on_one_line_keep_distinct_locations
as test_true_direct_calls_around_safe_comprehension_keep_distinct_locations,
so only the other five are added. Two of them reported the same call twice
on this branch:

- A variable window for a name spelled `true` (in any case) duplicated
  the case-insensitive direct `shell=True` owner at the same call. AST
  reconciliation removes that window only when the dataflow companion
  takes the call, so a file that does not parse, or a call after an
  unknown receiver effect, kept both findings. coalesce_path_findings
  now drops such a window when a direct lexical owner starts at the
  same call. A window for a call whose direct candidate was folded into
  an identical same-line call still reports that call.
- A call-anchored window for a longer true-prefixed name (`true_value`)
  used its raw text as identity, so the raw and normalized security
  views of one call did not collapse, and reconciliation then moved both
  to the call. The window now takes its identity from the normalized
  view text, as the direct owner already does.

Also cover both root causes with a test for a nested `true` call after
an unknown receiver effect and a normalized `true_value` call in a file
that does not parse.

Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(analyzer): report each true-prefixed shell call that #577 reports

The `\b`-bounded direct pattern does not match `shell=true_value`, so on
this branch such a call was reported only through a variable window. That
window pairs an assignment with a single call, skips an assignment
inside an earlier window, and was dropped outside Python. Compared with
#577 this lost:

- calls in JavaScript and Markdown files, fenced or plain;
- a later call that the AST companion abstains on, for example after an
  unknown receiver effect;
- all but the last same-line call in a file that does not parse, and the
  second of two identical same-line `shell=true` calls outside Python,
  which the direct candidate key folds into one;
- a call whose assignment sits inside the window of an earlier one.

_variable_shell_matches now pairs every assignment of a true-prefixed
name with each call its bounded window reaches, and the nearest
assignment owns each call. Candidate generation, analysis and AST
reconciliation share that list, and a window's candidate identity uses
its full text, so windows to different calls do not collapse on a shared
200-character preview. Outside Python, call-anchored windows for
true-prefixed names are kept, and coalescing still drops one when a
direct owner reports the same call. With an AST, each added window
passes the same visibility, counterevidence, cached replacement and
companion checks as before. Windows for other names are unchanged.

Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Signed-off-by: Christopher Kevin <christopherk@nvidia.com>
Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-authored-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>
Co-authored-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
rng1995 added a commit that referenced this pull request Oct 9, 2026
* fix(analyzer): stabilize TM1 window identity

Signed-off-by: Christopher Kevin <christopherk@nvidia.com>

* fix(analyzer): track eager receiver side effects

Signed-off-by: Christopher Kevin <christopherk@nvidia.com>

* feat(python): refresh execution surfaces on current main

Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>

* fix(analyzer): preserve lexical TM1 when dataflow abstains

Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>

* fix(analyzer): require a retained companion owner

Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>

* fix(analyzer): preserve execution scope in binding evidence

Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>

* fix(analyzer): preserve lexical TM1 across conservative Python analysis

Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>

* style(analyzer): simplify binding event deduplication

Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>

* fix: preserve runtime and retained shell evidence

Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>

* fix: preserve runtime and retained shell evidence

Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>

* fix(analyzer): preserve called-slot identity and cached method evidence

Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>

* fix(analyzer): preserve called-slot identity and cached method evidence

Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>

* fix: preserve cached subprocess slot state across scoped imports

Signed-off-by: Christopher Kevin <christopherk@nvidia.com>

* fix: preserve cached subprocess slot state across scoped imports

Signed-off-by: Christopher Kevin <christopherk@nvidia.com>

* fix(analyzer): reconcile explicit cached-slot evidence across scopes

Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>

* fix(analyzer): reconcile explicit cached-slot evidence across scopes

Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>

* fix: abstain from cached slot proof after eager effects

Signed-off-by: Christopher Kevin <christopherk@nvidia.com>

* fix: abstain from cached slot proof after eager effects

Signed-off-by: Christopher Kevin <christopherk@nvidia.com>

* test: retain lexical findings after unknown cached-slot effects

Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>

* test: retain lexical findings after unknown cached-slot effects

Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>

* fix: bound direct called-slot evidence to known execution order

Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>

* fix: bound direct called-slot evidence to known execution order

Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>

* fix: limit direct slot proof to single assignment targets

Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>

* fix: limit direct slot proof to single assignment targets

Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>

* fix: invalidate slot evidence on protocols and unsafe releases

Signed-off-by: Christopher Kevin <christopherk@nvidia.com>

* fix: invalidate slot evidence on protocols and unsafe releases

Signed-off-by: Christopher Kevin <christopherk@nvidia.com>

* fix: keep legacy shadow proof isolated from later effects

Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>

* fix: keep legacy shadow proof isolated from later effects

Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>

* fix: abstain from cached replacement proof after unknown effects

Signed-off-by: Christopher Kevin <christopherk@nvidia.com>

* fix: abstain from cached replacement proof after unknown effects

Signed-off-by: Christopher Kevin <christopherk@nvidia.com>

* fix: retain native callable detection across cached slot stores

Signed-off-by: Christopher Kevin <christopherk@nvidia.com>

* fix: retain native callable detection across cached slot stores

Signed-off-by: Christopher Kevin <christopherk@nvidia.com>

* fix: retain shell findings for native receiver assignments

Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>

* fix: retain shell findings for native receiver assignments

Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>

* fix: preserve native Popen replacement warnings

Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>

* fix: preserve native Popen replacement warnings

Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>

* test: retain occurrence metadata in prose shell reports

Signed-off-by: Christopher Kevin <christopherk@nvidia.com>

* fix(analyzer): require proven replacement before revoking lexical TM1

Address the three remaining review findings on lexical TM1 reconciliation,
where a bound shell=True call still ran natively but lost its HIGH finding:

- Untaken expression stores: a walrus in a short-circuited BoolOp operand,
  an untaken conditional-expression arm, or a comprehension/generator body
  is no longer recorded as replacement evidence. Only constant operands
  and tests prove that the store executed.
- Identity-preserving stores: unpacked targets are paired with their RHS
  elements. Self-stores are transparent, truthy constants keep the flag,
  and native receiver aliases (saved = subprocess, import subprocess as
  saved, Popen = subprocess.Popen) keep the native receiver.
- Deferred bodies: an outer store only proves replacement when every
  invocation the owner scope can reach sees it. Any reference to the
  function, method or class (compound statements, main guards, callbacks,
  other bodies) can invoke it from that point on, and decorators or
  lambdas can run once defined. This also covers a re-import before a
  later invocation.

The release-trigger test for a called function now matches its
module-level twin: the companion abstains and the lexical HIGH remains.

Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(analyzer): require proven replacement before revoking lexical TM1

Address the three remaining review findings on lexical TM1 reconciliation,
where a bound shell=True call still ran natively but lost its HIGH finding:

- Untaken expression stores: a walrus in a short-circuited BoolOp operand,
  an untaken conditional-expression arm, or a comprehension/generator body
  is no longer recorded as replacement evidence. Only constant operands
  and tests prove that the store executed.
- Identity-preserving stores: unpacked targets are paired with their RHS
  elements. Self-stores are transparent, truthy constants keep the flag,
  and native receiver aliases (saved = subprocess, import subprocess as
  saved, Popen = subprocess.Popen) keep the native receiver.
- Deferred bodies: an outer store only proves replacement when every
  invocation the owner scope can reach sees it. Any reference to the
  function, method or class (compound statements, main guards, callbacks,
  other bodies) can invoke it from that point on, and decorators or
  lambdas can run once defined. This also covers a re-import before a
  later invocation.

The release-trigger test for a called function now matches its
module-level twin: the companion abstains and the lexical HIGH remains.

Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: account for TM1 Python parsing in marker ownership ledger

Main's #784 test expects the tool-misuse ledger event for an invalid
Python file to carry obfuscated_instruction_text. This PR makes tool
misuse parse Python for TM1 reconciliation, so the runner's
higher-precedence syntax_error now names that partial event. Keep the
fail-closed contract: assert the marker reading through the lexical-only
anti-refusal analyzer and keep tool misuse partial with syntax_error.

Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: account for TM1 Python parsing in marker ownership ledger

Main's #784 test expects the tool-misuse ledger event for an invalid
Python file to carry obfuscated_instruction_text. This PR makes tool
misuse parse Python for TM1 reconciliation, so the runner's
higher-precedence syntax_error now names that partial event. Keep the
fail-closed contract: assert the marker reading through the lexical-only
anti-refusal analyzer and keep tool misuse partial with syntax_error.

Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(analyzer): keep one TM1 owner for true-prefixed names without AST proof

Port the six true-prefixed TM1 regression tests from #577. This branch
already carries test_true_direct_calls_on_one_line_keep_distinct_locations
unchanged, so only the other five are added. Two of them reported the
same call twice on this branch:

- A variable window for a name spelled `true` (in any case) duplicated
  the case-insensitive direct `shell=True` owner at the same call. AST
  reconciliation removes that window only when the dataflow companion
  takes the call, so a file that does not parse, or a call after an
  unknown receiver effect, kept both findings. coalesce_path_findings
  now drops such a window when a direct lexical owner starts at the
  same call. A window for a call whose direct candidate was folded into
  an identical same-line call still reports that call.
- A call-anchored window for a longer true-prefixed name (`true_value`)
  used its raw text as identity, so the raw and normalized security
  views of one call did not collapse, and reconciliation then moved both
  to the call. The window now takes its identity from the normalized
  view text, as the direct owner already does.

Also cover both root causes with a test for a nested `true` call after
an unknown receiver effect and a normalized `true_value` call in a file
that does not parse.

Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(analyzer): keep one TM1 owner for true-prefixed names without AST proof

Port the six true-prefixed TM1 regression tests from #577. This branch
already carries test_true_direct_calls_on_one_line_keep_distinct_locations
as test_true_direct_calls_around_safe_comprehension_keep_distinct_locations,
so only the other five are added. Two of them reported the same call twice
on this branch:

- A variable window for a name spelled `true` (in any case) duplicated
  the case-insensitive direct `shell=True` owner at the same call. AST
  reconciliation removes that window only when the dataflow companion
  takes the call, so a file that does not parse, or a call after an
  unknown receiver effect, kept both findings. coalesce_path_findings
  now drops such a window when a direct lexical owner starts at the
  same call. A window for a call whose direct candidate was folded into
  an identical same-line call still reports that call.
- A call-anchored window for a longer true-prefixed name (`true_value`)
  used its raw text as identity, so the raw and normalized security
  views of one call did not collapse, and reconciliation then moved both
  to the call. The window now takes its identity from the normalized
  view text, as the direct owner already does.

Also cover both root causes with a test for a nested `true` call after
an unknown receiver effect and a normalized `true_value` call in a file
that does not parse.

Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(analyzer): report each true-prefixed shell call that #577 reports

The `\b`-bounded direct pattern does not match `shell=true_value`, so on
this branch such a call was reported only through a variable window. That
window pairs an assignment with a single call, skips an assignment
inside an earlier window, and was dropped outside Python. Compared with
#577 this lost:

- calls in JavaScript and Markdown files, fenced or plain;
- a later call that the AST companion abstains on, for example after an
  unknown receiver effect;
- all but the last same-line call in a file that does not parse, and the
  second of two identical same-line `shell=true` calls outside Python,
  which the direct candidate key folds into one;
- a call whose assignment sits inside the window of an earlier one.

_variable_shell_matches now pairs every assignment of a true-prefixed
name with each call its bounded window reaches, and the nearest
assignment owns each call. Candidate generation, analysis and AST
reconciliation share that list, and a window's candidate identity uses
its full text, so windows to different calls do not collapse on a shared
200-character preview. Outside Python, call-anchored windows for
true-prefixed names are kept, and coalescing still drops one when a
direct owner reports the same call. With an AST, each added window
passes the same visibility, counterevidence, cached replacement and
companion checks as before. Windows for other names are unchanged.

Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(analyzer): report each true-prefixed shell call that #577 reports

The `\b`-bounded direct pattern does not match `shell=true_value`, so on
this branch such a call was reported only through a variable window. That
window pairs an assignment with a single call, skips an assignment
inside an earlier window, and was dropped outside Python. Compared with
#577 this lost:

- calls in JavaScript and Markdown files, fenced or plain;
- a later call that the AST companion abstains on, for example after an
  unknown receiver effect;
- all but the last same-line call in a file that does not parse, and the
  second of two identical same-line `shell=true` calls outside Python,
  which the direct candidate key folds into one;
- a call whose assignment sits inside the window of an earlier one.

_variable_shell_matches now pairs every assignment of a true-prefixed
name with each call its bounded window reaches, and the nearest
assignment owns each call. Candidate generation, analysis and AST
reconciliation share that list, and a window's candidate identity uses
its full text, so windows to different calls do not collapse on a shared
200-character preview. Outside Python, call-anchored windows for
true-prefixed names are kept, and coalescing still drops one when a
direct owner reports the same call. With an AST, each added window
passes the same visibility, counterevidence, cached replacement and
companion checks as before. Windows for other names are unchanged.

Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(supply-chain): opt into Python source-type propagation

SC2's literal-XOR decoding and SC3 branch on file_type, but the
supply-chain analyzer was not in the USES_PYTHON_SOURCE_TYPE set, so
Python executed through an extensionless shebang file or a shebang
Markdown file received its suffix type and silently lost the decoded
SC2 HIGH while completeness stayed true. Opt the module in and add
graph-level parity tests for .py, .pyw, extensionless and .md surfaces.

Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(build-context): decode PEP 263 primary Python before UTF-8 rejection

A selected primary file (direct .py/.pyw or extensionless Python
shebang) declaring a supported PEP 263 encoding such as latin-1 was
rejected as unsupported_primary_content because the generic primary
check only accepts UTF-8. The later source-decoding pass then decoded
it correctly, but the artifact stayed FAILED with a fatal event and 0%
coverage. Waive the generic UTF-8 rejection when the bytes classify as
Python and decode under their declared encoding; an unknown codec or
bytes invalid for the declared codec remain fatal, and format checks
(BOMs, archives, NUL density) are unchanged. The waiver is evaluated
lazily and not after the shared deadline, so no post-cache Python work
starts once the budget is spent.

Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Signed-off-by: Christopher Kevin <christopherk@nvidia.com>
Signed-off-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-authored-by: Christopher Kevin <256191862+chrisknvidia@users.noreply.github.com>
Co-authored-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant