Repository navigation
Static v2.11.1: distinguish inert security-test strings and taxonomy labels from executable behavior in findings #523
Description
Activity
Opened #530 to address the
uvicorn/gunicornSC6 false positive, with regression tests. This only addresses the typosquatting observation; the other findings in this issue remain out of scope.A few more examples of findings that infer behavior from vocabulary alone:
API reference: https://api.example.com/emits E1 External Transmission although it supplies no send operation.- Explicit sibling-skill workflow references are reported as AS3 snooping. The cited public example names the skills needed by that workflow; it does not establish a search for secrets or an authorization violation.
- A function assigning
prompt = "A blue circle"and returning that local value emits P6. The variable name alone does not establish protected system-prompt extraction. Using Chrome exposes cookies. Use a separate Chrome profile.emits HIGH YR1 through the built-ininfo_stealersignature. This is safety advice, not credential theft.
The E1/P6/YARA examples reproduce in direct detector tests. The sibling-skill case was checked against the full source linked below. A skill can still describe a risky capability, and its own claim of authorization is not enough to establish permission.
Regression coverage should pair these with instructions to send data, search credential directories, reveal the system prompt or steal credentials. Documentation, comments and test files still need inspection.
PR #658 covers narrower prohibition and install cases; PR #608 covers ransomware-payment prose. The examples above need additional handling. E2 environment pass-through is tracked separately in #441 / PR #492.
Related examples in public skills:
- sibling-skill workflow: explicit workflow references reported as snooping.
- skill-authoring workflow: authoring the requested skill described as unrequested self-modification.
- optional review-provider selection: an explicitly selected optional review path described as a covert provider change.
- security-hardening guidance: combined threat-model and disable-control vocabulary reported as a hack-tool indicator despite the local prohibitions.
- API tool-schema migration reference: combined tool-schema vocabulary reported as a hack-tool indicator.
- tool configuration instructions: configuration/profile setup reported as a malware/backdoor indicator.
Some examples do include real capabilities or configuration changes. Findings should describe those accurately and distinguish them from claims of covert or malicious behavior that the source does not support.
Relevant code
static_patterns_data_exfiltration.py:45, static_patterns_agent_snooping.py:107, static_patterns_system_prompt_leakage.py:58, static_yara.py:445.
Merge-state update checked on 2026-10-08: PR #608 has merged its narrow ransomware-payment-prose fix. PR #787 also merged handling for proven path exclusions in ignore files.
Keeping this broader report open. Fresh offline detector checks on main
8f4ca7d6still report E1 for a bare API-reference URL and P6 for a function returning an innocuous local prompt value. The inert-input, taxonomy, and evidence-context concerns are not collectively resolved by these narrow fixes. Existing unrelated findings are not being suppressed to mark the issue complete.
We integrated SkillSpector as an attributed independent evidence provider in
AI SAFE². Thank you for exposing coverage and degraded-analysis status explicitly;
those fields are important to our downstream review process.
Environment and scope
704bc9544260c2f41222dc0f92982521709496ab.scan TARGET --format json --no-llm.network access: supply-chain lookups can contact OSV.
skills/tree at6455eec76caec259ba3f9b902a4ab24f27b34263.upstream revision is affected.
Observed result
The tracked-source snapshot produced 50 issues and a risk assessment of
100 / CRITICAL / DO_NOT_INSTALL. A separate development checkout also produced
50 issues. Our clean and inert-hostile controls produced zero and seven issues,
respectively. These counts are not a calibrated precision or recall benchmark.
Several explanations appear to infer executable behavior from inert content:
assertions about sanitization. A lexical match should not itself establish
that the package follows those instructions.
self_modificationrisk-taxonomy key describes a risk being assessed;it is not an operation that modifies the program or its policy.
on its own to establish typosquatting.
exfiltration explanation should identify the actual sensitive source and sink.
Important caveat: some smoke tests send bearer tokens to a configured URL.
Their transport/destination validation merits separate review. We are not asking
for all test-path findings to be suppressed or claiming every finding is false.
Coverage warnings, degraded OSV checks, and unresolved references are genuine
limitations and should remain visible.
Reproduction
Recommend using a clean, disposable directory and isolated environment. Install the pinned scanner from its official source, not an unpinned dependency set. The source pin
does not itself pin all transitive dependencies.
Extract that public-source archive into a new
specimendirectory using yourplatform's archive utility. Scan only; do not install or execute specimen code:
The historical count is environment-dependent, particularly for network checks;
the useful reproduction target is the particular finding explanation and source
context, not an invariant total of 50. A fresh minimal specimen and current-main
comparison have not yet been run for this upstream submission.
Requested improvement / questions
executable behavior in the explanation and confidence semantics.
including malicious content placed under a test-like path.
tests/, private IPs, or loopback wholesale: those are not trustboundaries. Baselines should be reviewer-controlled and narrowly scoped.
or a newer analyzer revision? We can help supply focused specimens.
Our integration preserves SkillSpector's original score and recommendation. We
do not rewrite those into approval or claim NVIDIA endorsement or conformance.
Public supporting analysis
AI SAFE² second-pass findings and limitations
Related reports
Related: #37 (text-content false positives), #314 (security-test patterns flagged), and #103 (earlier static documentation findings). This report supplies a pinned v2.11.1 AI SAFE2 specimen and distinguishes specific explanation/precision concerns from genuine coverage limitations. Please consolidate if an existing issue is the preferred tracking location.