Skip to content

Trigger and quality-policy checks overinterpret descriptions and tool catalogs #669

Description

@Spectorian

Problem

Some trigger and quality-policy checks report instructions that the skill does not actually give.

Evidence

  • Static TR2 flags a description containing the noun “ask” as shadowing the built-in command, despite the actual quoted invocation being different.
  • Static TR3 flags “anything with PDF files” as universal activation, losing the PDF domain constraint.
  • An SQP-2 result demanded confirmation for a tool-catalog entry naming a bulk-execution tool, even though the text did not select an action to perform.
  • An SQP-3 result treated “in plain English” in a user-invoked clarification description as a forced output-language rule.

The TR2 and TR3 cases reproduce in deterministic detector tests. The SQP-2 and SQP-3 cases were observed in model output for the linked examples; their results can vary by model.

Expected behavior

Findings should identify the actual trigger, action or language requirement. A noun does not imply a command invocation, “PDF files” limits the trigger's scope, and listing a tool does not select an operation. A readability idiom alone does not establish an organization-policy violation.

Continue detecting real command interception, universal triggers, unrequested external actions, undisclosed destructive changes and explicit incompatible language requirements. Missing warnings alone should not establish unsafe behavior. These corrections should use the existing scoring policy and available evidence about host policy.

Use benign and risky pairs, with a small live-model check for semantic changes.

Examples: tool catalog, clarification-command description.

Relevant code

static_patterns_supply_chain.py:2498, semantic_quality_policy.py:86, semantic_quality_policy.py:108, semantic_quality_policy.py:134.

Activity

  1. Zhuoxi2000 commented on Oct 1, 2026

    @Zhuoxi2000
    Contributor

    I can reproduce the TR2 and TR3 cases deterministically on main (2226747) and would like to take the static part of this issue.

    • TR2: the shadowed command is taken from any built-in word in the clause, and any slash token satisfies the interception gate, so Use when the user says /ask-matt or wants an ask answered by Matt. reports built-in ask. Plan: match slash tokens by their whole name, and let an interception verb (intercepts/overrides/shadows/invokes) name a built-in only in the few words right after it.
    • TR3: _DESCRIPTION_UNIVERSAL_SCOPE_RE only exempts a following about, so the pdf skill's whenever the user wants to do anything with PDF files is reported as keyword baiting. Plan: also accept with/involving/regarding/concerning/related to as qualifiers, unless the object is only a pronoun (anything with it). in/for/on stay out, so any message in the chat still fires.

    SQP-2/SQP-3 depend on the model, so I'd leave them for a follow-up and reference this issue from the PR rather than close it. I'll open the PR shortly; happy to adjust if you'd prefer a different split.

    AI assistance: drafted with an AI coding assistant (Claude).

  2. Xnadir commented on Oct 1, 2026

    @Xnadir
    Contributor

    I'd like to take the two deterministic cases here (TR2 and TR3), and leave the SQP-2/SQP-3 model-dependent ones out of scope.

    From reading static_patterns_supply_chain.py, I think the causes are:

    • TR2: in the description path, shadowed collects any built-in command that appears as a bare word anywhere in the clause, and _DESCRIPTION_COMMAND_INTERCEPTION_RE matches on any slash token. So a description that mentions the skill's own /ask-matt command and also uses the word "ask" in prose gets reported as shadowing the built-in ask. Proposed fix: when the interception signal is a slash token, only count a built-in that appears as that slash token itself (/ask), not as a bare word elsewhere in the clause. Interception verbs (override, intercept, shadow) keep the current behavior.

    • TR3: _DESCRIPTION_UNIVERSAL_SCOPE_RE only treats about as a subject qualifier ((?!\s+about\b)), so "anything with PDF files" still counts as universal scope. Proposed fix: extend the qualifier lookahead to the other scoping prepositions (with, involving, regarding, concerning, related to, containing, in).

    I'd add benign/risky test pairs for each: "/ask-matt ... ask the user" stays negative while "intercepts /ask" stays positive; "anything with PDF files" stays negative while "whenever the user says anything" stays positive.

    Does that approach work for you? Happy to adjust before opening a PR.

  3. Zhuoxi2000 commented on Oct 1, 2026

    @Zhuoxi2000
    Contributor

    Hi @nadirali1350, heads-up so we don't duplicate work: the static TR2/TR3 part was claimed above and is implemented in #709 (opened a couple of hours ago), along the same lines you describe. SQP-2/SQP-3 are still open if you'd like to take those. Happy to coordinate if the maintainers prefer a different split.

    AI assistance: drafted with an AI coding assistant (Claude).

  4. Xnadir commented on Oct 2, 2026

    @Xnadir
    Contributor

    Thanks @Zhuoxi2000, I missed your claim when I posted. #709 looks good, and the TR2/TR3 part is yours. I'll take SQP-2 and SQP-3, if the maintainers are OK with that split.

    Plan (both are wording changes to the rules in semantic_quality_policy.py, plus fixtures):

    • SQP-2: add a "Do NOT flag" case for tool catalogs and reference tables. Naming a tool (for example a bulk-execution tool in a list of available tools) doesn't select an operation; only flag when the skill instructs the agent to perform it or the code performs it.
    • SQP-3: add a "Do NOT flag" case for readability idioms. "In plain English", "in simple terms", and "in layman's terms" ask for clarity, not a language. Only flag explicit forced-language instructions such as "always respond in Japanese".
    • Fixtures: benign/risky pairs next to the existing ones in tests/fixtures/sqp/ (a tool catalog vs. an instruction to run the tool; "explain in plain English" vs. "always answer in French").
    • Live check: run skillspector scan --llm with the claude_cli provider on the two linked examples (baselinker-automation and ask-matt) and the new fixtures, before and after the change, and post the results in the PR.

    Does that work?

    AI assistance: drafted with Claude. I review and test everything I submit.

  5. Zhuoxi2000 commented on Oct 2, 2026

    @Zhuoxi2000
    Contributor

    Thanks @nadirali1350, that split works for me.

  6. added 2 commits that reference this issue on Oct 4, 2026
  7. rng1995 commented on Oct 4, 2026

    @rng1995
    Collaborator

    Merge-state update checked on 2026-10-06: PR #709 has merged, and GitHub automatically closed this issue as completed.

    Both parts of this issue are now merged: static TR2/TR3 description matching in #709 and SQP guidance in #710. Offline trigger regressions pass on main; this is not a fresh live-model precision evaluation.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions