Repository navigation
Trigger and quality-policy checks overinterpret descriptions and tool catalogs #669
Description
Activity
I can reproduce the TR2 and TR3 cases deterministically on main (2226747) and would like to take the static part of this issue.
- TR2: the shadowed command is taken from any built-in word in the clause, and any slash token satisfies the interception gate, so
Use when the user says /ask-matt or wants an ask answered by Matt.reports built-inask. Plan: match slash tokens by their whole name, and let an interception verb (intercepts/overrides/shadows/invokes) name a built-in only in the few words right after it. - TR3:
_DESCRIPTION_UNIVERSAL_SCOPE_REonly exempts a followingabout, so the pdf skill'swhenever the user wants to do anything with PDF filesis reported as keyword baiting. Plan: also acceptwith/involving/regarding/concerning/related toas qualifiers, unless the object is only a pronoun (anything with it).in/for/onstay out, soany message in the chatstill fires.
SQP-2/SQP-3 depend on the model, so I'd leave them for a follow-up and reference this issue from the PR rather than close it. I'll open the PR shortly; happy to adjust if you'd prefer a different split.
AI assistance: drafted with an AI coding assistant (Claude).
- TR2: the shadowed command is taken from any built-in word in the clause, and any slash token satisfies the interception gate, so
I'd like to take the two deterministic cases here (TR2 and TR3), and leave the SQP-2/SQP-3 model-dependent ones out of scope.
From reading static_patterns_supply_chain.py, I think the causes are:
-
TR2: in the description path,
shadowedcollects any built-in command that appears as a bare word anywhere in the clause, and_DESCRIPTION_COMMAND_INTERCEPTION_REmatches on any slash token. So a description that mentions the skill's own/ask-mattcommand and also uses the word "ask" in prose gets reported as shadowing the built-inask. Proposed fix: when the interception signal is a slash token, only count a built-in that appears as that slash token itself (/ask), not as a bare word elsewhere in the clause. Interception verbs (override, intercept, shadow) keep the current behavior. -
TR3:
_DESCRIPTION_UNIVERSAL_SCOPE_REonly treatsaboutas a subject qualifier ((?!\s+about\b)), so "anything with PDF files" still counts as universal scope. Proposed fix: extend the qualifier lookahead to the other scoping prepositions (with, involving, regarding, concerning, related to, containing, in).
I'd add benign/risky test pairs for each: "/ask-matt ... ask the user" stays negative while "intercepts /ask" stays positive; "anything with PDF files" stays negative while "whenever the user says anything" stays positive.
Does that approach work for you? Happy to adjust before opening a PR.
-
Hi @nadirali1350, heads-up so we don't duplicate work: the static TR2/TR3 part was claimed above and is implemented in #709 (opened a couple of hours ago), along the same lines you describe. SQP-2/SQP-3 are still open if you'd like to take those. Happy to coordinate if the maintainers prefer a different split.
AI assistance: drafted with an AI coding assistant (Claude).
Thanks @Zhuoxi2000, I missed your claim when I posted. #709 looks good, and the TR2/TR3 part is yours. I'll take SQP-2 and SQP-3, if the maintainers are OK with that split.
Plan (both are wording changes to the rules in semantic_quality_policy.py, plus fixtures):
- SQP-2: add a "Do NOT flag" case for tool catalogs and reference tables. Naming a tool (for example a bulk-execution tool in a list of available tools) doesn't select an operation; only flag when the skill instructs the agent to perform it or the code performs it.
- SQP-3: add a "Do NOT flag" case for readability idioms. "In plain English", "in simple terms", and "in layman's terms" ask for clarity, not a language. Only flag explicit forced-language instructions such as "always respond in Japanese".
- Fixtures: benign/risky pairs next to the existing ones in tests/fixtures/sqp/ (a tool catalog vs. an instruction to run the tool; "explain in plain English" vs. "always answer in French").
- Live check: run
skillspector scan --llmwith the claude_cli provider on the two linked examples (baselinker-automation and ask-matt) and the new fixtures, before and after the change, and post the results in the PR.
Does that work?
AI assistance: drafted with Claude. I review and test everything I submit.
Thanks @nadirali1350, that split works for me.
- added a commit that references this issue
on Oct 3, 2026 - added a commit that references this issue
on Oct 3, 2026 Merge-state update checked on 2026-10-06: PR #709 has merged, and GitHub automatically closed this issue as completed.
Both parts of this issue are now merged: static TR2/TR3 description matching in #709 and SQP guidance in #710. Offline trigger regressions pass on main; this is not a fresh live-model precision evaluation.
- added a commit that references this issue
on Oct 6, 2026
Problem
Some trigger and quality-policy checks report instructions that the skill does not actually give.
Evidence
The TR2 and TR3 cases reproduce in deterministic detector tests. The SQP-2 and SQP-3 cases were observed in model output for the linked examples; their results can vary by model.
Expected behavior
Findings should identify the actual trigger, action or language requirement. A noun does not imply a command invocation, “PDF files” limits the trigger's scope, and listing a tool does not select an operation. A readability idiom alone does not establish an organization-policy violation.
Continue detecting real command interception, universal triggers, unrequested external actions, undisclosed destructive changes and explicit incompatible language requirements. Missing warnings alone should not establish unsafe behavior. These corrections should use the existing scoring policy and available evidence about host policy.
Use benign and risky pairs, with a small live-model check for semantic changes.
Examples: tool catalog, clarification-command description.
Relevant code
static_patterns_supply_chain.py:2498, semantic_quality_policy.py:86, semantic_quality_policy.py:108, semantic_quality_policy.py:134.