Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
21 commits
Select commit Hold shift + click to select a range
21528e0
fix: refresh NVIDIA Build default model and reasoning budget
yashrajp22 Oct 5, 2026
ef6145c
docs: update the OpenCode NVIDIA model example
yashrajp22 Oct 5, 2026
ad276c6
Merge branch 'main' into yashraj/refresh-nvidia-build-default
github-actions[bot] Oct 5, 2026
3ebdaf0
Merge branch 'main' into yashraj/refresh-nvidia-build-default
github-actions[bot] Oct 5, 2026
650730b
Merge branch 'main' into yashraj/refresh-nvidia-build-default
github-actions[bot] Oct 5, 2026
a31c6f0
Merge branch 'main' into yashraj/refresh-nvidia-build-default
github-actions[bot] Oct 5, 2026
15a1a41
Merge branch 'main' into yashraj/refresh-nvidia-build-default
github-actions[bot] Oct 5, 2026
a96b7c6
Merge branch 'main' into yashraj/refresh-nvidia-build-default
github-actions[bot] Oct 5, 2026
8d54ecc
Merge branch 'main' into yashraj/refresh-nvidia-build-default
github-actions[bot] Oct 5, 2026
39b2e3b
Merge branch 'main' into yashraj/refresh-nvidia-build-default
github-actions[bot] Oct 5, 2026
5dbc208
Merge branch 'main' into yashraj/refresh-nvidia-build-default
github-actions[bot] Oct 5, 2026
c43ef27
Merge branch 'main' into yashraj/refresh-nvidia-build-default
github-actions[bot] Oct 5, 2026
6f25ccc
fix(provider): clarify reasoning defaults and provenance
yashrajp22 Oct 5, 2026
c7af4cc
Merge branch 'main' into yashraj/refresh-nvidia-build-default
github-actions[bot] Oct 5, 2026
adf305d
Merge branch 'main' into yashraj/refresh-nvidia-build-default
github-actions[bot] Oct 5, 2026
2f2d5b7
Merge branch 'main' into yashraj/refresh-nvidia-build-default
github-actions[bot] Oct 5, 2026
8f8b6da
Merge branch 'main' into yashraj/refresh-nvidia-build-default
github-actions[bot] Oct 5, 2026
6eeeb74
Merge main while preserving the GLM-5.3 default and Gemini guidance
yashrajp22 Oct 6, 2026
9bfa428
Merge branch 'main' into yashraj/refresh-nvidia-build-default
github-actions[bot] Oct 6, 2026
7fce10d
Merge branch 'main' into yashraj/refresh-nvidia-build-default
github-actions[bot] Oct 6, 2026
6852526
Merge branch 'main' into yashraj/refresh-nvidia-build-default
github-actions[bot] Oct 6, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 2 additions & 1 deletion .env.example
Original file line number Diff line number Diff line change
Expand Up @@ -29,7 +29,8 @@ NVIDIA_INFERENCE_KEY=
OPENAI_API_KEY=
OPENAI_BASE_URL=
# Optional provider- and model-dependent reasoning-effort setting. Non-empty values
# are trimmed and passed through unchanged; unset or blank uses the provider default.
# are trimmed and passed through unchanged. Unset or blank sends high for nv_build
# with z-ai/glm-5.3; other provider/model combinations keep their endpoint defaults.
SKILLSPECTOR_REASONING_EFFORT=
# Optional language for human-readable LLM finding text. Machine-readable values
# such as rule IDs and severity values remain unchanged.
Expand Down
9 changes: 7 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -290,7 +290,7 @@ inference gateways.
| `anthropic` | `ANTHROPIC_API_KEY` | api.anthropic.com | `claude-opus-4-6` |
| `anthropic_proxy` | `ANTHROPIC_PROXY_API_KEY` + `ANTHROPIC_PROXY_ENDPOINT_URL` | Any Vertex-style raw-predict proxy | `claude-sonnet-4-6` |
| `bedrock` | `AWS_PROFILE` (optional) + `AWS_REGION` — SigV4 via boto3 | AWS Bedrock Runtime | `us.anthropic.claude-sonnet-4-6-20250915-v1:0` |
| `nv_build` | `NVIDIA_INFERENCE_KEY` | build.nvidia.com | `z-ai/glm-5.2` |
| `nv_build` | `NVIDIA_INFERENCE_KEY` | build.nvidia.com | `z-ai/glm-5.3` |
| `gemini` | `GOOGLE_CLOUD_PROJECT` (+ optional `GOOGLE_CLOUD_LOCATION`) via ADC | Google Cloud OpenAI-compatible Gemini endpoint | `gemini-3.8-flash` |
| `ollama` | _(none)_ | `OLLAMA_BASE_URL` (default `http://localhost:11434/v1`) | `llama3.1:8b` |
| `azure_openai` | `AZURE_OPENAI_API_KEY` + `AZURE_OPENAI_ENDPOINT` | Azure OpenAI Service | `gpt-4o` (deployment defaults to the model label) |
Expand All @@ -300,6 +300,11 @@ inference gateways.
| `gemini_cli` | _(none — uses local CLI auth)_ | local `gemini` binary | local Gemini runtime fallback, or `SKILLSPECTOR_MODEL` |
| `opencode_cli` | _(none — uses local CLI auth)_ | local `opencode` 1.18.33 binary | local OpenCode runtime fallback, or `SKILLSPECTOR_MODEL` |

For NVIDIA Build's `z-ai/glm-5.3`, SkillSpector requests `high` reasoning effort.
Set `SKILLSPECTOR_REASONING_EFFORT` to override it with `low`, `high`, or `max`.
The bundled 128,000-token context and 32,000-token output budgets are conservative
application limits; they do not claim the hosted endpoint's maximum capacity.

Structured output is requested through LangChain's `with_structured_output`,
whose default forces a tool call. Some models reject a forced tool call with
HTTP 400 (`tool_choice: type "tool" and "any" are not supported for this
Expand Down Expand Up @@ -714,7 +719,7 @@ Issues (2)
| `NVIDIA_INFERENCE_KEY` | Credential for the `nv_build` provider (build.nvidia.com). | Required for LLM analysis when `SKILLSPECTOR_PROVIDER=nv_build` |
| `OPENAI_API_KEY` | Credential for the OpenAI provider (`SKILLSPECTOR_PROVIDER=openai`). Also serves as the tier-2 fallback in the credential waterfall when the active provider returns no credentials. | Required for LLM analysis when `SKILLSPECTOR_PROVIDER=openai` |
| `OPENAI_BASE_URL` | Override the OpenAI endpoint (e.g. point at Ollama). | Optional |
| `SKILLSPECTOR_REASONING_EFFORT` | Optional provider- and model-dependent reasoning-effort setting. Non-empty values are trimmed and passed through unchanged; unset or blank preserves provider-default behavior. | Optional |
| `SKILLSPECTOR_REASONING_EFFORT` | Optional provider- and model-dependent reasoning-effort setting. Non-empty values are trimmed and passed through unchanged. When unset or blank, SkillSpector sends `high` for `nv_build` with `z-ai/glm-5.3`; other provider/model combinations keep their endpoint defaults. | Optional |
| `SKILLSPECTOR_OUTPUT_LANGUAGE` | Short, single-line language label (letters, numbers, spaces, `_`, or `-`; maximum 64 characters) for human-readable LLM finding text such as messages, explanations, and remediation. Rule IDs, severity values, paths, code, and other machine-readable values remain unchanged. Unset, blank, or invalid values preserve the default output language. | Optional |
| `SKILLSPECTOR_TEMPERATURE` | Optional sampling temperature from `0` to `1` for hosted providers. Unset or blank preserves the provider default. Lower values can reduce run-to-run variation but do not guarantee identical output. | Optional |
| `SKILLSPECTOR_SEED` | Optional integer sampling seed for OpenAI-compatible and Azure OpenAI providers. Other hosted providers and CLI providers do not receive it. Provider support remains model-dependent. | Optional |
Expand Down
2 changes: 1 addition & 1 deletion docs/DEVELOPMENT.md
Original file line number Diff line number Diff line change
Expand Up @@ -300,7 +300,7 @@ Copy [.env.example](../.env.example) to `.env` in the project root and set value
| `NVIDIA_INFERENCE_KEY` | Credential for `nv_build`. | `nvapi-...` |
| `OPENAI_API_KEY` | Credential for `SKILLSPECTOR_PROVIDER=openai`. Also tier-2 fallback for non-OpenAI providers. | `sk-...` |
| `OPENAI_BASE_URL` | Override the OpenAI endpoint (e.g. point at Ollama). | `http://localhost:11434/v1` |
| `SKILLSPECTOR_REASONING_EFFORT` | Optional provider- and model-dependent reasoning-effort setting. Non-empty values are trimmed and passed through unchanged; unset or blank preserves provider-default behavior. | `high` |
| `SKILLSPECTOR_REASONING_EFFORT` | Optional provider- and model-dependent reasoning-effort setting. Non-empty values are trimmed and passed through unchanged. When unset or blank, SkillSpector sends `high` for `nv_build` with `z-ai/glm-5.3`; other provider/model combinations keep their endpoint defaults. | `high` |
| `SKILLSPECTOR_OUTPUT_LANGUAGE` | Optional short, single-line language label (letters, numbers, spaces, `_`, or `-`; maximum 64 characters) for human-readable LLM finding text. Rule IDs, severity values, paths, code, and other machine-readable values remain unchanged. Unset, blank, or invalid values preserve the default output language. | `Japanese` |
| `SKILLSPECTOR_TEMPERATURE` | Optional sampling temperature from `0` to `1` for hosted providers. Unset or blank preserves provider defaults. Lower values reduce variation but do not guarantee identical output. | `0` |
| `SKILLSPECTOR_SEED` | Optional integer sampling seed for OpenAI-compatible and Azure OpenAI providers. Provider/model support is best-effort; CLI providers ignore it. | `42` |
Expand Down
2 changes: 1 addition & 1 deletion docs/OPENCODE_EXTENSION.md
Original file line number Diff line number Diff line change
Expand Up @@ -75,7 +75,7 @@ Use skillspector_scan on ./my-skill with noLlm=false.
export SKILLSPECTOR_PROVIDER=nv_build
export NVIDIA_INFERENCE_KEY=nvapi-...
# Optional; omit to use nv_build's bundled default model.
# export SKILLSPECTOR_MODEL=z-ai/glm-5.2
# export SKILLSPECTOR_MODEL=z-ai/glm-5.3
```

Other valid providers and their credential variables are listed in the main
Expand Down
5 changes: 3 additions & 2 deletions src/skillspector/providers/chat_models.py
Original file line number Diff line number Diff line change
Expand Up @@ -113,6 +113,7 @@ def create_openai_compatible_chat_model(
timeout: float | None = 120,
default_headers: dict[str, str] | None = None,
disabled_params: dict[str, object] | None = None,
default_reasoning_effort: str | None = None,
) -> BaseChatModel | None:
"""Create ``ChatOpenAI`` for providers serving OpenAI-compatible endpoints.

Expand All @@ -135,8 +136,8 @@ def create_openai_compatible_chat_model(
if disabled_params:
kwargs["disabled_params"] = disabled_params
reasoning_effort = resolve_reasoning_effort()
if reasoning_effort:
kwargs["reasoning_effort"] = reasoning_effort
if reasoning_effort or default_reasoning_effort:
kwargs["reasoning_effort"] = reasoning_effort or default_reasoning_effort
sampling_parameters = resolve_sampling_parameters(include_seed=True)
kwargs.update(sampling_parameters)
chat_model = ChatOpenAI(**kwargs)
Expand Down
8 changes: 8 additions & 0 deletions src/skillspector/providers/nv_build/model_registry.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -14,6 +14,14 @@
# can be added as needed.

models:
# Conservative application budgets, retained from the unknown-model fallback.
# The model card advertises 1M context, but the hosted endpoint's combined
# input/output limit has not been verified. These are not its capacity claims.
# https://build.nvidia.com/z-ai/glm-5-3/modelcard
"z-ai/glm-5.3":
context_length: 128000
max_output_tokens: 32000

# NVIDIA-curated NIMs on build.nvidia.com.
# 202749 e' il tetto COMBINATO ingresso+uscita che l'endpoint impone, ed e' riportato
# alla lettera nel 400 che restituisce quando lo si supera. Dichiarando la finestra
Expand Down
21 changes: 9 additions & 12 deletions src/skillspector/providers/nv_build/provider.py
Original file line number Diff line number Diff line change
Expand Up @@ -38,18 +38,12 @@
class NvBuildProvider:
"""build.nvidia.com credentials + bundled-YAML metadata provider."""

# General default. The previous defaults (deepseek-v4-flash and, for the
# meta_analyzer slot, deepseek-v4-pro) are no longer served: the former
# returns 410 Gone (end of life 2026-08-07) and neither appears in
# GET /v1/models, so the out-of-the-box path failed every call.
#
# The replacement is chosen for DETECTION, not latency. On a bait skill
# carrying prose-disguised credential exfiltration, the fast served model
# (deepseek-v4-flash-0731, ~1.6 s/call) completed every call, reported no
# degradation, and returned a clean verdict — a confident false negative,
# which on a security scanner is the worst possible failure. glm-5.2 costs
# ~16 s/call and flags it CRITICAL.
DEFAULT_MODEL = "z-ai/glm-5.2"
# GLM-5.2 is retired on the hosted endpoint (HTTP 410).
# Choose defaults for detection, not latency alone: a faster replacement
# previously missed the credential-disclosure control. GLM-5.3 detected
# SSD-3 on that control; the prompt-injection control remains unverified
# because its hosted requests timed out. Recheck both before replacing it.
DEFAULT_MODEL = "z-ai/glm-5.3"
SLOT_DEFAULTS: dict[str, str] = {}

def resolve_credentials(self) -> tuple[str, str | None] | None:
Expand All @@ -72,6 +66,9 @@ def create_chat_model(
credentials=self.resolve_credentials(),
max_tokens=max_tokens,
timeout=timeout,
# GLM-5.3 otherwise defaults to maximum reasoning effort.
# https://build.nvidia.com/z-ai/glm-5-3/modelcard
default_reasoning_effort="high" if model == "z-ai/glm-5.3" else None,
)

def get_context_length(self, model: str) -> int | None:
Expand Down
13 changes: 13 additions & 0 deletions tests/unit/test_constants.py
Original file line number Diff line number Diff line change
Expand Up @@ -131,6 +131,19 @@ def test_configured_provider_precedes_openai_fallback(

assert config["default"] == NvBuildProvider.DEFAULT_MODEL

def test_nv_build_slot_override_does_not_inherit_glm_reasoning(
self, monkeypatch: pytest.MonkeyPatch
) -> None:
monkeypatch.setenv("SKILLSPECTOR_PROVIDER", "nv_build")
monkeypatch.setenv("NVIDIA_INFERENCE_KEY", "nvapi-test")
monkeypatch.setenv("SKILLSPECTOR_MODEL_META_ANALYZER", "openai/gpt-oss-120b")
config = _reload_constants().build_model_config()
provider = NvBuildProvider()
default_model = provider.create_chat_model(config["default"], max_tokens=123)
meta_model = provider.create_chat_model(config["meta_analyzer"], max_tokens=123)
assert default_model._get_request_payload("hello")["reasoning_effort"] == "high"
assert "reasoning_effort" not in meta_model._get_request_payload("hello")

def test_cli_provider_precedes_openai_fallback(self, monkeypatch: pytest.MonkeyPatch) -> None:
monkeypatch.setenv("SKILLSPECTOR_PROVIDER", "codex_cli")
monkeypatch.setenv("OPENAI_API_KEY", "sk-test-openai")
Expand Down
41 changes: 41 additions & 0 deletions tests/unit/test_llm_provenance.py
Original file line number Diff line number Diff line change
Expand Up @@ -18,6 +18,7 @@
)
from skillspector.llm_utils import new_inference_usage_collector
from skillspector.providers.chat_models import create_openai_compatible_chat_model
from skillspector.providers.nv_build import NvBuildProvider


class OpenAIProvider:
Expand Down Expand Up @@ -262,6 +263,46 @@ def test_constructor_observation_supersedes_stale_configuration_capture(
assert result["sampling"]["reasoning_effort"]["forwarded_to_client"] == "high"


@pytest.mark.parametrize("effort", [None, " "])
def test_nv_build_adapter_default_is_recorded_without_a_user_override(
monkeypatch: pytest.MonkeyPatch, effort: str | None
) -> None:
provider = NvBuildProvider()
for name in ("SKILLSPECTOR_TEMPERATURE", "SKILLSPECTOR_SEED", "SKILLSPECTOR_REASONING_EFFORT"):
monkeypatch.delenv(name, raising=False)
if effort is not None:
monkeypatch.setenv("SKILLSPECTOR_REASONING_EFFORT", effort)
monkeypatch.setenv("NVIDIA_INFERENCE_KEY", "nvapi-test")
monkeypatch.setattr("skillspector.llm_provenance.get_active_provider", lambda: provider)
monkeypatch.setattr("skillspector.llm_provenance.get_model_config_provider", lambda: provider)
monkeypatch.setattr("skillspector.llm_utils.get_active_provider", lambda: provider)
model = provider.DEFAULT_MODEL
chat_model = provider.create_chat_model(model, max_tokens=128)
collector = new_inference_usage_collector(
node="semantic_security_discovery",
request_kind="structured_output",
model=model,
chat_model=chat_model,
)
collector.mark_response_received()

result = sanitize_llm_provenance(
capture_llm_provenance(_models(model)),
use_llm=True,
inference_usage=collector.snapshot(),
)

assert result["sampling"]["reasoning_effort"] == {
"requested": None,
"source": "provider_default", # SkillSpector's adapter default.
"forwarded_to_client": "high",
"adapter_support": True,
"provider_support": "unknown",
}
assert result["determinism"]["control_status"] == "provider_defaults"
assert result["determinism"]["provider_guarantee"] is False


def test_gpt_5_4_reports_only_controls_retained_by_request_payload(
monkeypatch: pytest.MonkeyPatch,
) -> None:
Expand Down
47 changes: 47 additions & 0 deletions tests/unit/test_providers.py
Original file line number Diff line number Diff line change
Expand Up @@ -33,6 +33,7 @@

import skillspector.providers as providers_module
import skillspector.providers.anthropic.provider as anthropic_provider_module
from skillspector.inference_usage import chat_model_controls, chat_model_requested_controls
from skillspector.providers import (
NO_LLM_API_KEY_MESSAGE,
chat_models,
Expand Down Expand Up @@ -143,6 +144,7 @@ class TestNvBuildProvider:
@pytest.mark.parametrize(
("model", "context_length"),
[
("z-ai/glm-5.3", 128_000),
("z-ai/glm-5.2", 202_749),
("moonshotai/kimi-k2.6", 256_000),
],
Expand All @@ -162,6 +164,51 @@ def test_glm_declares_both_limits(self) -> None:
assert provider.get_context_length("z-ai/glm-5.2") == 202_749
assert provider.get_max_output_tokens("z-ai/glm-5.2") == 32_768

def test_default_model_keeps_conservative_token_budgets(self) -> None:
provider = NvBuildProvider()
assert provider.DEFAULT_MODEL == "z-ai/glm-5.3"
assert provider.get_context_length(provider.DEFAULT_MODEL) == 128_000
assert provider.get_max_output_tokens(provider.DEFAULT_MODEL) == 32_000

@pytest.mark.parametrize(
("model", "configured_effort", "expected_effort"),
[
("z-ai/glm-5.3", None, "high"),
("z-ai/glm-5.3", " ", "high"),
("z-ai/glm-5.3", " low ", "low"),
("z-ai/glm-5.3", "high", "high"),
("z-ai/glm-5.3", "max", "max"),
("z-ai/glm-5.2", None, None),
("z-ai/glm-5.3-flash", None, None),
("another/model", None, None),
],
)
def test_reasoning_default_is_model_specific_and_preserves_user_override(
self,
monkeypatch: pytest.MonkeyPatch,
model: str,
configured_effort: str | None,
expected_effort: str | None,
) -> None:
monkeypatch.setenv("NVIDIA_INFERENCE_KEY", "nvapi-test")
if configured_effort is not None:
monkeypatch.setenv("SKILLSPECTOR_REASONING_EFFORT", configured_effort)
llm = NvBuildProvider().create_chat_model(model, max_tokens=123)
assert isinstance(llm, ChatOpenAI)
assert llm._get_request_payload("hello").get("reasoning_effort") == expected_effort
assert chat_model_controls(llm)["reasoning_effort"] == expected_effort
assert chat_model_requested_controls(llm)["reasoning_effort"] == (
configured_effort.strip() or None if configured_effort else None
)

def test_glm_reasoning_default_does_not_apply_to_other_providers(
self, monkeypatch: pytest.MonkeyPatch
) -> None:
monkeypatch.setenv("OPENAI_API_KEY", "sk-test")
llm = OpenAIProvider().create_chat_model("z-ai/glm-5.3", max_tokens=123)
assert isinstance(llm, ChatOpenAI)
assert "reasoning_effort" not in llm._get_request_payload("hello")

@pytest.mark.parametrize("model", ["glm-5.2", "z-ai/glm-5.2 "])
def test_nv_build_model_near_match_stays_unresolved(self, model: str) -> None:
provider = NvBuildProvider()
Expand Down
Loading