Skip to content

ci: run Strix behavioral evals through OrchestrAI - #247

Open
saman-amd wants to merge 14 commits into
mainfrom
saman/orchestrai-evals
Open

saman-amd wants to merge 14 commits into
mainfrom
saman/orchestrai-evals

Conversation

@saman-amd

@saman-amd saman-amd commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

Run selected Strix behavioral evals through OrchestrAI at the exact candidate commit. Keep the Skillscope release, test expectations, Windows provisioning sequence, redacted test streams and evals / results check.

Protect privileged jobs with reviewed controller code, protected Environment checks, same-repository restrictions and least-privilege tokens. Validate installer paths and matrix arguments, strengthen encoded-secret/URL redaction, and keep verdict processing on hosted runners. Add a credential-free Linux/Windows security preview.

Rollout: eligible graded runs automatically verify Environment protections and the trusted controller; no enable/disable variable is required. The evaluation gate fails when required policy checks or grading do not pass. Administrators must configure approvals, move credentials into protected Environments, verify runner/federation isolation and establish the trusted controller. Manual runs remain restricted to the default branch; same-repository PRs must target it. YAML checks alone do not protect a workflow that a PR rewrites. See the setup checklist.

Validation: 253 repository tests (182 OrchestrAI tests), workflow/Python lint, structural/federation/manifest checks and the synthetic Linux/Windows preview pass locally. No manual live hardware run or administrative settings change was performed for this update. The additional Windows GPU-readiness check remains deferred.

Previous-revision GitHub security preview passed: 178 regression tests and both synthetic OS flows. The latest local regressions additionally verify that removing the enable variable retains policy enforcement and failure propagation.

@saman-amd saman-amd added the run_behavioral Run behavioral tests on PR label Sep 29, 2026
@@ -0,0 +1,282 @@
#!/usr/bin/env python3

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

https://github.com/amd/skills/actions/runs/36598752509
Image
Looks like failing tests are not showing any logs. Those are potentially part of a log file now, but this is not great for visibility. Is there any way in which we can show the logs here as they used to show on other machines?

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated the behavioral jobs to show sanitized grader output directly in GitHub, including case/expectation PASS/FAIL lines, totals, and unmet expectations with judge explanations. These also appear in the job summary.
Full logs remain available through ReportPortal

Comment thread docs/evals.md Outdated
Comment on lines +130 to +131
uv tool install --system-certs git+https://github.com/amd/skillscope@v0.1.2
SKILLSCOPE_SHA="$(sed -n 's/^ SKILLSCOPE_SHA: //p' .github/workflows/evals.yml)"
uv tool install --system-certs "git+https://github.com/amd/skillscope@$SKILLSCOPE_SHA"

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Recommend simply pinning it to a version explicitly

Comment thread .github/workflows/evals.yml Outdated

env:
SKILLSCOPE_REPOSITORY: amd/skillscope
SKILLSCOPE_SHA: 0f734b2dd344cd72b5e016ad6c3b1ac2ce36f6f3

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We shouldn't stick with sha versions. Only releases like it was before.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done the local installation example now explicitly uses amd/skillscope@v0.1.3, matching CI. It no longer extracts a SHA from the workflow.

@saman-amd
saman-amd deployed to behavioral-instinct October 1, 2026 18:47 — with GitHub Actions Active
@saman-amd
saman-amd deployed to behavioral-instinct October 1, 2026 18:47 — with GitHub Actions Active
@saman-amd
saman-amd force-pushed the saman/orchestrai-evals branch from 14e7c3b to bd879a9 Compare October 2, 2026 04:45
- Treat enable_mi_ci as a relevant label again, so adding it starts the
  Instinct leg as it did before the OrchestrAI change.
- Fail evals / results for mock or single-OS dispatches; the check
  attaches to a commit and must not report a pass without full hardware.
- Pass the workflow's Skillscope pin to each OrchestrAI session as
  SKILLSCOPE_REF so Strix legs move with the same version.
- Add orchestrai_logs.py to infra paths; it shapes verdicts.
@saman-amd
saman-amd deployed to behavioral-instinct October 2, 2026 14:41 — with GitHub Actions Active
@saman-amd
saman-amd deployed to behavioral-instinct October 2, 2026 14:41 — with GitHub Actions Active
@saman-amd
saman-amd deployed to behavioral-instinct October 2, 2026 18:54 — with GitHub Actions Active
@saman-amd
saman-amd deployed to behavioral-instinct October 2, 2026 18:54 — with GitHub Actions Active
@saman-amd
saman-amd deployed to behavioral-instinct October 5, 2026 14:42 — with GitHub Actions Active
@saman-amd
saman-amd deployed to behavioral-instinct October 5, 2026 14:42 — with GitHub Actions Active
@johnl-amd

johnl-amd commented Oct 7, 2026 •

Copy link
Copy Markdown
Collaborator

Really solid work, the redaction design is more careful than most things I've read for this purpose. Three things I'd want fixed before this goes in, plus one question.

Blocker 1: a password can reach a public log

orchestrai_stdout.py:41 matches only https, http, ssh, ftp and git. _url() at :117 is the only thing that strips credentials, so any other scheme never gets there. Ran it:

in:  could not connect to postgres://admin:$PASSWORD@internal-db.example.com:5432/prod
out: could not connect to postgres://admin:$PASSWORD@[HOST REDACTED]/prod

in:  mysql://svc:$PASSWORD@dbhost:3306/app
out: unchanged

(placeholders above stand in for the literal values I actually tested with)

Host redacted, password published verbatim. Same for redis. A DB connection error during dependency setup is an ordinary thing for this harness to capture. Adding postgres, mysql, mongodb, redis, amqp and sftp to the scheme list covers it, or invert to any scheme://.

Blocker 2: grader errors get reported as skill failures, and it's live on this PR

Two of the six red checks here are infrastructure, not skills:

tracelens, Linux
  [FAIL] (expected_behavior) Filled the Executive Summary ...
    -- llm_judge gave no JSON verdict: 'Prompt is too long'
  [FAIL] gemm-01-repeatability: 6/7 checks in 370.55s

magpie, Windows
  [FAIL] (unexpected_behavior) Run a benchmark, install or download anything ...
    -- llm_judge gave no JSON verdict: 'API Error: Connection dropped (ECONNRESET)'
  [FAIL] benchmark-optimization-plan: 2/3 checks in 198.08s

The judge died, skillscope marked that one expectation unmet, and orchestrai_run.py:349 mapped the case to "One or more behavioral expectations were not met". Worth flagging that this is easy to miss because the judge fails per expectation, not per case, so the case still reads 6 of 7 rather than 0 of 7.

The diagnostics you want already exist at orchestrai_run.py:255. They're just nested under if "produced no parseable stream-json output" in output:, which catches the agent CLI failing to emit output, not a judge failing after a good run. Hoisting them out and mapping to status = "error" should do it. There's also no error bucket in the summary at all, orchestrai_verdict.py:374 is "failed": 0 if ok else 1.

Worth a regression test too, there's exactly one assertion in the 3056 line test file that touches result categorisation (test_orchestrai_evals.py:2294).

I filed the skillscope half separately as amd/skillscope#25, _grade_with_llm returns False when it cannot grade, same value as a real negative verdict. Fixing that doesn't remove the need for the change here though, since this classifier pattern matches harness strings and will keep mislabelling whatever shape they take.

Blocker 3: the Windows leg isn't exercising the GPU path at all

Not in the diff, but you should know. provisioning in orchestrai-config.json has linux_install_scripts and nothing for Windows. One Windows box came back with driver error 28 and no video controllers, which is in the judge line verbatim:

reported concrete blocking evidence (driver error code 28, no video controllers, no Windows ROCm wheel) before asking how to proceed

quark-install passed on Windows because of that. The dead GPU routed the agent down the easy branch of the expectation ("if compute mode is unresolved, just report and ask"). On Linux the GPU was healthy, so it got the hard branch and failed. That means every Windows result in this PR is suspect, not just this one. A pass that tested nothing bothers me more than any of the red ones here.

Question, not a change request: is ORCHESTRAI_CONTROL_RUNNER self hosted?

I noticed routing keeps if: needs.discover.outputs.routing == 'true' with no same repo check while the new behavioral jobs got one. Then I checked skillscope v0.1.3 and line 515 has the identical condition, and line 596 shows behavioral had no guard upstream either. So there was nothing here to drop, and you've actually tightened this by adding the guards and moving routing off the Strix boxes. Not asking you to change it.

But if that control runner is self hosted, fork PRs touching skills still execute on it, and nothing resets the control runner between jobs. Secrets aren't exposed since the trigger is pull_request, so this is about code execution, not credential theft. If it is self hosted, maybe worth a follow up issue.

Smaller stuff, not blocking

  • opaque value heuristic at orchestrai_stdout.py:254 needs upper, lower and digits all present, so an all lowercase 32 char key passes straight through
  • hostname denylist at :42 covers only amd.com, internal, local, corp and lan, so something like build-node-42 survives. X-Amz-Security-Token also slips the lookbehind at :87, and those two compound: the header name isn't matched and an all lowercase value misses the mixed class test, so the whole thing goes out in the clear
  • IDENTITY at :45 eats to end of line, so {"user": "x", "result": "pass"} becomes {"[IDENTITY REDACTED]. Fails safe, but docs/evals.md saying "every received test-output line is retained" reads a bit optimistically
  • {"items": 5} throws a raw TypeError instead of the "could not be verified" path, orchestrai_verdict.py:401 only catches ValueError
  • check.sh:8 says green here means the same as CI, but CI now runs the unittest discover at evals.yml:127 and check.sh doesn't
  • PR body says 195 unit tests, I get 125 running CI's exact command

On the rest of the red

The other three are ordinary agent variance, not your machines. lemonade stops mid sentence before printing its curl commands, local-ai-use installs TTS and STT backends during an image task. The work itself was done correctly in both. The self hosted boxes do this too, rocm-doctor on Windows went fail, pass, fail across three consecutive runs on 06 and 07 Oct.

@saman-amd saman-amd removed the run_behavioral Run behavioral tests on PR label Oct 7, 2026
@saman-amd
saman-amd deployed to behavioral-instinct October 7, 2026 18:22 — with GitHub Actions Active
@saman-amd
saman-amd deployed to behavioral-instinct October 7, 2026 18:22 — with GitHub Actions Active

This branch was successfully deployed

1 active (outdated) deployment
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants