Repository navigation
Conversation
| @@ -0,0 +1,282 @@ | |||
| #!/usr/bin/env python3 | |||
There was a problem hiding this comment.
https://github.com/amd/skills/actions/runs/36598752509

Looks like failing tests are not showing any logs. Those are potentially part of a log file now, but this is not great for visibility. Is there any way in which we can show the logs here as they used to show on other machines?
There was a problem hiding this comment.
Updated the behavioral jobs to show sanitized grader output directly in GitHub, including case/expectation PASS/FAIL lines, totals, and unmet expectations with judge explanations. These also appear in the job summary.
Full logs remain available through ReportPortal
| uv tool install --system-certs git+https://github.com/amd/skillscope@v0.1.2 | ||
| SKILLSCOPE_SHA="$(sed -n 's/^ SKILLSCOPE_SHA: //p' .github/workflows/evals.yml)" | ||
| uv tool install --system-certs "git+https://github.com/amd/skillscope@$SKILLSCOPE_SHA" |
There was a problem hiding this comment.
Recommend simply pinning it to a version explicitly
|
|
||
| env: | ||
| SKILLSCOPE_REPOSITORY: amd/skillscope | ||
| SKILLSCOPE_SHA: 0f734b2dd344cd72b5e016ad6c3b1ac2ce36f6f3 |
There was a problem hiding this comment.
We shouldn't stick with sha versions. Only releases like it was before.
There was a problem hiding this comment.
Done the local installation example now explicitly uses amd/skillscope@v0.1.3, matching CI. It no longer extracts a SHA from the workflow.
14e7c3b to
bd879a9
Compare
- Treat enable_mi_ci as a relevant label again, so adding it starts the Instinct leg as it did before the OrchestrAI change. - Fail evals / results for mock or single-OS dispatches; the check attaches to a commit and must not report a pass without full hardware. - Pass the workflow's Skillscope pin to each OrchestrAI session as SKILLSCOPE_REF so Strix legs move with the same version. - Add orchestrai_logs.py to infra paths; it shapes verdicts.
|
Really solid work, the redaction design is more careful than most things I've read for this purpose. Three things I'd want fixed before this goes in, plus one question. Blocker 1: a password can reach a public log
Host redacted, password published verbatim. Same for redis. A DB connection error during dependency setup is an ordinary thing for this harness to capture. Adding postgres, mysql, mongodb, redis, amqp and sftp to the scheme list covers it, or invert to any Blocker 2: grader errors get reported as skill failures, and it's live on this PR Two of the six red checks here are infrastructure, not skills: The judge died, skillscope marked that one expectation unmet, and The diagnostics you want already exist at Worth a regression test too, there's exactly one assertion in the 3056 line test file that touches result categorisation ( I filed the skillscope half separately as amd/skillscope#25, Blocker 3: the Windows leg isn't exercising the GPU path at all Not in the diff, but you should know.
quark-install passed on Windows because of that. The dead GPU routed the agent down the easy branch of the expectation ("if compute mode is unresolved, just report and ask"). On Linux the GPU was healthy, so it got the hard branch and failed. That means every Windows result in this PR is suspect, not just this one. A pass that tested nothing bothers me more than any of the red ones here. Question, not a change request: is I noticed routing keeps But if that control runner is self hosted, fork PRs touching skills still execute on it, and nothing resets the control runner between jobs. Secrets aren't exposed since the trigger is Smaller stuff, not blocking
On the rest of the red The other three are ordinary agent variance, not your machines. lemonade stops mid sentence before printing its curl commands, local-ai-use installs TTS and STT backends during an image task. The work itself was done correctly in both. The self hosted boxes do this too, rocm-doctor on Windows went fail, pass, fail across three consecutive runs on 06 and 07 Oct. |
Run selected Strix behavioral evals through OrchestrAI at the exact candidate commit. Keep the Skillscope release, test expectations, Windows provisioning sequence, redacted test streams and
evals / resultscheck.Protect privileged jobs with reviewed controller code, protected Environment checks, same-repository restrictions and least-privilege tokens. Validate installer paths and matrix arguments, strengthen encoded-secret/URL redaction, and keep verdict processing on hosted runners. Add a credential-free Linux/Windows security preview.
Rollout: eligible graded runs automatically verify Environment protections and the trusted controller; no enable/disable variable is required. The evaluation gate fails when required policy checks or grading do not pass. Administrators must configure approvals, move credentials into protected Environments, verify runner/federation isolation and establish the trusted controller. Manual runs remain restricted to the default branch; same-repository PRs must target it. YAML checks alone do not protect a workflow that a PR rewrites. See the setup checklist.
Validation: 253 repository tests (182 OrchestrAI tests), workflow/Python lint, structural/federation/manifest checks and the synthetic Linux/Windows preview pass locally. No manual live hardware run or administrative settings change was performed for this update. The additional Windows GPU-readiness check remains deferred.
Previous-revision GitHub security preview passed: 178 regression tests and both synthetic OS flows. The latest local regressions additionally verify that removing the enable variable retains policy enforcement and failure propagation.