Repository navigation
feat(federation): add rocm-doctor from ROCm/rocm-cli - #259
Conversation
Diagnoses why ROCm, the HIP SDK, PyTorch or llama.cpp is broken on an AMD GPU across Linux, Windows and WSL2, against a closed catalog of 25 known misconfigurations, and either applies a low-risk fix with consent or hands back the exact next step. The probe, catalog and fixes live in the rocm CLI and are versioned with the binary; the skill is a driver over it. Product repo approval is recorded in .github/skill_owners.json. Signed-off-by: Eugene Volen <Eugene.Volen@amd.com>
Picks up the WSL2 eval reword from ROCm/rocm-cli#561, which drops an expectation a Windows grading host cannot satisfy: there is no WSL2 guest there for `rocm diagnose` to report on. Also carries upstream movement since the first import: the catalog grew to 26 entries with fix-16-vllm-oom, which fills what was a reserved handle. Signed-off-by: Eugene Volen <Eugene.Volen@amd.com>
|
Re-imported at I went with the reword rather than Both The import also picked up upstream movement since the first one: the catalog is now 26 entries,
On #260 — I would merge them separately, this one first. It is already shaped right as a dependent draft, and the two changes have different blast radii: this adds a federated skill, that deletes a staged draft and rewrites a README row. Keeping them apart means either can be reverted without the other. Your point about the copies having diverged is right and worth doing, just after rather than inside. |
danielholanda
left a comment
There was a problem hiding this comment.
Tested locally and works as expected. The skill is also tech preview on rocm-cli, which makes it eligible here. Happy to approve!
danielholanda
left a comment
There was a problem hiding this comment.
Minor change requested before merging: The evals look flaky. rocm-under-wsl2 failed this run (log) because the judge counted "offer to install the rocm CLI" as a non-CLI remediation, even though its own reasoning says that's not what the check is about.
Can we fix this before merging? Either make the skill behave more consistently or reword the expectation so it's clear what counts as a failure. Thanks!
Picks up ROCm/rocm-cli#596, which rewords two judged eval items so that installing the rocm CLI with consent, as Phase 0 of the skill instructs, no longer reads as a remediation the CLI did not return. Signed-off-by: Eugene Volen <Eugene.Volen@amd.com>
6346a16 to
39da6f1
Compare
|
@danielholanda fixed by rewording the expectation, which was the second option you offered. Re-imported at Cause. The case failed on wording, not on behaviour. The check said "a remediation the CLI did not return", and installing the Change. The item now names the class of remedy it is about: a driver or permissions change such as loading Results on this head. One green run proves little against a 2-in-6 failure rate, so I re-ran the Windows leg until there were three results:
That is 20 of 20 case runs, and |
Picks up ROCm/rocm-cli#611. The `rocm-nothing-established-routes-upstream` case asked the agent for the exact error text, although its prompt says the user cannot paste one and `rocm diagnose` probes the host without `--symptom`. The judge passed that behaviour on some runs and failed it on others. The item now grades the guard it stood for, and the `has_match` rule moves into the existing no-remediation item. Eval data only; the skill itself is unchanged. Signed-off-by: Eugene Volen <Eugene.Volen@amd.com>
Adds
rocm-doctorto the catalog. Product repo approval is recorded in.github/skill_owners.json(ROCm/rocm-cli, engineering owner and product release owner both signed off).The skill
Diagnoses why ROCm, the HIP SDK, PyTorch or llama.cpp is broken on an AMD GPU across Linux, Windows and WSL2, against a closed catalog of 26 known misconfigurations, then either applies a low-risk fix with consent or hands back the exact next step. It also routes Lemonade, LM Studio and Ollama problems to the right upstream channel rather than guessing.
It is a thin driver over the
rocmCLI —rocm examine/rocm diagnose/rocm fix. The probe, the catalog and the fixes live in the binary and are versioned with it, so the skill does not carry a second copy of the knowledge that can drift.Product status is Tech Preview.
What I verified before opening this
Ran the catalog's own tooling against a real import of this skill:
federate_skills.py --only rocm-doctorSKILL.md,reference.md,skill-card.md,evals/evals.jsoncheck.sh0 error(s), checked against 11 skillsskillscope structural--skill-files skill-card.md --skill-sections Description,Owner,Licensepublish.sh.claude-plugin/marketplace.jsonuntouchedAgainst the intake checklist: name is lowercase-hyphenated, matches the folder, unique in the catalog, no forbidden substrings. Description is 962 characters.
SKILL.mdis 208 lines with reference material in a sibling file one level deep. All paths use forward slashes. Prerequisites state explicitly that no fixed ROCm version,gfxtarget or container image is assumed — the CLI detects each — and warn against hand-settingHSA_OVERRIDE_GFX_VERSION.Evals carry 5 positive, 2 negative, and 4 positives with behavioural expectations, against minimums of 3 / 2 / 2.
One checklist item I have not met
Routing and behavioural have not been run in
ROCm/rocm-cli. Both need a model key that repo does not hold — I opened a PR there to wire up the reusable workflow and closed it again, because atv0.1.3it takes noapi_base_urlorapi_custom_headersand so cannot reach the AMD gateway at all. A follow-up is going into that repo's existing CI job, which calls the skillscope action directly, onceORCHESTR_API_KEYis provisioned there.So this PR's routing and behavioural legs will be their first real run. Flagging it rather than leaving you to find out — if you would rather wait until they are green on our side first, say so and I will hold this until the secret lands.
Not included
No walkthrough yet. It is listed as recommended rather than required, and I would rather add one that reflects a real session than ship a sketch.