MIOpen Conv3d Failure on AMD Strix Halo under Windows: Reproduction, Isolation and a Scoped Workaround
This technical white paper documents investigation of a reproducible MIOpen Conv3d failure on an AMD Ryzen AI MAX+ 395 “Strix Halo” system running Windows 11. A bounded stage-3 video-VAE workload reproduced a size-dependent failure on the tile-512 MIOpen path, while the corresponding tile-256 control completed successfully. The same tile-512 geometry also completed when the MIOpen/cuDNN-compatible backend path was disabled using torch.backends.cudnn.enabled = False, providing evidence that the failure is associated with the MIOpen execution path rather than the tensor geometry alone. The paper describes the reproduction boundary, isolation methodology, evidence obtained, and a narrowly scoped application-level workaround candidate that can allow affected users to bypass the failing MIOpen path while leaving the wider application unchanged. The workaround should be considered experimental rather than a confirmed mitigation. The ACE evaluation terminated as INCONCLUSIVE because repeated validation and numerical-correctness comparison against a declared reference tolerance were not completed. An upstream correction to the affected AMD/MIOpen path remains the preferred resolution. This work was conducted independently by NH Applied on consumer AMD Strix Halo hardware. GitHub discussion and implementation context:[Issue]: gfx1151 / Windows — MIOpen Conv3d aborts during large video VAE decode (tile 512), tile 256 succeeds · Issue #5845 · ROCm/TheRock This GitHub issue contains the upstream discussion, reproduction context, and any subsequent updates relevant to the failure and workaround described in this paper.
Authors
- Nigel Hutchinson
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-05
- DOI
- https://doi.org/10.5281/zenodo.23158751
- Primary Topic
- Parallel Computing and Optimization Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00