Name and Version
last version of atomic llama.cpp turboquant (b10269-1.6.0)
Operating systems
Windows
GGML backends
HIP
Hardware
RX 9070 XT
Models
Qwen 3.8 27b, Qwen 3.6 35b A3B
Problem description & steps to reproduce
Summary
The shipped ROCm runtime package (b10269-1.6.0) is missing the rocBLAS Tensile data and the precompiled HIP kernels for GPU arch gfx1201 (RDNA 4 / RX 9000 series). The server loads and listens fine, but the first request kills the process ~0.5 s after inference starts, with no graceful fallback. The failure is fully reproducible in the log on two different models/templates, while the same model works perfectly on the Vulkan backend.
Environment
Backend: A:\AI\Backends\atomic-llama-cpp-turboquant\llama-server.exe (bundled ROCm build b10269-1.6.0)
GPU: AMD gfx1201
Models: Qwen3.8-27B-GSQ-RCO-IQ3_XXS-mtp.gguf (ctx 131072, gpu-layers 65), Hermes3.6-35B-A3B-…-MTP-APEX-Compact.gguf + mmproj F16 (ctx 262144, gpu-layers auto)
Reproduction
Start any template with the ROCm backend on gfx1201.
Send the first chat message.
launch_slot_ starts → rocBLAS error cascade → process exits immediately.
Root cause
The package simply doesn't contain the gfx1201 runtime files:
…\rocm\b10269-1.6.0\bin\rocblas\library\ — directory missing (empty Tensile file list), so no TensileLibrary.dat / TensileLibrary_lazy_gfx1201.dat
Kernels.so-000-gfx1201.hsaco (+ -xnack-/-xnack+) — not shipped
The first GEMM in inference triggers Tensile host init and kernel load → both fail → process death. It's a packaging/build-coverage gap for gfx1201, not a model quirk.
Secondary observations
llama_sampler_backend_support: device 'ROCm0' does not have support for op TOP_K needed for sampler 'top-k' — emitted per slot (×4 for n_slots=4) on every ROCm run: top-k sampling is unsupported on the ROCm device (CPU fallback).
Expected vs actual
Expected: server loads, first message is processed on GPU; at worst a startup warning + CPU fallback.
Actual: hard process exit ~0.5 s into the first inference.
Suggested fixes / verification
Ship the complete rocblas/library directory (Tensile TensileLibrary.dat incl. gfx1201 lazy variants) in the b10269-1.6.0 package.
Build and ship Kernels.so-000-gfx1201{,-xnack-,-xnack+}.hsaco.
Verify with llama-cli on gfx1201: list of available TensileLibrary Files should be non-empty, and a prompt should complete without the process exiting.
Follow-up: ROCm TOP_K sampler support
Assisted by Qwen 3.8 27b
First Bad Commit
No response
Relevant log output
15:24:45.493 I slot launch_slot_: id 0 | task 0 | processing task, is_child = 0
15:24:45.646 rocblaslt error: Cannot read "TensileLibrary_lazy_gfx1201.dat" (or .zlib variant): No error
15:24:45.647 rocBLAS error: Cannot read A:\AI\Backends\atomic-llama-cpp-turboquant\rocm\b10269-1.6.0\bin\rocblas\library\TensileLibrary.dat:
No such file or directory for GPU arch : gfx1201
List of available TensileLibrary Files : ← empty
15:24:45.646 hipModuleLoad failed: Kernels.so-000-gfx1201.hsaco / -xnack- / -xnack+ : file not found
15:24:45.651 rocBLAS error: Could not initialize Tensile host:
directory_iterator: The system cannot find the path specified: "…\bin\rocblas\library"
15:24:46.162 ■ Process exited
Full log
xlm-studio-logs-1790177241806.txt
Note
Currently upstream official llama.cpp has an issue of not using GPU at all with ROCm backend ggml-org#26964
Name and Version
last version of atomic llama.cpp turboquant (b10269-1.6.0)
Operating systems
Windows
GGML backends
HIP
Hardware
RX 9070 XT
Models
Qwen 3.8 27b, Qwen 3.6 35b A3B
Problem description & steps to reproduce
Summary
The shipped ROCm runtime package (b10269-1.6.0) is missing the rocBLAS Tensile data and the precompiled HIP kernels for GPU arch gfx1201 (RDNA 4 / RX 9000 series). The server loads and listens fine, but the first request kills the process ~0.5 s after inference starts, with no graceful fallback. The failure is fully reproducible in the log on two different models/templates, while the same model works perfectly on the Vulkan backend.
Environment
Backend: A:\AI\Backends\atomic-llama-cpp-turboquant\llama-server.exe (bundled ROCm build b10269-1.6.0)
GPU: AMD gfx1201
Models: Qwen3.8-27B-GSQ-RCO-IQ3_XXS-mtp.gguf (ctx 131072, gpu-layers 65), Hermes3.6-35B-A3B-…-MTP-APEX-Compact.gguf + mmproj F16 (ctx 262144, gpu-layers auto)
Reproduction
Start any template with the ROCm backend on gfx1201.
Send the first chat message.
launch_slot_ starts → rocBLAS error cascade → process exits immediately.
Root cause
The package simply doesn't contain the gfx1201 runtime files:
…\rocm\b10269-1.6.0\bin\rocblas\library\ — directory missing (empty Tensile file list), so no TensileLibrary.dat / TensileLibrary_lazy_gfx1201.dat
Kernels.so-000-gfx1201.hsaco (+ -xnack-/-xnack+) — not shipped
The first GEMM in inference triggers Tensile host init and kernel load → both fail → process death. It's a packaging/build-coverage gap for gfx1201, not a model quirk.
Secondary observations
llama_sampler_backend_support: device 'ROCm0' does not have support for op TOP_K needed for sampler 'top-k' — emitted per slot (×4 for n_slots=4) on every ROCm run: top-k sampling is unsupported on the ROCm device (CPU fallback).
Expected vs actual
Expected: server loads, first message is processed on GPU; at worst a startup warning + CPU fallback.
Actual: hard process exit ~0.5 s into the first inference.
Suggested fixes / verification
Ship the complete rocblas/library directory (Tensile TensileLibrary.dat incl. gfx1201 lazy variants) in the b10269-1.6.0 package.
Build and ship Kernels.so-000-gfx1201{,-xnack-,-xnack+}.hsaco.
Verify with llama-cli on gfx1201: list of available TensileLibrary Files should be non-empty, and a prompt should complete without the process exiting.
Follow-up: ROCm TOP_K sampler support
Assisted by Qwen 3.8 27b
First Bad Commit
No response
Relevant log output
15:24:45.493 I slot launch_slot_: id 0 | task 0 | processing task, is_child = 0
15:24:45.646 rocblaslt error: Cannot read "TensileLibrary_lazy_gfx1201.dat" (or .zlib variant): No error
15:24:45.647 rocBLAS error: Cannot read A:\AI\Backends\atomic-llama-cpp-turboquant\rocm\b10269-1.6.0\bin\rocblas\library\TensileLibrary.dat:
No such file or directory for GPU arch : gfx1201
List of available TensileLibrary Files : ← empty
15:24:45.646 hipModuleLoad failed: Kernels.so-000-gfx1201.hsaco / -xnack- / -xnack+ : file not found
15:24:45.651 rocBLAS error: Could not initialize Tensile host:
directory_iterator: The system cannot find the path specified: "…\bin\rocblas\library"
15:24:46.162 ■ Process exited
Full log
xlm-studio-logs-1790177241806.txt
Note
Currently upstream official llama.cpp has an issue of not using GPU at all with ROCm backend ggml-org#26964