Skip to content

Eval bug: ROCm backend crashes on first inference (gfx1201) — missing Tensile library & precompiled kernels #83

Description

@RenZekta

Name and Version

last version of atomic llama.cpp turboquant (b10269-1.6.0)

Operating systems

Windows

GGML backends

HIP

Hardware

RX 9070 XT

Models

Qwen 3.8 27b, Qwen 3.6 35b A3B

Problem description & steps to reproduce

Summary

The shipped ROCm runtime package (b10269-1.6.0) is missing the rocBLAS Tensile data and the precompiled HIP kernels for GPU arch gfx1201 (RDNA 4 / RX 9000 series). The server loads and listens fine, but the first request kills the process ~0.5 s after inference starts, with no graceful fallback. The failure is fully reproducible in the log on two different models/templates, while the same model works perfectly on the Vulkan backend.

Environment

Backend: A:\AI\Backends\atomic-llama-cpp-turboquant\llama-server.exe (bundled ROCm build b10269-1.6.0)
GPU: AMD gfx1201
Models: Qwen3.8-27B-GSQ-RCO-IQ3_XXS-mtp.gguf (ctx 131072, gpu-layers 65), Hermes3.6-35B-A3B-…-MTP-APEX-Compact.gguf + mmproj F16 (ctx 262144, gpu-layers auto)

Reproduction

Start any template with the ROCm backend on gfx1201.
Send the first chat message.
launch_slot_ starts → rocBLAS error cascade → process exits immediately.

Root cause

The package simply doesn't contain the gfx1201 runtime files:

…\rocm\b10269-1.6.0\bin\rocblas\library\ — directory missing (empty Tensile file list), so no TensileLibrary.dat / TensileLibrary_lazy_gfx1201.dat
Kernels.so-000-gfx1201.hsaco (+ -xnack-/-xnack+) — not shipped
The first GEMM in inference triggers Tensile host init and kernel load → both fail → process death. It's a packaging/build-coverage gap for gfx1201, not a model quirk.

Secondary observations

llama_sampler_backend_support: device 'ROCm0' does not have support for op TOP_K needed for sampler 'top-k' — emitted per slot (×4 for n_slots=4) on every ROCm run: top-k sampling is unsupported on the ROCm device (CPU fallback).

Expected vs actual

Expected: server loads, first message is processed on GPU; at worst a startup warning + CPU fallback.
Actual: hard process exit ~0.5 s into the first inference.

Suggested fixes / verification

Ship the complete rocblas/library directory (Tensile TensileLibrary.dat incl. gfx1201 lazy variants) in the b10269-1.6.0 package.
Build and ship Kernels.so-000-gfx1201{,-xnack-,-xnack+}.hsaco.
Verify with llama-cli on gfx1201: list of available TensileLibrary Files should be non-empty, and a prompt should complete without the process exiting.
Follow-up: ROCm TOP_K sampler support

Assisted by Qwen 3.8 27b

First Bad Commit

No response

Relevant log output

15:24:45.493 I slot launch_slot_: id 0 | task 0 | processing task, is_child = 0
15:24:45.646 rocblaslt error: Cannot read "TensileLibrary_lazy_gfx1201.dat" (or .zlib variant): No error
15:24:45.647 rocBLAS error: Cannot read A:\AI\Backends\atomic-llama-cpp-turboquant\rocm\b10269-1.6.0\bin\rocblas\library\TensileLibrary.dat:
No such file or directory for GPU arch : gfx1201
List of available TensileLibrary Files : ← empty
15:24:45.646 hipModuleLoad failed: Kernels.so-000-gfx1201.hsaco / -xnack- / -xnack+ : file not found
15:24:45.651 rocBLAS error: Could not initialize Tensile host:
directory_iterator: The system cannot find the path specified: "…\bin\rocblas\library"
15:24:46.162 ■ Process exited

Full log

xlm-studio-logs-1790177241806.txt

Note

Currently upstream official llama.cpp has an issue of not using GPU at all with ROCm backend ggml-org#26964

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions