Evidence-backed AMD Strix Halo local-AI setup and benchmarks: Qwen3.8, Ollama, llama.cpp, Vulkan/ROCm, large GGUFs, and cross-OEM results.
-
Updated
Oct 2, 2026 - Python
Evidence-backed AMD Strix Halo local-AI setup and benchmarks: Qwen3.8, Ollama, llama.cpp, Vulkan/ROCm, large GGUFs, and cross-OEM results.
Performance-tuned llama.cpp for AMD Strix Halo (gfx1151): FA + MoE-prefill fixes with a bundled current Mesa driver. Vulkan and HIP; portable dir, Docker, and distrobox.
vLLM Qwen 3.6-27B (AWQ-INT4) + DFlash speculative decoding on AMD Strix Halo (gfx1151 iGPU, 128 GB UMA, ROCm 7.13). 24.8 t/s single-stream, vision, tool calling, 256K context, OpenAI-compatible, Docker. Matches DGX Spark FP8+DFlash+MTP at a third of the cost. No CUDA.
Production-ready, reproducible Ansible for DeepSeek V4 Flash on 128 GiB AMD Strix Halo, with two qualified Vulkan/ROCmFPX stacks, matched quality and throughput benchmarks, and 512K context validation.
Direct EC fan control for AMD Strix Halo (Ryzen AI Max+ 395) mini-PCs — validated on the Bosgame M5. Fixes sustained-load thermal throttling.
GNOME top-bar indicator for the AMD Strix Halo power level on the Bosgame M5 (Sixunited AXB35-02, Ryzen AI Max+ 395) — reads the hardware power button's EC register and adds an ACPI platform_profile so GNOME's own Power Mode menu drives the real TDP.
ROCmFPX llama.cpp fork for Windows 🏆 — native build, headless OpenAI-compatible server & benchmarks. Tested on AMD Strix Halo (gfx1151), runs on other GPUs too.
Reproducible local-LLM benchmark harness: llama.cpp on AMD Strix Halo (gfx1151, Ryzen AI Max+ 395) and NVIDIA DGX Spark — frozen corpora, quality gates with unit tests, sealed run bundles. Apache-2.0
A turnkey, fully-local AI workstation engineered for the AMD Ryzen AI Max+ 395. LLM inference, voice, document parsing, browser automation, agents — all on-device.
Claude Code skill for AMD Strix Halo (Ryzen AI MAX+ 395) ML setup. Handles PyTorch installation (official wheels don't work with gfx1151), GTT memory config, and environment setup. Enables 30B parameter models.
Playtest-graded benchmark for local LLM inference stacks — a coding agent builds a 10-file game, static + runtime gates + human playtest grade it. Validated on llama.cpp Vulkan / AMD Strix Halo.
Measured inference stack for Strix Halo (Ryzen AI Max+ 395, 128 GB) — llama.cpp serving plus TTS/TTI/TTV media workloads under one memory authority. Not a fork: master plus a short list of patches. Memory budget, defect registry, prefix-cache gateway.
Optimized dual AMD Strix Halo (gfx1151) vLLM MoE inference: TP=2 over USB4, tuned int4 MoE kernel, ~8us interconnect, serialized serving
Runnable local coding-agent stack for AMD Strix Halo: llama.cpp (Vulkan) server configs for Qwen3.8-Flash-Next 125B + 27B, pi agent wiring, spec decoding (MTP / DFlash2).
Linux fan control for the Minisforum MS-S1 MAX (Ryzen AI Max+ 395 / Strix Halo)
Talos-O (Omni): A sovereign, embodied agentic organism forged on AMD Strix Halo. Integrating the Chimera Kernel (Linux 7.0), Zero-Copy Introspection, and the Phronesis Engine. Built from First Principles.
Run Qwen 3.6-27B AWQ-INT4 models with DFlash speculative decoding on AMD Strix Halo hardware using vLLM for high-throughput inference.
Drop-in recipe for running faster-whisper on AMD Strix Halo (Ryzen AI Max+ 395, gfx1151) with Ubuntu 26.04 + ROCm 7.2.2 — no source build required
Docker stack: Ollama v0.21.0 built from source against ROCm 7.2.2 with native gfx1151 (Strix Halo) — serves Gemma 4 up to 256K context on AMD Ryzen AI MAX+ 395 / Radeon 8060S. Includes a 9-layer make validate ladder for the host firmware, ROCm runtime, container, and long-context inference.
Fan Control plugin for the Bosgame M5 / Sixunited AXB35 (AMD Strix Halo) embedded controller: EC temperature, fan RPM and 0-100% fan control via PawnIO, Secure Boot stays on
To associate your repository with the ryzen-ai-max topic, visit your repo's landing page and select "manage topics."