vibescoder

Turning Strix Halo Into a Headless Local AI Node

·9 min read

The Strix Halo box arrived with Windows 11 Pro, a shiny desktop login, and 128GB of unified memory. I immediately wanted to hand all that to Linux.

Originally, I envisioned the Strix Halo as my daily desktop. Runs Windows for work apps, does LLM as a side hustle. Then I remembered something important. I hate Windows. I’ll suffer through it for gaming, but not work. And even then, I can’t wait for SteamOS to support NVIDIA GPUs and Xbox Game Pass. Then it’s goodbye Windows everywhere.

Until then, I don’t want this machine to be a desk ornament. It was the missing AMD leg in the homelab I sketched in the local AI hardware post: Spark for the giant model, the RTX 5090 workstation for fast local experiments, and Strix Halo for a third always-reachable local AI node with a huge memory pool.

So I found a good candidate and wiped it.

The Windows Box Became a Headless Ubuntu Node

I built the Ubuntu installer from the AI workstation, wrote it directly to a USB stick, and booted the new box into the installer. The only Windows drama was predictable. Changing boot settings triggered a BitLocker recovery prompt. Since I planned to erase the disk, I ignored Windows Boot Manager and booted the USB instead.

I kept the install boring on purpose. No ZFS. No LVM. No full-disk encryption. A homelab inference node that needs a keyboard after every reboot is not a server. It is a desktop with aspirations.

The first boot turned into the usual server checklist:

sudo apt update
sudo apt full-upgrade -y
sudo apt install -y openssh-server curl git vim htop btop nvme-cli lm-sensors
curl -fsSL https://tailscale.com/install.sh | sh
sudo tailscale up --ssh=false

Then I added the workspace SSH key, confirmed the Tailnet address, and switched the box to headless boot. The machine now lands on a local role-based login prompt after reboot, while SSH and Tailscale come up automatically.

That last part mattered more than the install. If a local AI node needs me to log into GNOME before it can serve a model, it is not part of the homelab yet.

Key-Only SSH Made the Box Feel Like the Rest of the Homelab

The box joined the fleet with a role-based hostname, a Tailnet address, and a normal non-root admin user. I am not publishing the literal hostname, LAN IP, Tailnet IP, or Linux username here. Those details do not help anyone reproduce the setup. They only turn a useful post into future cleanup work.

So the public version uses role names and placeholders:

Hostname: strix-ai-node.internal.example
Tailnet: <tailnet-address>
LAN: 192.0.2.41
User: <admin-user>

192.0.2.0/24 is documentation space, not my LAN. The structure is real. The address book is not.

I hardened SSH the same way I did the other machines. Password auth off. Root login off. Key auth on. Sleep and hibernate masked. NetworkManager, ssh, tailscaled, unattended-upgrades, and fstrim.timer all enabled at boot.

I also switched the default target from the desktop to multi-user.target and masked GDM. That made the box boring in the best possible way. Reboot it, wait a minute, SSH in.

The first clean reboot proved the core path:

ssh: active
Tailscale: active
NetworkManager: active
gdm: masked / inactive
sleep targets: masked

At that point it was a homelab machine. It was not yet a useful AI machine.

Vulkan Saw the GPU Only After I Could Reach Render Devices

The hardware showed up exactly how I hoped:

CPU: AMD Ryzen AI Max+ 395 w/ Radeon 8060S
RAM: 123GiB visible
GPU: Radeon 8060S Graphics (RADV GFX1151)

But my first headless Vulkan check only saw llvmpipe. The fix was simple and easy to forget on a fresh Ubuntu install. I needed access to the render and video groups:

sudo usermod -aG render,video <admin-user>
sudo reboot

After that, vulkaninfo reported the actual GPU:

Radeon 8060S Graphics (RADV GFX1151)
Mesa 25.2.8

That unlocked the practical path. I built llama.cpp with Vulkan, not ROCm. That was a deliberate choice, not a reflex.

I looked at the usual suspects before committing the box to a serving stack:

StackWhy It Was TemptingWhy I Did Not Start There
llama.cpp VulkanIt runs GGUFs directly, exposes an OpenAI-compatible server, and uses Mesa RADV on the Radeon 8060S without a full ROCm userspace stack.This was the boring path, which is exactly why I chose it first.
ROCm / HIPNative AMD compute should eventually be the clean answer for this APU.Strix Halo support is still moving fast. I did not want the first bring-up blocked on ROCm version pinning, kernel edge cases, or model-specific HIP bugs.
LemonadeAMD is building it for local AI on Ryzen and Radeon hardware, and it can present a friendlier server layer.It wraps the same underlying runtime choices. I wanted to prove the raw backend first before adding an orchestration layer.
OllamaIt is easy to operate and already has a good service model.It hides too many knobs. Context size, KV cache types, YaRN settings, and exact GGUF choices mattered for this test.
LM Studio headlessIt is pleasant for desktop model testing and now has headless support.This box was becoming infrastructure, not a desktop app host. I wanted systemd, logs, and a plain server binary.
vLLMIt is the right shape for high-throughput OpenAI-compatible serving.On this machine, the near-term goal was large GGUFs on unified memory. vLLM is a better future experiment than first boot path.

That left Vulkan llama-server as the first production candidate. It gave me the control I needed and the fewest new moving parts.

The build was uneventful after installing the missing SPIR-V packages. The memory tuning was not.

Unified Memory Still Needed a Kernel Limit Raised

The box has 128GB of unified memory. Vulkan did not get 128GB by default.

That was the first real trap. AMD’s Strix Halo guide explains why. On RDNA3.5 APUs, the GTT/TTM limit controls how much system memory GPU processes can map. The default is roughly half of RAM. That is fine for games. It is not fine when I ask a local model runtime to map a 78GB GGUF.

At first, RADV exposed about 42GB of device-local heap. I raised TTM to 110GB and rebooted. RADV exposed about 74.5GB. Still not enough for GLM-4.5-Air Q5.

The setting that worked was 120GB:

TTM pages_limit: 31457280
TTM page_pool_size: 31457280
RADV device-local heap: ~81.2GB

The persistent config lives here:

/etc/modprobe.d/99-strix-halo-llm-ttm.conf
/etc/default/grub.d/99-strix-halo-llm.cfg

A live sysfs write changed the reported TTM value, but Vulkan did not see the new heap until reboot. That is the kind of detail that turns a five-minute model test into an evening.

GLM Proved the Stack Worked, Then Lost the Job

My first serious model was GLM-4.5-Air Q5. The Strix Halo enthusiast community raved about it. Makes sense. It fit after the 120GB TTM change. It answered the smoke test. It ran around 20 tokens per second on a tiny generation.

Then I checked the release date.

GLM-4.5-Air landed on Hugging Face in July 2025. The GGUF was from August 2025. That is ancient in local model years. The model proved the box could run a big GGUF through Vulkan, but it did not match the job anymore. I wanted a current coding model with a large context window, not a museum piece that happened to fit.

Don’t believe everything you or your agent reads on the internet.

I switched targets.

Qwen 3.8 Flash Next Gave the Box Its Actual Job

The better fit was Qwen3.8 Flash Next GSQ-RCO Coder, a September 2026 compressed coding variant of Qwen3.8 Flash Next. The model card made the hardware story much more interesting. The download is about 58.4GB, but only about 29.6GB must be resident. The second shard is an n-gram table that can be served from disk.

That is exactly the kind of weird model format a homelab exists to test.

The first 512K attempt loaded but capped itself to 262K. Then I added YaRN settings:

--ctx-size 524288
--rope-scaling yarn
--rope-scale 2
--yarn-orig-ctx 262144
--cache-type-k q8_0
--cache-type-v q8_0

This time the server reported:

n_ctx: 524288
n_ctx_train: 524288

It still warned that the base model trained at 262K before the scaling took effect. That is fine. The point of the test was to see whether the box could load and serve the 512K configuration without falling over.

It did.

The smoke test was intentionally dumb:

Prompt: What is 17*19? Return only the final integer.
Answer: 323

The timings were better than I expected:

Prompt eval: ~29 tokens/sec
Generation:  ~25 tokens/sec

While loaded, the machine still had about 50GB available. That matters. The GLM experiment felt like I was wedging a model into the box. The Qwen 3.8 experiment felt like the box found its shape.

The Strix Box Became the Fresh Local Coding Node

The final service is now a normal systemd unit:

strix-qwen38-coder.service

It listens on port 8080 and advertises three aliases:

qwen3.8-flash-next-coder
qwen38-flash-coder
strix-coder

The Strix Halo box now fills the role I wanted from the beginning. It is not the biggest model in the house. The Spark cluster owns that job. It is not the fastest local GPU box. The RTX 5090 workstation still has that lane.

It is the freshest local coding endpoint with a 512K context window, running on a headless machine I can reach from Coder Agents.

That is more useful than the spec sheet.

By the Numbers

  • 1 Windows install — wiped and replaced with Ubuntu 24.04.5.
  • 128GB unified memory — the reason Strix Halo was worth testing at all.
  • 120GB TTM limit — the kernel-side setting needed before large Vulkan allocations became practical.
  • ~81GB RADV heap — the Vulkan memory budget after the TTM reboot.
  • 58.4GB — the Qwen3.8 Flash Next GSQ-RCO Coder download size.
  • 512K tokens — the context window that loaded with YaRN scaling.
  • ~25 tokens/sec — smoke-test generation speed on the Qwen endpoint.

Comments