Teaching Coder Agents About the Local Model Fleet
Once the Strix Halo box could serve a model, I had a new problem. The homelab had three serious local model endpoints, and Coder Agents only cared if I could pick them from the model dropdown.
That sounds like a small UI problem. It was not. It forced the local AI stack to grow up from “a pile of ports I know by memory” into something closer to a capability graph.
This is a direct continuation of the Homelab Router series. Part 1 made a static router. Part 2 proved it could control a real Home Assistant light. Part 3 split fast voice control from agentic control. This session moved the same idea into Coder Agents. The router still owns policy. Coder also gets explicit model choices when I want to force a specific machine.
The Homelab Had Three Different Model Jobs
By the end of the Strix bring-up, the local fleet looked like this:
Strix Halo
Qwen3.8 Flash Next Coder
512K context
local coding endpoint
Spark cluster
DeepSeek V4 Flash Vision-Exp
1M context
multimodal heavyweight endpoint
AI Workstation
Qwen 3.6 35B-A3B
131K context
RTX 5090 incumbent local endpointThose machines do not solve the same problem. Spark is the giant multimodal endpoint. The RTX 5090 workstation is the known-good local baseline. Strix is the fresh coding endpoint with a big context window.
The mistake would be pretending one endpoint should hide all of that all the time.
The Router Stayed the Default Local Path
The local-agent-router still does useful work. It sits on the always-on Proxmox box and keeps clients from learning every backend address.
The current policy is simple:
vision requests -> Spark DeepSeek
home-assistant tools -> Granite on the RTX 5090 path
long-context requests -> Spark DeepSeek
plain coding/default -> Strix Qwen3.8That makes the router a good default provider for Coder. If I start a normal local coding chat, Strix should get the work. If the request includes images or truly huge context, the router should prefer Spark.
But I also wanted explicit overrides. Sometimes I do not want policy. I want to say, “run this on DeepSeek” or “run this on the RTX 5090 Qwen model” and know exactly which machine I am testing.
That meant adding direct Coder model entries too.
Strix Needed the Router, Not Direct DNS
A redaction note before the config snippets: this post uses role-based names and documentation IPs. The real hostnames, LAN addresses, Tailnet addresses, and Linux usernames stay out of the post for the same reason they stayed out of the Strix bring-up. They make the piece less safe and not more useful.
The Strix endpoint works locally as a normal llama.cpp server:
http://strix-ai-node.internal.example:8080/v1The Coder host could resolve local names sometimes. The Coder server process could not rely on them. I hit a DNS failure from inside the Coder AI provider path:
lookup local-router.internal.example on 127.0.0.53: server misbehavingSo I stopped being cute and used the router’s stable provider endpoint for Coder. The public version uses a documentation address:
http://192.0.2.32:8088/v1/That gave Coder a durable OpenAI-compatible provider while keeping Strix behind the router policy. 192.0.2.32 is an example address, not my LAN.
The Strix Coder model config now looks like this:
Display name: Strix Coder (Qwen 3.8 Flash Next)
Model: qwen3.8-flash-next-coder
Context: 524288
Compression threshold: 85%That last number matters. The first GLM attempt had Coder set to 32768 tokens with a 70% compression threshold. The agent started summarizing around 23K tokens. That felt broken, but it was just configured exactly that way. With the Qwen endpoint at 512K and compression at 85%, Coder can run much longer before it has to compress.
Spark Needed a Direct Model Entry
The policy router does not honor a user’s requested model as an override. That is intentional. If I send a plain request through the router and set model: deepseek-v4-flash-vision-exp, the default route still chooses Strix, because the router routes by request facts rather than by vibes in the model field.
That is exactly what I want for the default local path. It is not what I want when I explicitly choose DeepSeek from a dropdown.
So I added Spark as a direct Coder provider:
Provider: Spark DeepSeek
Base URL: http://192.0.2.4:8888/v1/
Model: deepseek-v4-flash-vision-exp
Display name: DeepSeek Spark (Vision-Exp)
Context: 1048576This bypasses the router and talks straight to the vLLM OpenAI-compatible server on the Spark head node. The example address is documentation space, not the real Spark address.
That does not break the policy router. It gives me two tools:
Default local routing: use Strix unless policy says otherwise
Explicit DeepSeek test: pick Spark DeepSeek from the dropdownBoth are useful. They are just different control surfaces.
The RTX 5090 Model Was Already There, But Not Explicit Enough
The AI Workstation already ran llama-swap on port 8080. The currently loaded model was Qwen 3.6 35B-A3B:
Model ID: qwen
Context: 131072
GPU: RTX 5090
VRAM used: ~24GB of 32GBThat model has been the incumbent daily driver in the homelab. It deserved a named dropdown entry instead of hiding behind an old generic local provider.
So I added:
Display name: AI Workstation Qwen 3.6 35B-A3B
Model: qwen
Context: 131072
Compression threshold: 85%Now I can pick the RTX 5090 model directly when I want the older baseline instead of the new Strix endpoint.
The Final Dropdown Matches the Actual Homelab
The Coder model list now has three local choices that map to three physical roles:
Strix Coder (Qwen 3.8 Flash Next)
Fresh coding model, 512K context, AMD Strix Halo
DeepSeek Spark (Vision-Exp)
Heavy multimodal model, 1M context, dual Spark cluster
AI Workstation Qwen 3.6 35B-A3B
Incumbent local baseline, 131K context, RTX 5090This is the part that feels different from earlier local model experiments. I am not picking between random GGUF files anymore. I am choosing the machine and role I want for the task.
That is the point of the router series. The homelab should not only run models. It should know what each one is for.
The Gotchas Were Mostly Boring and Important
The biggest problems were not model problems.
The first was path construction. My first Strix provider attempt made Coder call:
POST /strix-halo/v1/chat/completionsThat failed with a 404. Coder needed a provider base URL it could treat like a normal OpenAI-compatible endpoint, not a name that turned into an internal route prefix.
The second was model health. The router checks /v1/models and only marks a configured model healthy if the configured ID appears in the backend response. I first configured Strix as strix-coder, but the backend advertised glm-4.5-air-q5. The node was healthy. The model was not. Matching IDs fixed it.
The third was DNS. Local names are convenient in shells and brittle inside services. Stable, redacted provider endpoints won for Coder configuration.
None of this is glamorous. All of it matters if I want Coder Agents to treat local models like real first-class options.
By the Numbers
- 3 local model backends — Strix Halo, Spark, and the RTX 5090 workstation now appear in Coder.
- 524288 tokens — the Strix Coder context window after the Qwen3.8 swap.
- 1048576 tokens — the DeepSeek Spark context window exposed as a direct Coder model.
- 131072 tokens — the AI Workstation Qwen context window.
- 85% compression threshold — the setting now used for the local Coder model configs added in this session.
- 1 policy router — still the default local path, but no longer the only way to reach a model.
- 2 direct overrides — Spark DeepSeek and RTX 5090 Qwen can be selected explicitly.