Homelab Router Part 4: The Router Picks the Model, Not You
I opened the Coder model picker and saw a list of models I did not want to choose between.
There was DeepSeek. There was Qwen. There was Granite. There was a Qwen variant on Strix, a Qwen variant on the workstation, a DeepSeek on the Spark pair. Each one was a real backend, and each one looked like a decision I had to make.
That list was wrong. Not because the backends were wrong, but because it put the choice in the wrong place. The policy router already decides which model gets a request. By exposing every backend as a separate model in the Coder picker, I was asking the model selection question twice, and giving the human the losing half of it.
This is the part of the router series where the abstraction finally stopped being a proxy and became a policy layer.
I Tested Whether the Model Name Mattered at All
Before I consolidated anything, I wanted to know one thing. Did the model name I sent to the router change what it did?
The router routes by request facts. It checks whether the request has an image, whether it declares Home Assistant tools, and how large the input is. It does not read the model field on the request. I confirmed that by sending nonsense model names straight at the endpoint.
| Request fact | Router checks |
|---|---|
has_image | Does the request contain an image? |
tools_include | Does it declare HassGetState, HassTurnOn, HassTurnOff? |
input_tokens_gt | Is the prompt bigger than the threshold? |
A request with model set to local-router-default came back with the right answer. A request with model set to anything-at-all came back with the right answer. The routing decision did not change.
That is the whole point. The router is not a proxy that forwards whatever model you ask for. It is the decision maker. The name I send is cosmetic, and acting like it was meaningful was creating a false sense of control.
The Fleet Did Not Agree on Model Names, and That Broke the Router
The second problem was that my four backends did not name their models the same way, and the router’s health check only trusts names a backend advertises.
A vLLM endpoint exposes a model ID. A llama-swap endpoint exposes canonical model IDs and keeps alternate names as aliases. A llama.cpp server does its own thing. Three different conventions on four backends.
The router checks whether a model is healthy by asking the backend for its model list and looking for the model ID it is about to forward. That worked for the names the backends published as canonical IDs. My unified names were aliases, and aliases did not show up as IDs until I turned on includeAliasesInList in the llama-swap config.
# llama-swap-config.yaml
includeAliasesInList: true # without this, aliases never appear as model IDsWith that flipped, qwen-3.8-27b finally appeared as a real model ID, and the router could see it was healthy. Without it, the model looked dead even though it served requests fine.
I landed on one consistent naming scheme across the fleet so the router, the Coder config, and the backends all agree:
deepseek-v4-flash-vision-exp Spark cluster
qwen-3.8-27b RTX 5090 workstation
granite-4.2-nothink RTX 5090 workstation
qwen-3.8-flash-next-coder-262k Strix Halo
qwen-3.8-flash-next-coder-512k Strix Halo| Backend | Canonical id | Unified alias used everywhere |
|---|---|---|
| Spark vLLM | deepseek-v4-flash-vision-exp | deepseek-v4-flash-vision-exp |
| RTX llama-swap | qwen38 | qwen-3.8-27b |
| RTX llama-swap | granite-4.2-nothink | granite-4.2-nothink |
| Strix llama-swap | qwen38-coder-262k | qwen-3.8-flash-next-coder-262k |
| Strix llama-swap | qwen38-coder-512k | qwen-3.8-flash-next-coder-512k |
The model IDs match the model servers, the aliases I added, and the Coder config.
Coder Would Not Show the Router Until It Had a Key
The router was healthy. Its models were correct. Coder still would not list them in the model picker.
The cause was not the router. The local-agent-router provider in Coder had zero API keys, and Coder does not surface models for a provider with no key. The router does not require a key, so a placeholder value works fine. It forwards any authorization header upstream and never validates it.
# Coder AI provider for the router
base_url: http://local-agent-router.internal:8088/v1
api_keys:
- dummy-local-agent-router-key # the router ignores it, Coder just needs oneThat is a small gotcha with a big consequence. The router models sat invisible in the admin forever because the provider looked incomplete. One key entry and they showed up.
The Router Config Was Drifting Away From What Systemd Loaded
I found a deeper problem while I was in there. The router config that the documentation pointed at was not the config the service read.
# the doc showed this path
/opt/local-agent-router/examples/config.yaml
# what the systemd unit loaded
/etc/local-agent-router/config.yamlThe systemd unit sets LOCAL_AGENT_ROUTER_CONFIG to a config path. The router docs and the first deployment edits pointed at a different file. So edits I made to the documented file never changed the running router. I was tuning a file nobody loaded.
I moved the real config to the path the service reads, and restarted the router. The health endpoint finally reported the two models I expected instead of a stale set.
I Collapsed Everything Into One Default Model
With the mechanics sorted, the consolidation was obvious.
The router does the model selection. Coder just needs to point at the router. So instead of five Coder model entries for the local fleet, the router path got one entry called Local Agent Router, set as the default. I kept the direct backends as separate entries too, but only so I could force a specific machine when I am testing something.
| Before | After |
|---|---|
| DeepSeek on Local Agent Router | Local Agent Router (default), one entry |
| Qwen 3.8 27B on Local Agent Router | same entry, router decides |
| Granite 4.2 No-Think on RTX | stays as direct-only test entry |
| Qwen on RTX | stays as direct-only test entry |
| Strix coder on Strix Halo | stays as direct-only test entry |
The current policy is simple:
| Request | Router sends it to |
|---|---|
| Default / long context | DeepSeek on Spark |
| Home Assistant tools | Qwen on RTX 5090 |
| Has an image | DeepSeek on Spark |
| Strix, Granite | direct-only until tuned |
That matches what the interlude concluded. Let the router remember which model earned which lane. Do not make a human re-derive it every session.
The mental model that finally stuck is this. The Coder model picker is an endpoint picker, not a model picker, when a policy router is in front of it. So I keep one default entry that says the router decides, and I keep the explicit backends as escape hatches for when I want to test one machine.
The next question is whether the household agents should all funnel through that one policy entry too, or keep choosing a backend directly.
By the Numbers
- 4 local model backends unified behind one Coder default entry
- 3 routing rules that choose a backend by request facts, not by model name
- 1 config file that was drifting from the path the systemd service loaded
- 1 missing API key that kept the router models invisible in the Coder picker
- 2 nonsense model names I sent to the router to prove it ignores the field
- 5 model IDs that now match across the backends, the router, and Coder