Homelab Router Interlude: DeepSeek Wins the Local Coder Bakeoff
The homelab router did not need another diagram today. It needed a personnel decision.
After wiring the local model fleet into Coder Agents, I had four plausible ways to ask an agent to write code: GPT-5.5 in the cloud, Qwen on the Strix Halo node, Qwen on the RTX 5090 workstation, and DeepSeek V4 Flash Vision-Exp on the Spark pair.
This was not a test of the router. I did not route a smart-home command. I did not touch the Home Assistant path. I wanted a more basic answer before I harden the router policy: when I ask the local stack to do real Coder work, which model should get the job?
That answer matters because the Homelab Router is becoming a policy layer, not just a proxy. Part 1 made the static router. Part 2 proved the agentic path could turn off a real light. Part 3 split fast voice control from slower agentic control. This interlude asks which model belongs behind the coding lane before I settle the broader model policy.
I Gave Every Model the Same Coder Job
The task was intentionally ordinary. I asked each model to create a fresh Coder workspace, clone the blog engine, make a feature branch from the same baseline commit, and add a lightweight Draft Readiness Checklist to the admin preview and edit flow.
The feature was useful but small. The checklist had to inspect generated MDX and flag common draft problems before I saved or published a post:
- missing or malformed frontmatter
- missing title, description, date, tags,
published,type, orsyndicate - empty body
- missing
## By the Numbers - slug and title drift
The instruction was also deliberately annoying in one important way. I told the agents not to commit until I released them. That let each model work in isolation without leaking code into GitHub or contaminating another run. Once every model stopped, I told them to make a local commit only. No pushes.
The four contestants were:
| Candidate | Shape | Served from | Context exposed to Coder | Why it was in the race |
|---|---|---|---|---|
| GPT-5.5 | Cloud frontier model | Coder’s cloud provider path | about 1M tokens | The reference answer. If the local fleet cannot beat it on quality or privacy, at least I know the bar. |
| Qwen3.8 Flash Next Coder | MoE-style architecture, 125B parameters with 6B active, plus n-gram embeddings and MTP | Strix Halo through llama.cpp | 524K tokens | The fresh coding candidate on the new AMD node. Huge context, local control, unknown patience. |
| Qwen3.6 35B-A3B | MoE, 35B total with 3B active | RTX 5090 workstation through llama-swap | 131K tokens | The incumbent local baseline. Fast enough to feel practical and already wired into the homelab. |
| DeepSeek V4 Flash Vision-Exp | Multimodal DeepSeek V4 Flash variant with MoE and DSpark support | Dual Spark cluster through vLLM | about 1M tokens | The heavyweight local generalist. It already won the broad homelab agent benchmark. |
I used real Coder Agents, real top-level chats through the Coder API, and real fresh workspaces. Workspace creation counted as part of the test. That matters. A coding agent that cannot bootstrap its environment is not faster just because it writes a nice final paragraph.
DeepSeek Wrote the Best Merge Candidate
DeepSeek won the code-quality round.
It did not write the most ambitious solution. Strix did. It did not finish first. GPT-5.5 did. DeepSeek won because it landed the best balance: the right files, a compact diff, a pure helper, both admin flows wired up, and enough verification to trust the shape of the change.
The final scorecard looked like this:
| Model | Wall clock | Nudges | Files changed | Score |
|---|---|---|---|---|
| DeepSeek Spark | 24m 45s | 0 | 4 | 85 |
| Strix Qwen 3.8 | 71m 58s | 0 | 6 | 81 |
| GPT-5.5 | 4m 35s | 0 | 4 | 69 |
| Workstation Qwen 3.6 | 6m 37s | 1 | 3 | 57 |
The DeepSeek version added src/lib/draft-readiness.ts, a small DraftReadinessChecklist component, and integrations in both PostPreview and /admin/edit/[slug]. It kept the checklist advisory. Save and publish still worked. A valid fixture returned no issues in a helper probe.
That was the important part. I wanted a feature I could imagine merging after a little human polish, not a demo that only looked good in the final answer.
Strix Was Thorough Enough to Become a Warning Label
Strix was fascinating because it over-delivered.
It added a 433-line validator, a panel component, integrations in three edit surfaces, and a 275-line fixture script with eight cases. Its local endpoint used llama.cpp server, while Spark DeepSeek came through a vLLM-compatible endpoint. It caught nested frontmatter pitfalls, bad booleans, invalid post types, future dates, body length, date-prefixed slugs, and malformed fences. It even avoided pulling YAML parsing into the browser bundle and explained why.
That is good agent behavior in one sense. It found edge cases I would want in the final version.
It also took almost 72 minutes.
That is the router-policy lesson. The Strix node may still be useful as a local coding endpoint, especially for background work where thoroughness beats latency. But I should not make it the default fast coding lane just because it has a big context window and a model name with “Coder” in it. The router needs policy that includes time, not just capability.
This rhymes with the earlier DeepSeek smart-home benchmark. DeepSeek won the broad score there too, but Granite still owned the Home Assistant slice. One model winning one kind of work does not mean it should own the whole house.
The Fastest Runs Were Not the Best Runs
GPT-5.5 finished in 4 minutes and 35 seconds. It wired the feature into the right flows and produced a plausible implementation. It also did not test behavior deeply, and its checklist UI was noisier than I wanted. It felt like the fastest acceptable first draft, not the best merge candidate.
The RTX 5090 Qwen run finished in 6 minutes and 37 seconds, but it needed one nudge. It also missed the generated preview flow and mishandled inline tags like tags: ["meta", "agents"]. That is a common shape in this repo. Missing it matters.
This was the useful contrast. Speed alone would crown GPT-5.5. Thoroughness alone would crown Strix. The model I would actually build on was DeepSeek.
That is exactly the kind of decision the router needs to encode. Local policy should not ask, “Which model is best?” It should ask, “Which model is best for this job, under this latency budget, with this failure cost?”
This Was a Router Policy Test, Not a Router Test
I want to be precise about what this proves.
It does not prove the local-agent-router can route Coder traffic better than Coder’s model picker. It does not prove DeepSeek should control lights. It does not test live Home Assistant actions. It does not settle every local model question.
It does prove that my policy table needs a coding lane with real evidence behind it.
The current mental model now looks like this:
fast voice command -> Home Assistant Assist
smart-home agent path -> Granite-style specialist, with fallback
local coding default -> DeepSeek until Strix gets faster or more constrained
background deep review -> Strix may be worth the wait
vision or long context -> Spark DeepSeek
cloud control baseline -> GPT-5.5 when I want the reference answerThat is not final policy. It is a better starting point than vibes.
It also changes how I should tune the next bakeoff. I launched all four runs with the same nominal reasoning setting. That was fair enough for a first comparison, but it probably punished Strix. Future local runs should smoke-test model settings first. A local model in an expensive thinking mode is not the same product as the same model tuned for agentic coding speed.
The Router Is Becoming a Model Operations Layer
The early router posts focused on routing requests. That was the obvious problem. Hermes needed one OpenAI-compatible endpoint, and the homelab had multiple model backends.
The deeper problem is model operations.
Every local model now has a shape:
- capability
- latency
- reliability
- context window
- tool behavior
- setup friction
- best use case
The router should eventually know those shapes. Coder AI Gateway can expose explicit model choices. The homelab router can expose policy defaults. The two should reinforce each other instead of hiding the fleet behind one magic endpoint.
That is why this bakeoff belongs in the Home Lab Router series. It did not build the router. It gave the router better data.
DeepSeek gets the next coding-lane vote. Strix gets another run after I tune its settings. The RTX 5090 Qwen path stays useful as a fast local baseline. GPT-5.5 remains the cloud reference I measure against.
The router’s job is not to crown one model forever. It is to remember which model earned which lane.
By the Numbers
- 4 Coder Agents contestants: GPT-5.5, Strix Qwen, RTX 5090 Qwen, and Spark DeepSeek
- 4 fresh Coder workspaces created as part of the test
- 1 nudge required, for the RTX 5090 Qwen run after it stopped early
- 4m 35s GPT-5.5 wall time, the fastest run
- 6m 37s RTX 5090 Qwen wall time, the fastest local run
- 24m 45s DeepSeek Spark wall time, the winning implementation
- 71m 58s Strix Qwen wall time, the most thorough but slowest run
- 85 / 100 DeepSeek Spark’s final score
- 829 lines added by Strix, the largest diff
- 343 lines added by DeepSeek, the best merge candidate