vibescoder

Homelab Router Part 5: Seven Models, One Writing Job

·9 min read

The interlude settled the coding lane. DeepSeek on Spark got the agent work, Strix got a warning label, and the router policy table codified the coder.

But the blog writing was also defaulting to DeepSeek. Is that bad? Probably not. The writing has been serviceable. But it’s also also. At least the way I work. I like to iterate on the post with my agent. It’s pair writing to my vibe coding pair programming.

So I decided my writing model would be resident on my token-munching RTX 5090. I chose seven local models to throw at my one GPU. Each model received three writing tasks, and all scoring was blind. A sealed identity map stayed sealed until both scoring passes were finished.

This is Part 5 of the Homelab Router series. Parts 1 through 4 built the router, proved the agentic path, and refined the model picker logic. This bakeoff seeks to test the first of many use case specific models: writing.

I Put the Whole Fleet on the Same Page

The field was seven entries, all served through llama-swap on the RTX 5090 workstation, one resident at a time:

EntryModelWhat it brought to the race
qwen38Qwen 3.8 27B, dense with a Gated DeltaNet hybridThe incumbent blog writer. Fast KV, thinks a lot before it writes.
muse-glimmerMuse Glimmer 30B with a DFlash draft modelMeta’s new dense writer. Speculative decoding should make it fast.
gemma-4Gemma 4 31B, QAT q4_0, visionGoogle’s fresh 31B. Recommended quant, no thinking overhead.
granite-4.2Granite 4.2 30B, thinking onThe controlled variable.
granite-4.2-nothinkSame weights, thinking offThe other half of the controlled variable.
glm-4.7-flashGLM-4.7-FlashThe speed demon. 200 tok/s measured.
hermes4.3Hermes 4.3 36BThe wildcard I had been leaning on too much.

Three tasks, each written from real source material:

  • Task 1, a Friday Fixes post from fresh session fodder
  • Task 2, a build post for a new writing-intent route in the router
  • Task 3, an edition of my wife’s LinkedIn newsletter, using her own skill file

Twenty-one runs. Every model got the byte-identical prompt for a given task, a 16,000-token cap, and nothing that mentioned a competition. The harness logged timings, token counts, and thinking tokens separately, because a thinking model that writes 11,000 invisible tokens before the first visible word is not the same animal as a slow model. Consider that foreshadowing.

The Protocol Was Blind, and It Changed What I Trusted

The reading sheets labeled the drafts A through G. No model names. No ordering hints. I read all seven drafts for a task, scored them, and wrote my notes before the judge pass ran. The judge was a model that was not in the field, qwen3.8-flash-next-coder on the Strix Halo node, and it saw the same lettered sheets with no identities. Only after both passes were done did the sealed map get opened.

The blinding worked. Without it, I would have read qwen38 into draft B before I finished the first paragraph. I thought it would win. I daily-drive Qwen 3.6 on that RTX 5090, and I figured the newer 3.8 would just be better. Same architecture family, one generation newer, so the tidy story was already written before I read a word. The blind assessment forced me to look at secondary factors like speed to determine a winner.

The results:

TaskMy pickJudge’s pick
Friday Fixesqwen38 (79/100)qwen38
Router build postqwen38 (86/100)gemma-4
LinkedIn newslettermuse-glimmer (my wife picked)muse-glimmer

There were two human readers: me and my wife. And an AI reader: Qwen 3.8 Flash Next Coder with 512K context on the Strix. We agreed on the winner in two of the three tasks. On the router post we picked different winners but agreed on the top three set, qwen38, muse-glimmer, and gemma-4. On the newsletter we agreed completely, winner included.

The House Writer Was Not the Fastest Writer

qwen38 won both house-style blog tasks, and it did it the way I wanted. Draft B on Friday Fixes read like the blog. Draft F on the router post read like the blog, and I called it the strongest complete build narrative in the set. Both still needed a five-point edit pass before they were publishable.

But then there’s my preference to pair write with my agent. qwen38 thinks before it writes, and it thinks a lot. On the two blog tasks it generated 11,400 to 11,900 thinking tokens per draft. At 73 tok/s that is about three minutes of wall time per post, most of it invisible.

muse-glimmer is the interesting counterweight. It is nearly twice as fast, 133 to 147 tok/s, with 57 to 66 percent of its draft tokens accepted. Across all three tasks it finished in 82 seconds of generation time while qwen38 spent 435, most of it thinking. The fast drafts were also the shorter ones, 609 to 777 visible words against qwen38’s 891 to 1,319. That tradeoff cost it on the house-style posts. On the structured skill prompt the length was not the point, and it was simply the best in the room, a better all-round writer than I expected from a model that seemingly emerged from nowhere a few short weeks ago.

gemma-4 was the judge’s favorite on the blog tasks. It ranked first on the router post and second on Friday Fixes, while my read had it third and fifth. On Friday Fixes the gap was the widest in the bakeoff. I gave it 64, the judge called it the runner-up. I read it as competent, clean, and a little safe. The judge read the best voice fit of the set. One of us is wrong, and the blind protocol at least made sure we were wrong independently.

Then there was the speed check. glm-4.7-flash hit 200 tok/s and finished each draft in eleven to fifteen seconds. It also wrote 365 words on the Friday Fixes task, under half the band a real post needs, and it finished in the bottom half of all three judge rankings. It reads like a lazy student phoning in his homework.

hermes4.3 was last or near-last in both qualitative reads, and its router draft contained the bakeoff’s only outright factual reversal. It claimed a test no longer hung when the source clearly said it had. I’m a vibe coder, not a vibe writer. I strive for factual accuracy.

The Thinking Control Answered a Question I Was Avoiding

Granite 4.2 ran twice, same weights, thinking on and off. It was never a contender for the writing task. It’s an agentic tool call specialist. But it was the easiest model to A/B test for thinking.

The thinking run cost 5.5 times the wall time of the nothink run. 313 seconds against 57. The quality did not move. Thinking ranked third or fourth. Nothink ranked fourth or fifth. Adjacent in my read, within two in the judge’s.

That is the most expensive data point in the bakeoff, and the cheapest lesson in it. Thinking overhead is a real cost line on the router’s budget, and on a 30B dense writer it did not buy anything measurable. That also hurt in Home Assistant testing. I don’t need a model to debate the value of a light switch. I just need it on, off, or dimmed.

The Writing Lane Gets a Name

The policy table from the interlude and Part 4 had a coding lane and a vision lane (for seeing screenshots and websites). Now it has one writing lane, and it surprised me.

blog tool calls         -> muse-glimmer, ~30 s, the daily writer
long-context or vision   -> Spark DeepSeek, unchanged
coding                   -> Spark DeepSeek, unchanged

qwen38 won the house-style posts on the merits. It read the most like the blog, and two of my three top marks went to it. If the writing lane was decided by an objective score alone, qwen38 owns it.

But I do not run the writing lane on a score alone. I pair-write. I sit with the agent and iterate on a post, and I want a model that turns drafts around fast enough for that loop to feel like a conversation. muse-glimmer was nearly as good and roughly twice as fast, and its drafts landed the structured-skill task better than anything else in the field. That is the trade I want for a daily writer. The bakeoff picked qwen38 on merit, and I picked muse-glimmer for the way I work.

That is the difference between a bakeoff and a policy. The bakeoff tells you which model was best. The policy tells you which model you will use, and it is allowed to weigh something other than the score. I am not pretending muse-glimmer out-wrote qwen38. I am saying it wrote well enough, and it does it in a fraction of the time, and that is what a pair-writing loop needs.

The blog_draft route in the router now prefers muse-glimmer and falls back to deepseek. The route’s job is to remember which model I chose for the lane, not to re-litigate the bakeoff every time I write. That is the difference between a policy layer and a proxy. The proxy forwards whatever you ask for. The policy remembers the decision, and it only changes its mind when the evidence says so.

The next bakeoff will not be a model bakeoff. It will be a prompt bakeoff. Same model, three different versions of the house style, and I want to know how much of what I have been crediting to the model is the prompt doing the work.

Which of the seven would you have trusted with your voice before the map got opened?

By the Numbers

  • 7 models, 3 tasks, 21 runs, zero DNFs, one GPU
  • 1 sealed mapping, opened only after both scoring passes finished
  • 11,900 thinking tokens in qwen38’s best router draft, at 73 tok/s
  • ~2x faster, muse-glimmer versus qwen38, with 57 to 66% draft acceptance
  • 15 s wall for glm-4.7-flash’s 365-word draft, under half the target band
  • 5.5x wall-time cost for granite thinking, with no ranking improvement
  • 1 factual reversal in 21 drafts, in the draft that finished last or near-last in both reads

Comments