vibescoder

Homelab Router Part 3: Hey Jarvis, Turn Off the TV Lights

·9 min read

The best local smart-home voice assistant in my house is not Hermes.

That was the useful lesson from the last few days. Part 1 put a policy router between Hermes and the local model fleet. Part 2 proved Hermes could turn off a real Z-Wave light through Home Assistant MCP, once the Z-Wave dongle, Z-Wave JS, and entity mapping were all correct.

Then I measured the typed Hermes path. It worked, but it took about 23 seconds after optimization.

That is fine for an agentic assistant. It is not spouse-compatible for “turn off the lights.” The product requirement from my wife was simpler and more brutal: make it as fast as Google Home, or faster. No one wants to stand in a dark hallway waiting for an LLM to reflect on the ontology of lamps.

So I stopped trying to make Hermes the low-latency voice assistant. I made Home Assistant Assist the fast local voice path instead. It is faster than Google in the one test that matters: the light changes before the conversation becomes awkward.

The Split Became Obvious Once the Lights Worked

The router path is powerful because Hermes can reason, search, use tools, and make policy-heavy decisions. It is the right place for questions like:

  • Is anyone home before shutting this down?
  • Turn off the TV lights if nothing is playing.
  • Audit which Home Assistant entities are exposed to voice.
  • Choose the right local model for this request.

It is the wrong place to pay a 20-plus-second tax for a basic light toggle.

Home Assistant already has a purpose-built voice stack. Assist can take a command, match an intent, and call the device service directly. That is the right fast path for common household commands.

The architecture now has two lanes:

Fast local voice control:
Pixel Tablet -> HA Assist -> Z-Wave -> lights
 
Agentic local control:
Hermes -> local-agent-router -> HA MCP -> Z-Wave -> lights

That is not a retreat from the router. It is the router finding its actual job. Hermes should not be the thing that flips every switch in the house. It should be the thing that handles the weird requests where policy and context matter.

The ThinkCentre Had Enough Room for the Voice Brain

Before installing anything, I checked the always-on ThinkCentre. It is not a monster box. It is an i5-8400T Proxmox host with six cores and 16GB of RAM, running Home Assistant OS and a handful of containers.

The headroom was fine:

Host memory available: about 8GB
HAOS VM: 4 cores, 8GB RAM
Home Assistant Core: under 1GB RAM
Z-Wave JS: about 115MB RAM

That was enough for a conservative local voice pipeline.

I checked the Home Assistant side next. Assist was already present, and the default pipeline existed, but it was text-only:

assist_pipeline: yes
conversation: yes
stt_engine: null
tts_engine: null
Whisper: not installed
Piper: not installed
openWakeWord: not installed

So the work was not “turn on voice.” It was building the pieces Assist expects.

I installed two official Home Assistant add-ons:

Whisper -> speech-to-text
Piper   -> text-to-speech

Both installed cleanly, but that still did not make the pipeline usable. The add-ons expose voice services to Home Assistant through the Wyoming Protocol, a small protocol Home Assistant uses for STT, TTS, wake word, and satellite services. After the containers were healthy, Home Assistant had pending discovery flows for two Wyoming services:

Piper   -> tts.piper
Whisper -> stt.faster_whisper

Those flows had to be accepted. Only then did the service entities exist.

Then the preferred Assist pipeline still had to be updated. The final pipeline became:

STT: stt.faster_whisper
Language: en
Conversation: conversation.home_assistant
TTS: tts.piper
Voice: en_US-lessac-medium
Prefer local intents: true

That last line matters. For simple home-control sentences, I want Home Assistant’s local intent engine to win before anything tries to be clever.

I deliberately skipped openWakeWord on the ThinkCentre. The wake word in this setup runs on the Android tablet, not on the server. The tablet listens locally for the wake phrase. Only after it hears the wake word does it send the command to Home Assistant for transcription and intent handling.

That separation is important:

Pixel Tablet:
  wake-word detection
  microphone capture
 
ThinkCentre:
  Whisper transcription
  Home Assistant intent matching
  Piper response
  Z-Wave control

The gotcha is subtle but worth calling out. Installing voice add-ons is not the same as giving Assist a voice pipeline. The add-ons, Wyoming integrations, and preferred pipeline are three separate steps.

The Pixel Tablet Became the Voice Appliance

The hardware choice was a Pixel Tablet. It already has a microphone, speaker, battery, dock, screen, and Android. That makes it a much better first test device than a pile of ESP32 parts and optimism.

The Home Assistant Android docs call out exactly the path I needed: use the Companion App, set Home Assistant Assist as the Android default assistant, then enable wake word detection.

The wake word piece is local to the tablet. Android uses microWakeWord on-device. Audio is not sent to Home Assistant until the wake word is detected. After that, the command goes to the ThinkCentre for Whisper, intent handling, and Piper.

The tradeoff is battery. Google Assistant gets special low-power hardware access on supported devices. Third-party assistants do not. Home Assistant has to use normal audio processing, which costs more power.

For a docked Pixel Tablet that mostly lives as a home control panel, that is acceptable. For a phone in your pocket, I would be more careful.

The Wake Word Was Grayed Out Until Android Trusted Assist

The setup was not quite one click.

The Pixel app could use Assist manually, but the wake word toggle was grayed out. That was not a ThinkCentre problem. The server pipeline was already ready. The Android side had not made Home Assistant the default assistant app yet.

The fix was on the tablet:

Android Settings
  -> Default apps
  -> Digital assistant app
  -> Home Assistant

Then the Home Assistant app could enable wake word detection. The wake word became:

Hey Jarvis

That is a small thing, but it changes the shape of the project. The Hub Maxes are still good devices. I am just no longer required to route through Google for this path.

The First Voice Commands Worked but Naming Still Mattered

The first successful command felt exactly like the system I wanted:

Hey Jarvis, turn off TV lights.

The Pixel Tablet heard the wake word, Home Assistant Assist handled the command, and the Z-Wave dimmer turned off the physical lights.

Then it missed a word.

When I said “turn off the TV lights” too quickly, it heard or parsed the phrase as “turn off the TV.” Home Assistant then looked for a device called TV instead of the light entity. That is not a model failure. That is normal voice UX. Short names get clipped. Nouns matter.

The fix was to add aliases.

For the TV lights:

TV
TV light
TV lights
TV Room lights
television lights
media lights

For hallway:

Hallway
Hallway light
Hallway lights
hall lights

I also added area aliases for the TV Room and Hallway so Assist has multiple ways to map speech to the intended device.

The practical lesson is simple. A voice assistant is not done when the entity exists. It is done when the names people actually say work reliably.

The Timing Story Got Clearer

The latency numbers now make sense because the stack has separate lanes.

Direct Home Assistant service calls are nearly instant from HA’s point of view:

Direct HA service to observed state: about 0.04 seconds

Hermes through the router started at roughly 37 seconds for the no-entity command. After adding a Home Assistant skill, caching aliases, and switching from the kawaii personality to concise responses, it dropped to about 23 seconds.

That is a big improvement, but it is still an agent loop.

Home Assistant Assist is the right thing to measure next because it skips the model reasoning loop entirely. It should be much closer to the direct HA path plus STT/TTS overhead. That is the number that matters for a dedicated room voice assistant.

The important part is that both lanes converge on the same backend:

Home Assistant entities
Z-Wave JS
GE/Jasco dimmers

The control plane is shared. The user interfaces can specialize.

This Changed the Role of the Router

The router is still the project. It just is not the wake word.

The right job split now looks like this:

HA Assist:
  Fast commands.
  Local voice.
  Lights, switches, scenes, rooms.
 
Hermes + local-agent-router:
  Slower but smarter commands.
  Policy decisions.
  Cross-tool reasoning.
  Model selection.
  Home Assistant plus everything else.

That makes Part 4 clearer. The router still needs dynamic node heartbeat and self-advertisement. Machines should announce their capabilities instead of living forever in static YAML. But the voice-control layer no longer has to wait for that. The fast path already works.


The sentence that mattered was not complicated:

Hey Jarvis, turn off TV lights.

It worked. Locally. Through Home Assistant. Through Z-Wave. Without Google. Without the Hub Max. Without Hermes waiting 23 seconds to decide how to call a light service.

That is the shape I want for the house. Fast local control for the boring commands. Local agents for the interesting ones.

Oh, and not only is it spouse-approved. It’s spouse-endorsed. The local loop is more responsive than our old Google Home cloud relay. Winning!

By the Numbers

  • 1 Pixel Tablet promoted from general tablet to Home Assistant voice panel
  • 2 local voice add-ons installed on the ThinkCentre: Whisper and Piper
  • 1 Assist pipeline updated to use local STT, local TTS, and local intents
  • 1 Android wake word enabled: Hey Jarvis
  • 2 Z-Wave dimmers now controllable by voice through Home Assistant
  • 10 voice aliases added across the TV and hallway light paths
  • ~37 seconds was the original typed Hermes no-entity path before optimization
  • ~23 seconds was the optimized Hermes path after skill, aliases, and concise responses
  • ~0.04 seconds is the direct Home Assistant observed-state baseline
  • 0 Google Assistant or Home Assistant Cloud relay required for the Pixel Tablet path

Comments