Performance Tuning Hermes for Local AI Using a Two-Node DGX Spark Cluster
My wife’s Hermes agent was slow, and for a while I blamed the cluster.
That was the easy answer. The Spark pair is the heavy backend. The DeepSeek model is a big multimodal generalist. A slow-feeling agent against a big local model reads like a capacity problem. I went looking for the knob.
I was wrong. The cluster was fast. The slowness was one config value on the client, and that value was not just adding latency. It was breaking the agent’s tools.
Direct Against the Cluster Was Fast, So I Stopped Blaming It
I measured the cluster directly before touching anything. A single request against the DeepSeek endpoint, with thinking off and a tool available, finished in about 4 seconds and produced a correct tool call.
That turned out to be a best case. A tiny stateless prompt and one call is the easiest thing an endpoint can do. A real agent turn is not that. But even as a best case, it was far faster than what Hermes was producing.
The equivalent work through Hermes took about 23 seconds. Same model, same endpoint, roughly same request. That was a 7x gap on an apples-to-apples task, so the cluster was not the problem.
| Path | Wall time | Tool calls |
|---|---|---|
| Direct against the cluster, tiny prompt (best case) | 4.3s | 2 |
| Through Hermes | 23s | 2 |
Sometimes the answer to a slow system is a bigger system. Sometimes it is the config sitting between the agent and the model. This was the second kind.
reasoning_effort: medium Was Not Even a Real Value
I asked Hermes what settings it sent, and it reported reasoning_effort: medium at the agent level. I checked what the serving layer supports for this model, and medium was not in the list. The server maps low and high, and the model family uses its own thinking control, but medium was not a recognized value.
# ~/.hermes/config.yaml: what Hermes was sending
reasoning_effort: medium # <- not a value this model acceptsThat matters because an unrecognized value does not get ignored cleanly. It gets mapped onto something the model does not expect, and the result was worse than either setting on its own.
I tested every combination against the cluster for a tool-call request:
| Setting | Wall time | Tool calls | Content |
|---|---|---|---|
| Thinking off | 4.3s | 2 | 103 chars |
reasoning_effort: medium | 10.6s | 2 | 0 chars |
reasoning_effort: low | 11.2s | 2 | 96 chars |
Two real problems came out of that table.
The medium setting made a tool call take over twice as long. And in repeated runs it produced zero visible content, because the reasoning consumed the whole token budget and the request ended with a length finish reason instead of a tool call. The agent got nothing actionable back.
That is the failure mode that makes a home agent feel broken. Not a wrong answer. An empty answer, after waiting a long time, for a task that needed a tool.
Context Size Was the Second Half of the Story
Hermes corrected me on one thing before I locked in the conclusion. The 23 seconds was not one request. The full model time was about 114 seconds split across three sequential calls, each carrying around 240K input tokens.
| Model call | Input tokens | Time |
|---|---|---|
| Call 6 | 240,516 | 50.8s |
| Call 7 | 240,833 | 22.6s |
| Call 8 | 241,423 | 41.2s |
That is a long-running session history plus a heavy system prompt. My measurement used a tiny stateless prompt, so I was comparing a clean request to a long conversation. Hermes was right to call that out. The context size was a real contributor.
But it was not the whole story. The reasoning setting still made the requests slower and the tool behavior worse on the same prompt. The fix was on the config side, and it was the client that needed the change.
The Fix Was on the Client, Not the Cluster
The change that fixed it was on the Hermes side. Remove reasoning_effort, let the cluster’s thinking-off default apply, keep a low temperature and a sane max token ceiling.
# ~/.hermes/config.yaml: after
# reasoning_effort removed; let the cluster's thinking-off default apply
temperature: 0.2
max_tokens: 2048I kept temperature at 0.2 rather than dropping to zero. My measurements showed temperature zero gave the best acceptance rate, but 0.2 is close enough to greedy to keep most of that benefit while leaving a little variety. I kept a max token cap as a guard against the truncation and empty reply case.
| Setting | Why |
|---|---|
Remove reasoning_effort | The value was unrecognized and broke tool calls |
temperature: 0.2 | Close to greedy, keeps most of the acceptance benefit |
max_tokens: 2048 | Guards against truncation and empty replies |
The cluster itself was already tuned. Its default thinking behavior, its speculative decode length, and its concurrency were done separately. Once the client stopped overriding with a broken value, the fast backend stayed fast.
The Router Was Sending Her Traffic to the Wrong Model
The last finding was the sneakiest. I checked where Hermes traffic landed, and a generic household request was not going to DeepSeek at all.
The router had a default route that preferred the Strix Qwen model and fell back to DeepSeek only if Strix was unhealthy. A plain Hermes request with no image, no Home Assistant tools, and a modest prompt matched that default route. So it went to Strix.
# what a generic Hermes request matched
default -> prefer: strix-coder
fallback: deepseek-vision
# so plain household traffic hit Strix, not DeepSeek| Request shape | Route | Backend it hit |
|---|---|---|
| Plain prompt, under 50K tokens | default | Strix Qwen |
| Has an image | vision | DeepSeek |
| Declares Home Assistant tools | home_assistant | Granite, fallback DeepSeek |
| Over 50K tokens | long_context | DeepSeek |
That means I was tuning the cluster to make Hermes feel fast, while generic Hermes requests never hit the cluster. They hit a different machine entirely. I only caught it because I asked which route a generic request matched.
That is the lesson that outlived the tuning. The client config mattered, but the routing policy mattered too, and they were sending requests to a backend I was not tuning. I fixed the default route in the router work that followed.
The Compression Cut the Real Cost, and It Was Context
Once the client was clean, I had Hermes compress its context and rerun the same task. That isolated the remaining variable, and it was bigger than I expected.
| Metric | Before compression | After compression |
|---|---|---|
| Input tokens per call | about 240K | 63.6K |
| Model calls per turn | 3 | 2 |
| Call latency, typical | 22-50s | 15-16s |
| Turn wall-clock | about 120s | about 32s |
That is roughly a 4x turn speedup from context alone. The input size was about 64 percent of the original latency. Thinking was already off, so the binding constraint was context volume, not the model’s reasoning mode or temperature. The reasoning_effort fix removed the broken setting. The compression removed the tonnage.
I had a fresh-session direct measurement after that, a 47-token stateless prompt with no agent loop, to find the true floor.
| Call | Wall time | Input / output | Finish |
|---|---|---|---|
| Cold pass | 6.55s | 47 / 34 | stop |
| Warm pass | 6.56s | 47 / 31 | stop |
Even with a 47-token prompt and no agent loop, the endpoint returns in about 6.5 seconds, not 3.2. Warm and cold were identical, so connection churn is negligible. The 6.5 seconds is model and server time.
The Honest Breakdown Is 6.5s Times N Plus Context
Putting it together, the original 23 seconds was additive.
| Component | Contribution | Can config tuning remove it? |
|---|---|---|
| Server base per call | about 6.5s | No |
| Sequential calls per turn | 3 before, 2 after | Partly, via context |
| Context-size cost | 240K vs 64K | Yes |
| Tool-loop serialization | forces the multiple calls | No |
So the 3.2s benchmark I started with is achievable only as one call on a tiny prompt with no agent loop. Any real agent turn pays the base per call times the number of calls, plus the context cost. Config tuning removed the context share and the reasoning penalty. It cannot remove the 6.5-second base floor or the fact that an agent turn is several sequential model calls.
That is the honest limit of what tuning can do. The fixable part was real and worth doing. The rest is structural.
reasoning_effort: medium taught me to distrust an agent config that claims an option the model does not support. The router default taught me that the path a request takes decides which model you are testing. The follow-up taught me where the hard floor is. Measure where the work goes before you tune the thing you assume is running it, then check whether the rest is a knob or a structure.
The open question is whether a single parallel-safe tool call could collapse the two remaining round trips into one, or whether the compressed context is already the ceiling.
By the Numbers
- 7x the wall time of a Hermes tool call versus the same request against the cluster
- 114s the real model time split across three sequential calls before the fix
- 240K input tokens per call in a long Hermes session
- 64K input tokens per call after compression
- 4x turn speedup from compressing context (about 120s to about 32s)
- 64 percent of the original latency that was raw input-size cost
- 3 model calls per turn before compression, 2 after
- 6.5s the endpoint’s per-call floor on a 47-token prompt with no agent loop
- 2.5x slower tool call with
reasoning_effort: mediumversus thinking off - 0 tool calls produced when the reasoning ate the token budget
- 1 unrecognized
reasoning_effortvalue that the serving layer could not honor