vibescoder

Tuning Turns on a Two-Node DGX Spark Cluster: Fairness, Not Just Throughput

·7 min read

The Spark cluster is fast. That was never the problem. The problem was that two agents hitting it at once did not get fair turns.

I had a tuning session on the dual Spark pair and I expected to chase raw throughput. Instead I found something more useful. The one lever that mattered most was something the deployment’s own back-to-back A/B test had measured and already recommended, and production had quietly put it back to the wrong value.

This is a story about reading your own data instead of re-tuning from intuition.

My First Benchmark Number Was Noise

Before changing anything I re-ran the throughput matrix to get a clean baseline. The first run looked worrying. A concurrency case that had measured about 5.6 tokens per second suddenly showed about 3.1.

Concurrency caseFirst runRe-run, 5 trialsNotes
c=6 / 8K prompt3.1 tok/s6.7 tok/s (6.4 - 10.2)First run was a single-trial artifact

I almost chased it. A drop like that points at a mixed prefill and decode problem, and the deployment notes warned about exactly that. But a single bad trial is not a result. I re-ran with five trials and the same case measured about 6.7 tokens per second, with a range from 6.4 up to 10.2.

That was better than the number I had before the change. The 3.1 was a one-off artifact, not a regression.

The lesson was not about the specific number. It is that I went looking for a throughput win and almost tuned a ghost. Measure more than once before you turn a knob.

The A/B Test Already Told Me the Answer

The real find was in the deployment’s own notes on the distributed scheduler.

# the setting that decides how many prefills overlap
DSPARK_MAX_INFLIGHT_PREFILLS
 
# what the A/B concluded it should be: 2
# what production was running:        1

The cluster uses a DSpark stack through an inference server, and there is a setting that controls how many prefills can run at once. The notes contained a back-to-back A/B run from earlier that measured this exact knob with a four-request concurrency case at an 8K prompt size.

The A/B concluded the value should be two. With two overlapping prefills, the worst-case wait between the concurrent streams nearly halved, and decoding showed no regression at the concurrency levels that matter. The two-agent case is the household reality. One agent for the house, one for my work.

Production had it set to one.

That is the whole story. The measured answer was sitting in the repo and the running config had drifted from it. So the first thing I did was restore the value the A/B had already earned, then verify it was live in the running process and not just in the file.

The Fairness Win Showed Up in the Spread

The point of that knob is not to make one request faster. It is to keep a second request from waiting for the first to finish prefilling. That is the difference between throughput and fairness.

I measured the time to first token across four concurrent requests at the concurrency case the A/B used.

StreamSerialized (inflight=1)Restored (inflight=2)
First4.7s6.7s
Second9.4s9.6s
Third14.3s14.4s
Fourth19.3s15.7s
Worst-case spread14.6sabout 9s

The spread tightened compared with the serialized baseline, where the later streams waited much longer. The aggregate improved and the worst case got better. Trial zero was an outlier, so trials one and two are the consistent read.

That is exactly the multi-agent experience I wanted. Nobody waits for the whole family.

Thinking Off Was Already the Right Default

While I was in the serving config I confirmed the thinking behavior was still set to off.

default_chat_template_kwargs: {'thinking': False}
num_speculative_tokens: 5
DSPARK_MAX_INFLIGHT_PREFILLS: 2

A long reasoning budget on this model wastes latency on household agent work. The tool-call requests the house sends are better served with thinking off, and the model still reasons enough to call its tools. The deployment’s own bakeoff comparisons keep reaching the same place, so thinking off stayed as the default.

That pairs with the speculative decode length I had settled on. The model family gains most from a short speculation window, and the measured code quality held while the wall time dropped.

The combined shape was a fast default that does not overthink a tool call it can just make.

I Hit the Host Memory Gate and Stopped

Not every lever was worth pulling.

Host memoryValue
Used109 GiB
Total121 GiB
Freeabout 12 GiB
Hard stop rulebelow 3 GiB free

The deployment had a hard rule. If available host memory drops below 3 GiB, stop tuning. I was not at the stop line, but I was close enough that the remaining concurrency and KV tradeoffs were not worth the risk. Raising admission further would trade host memory headroom for a small fairness gain I could not verify was safe.

So I left that lever alone. The honest answer to whether more concurrent requests would help is that it might help a little, but I could not prove it without risking the host.

The Wrong Knob Can Still Cost You

The useful contrast came from the earlier tuning round on the same cluster. On face value, unlocking the model’s native speculative block size looked promising because it matches the model’s internal structure. The stock serving path uses a different size because of a runtime divisibility rule.

KnobReasoningResult on the real workload
Native spec-decode block sizeMatches the model’s internalsRegressed smart-home, portfolio, to-do, and made a coding task slower
DSPARK_MAX_INFLIGHT_PREFILLS=2Matched the A/B’s two-agent caseFairness spread tightened, no decode regression

When I tested the block-size knob against the actual household workload, it regressed. It hurt the smart-home slice, the portfolio task, and the to-do task, and it made a coding conflict task slower.

That is why the lesson here is not trust the measurement, it is trust the measurement for the workload you run. A knob that sounds closer to the model’s internals can still be the wrong move for your real traffic. The fairness knob from the A/B was the right move because the A/B tested the same two-agent scenario the house runs every day.

The cluster now runs with the A/B-confirmed fairness setting, thinking off, the measured speculation length, and a lot of host memory already committed. It is faster where it matters and fairer where it counts.

The next tuning question is whether more concurrent requests can fit once I understand what is holding that 12 GiB of headroom.

By the Numbers

  • 6.7 tokens per second measured at the concurrency case, over five trials
  • 3.1 the single-trial number that looked like a regression and was not
  • 2 prefill slots restored after the A/B test recommended them
  • 9s the worst-case wait between concurrent streams after the change
  • 109 GiB of host memory used out of 121 GiB on the head node
  • 1 A/B-confirmed lever that production had silently reverted

Comments