Tuning the Hermes Context Window: How I Burned 72.7M Tokens So You Don’t Have To
A homelab experiment set out to find the optimal context window for a local Qwen 3.6 daily driver. It failed at that job, and revealed a more useful one: three real bugs, a clean speed-versus-depth curve, a prompt cache that hid 93% of the real token count, and a config change worth making in Hermes Agent’s compaction settings. Six tables. Zero cloud tokens spent.