Splitting the Bakeoff: Coding Gets Its Own Test, and It’s Terminal-Bench Over SWE-bench
Round 2 of the local agent bakeoff saturated its coding domain completely, zero cracks across 18 fresh evals from two coding-focused models. That’s the moment to stop patching a broken test and split the question in two: what’s the best Home Assistant butler, and separately, what’s the best local coding model that fits a 32GB card. Here’s the research behind picking Terminal-Bench 2.1 over SWE-bench, and the plan for the actual run.