The Homelab Grand Prix: Racing 14 Quantized LLMs on Two DGX Sparks
I put fourteen quantized LLMs on the grid of my two-node DGX Spark cluster and timed every one, top speed measured in tokens per second. A 35B model took pole at 57.1 tok/s, but the real finding sits behind it: a 122B model lands within 29 percent of that, and speculative decoding nearly doubled a 120B into the same class. Total size predicts very little. Here is the full grid, the configuration each car ran, and what actually moves the needle.