INSIGHTS

Exploring the intersection of AI, SAP, and enterprise automation through research-driven perspectives and practical experience.

AI ENGINEERING
2026-07-267 min read

The Homelab Grand Prix: Racing 14 Quantized LLMs on Two DGX Sparks

I put fourteen quantized LLMs on the grid of my two-node DGX Spark cluster and timed every one, top speed measured in tokens per second. A 35B model took pole at 57.1 tok/s, but the real finding sits behind it: a 122B model lands within 29 percent of that, and speculative decoding nearly doubled a 120B into the same class. Total size predicts very little. Here is the full grid, the configuration each car ran, and what actually moves the needle.

READ MORE

Stay Informed

Follow my research and insights on AI-driven enterprise automation