Recent measured numbers

~70 tok/sGemma 4 12B · RTX 5080 · coding bake-off · ~10.4GB VRAM
~28 tok/sQwen 27B Unsloth IQ4_XS · same tasks · ~15.6GB (tight)
~43 tok/s27B IQ4_XS full-GPU · Ollama ctx 4096 · vs ~18 / ~7 when spilled
~60–90 wallFreeToken 35B-A3B MoE on 16GB · Flash-Next did not fit

Single-stack readings. Quant, ctx, and thinking mode move the needle. Full write-ups live on llmlanes.com/builds.

What this site is

Contact