Recent measured numbers
~70 tok/sGemma 4 12B · RTX 5080 · coding bake-off · ~10.4GB VRAM
~28 tok/sQwen 27B Unsloth IQ4_XS · same tasks · ~15.6GB (tight)
~43 tok/s27B IQ4_XS full-GPU · Ollama ctx 4096 · vs ~18 / ~7 when spilled
~60–90 wallFreeToken 35B-A3B MoE on 16GB · Flash-Next did not fit
Single-stack readings. Quant, ctx, and thinking mode move the needle. Full write-ups live on llmlanes.com/builds.
What this site is
- Not a redirect to LLM Lanes — sister site with a tok/s review-bench focus.
- Hard cross-links to the lanes map and Builds on llmlanes.com.
- Numbers over hype — wall-clock and engine tok/s labeled as such when they differ.