Skip to content
Andrew Voirol
WorkLogAboutContact
HomeWorkLog

AboutContact

✦ Just one prompt away from figuring it all out.

Thursday, April 9, 2026gemma-4-benchmarks

The inverted ladder. Q4 > Q8 > F16.

Gemma 4QuantizationQ4

We ran the complete E2B quantization ladder — Q4, Q8, and BF16 — through 39 identical tests on the same hardware. Q4 scored 92.2%. Q8 scored 91.8%. F16 scored 91.2%. The relationship is perfectly inverted: lower precision = higher score. This shouldn't happen. Quantization is lossy compression. But on DDR4 bandwidth-constrained hardware, smaller weights mean more of the model stays in CPU cache. Fewer cache misses. More consistent throughput. And the precision you 'lose' at Q4? It's apparently noise, not signal — at least for the tasks that matter. The daily driver isn't the most precise model. It's the smallest one. And it's also the fastest.

← Previous

Dark mode toggle with jelly physics

Next →

F16 is the ceiling. Q8 already passed it.


Andrew Voirol

Builder, hacker, shipper. Currently leaving localhost.

Navigate

WorkBuilder’s LogAboutContactRSS Feed

Connect

X / TwitterGitHubLinkedIn

© 2026 Andrew Voirol✦Just one prompt away from figuring it all out.