The quantization showdown: Unsloth wins on efficiency, E4B is dead weight
Ran 4 models through 39 core tests each: E2B Q8, E4B Q8, Unsloth DQ4, and 31B Q4. The E4B — Google's 'recommended' mid-range model — is the worst value proposition in the lineup. It uses 2× the memory of the E2B Q8 for marginally better scores, and it's slower. Meanwhile Unsloth's dynamic Q4 of the E2B matches the E4B's quality at half the size and 38 tok/s. The 31B Q4 is the accuracy king at 94% but runs at 7.5 tok/s. The practical daily driver is Unsloth UD-Q4_K_XL: 2.94 GB, 38 tok/s, and the efficiency curve tells a story the spec sheets don't — bigger isn't always better when the quantization is smart enough.