The bartowski experiment: confirming the impossible
The HuggingFace community said 'bartowski quants fix the <unused50> bug.' Downloaded 12 GB of bartowski's Gemma 4 26B-A4B Q3_K_M. Loaded it on a separate port (8081) so the working E4B stayed live on 8080. First test with --cpu-moe: <unused50> at 10 tok/s. Faster garbage, but still garbage. Without --cpu-moe: <unused50> at 0.5 tok/s. Even raw /completion endpoint (no chat template whatsoever): <unused50>. Seven configurations tested, zero success. The community was wrong — it's not a quant issue at all, it's a Vulkan MoE shader bug. My E4B pipeline stays as the daily driver. Sometimes the answer is: the model you have, configured well, beats the model you want.