Ran Think Mode A/B tests today and the results are not what I expected. On the E2B model, enabling Think Mode on a math task dropped the score from 100% to 20%. On the 26B, it dropped logic from 60% to 20%. My hypothesis: on memory-constrained hardware, the extra tokens consumed by 'thinking' crowd out the tokens needed for a good answer. The model literally thinks itself into a worse response. This is the kind of finding you only get from testing on real hardware with real constraints.