We were taking one step forward and two steps back tuning the subjective feel of the brush. After repeatedly losing the "magic" of a good stroke to regression bugs, I realized we needed a methodical way to evaluate tweaks. I built an automated visual comparison suite that captures macro and micro snapshots of kanji character strokes across a speed ladder. By running side-by-side A/B baselines, we can finally prove whether a new physical tip lag or viscosity adjustment is actually an improvement or just a different flavor of wrong.