Skip to content
Andrew Voirol
Work → MUJŌ: WebGPU Watercolor Engine → Building MUJŌ: WebGPU Watercolor Engine
ThreadActive

Building MUJŌ: WebGPU Watercolor Engine

A physically-based Japanese watercolor and sumi-e simulation running entirely on WebGPU compute shaders — Kubelka-Munk optics, fluid dynamics, 3D elastic bristle physics, and Zen impermanence.

Code
Started Aug 15, 2026·Latest Aug 20, 2026·12 entries

MUJŌ watercolor app showing five Nihonga mineral pigment brush strokes — Sumi (black), Shu (red), Ai (indigo), Ōdo (ochre), and Rokushō (green) — on Kizuki Kōzo washi paper, with the Suzuri Palette, brush selector, water dilution slider, and canvas tilt controls

Five authentic Nihonga mineral pigments — Kubelka-Munk spectral mixing on procedural washi

MUJŌ tarashikomi wet blending technique — Ai indigo ink flowing into a wet wash with Shu red bleeding at the edges, on Kobishi paper, demonstrating Navier-Stokes fluid dynamics and porous media absorption

MUJŌ (無常) started as a question: can you make a digital brush that feels like real sumi ink on handmade washi paper? Not "looks like" in a screenshot — feels like under your hand. The answer required going deep: Navier-Stokes fluid mechanics for the ink flow, 2-flux Kubelka-Munk spectral radiative transfer for physically correct pigment mixing, Darcy tensor porous flow through virtual paper fibers, and 48 elastic guide bristle rods simulated via Position-Based Dynamics entirely in WGSL compute shaders.

No Three.js. No React. No physics libraries. Every simulation pass — from the Jacobi pressure solver to the coffee-ring evaporation pinning — is a custom WGSL compute shader running on the GPU. The result: a meditative painting canvas with five authentic Nihonga mineral pigments, three Japanese brush types, three washi paper varietals, and a procedural sound engine that responds to your brush pressure.

This thread tracks the engineering from first ink to artifact elimination, and the Atelier testbed that proves the physics hold up.

Latest Update

Pushing the Apple M4 Pro with half-precision compute

Thu, Aug 20, 2026

Timeline

Built a complete WebGPU watercolor simulation engine in a single session. MUJŌ (無常) — a meditative Japanese sumi-e painting canvas running entirely on GPU compute shaders. No Three.js, no React, no external physics libraries. Custom Navier-Stokes fluid solver, 2-flux Kubelka-Munk spectral color mixing, and Darcy porous flow through virtual paper fibers. Five authentic Nihonga mineral pigments — sumi carbon black, cinnabar vermilion, fermented indigo, yellow ochre, malachite verdigris — each with real spectral absorption/scattering tables. The ink bleeds through virtual Kozo mulberry fibers and darkens along evaporating edges. 1024×1024 simulation grid running at interactive framerates on M4 Pro.

WebGPUMUJŌShaders
Permalink →

Getting the physical tip lag right was a massive battle against stiffness and chaotic particles. I spent hours dialing in a discrete 3D Cosserat elastic rod simulation for the brush bristles, implemented in a 287-line CosseratBristleCluster.ts file. Early iterations were either too rigid or produced bizarre, sperm-like particle trails that broke the meditative vibe. By tuning the bending stiffness and stepping the physics at 0.016s intervals, we finally achieved a natural viscoelastic lag that responds organically to trackpad pressure and momentum.

WebGPUSimulationPhysics
Permalink →

Lattice Boltzmann provides incredible fluid mechanics, but it devours memory bandwidth on Apple Silicon. I set up an A/B test in a 696-line FluidBleedExperiment.ts harness to compare an LBM D2Q9 model against an anisotropic Darcy tensor for simulating the lush capillary tendrils (Hige-nijimi) of watercolor bleed. While LBM is mathematically elegant, it bottlenecked the unified memory architecture. We pivoted to a 2-layer Darcy porous flow with saturation gating—yielding a stable WebGPU compute pipeline that visually matches the organic wicking without the performance penalty.

Fluid DynamicsWebGPUResearch
Permalink →

Standard RGB alpha blending completely ruins the luminosity of digital watercolor. Layering yellow over indigo usually produces a muddy, desaturated grey in traditional digital canvases. To fix this, I wrote a 140-line render_km.wgsl shader implementing a robust 2-flux Kubelka-Munk radiative transfer model. By using true physical spectral values (absorption K and scattering S), the pigments now interact subtractively just like real minerals—yielding deep, authentic celadon greens and vibrant glazes that maintain the optical reality of actual paint.

Color ScienceShadersWebGPU
Permalink →

A single Kubelka-Munk layer isn't enough to capture internal light bounces inside the pigment pool. I built a 659-line experimental lab to compare our baseline Kubelka-Munk output against a full Saunderson 3-layer glaze matrix. The goal was to accurately render the Fuchidori (coffee-ring borders) and refractive-index matching wet-darkening. Factoring in Fresnel internal reflections proved computationally heavy, but the optical depth it adds to the drying decay of the watercolor pools is unmistakable.

OpticsColor ScienceResearch
Permalink →

The capillary mechanics of fluid change drastically depending on the paper's microscopic fiber skeleton. I expanded the parchment_gen.wgsl shader to 141 lines to procedurally generate three distinct traditional substrates: Kizuki Kōzo (unbleached mulberry), Torinoko (sized eggshell), and Kobishi (vintage tea-tannin). By tracking physical tangents along the sinuous bast fibers, the fluid dynamics solver now knows exactly where to pool and where to bleed aggressively, matching the distinct behaviors of each paper type.

Procedural GenerationShadersDesign
Permalink →

Replaced the simple swept-capsule brush with 48 Position-Based Dynamics elastic guide rods running entirely in WGSL compute shaders. Each bristle has distance constraints, bending rigidity, Coulomb friction against the paper surface, and capillary clumping when wet. The brush fans out under pressure, pinches on turnarounds, and flicks on lift-off with real inertia. Three brush types: Maru-fude (round calligraphy), Menso (fine liner), and Hake (broad wash). The difference is visceral — strokes went from "digital stamp" to "I can feel the paper" in one refactor. Also added a procedural Web Audio soundscape: bamboo water drops, singing bowls, and paper friction acoustics that respond to brush pressure.

WebGPUMUJŌShadersAnimation
Permalink →

Fast brush flicks were starving the interpolation curve, leaving ugly jagged staircase edges on the canvas. When relying strictly on raw coordinate deltas, rapid trackpad swipes didn't provide enough data points for smooth splines. The fix was moving to a timestamp-normalized kinematic model. I implemented a dynamic Centripetal Catmull-Rom ingester that scales subdivision steps based on true elapsed millisecond velocity. Now, even a 2,500px/s flick renders with perfect tangent continuity and natural taper.

InputKinematicsMath
Permalink →

Spent two days hunting and killing every artifact that made the strokes look digital instead of organic. The "caterpillar beads" — visible blob spacing along fast strokes — were eliminated by switching to Catmull-Rom C1 spline interpolation with lag-by-one future-point basis. The "reed-screen grating" came from frame-boundary discontinuities in bristle injection, fixed by continuous sub-frame interpolation. Then the needle artifact at stroke ends: a kinematic speed/pressure gate smoothly tapers the contact area to zero. Finally, added fude-ashi bristle fringes and tooth gating so even straight lines have the organic irregularity of real ink on handmade paper. The Shisho benchmark suite (永 永字八法, 一 ichi, 心 kokoro, 円 ensō) now produces strokes that would fool a calligrapher at thumbnail scale.

WebGPUMUJŌDebuggingShaders
Permalink →

We were taking one step forward and two steps back tuning the subjective feel of the brush. After repeatedly losing the "magic" of a good stroke to regression bugs, I realized we needed a methodical way to evaluate tweaks. I built an automated visual comparison suite that captures macro and micro snapshots of kanji character strokes across a speed ladder. By running side-by-side A/B baselines, we can finally prove whether a new physical tip lag or viscosity adjustment is actually an improvement or just a different flavor of wrong.

TestingEngineeringTooling
Permalink →

Built a companion scientific testbed — MUJŌ Atelier — with 5 standalone lab experiments. Each experiment isolates and compares one physics layer: Darcy tensor diffusion vs. Lattice Boltzmann D2Q9 fluid bleed, single-layer Kubelka-Munk vs. Saunderson 3-layer glaze stacks, 6 botanical washi substrates with different fiber structures, and an M4 Pro hardware profiler measuring per-pass GPU timings. The Atelier is the "show your work" for the simulation — same ethos as the Gemma 4 benchmark suite, different domain. If someone asks "how does your Kubelka-Munk compare to LBM?" the answer is a side-by-side live experiment, not a paragraph.

WebGPUMUJŌBenchmarks
Permalink →

Sustaining 120 FPS on a 9-million cell fluid grid requires militant GPU memory management. To hit our performance targets on Apple Silicon without dropping frames, we moved the entire simulation pipeline to rgba16float (f16) half-precision math. The real breakthrough was consolidating Navier-Stokes advection, Jacobi pressure projection, and capillary diffusion into a single zero-allocation GPU command submission per frame. By eliminating runtime bind group creation, the M4 Pro chews through 3072² grids with almost zero latency.

WebGPUPerformanceOptimization
Permalink →

Andrew Voirol

Builder, hacker, shipper. Currently leaving localhost.

Navigate

WorkThreadsBuilder's LogAboutContactRSS Feed

Connect

X / TwitterGitHubLinkedIn

© 2026 Andrew Voirol·Back to top ↑
✦Just one prompt away from figuring it all out.