Pushing the Apple M4 Pro with half-precision compute
Sustaining 120 FPS on a 9-million cell fluid grid requires militant GPU memory management. To hit our performance targets on Apple Silicon without dropping frames, we moved the entire simulation pipeline to rgba16float (f16) half-precision math. The real breakthrough was consolidating Navier-Stokes advection, Jacobi pressure projection, and capillary diffusion into a single zero-allocation GPU command submission per frame. By eliminating runtime bind group creation, the M4 Pro chews through 3072² grids with almost zero latency.