Published the full benchmark source code to GitHub. Every prompt. Every grading function. Every raw JSON result. Because if you're going to claim '94% on consumer hardware' on a fancy website and then compare yourself to Google's published benchmarks, you better show how you got those numbers. The repo has the Python orchestrator, the Ollama client wrapper, the test definitions across 7 categories (reasoning, coding, tool calling, creative, context, agentic, performance), and the raw results from all 414 test runs. MIT licensed. Anyone can clone it, pull a Gemma model, and run the same suite on their own hardware. That's the difference between a benchmark and a blog post.