AI InfrastructureWIP
OpenClaw + Gemma 4: Local Agent Pipeline
A 2017 iMac running an autonomous AI agent via Telegram. Vulkan GPU acceleration, DuckDuckGo web search, tool-calling chains.
7.5 tok/s · web search · Telegram bot
Started Apr 6, 2026

The full local agent pipeline — Telegram to tool-calling on consumer hardware
Building an autonomous AI agent that texts me through Telegram, searches the web, and runs entirely on consumer hardware. No cloud. No API keys. Just a 9-year-old desktop with a 4 GB GPU and some stubbornness.
The Stack
- Model: Gemma 4 E4B (7.5B dense · Q4_K_M · 4.5 GB)
- Inference: llama.cpp v38 (Vulkan/MoltenVK · 7.46 tok/s)
- Orchestration: OpenClaw (Gateway + tool-calling + compaction)
- Interface: Telegram Bot (ClawdyDawdy · web search · always on)
The Hardware
2017 iMac 27"
- CPU: i7-7700K
- RAM: 40 GB DDR4
- GPU: Radeon Pro 575
- VRAM: 4 GB GDDR5
The GPU runs via Vulkan/MoltenVK — not Metal (crashes on discrete AMD), not ROCm (Linux-only), not Ollama (can't see it). Three hours of empirical testing turned "worthless for LLMs" into "conversational speed for 7.5B models."
What Actually Works
- Validated Capabilities: Web search via DuckDuckGo, tool-calling chains, context compaction, Telegram bot, auto-restart via launchd.
- What Doesn't Work (Yet): 26B MoE models on Vulkan, Flash attention on AMD, Models over 4 GB VRAM.
Myths Debunked
- "You need Apple Silicon for local LLMs" — A 2017 Intel iMac with Vulkan runs Gemma 4 at 7.5 tok/s — fast enough for an autonomous agent.
- "20+ tok/s on 26B with --cpu-moe" — We measured 0.5–10 tok/s. And the output was garbage tokens. Performance claims without hardware specs are worthless.
- "Just use the bartowski quants, they fix the bug" — Tested bartowski Q3_K_M. Same
<unused50>tokens. The bug is in Vulkan's MoE shaders, not the quantization. - "--flash-attn works on AMD" — Flash attention requires Metal (Apple Silicon) or CUDA (NVIDIA). Not available on Vulkan.
OpenClawGemma 4VulkanAgentic
