Skip to content
Andrew Voirol
Work → OpenClaw + Gemma 4: Local Agent Pipeline
AI InfrastructureWIP

OpenClaw + Gemma 4: Local Agent Pipeline

A 2017 iMac running an autonomous AI agent via Telegram. Vulkan GPU acceleration, DuckDuckGo web search, tool-calling chains.

7.5 tok/s · web search · Telegram bot

Started Apr 6, 2026
Architecture diagram showing the OpenClaw pipeline: Telegram → Gateway → Gemma 4 LLM (Vulkan GPU) → Tool Router → DuckDuckGo and Memory Store

The full local agent pipeline — Telegram to tool-calling on consumer hardware


Building an autonomous AI agent that texts me through Telegram, searches the web, and runs entirely on consumer hardware. No cloud. No API keys. Just a 9-year-old desktop with a 4 GB GPU and some stubbornness.

The Stack

  • Model: Gemma 4 E4B (7.5B dense · Q4_K_M · 4.5 GB)
  • Inference: llama.cpp v38 (Vulkan/MoltenVK · 7.46 tok/s)
  • Orchestration: OpenClaw (Gateway + tool-calling + compaction)
  • Interface: Telegram Bot (ClawdyDawdy · web search · always on)

The Hardware

2017 iMac 27"

  • CPU: i7-7700K
  • RAM: 40 GB DDR4
  • GPU: Radeon Pro 575
  • VRAM: 4 GB GDDR5

The GPU runs via Vulkan/MoltenVK — not Metal (crashes on discrete AMD), not ROCm (Linux-only), not Ollama (can't see it). Three hours of empirical testing turned "worthless for LLMs" into "conversational speed for 7.5B models."

What Actually Works

  • Validated Capabilities: Web search via DuckDuckGo, tool-calling chains, context compaction, Telegram bot, auto-restart via launchd.
  • What Doesn't Work (Yet): 26B MoE models on Vulkan, Flash attention on AMD, Models over 4 GB VRAM.

Myths Debunked

  • "You need Apple Silicon for local LLMs" — A 2017 Intel iMac with Vulkan runs Gemma 4 at 7.5 tok/s — fast enough for an autonomous agent.
  • "20+ tok/s on 26B with --cpu-moe" — We measured 0.5–10 tok/s. And the output was garbage tokens. Performance claims without hardware specs are worthless.
  • "Just use the bartowski quants, they fix the bug" — Tested bartowski Q3_K_M. Same <unused50> tokens. The bug is in Vulkan's MoE shaders, not the quantization.
  • "--flash-attn works on AMD" — Flash attention requires Metal (Apple Silicon) or CUDA (NVIDIA). Not available on Vulkan.
OpenClawGemma 4VulkanAgentic

Related Threads

OpenClaw agent architecture — Telegram App sends messages through the OpenClaw Gateway to a Vulkan-accelerated Gemma 4 LLM, which routes tool calls to DuckDuckGo web search and a Memory Store, with results flowing back to Telegram

OpenClaw: Building a Local AI Agent

From benchmark suite to autonomous Telegram bot — deploying Gemma 4 as an agentic pipeline on a 2017 iMac with Vulkan GPU acceleration.


Andrew Voirol

Builder, hacker, shipper. Currently leaving localhost.

Navigate

WorkThreadsBuilder's LogAboutContactRSS Feed

Connect

X / TwitterGitHubLinkedIn

© 2026 Andrew Voirol·Back to top ↑
✦Just one prompt away from figuring it all out.