Subzero

2026-10-11

Run Gemma 4 Offline on a Raspberry Pi 5

Build a fully offline AI assistant on Raspberry Pi 5 with Gemma 4. Test three runtimes, cool it properly, and add tools for reliable answers—no internet needed.

In short

  • A Raspberry Pi 5 with 8GB RAM and NVMe SSD can run Gemma 4 locally with no Wi-Fi or subscription.
  • Cooling matters: without a fan the Pi throttles and loses 25-35% speed; with a heatsink and fan it runs 29-55% faster.
  • Three runtimes work on the Pi—llama.cpp, LiteRT-LM (fastest), and LM Studio—each with different trade-offs for speed and ease of use.
  • Add three open-source tools (grammars, Outlines, Instructor) to force the model into clean, structured answers for reliable automation.

Why Run AI Offline on Your Pi

Most people think AI needs the cloud. It doesn't. A Raspberry Pi 5 can run Gemma 4, a capable language model, entirely offline with no Wi-Fi and no monthly bill. Your data stays on the device, your assistant works without internet, and you control everything.

The catch: it's slower than cloud AI, and you need to set it up yourself. But if you want privacy, reliability, or to avoid subscription costs, this is worth the effort.

What You Need to Build It

Start with a Raspberry Pi 5 with 8 gigabytes of RAM. Add an NVMe SSD (the Q4_0 quantization of Gemma 4 is 2.84 gigabytes, so it fits in 8GB), an NVMe HAT to connect it, a heatsink, an external fan, and a power supply. You don't need an AI HAT—the Pi 5 CPU runs Gemma 4 on its own.

Skip the extra hardware. The AI HAT+ only does vision, and the AI HAT+ 2 only runs tiny models. The Pi 5 is powerful enough without them.

Assemble and Boot from NVMe

Mount the NVMe HAT onto the Pi 5 using standoffs and a screwdriver. Insert the NVMe SSD into the HAT's M.2 slot at a 45-degree angle, then press it down flat and screw it in place. Connect the thin PCIe ribbon cable between the Pi 5 and the NVMe HAT by lifting the latches on both ends, inserting the cable, and closing the latches.

Download Raspberry Pi Imager and flash Raspberry Pi OS 64-bit onto the NVMe drive. Select Raspberry Pi 5 as the device, Raspberry Pi OS 64-bit as the operating system, and your NVMe drive as the storage device. Click Write and wait a few minutes.

Once it's done, eject the drive, plug it into the NVMe HAT on the Pi, and power on. On first boot, open a terminal and update the system. Then configure the Pi to boot from the NVMe drive by default: open the bootloader configuration, go to Advanced Options, then Boot Order, and select NVMe SSD. Save and reboot. The Pi will now boot from the NVMe drive every time.

Cooling Makes a Huge Difference

This is the part most people skip, and it costs them real speed. Without a fan, idle temperature hits 75 degrees Celsius. Within five minutes it climbs to 86 degrees and the Pi throttles itself. With a heatsink and external fan, idle is 34.5 degrees and peak is about 65 to 73 degrees, with no throttling.

The speed gain is striking. Writing speed went from 5.7 to 7.35 tokens per second, about 29 percent faster. Reading the prompt went from 31 to 48 tokens per second, about 55 percent faster. If you skip cooling, you lose about a quarter to a third of your speed. Get a heatsink and a fan.

Three Runtimes: Speed, Ease, and Control

Gemma 4 on a Pi 5 can run three different ways. Each trades off speed, ease of use, and how much control you get. The video above goes deeper into the benchmarks, but here's what matters: use four CPU threads for speed, and turn reasoning off if you need fast answers.

llama.cpp is fast and open source. Install the build tools, clone the repository from GitHub, prepare the build with cmake, then compile it with cmake build using four CPU cores. On a Pi, compilation took 13 minutes and 33 seconds. Download the Gemma 4 E2B model in GGUF format (the Q4_0 quantization is 2.84 gigabytes). Run it in chat mode with two threads and reasoning off, then test with four threads. With four threads, writing stayed about the same but reading the prompt nearly doubled. Use four threads.

LiteRT-LM is Google's lightweight runtime for ARM devices like the Pi 5. Install it with pip (7 seconds on a Pi), download and import the model (37 seconds total), then benchmark it with four threads. LiteRT-LM is the fastest of the three: 8.59 tokens per second writing, 116 tokens per second reading the prompt, and the first word came after 1.2 seconds. That's about 17 percent faster writing and 2.4 times faster reading than llama.cpp.

LM Studio also has a command-line interface, but there's a gotcha on Raspberry Pi OS. The official installer fails with 'ldconfig must be available on your PATH'. Fix it by setting the path first in the same terminal, then run the official install command. Once you set the path, LM Studio installed in 55 seconds on a Pi. Download the Gemma 4 model (2 minutes for a 4.41 gigabytes file). LM Studio's default file is bigger than llama.cpp's, and reasoning is on by default, so answers take longer. Always check the file size and reasoning setting before comparing tools.

Build a Real Offline Assistant

Serve Gemma with llama-server because it gives a standard chat API that Python scripts and other tools can all talk to. Create a Python file called assistant.py that points at the model server and at Kiwix (offline Wikipedia), keeps your chat history, and has two helpers: ask sends the conversation to Gemma, and wiki looks a topic up in offline Wikipedia.

The script reads what you type. A normal question goes to Gemma with the whole conversation, so it remembers what you said. Start a line with slash wiki, and it answers from the Wikipedia article instead. Start the model server in the background with four threads and reasoning off, then run the script. It works, and it never touches the internet.

Sometimes the answers are messy or incomplete. Add three open-source tools to fix that. llama.cpp has a built-in feature called grammars. You give it a grammar file that defines the exact format the model must follow. Create a grammar file that forces the model to output JSON. Run llama.cpp with the grammar file. On a Pi, classifying a command took 4.0 seconds and always produced valid JSON. Without the grammar, the model rambles. With it, you get valid JSON every time.

Outlines is a library that restricts the model to only certain outputs. Install it, but remember to install numba first, or it won't work. On a Pi, Outlines picked the right action for all five test requests in about 2.5 seconds each. Now the model can only answer with one of those allowed choices. No rambling, no typos.

Instructor is the most powerful tool. It validates the model's answer against a shape you define, and if the answer doesn't match, it retries automatically. Install it and create a Python script that defines the shape of the answer you want, then ask the model for it. Instructor talks to your local llama.cpp server. On a Pi, Instructor pulled a name and age out of a sentence in 7.3 seconds, in exactly the shape asked for. Use Instructor when you need to extract structured data from unstructured text.

Add Offline Wikipedia as a Brain

A small model forgets facts. But give it the right article and it can answer from a source, fully offline. Kiwix serves Wikipedia from one local file, no internet. Download kiwix-tools for ARM64 and the English Wikipedia file. The full version with pictures is 119 gigabytes and took about 26 minutes on a typical connection. Smaller versions exist: 13 gigabytes intros only, or 49 gigabytes full text without pictures.

Run kiwix-serve with the file and open it in a browser on any device at home. Searching 'Raspberry Pi' in the offline Wikipedia took 0.2 seconds. Then Gemma answered 'the first Raspberry Pi was introduced in 2012' from that article in 17.4 seconds, correct and fully offline. Now your assistant has a brain. It can search Wikipedia and answer from real sources, all on the Pi.

When Speed Matters: Cloud Decisions

Everything so far runs offline on the Pi. But what if you wanted faster decisions? JEV is TypeSafe's decision model. It doesn't chat like Gemma. It answers typed questions: yes or no, pick one from a list, or give a score. It runs in the cloud through an API, so it needs internet and an account.

The offline assistant makes small decisions, like whether something is a question or a command, or which action to take. On the Pi, each local answer with Outlines takes about 2.5 seconds. In the cloud, JEV answers in about 0.23 seconds, roughly 11 times faster. TypeSafe charges roughly 0.04 cents per million input tokens, with output free.

Here's what this means for your project. If you keep everything local, the assistant is fully offline and your decisions stay on the Pi, but the decisions are slower. If you send only the quick decisions to JEV, the assistant feels faster, but it needs internet and an account, and those decisions leave the Pi. It depends on what matters more to you: speed, privacy, or staying offline.

Watch "I Built an AI That Knows Everything (And It’s Offline)" on YouTube

Sources

  1. raspberrypi.com
  2. github.com
  3. huggingface.co
  4. github.com
  5. huggingface.co
  6. lmstudio.ai
  7. github.com
  8. github.com
  9. docs.typesafe.ai
  10. download.kiwix.org
  11. download.kiwix.org
  12. Pi Pico W
  13. Raspberry Pi 5 8GB
  14. Raspberry Pi 5 Starter Kit - 16GB RAM
  15. NVMe Base for Raspberry Pi 5 M.2 HAT
  16. ELECROW 5 Inch Mini Touchscreen Monitor
  17. NIMO Powerful AI Mini PC, AMD Ryzen AI Max+ 395
  18. NIMO Mini PC, AMD Ryzen AI Max+ 395 128GB LPDDR5X
  19. AX9 Max Mini PC (vapor chamber), Ryzen AI 9 HX 470
  20. EVGA GeForce RTX 3090
  21. Raspberry Pi Zero 2W Kit
  22. Pi Zero Aluminum Case Kit with Pin Header
  23. Crucial MX500 2TB 3D NAND SATA 2.5 Inch
  24. Corsair Vengeance RGB RS DDR5 32GB (2 x 16GB)

More videos