Deploy a Personal AI Assistant on a Raspberry Pi: A Complete Step‑by‑Step Guide for Developers
Read this article in clean Markdown format for LLMs and AI context.Ever thought about having a tiny AI buddy that lives on a cheap board you can stick on a shelf? It’s a fun project, and it actually helps you learn a lot about how AI works on low‑power devices. At Tech Brew we love taking big ideas and shrinking them down to something you can hold in your hand. This guide walks you through the whole process, no fluff, just clear steps you can follow tonight.
Why you might want an AI assistant on a Pi
- Always on, no cloud fees – Your Pi runs 24/7, so the assistant is ready whenever you are.
- Privacy – All data stays on your own hardware. No big tech company watching your voice.
- Learning – You get to see how models are loaded, how audio works, and how to debug on a tiny Linux box.
I built my first Pi assistant last winter just to control the lights in my living room. The first time it turned on the lamp after I said “Hey Pi, lights on,” I felt like a sci‑fi director. That moment made me realize how rewarding this kind of project can be, and I wanted to share the recipe on Tech Brew.
What you’ll need
| Item | Reason |
|---|---|
| Raspberry Pi 4 (4 GB or more) | Enough RAM to load a small language model |
| Micro‑SD card (32 GB, Class 10) | Fast storage for the OS and model files |
| Power supply (5 V 3 A) | Keeps the Pi stable under load |
| USB microphone and speaker (or a USB sound card) | For voice input and output |
| Internet connection (Wi‑Fi or Ethernet) | To download software and models |
| Optional: case and heat sink | Keeps the Pi cool during long runs |
All of these items are easy to find on Amazon or your local electronics store. If you already have a Pi lying around, you can skip the purchase step and jump straight to the setup.
Setting up the Pi
- Flash Raspberry Pi OS – Download the “Raspberry Pi OS Lite” image from the official site. Use the Raspberry Pi Imager (or balenaEtcher) to write it to the SD card. Choose the “Lite” version because we don’t need a desktop for this project.
- Boot and SSH – Insert the card, power up the Pi, and find its IP address (you can check your router). Enable SSH by creating an empty file named
sshon the boot partition before you insert the card. Then SSH in:ssh pi@<your_pi_ip>The default password is
raspberry. Change it right away withpasswd. - Update the system – Run:
sudo apt update && sudo apt upgrade -yThis makes sure you have the latest packages and security patches.
Installing the software
1. Install Python and essential tools
sudo apt install -y python3 python3-pip python3-venv git build-essential libatlas-base-dev
2. Set up a virtual environment
python3 -m venv ~/assistant_venv
source ~/assistant_venv/bin/activate
Using a virtual environment keeps the libraries separate from the system Python, which avoids version clashes.
3. Get the Whisper speech‑to‑text engine
For voice input we’ll use OpenAI’s Whisper model. The tiny version works well on a Pi.
pip install --upgrade pip
pip install git+https://github.com/openai/whisper.git
pip install torch==2.0.0+cpu -f https://download.pytorch.org/whl/torch_stable.html
The torch line installs the CPU‑only build of PyTorch, which is what the Pi can run.
4. Install a small language model
We’ll use the “tiny” version of the LLaMA‑like model from the ggml community. It’s a few hundred MB and runs fast enough for simple commands.
mkdir -p ~/models && cd ~/models
wget https://huggingface.co/ggml/llama-2-7b-chat/resolve/main/ggml-model-q4_0.bin
If the link changes, just look for the latest ggml-model-q4_0.bin on Hugging Face.
If you’re curious about bringing similar LLM capabilities to web applications, see our guide on integrating open‑source LLMs into JavaScript projects.
5. Install audio handling tools
sudo apt install -y sox libsox-fmt-all
pip install pyaudio sounddevice
sox helps us record and play audio, while pyaudio lets Python talk to the microphone.
Getting the model running
Create a new Python file called assistant.py in your home folder:
#!/usr/bin/env python3
import os
import subprocess
import sounddevice as sd
import numpy as np
import whisper
import torch
# Load Whisper model
whisper_model = whisper.load_model("tiny")
# Load LLaMA model via llama.cpp wrapper (assume you have llama.cpp installed)
# For simplicity we call a shell command that runs the model
def run_llama(prompt):
cmd = [
"./llama-cli",
"-m", os.path.expanduser("~/models/ggml-model-q4_0.bin"),
"-p", prompt,
"-n", "128"
]
result = subprocess.run(cmd, capture_output=True, text=True)
return result.stdout.strip()
def record_audio(duration=5, fs=16000):
print("Listening...")
audio = sd.rec(int(duration * fs), samplerate=fs, channels=1, dtype='int16')
sd.wait()
return audio.squeeze()
def main():
while True:
audio = record_audio()
# Save to wav for Whisper
wav_path = "input.wav"
sd.write(wav_path, audio, 16000)
# Transcribe
result = whisper_model.transcribe(wav_path)
text = result["text"].strip()
if not text:
continue
print(f"You said: {text}")
# Simple exit command
if "stop" in text.lower():
print("Goodbye!")
break
# Get response from LLaMA
response = run_llama(text)
print(f"Assistant: {response}")
# Speak the response (using espeak)
subprocess.run(["espeak", "-v", "en-us", response])
if __name__ == "__main__":
main()
A few notes:
- The script records 5 seconds of audio, sends it to Whisper, then feeds the transcription to the LLaMA model.
- We use
espeak(installed viasudo apt install espeak) to read the answer out loud. - The
run_llamafunction assumes you have compiledllama-clifrom thellama.cpprepo. That compile step is simple:
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
make
Now you have an executable llama-cli in the folder.
Testing your assistant
- Make the script executable:
chmod +x assistant.py - Run it:
./assistant.py - Say something like “What’s the weather today?” or “Turn on the fan”. The assistant will print the transcription, call the tiny LLaMA model, and speak the answer.
If you hear garbled audio, check that your microphone and speaker are correctly set up. Run arecord -l and aplay -l to list devices, then edit the sd.rec line with the proper device index.
Next steps and tips
- Add custom commands – You can add a simple
ifblock that looks for “lights on” and then runs a GPIO command to toggle a relay. - Swap models – If you have a more powerful Pi (like a Pi 5) you can try the “base” Whisper model for better accuracy.
- Run headless – Set the Pi to start the script on boot by adding a line to
~/.bashrcor creating a systemd service. For automated deployments you might explore a personal CI/CD pipeline with GitHub Actions. - Keep it tidy – Delete the large model files you’re not using. Space is cheap, but the SD card can fill up fast.
That’s it! You now have a personal AI assistant humming away on a Raspberry Pi. At Tech Brew we love seeing how a few lines of code can turn a cheap board into a useful tool. Play around, break things, and then fix them – that’s the best way to learn.
- →
- →
- →
- →
- →