How to Run AI Locally: Vision, Speech, Images, and Agents

TLDR: To run AI locally beyond a chatbot, give each workload its own engine: Ollama for chat, tool calling, and image understanding, whisper.cpp for speech-to-text, Piper for text-to-speech, and ComfyUI or stable-diffusion.cpp for image generation. Then budget memory across all of them, check the license of every engine, model, and voice, and switch off the cloud features that local tools now ship with.
Local AI is several engines, not one. Each has its own model format, default port, memory footprint, and license terms, and the trouble starts where they meet: two servers competing for port 8080, a GPU that cannot hold the chat model and the image model at once, a voice whose training data carries a research-only license. This guide covers vision, speech, image generation, and a local voice agent that ties them together. If you only need a chat model, the walkthrough for running an LLM locally covers install and testing, and the local AI models guide compares LLM families, sizes, and licenses. Neither is repeated here.
What does running AI locally involve beyond a chat model?
Start with a map. Each row is a separate process you install, start, and monitor. Sizes and addresses come from each project's documentation as of September 2026.
| Workload | Engine (license) | Starter model | Download or memory | Default address |
|---|---|---|---|---|
| Chat and tool calling | Ollama (MIT) | qwen3:8b | 5.2 GB | 127.0.0.1:11434 |
| Image understanding | Ollama (MIT) | qwen3-vl:4b | 3.3 GB | 127.0.0.1:11434 |
| Speech-to-text | whisper.cpp (MIT) | Whisper base.en | 142 MiB on disk, about 388 MB in memory | 127.0.0.1:8080 |
| Text-to-speech | Piper (GPL-3.0) | en_US-lessac-medium | 63 MB | 0.0.0.0:5000 |
| Image generation | ComfyUI (GPL-3.0) or stable-diffusion.cpp (MIT) | FLUX.2 [klein] 4B | About 13 GB of VRAM | 127.0.0.1:8188 (ComfyUI) |
Sources: the Ollama library and FAQ, the whisper.cpp README, Piper's voice repository, the FLUX.2 [klein] 4B card, and ComfyUI's command-line options.
Two of those defaults bite on day one. llama.cpp's llama-server (launched as llama serve in its current quick start) and whisper.cpp's whisper-server both listen on 127.0.0.1:8080, per the llama.cpp server README and the whisper.cpp server README, so whichever starts second fails to bind. And Piper's HTTP server binds 0.0.0.0, meaning every network interface, unless you pass --host, according to its server source. The other engines default to loopback. Assign ports up front and bind everything to 127.0.0.1.
How do you add vision to a local model?
Image understanding runs through the same Ollama server as chat. Send base64-encoded images in the images array of a /api/chat message, as the Ollama vision docs show, and the model can read receipts, screenshots, charts, and photos.
ollama pull qwen3-vl:4b
IMG=$(base64 < receipt.jpg | tr -d '\n')
curl -s http://127.0.0.1:11434/api/chat -d '{
"model": "qwen3-vl:4b",
"stream": false,
"messages": [{"role": "user", "content": "List the line items and the total.", "images": ["'"$IMG"'"]}]
}'
Choose the vision model on three properties. Size decides whether it fits. The capability tags on its Ollama library page decide whether it can drive an agent. The license decides whether you can ship it.
| Model | Sizes on Ollama | Context | Tool calling tag | License |
|---|---|---|---|---|
| qwen3-vl | 2B 1.9 GB, 4B 3.3 GB, 8B 6.1 GB, 32B 21 GB | 256K | Yes | Apache 2.0 |
| gemma4 | E4B 9.6 GB, 12B 7.6 GB, 31B 20 GB | 128K to 256K | Yes | Apache 2.0 |
| gemma3 | 4B 3.3 GB, 12B 8.1 GB, 27B 17 GB (1B and 270M are text-only) | 128K | Not listed | Gemma Terms of Use |
| llama3.2-vision | 11B 7.8 GB, 90B 55 GB | 128K | Not listed | Llama 3.2 Community License, with an EU carve-out |
The Llama row carries a trap. Llama 3.2's acceptable use policy withholds the license grant for its multimodal models from individuals domiciled in the European Union and from companies with their principal place of business there. End users of a product built on those models are not affected, but an EU team building one is. The tag column matters later: qwen3-vl and gemma4 list tool calling, so one model can both read an image and call a function, while gemma3 and llama3.2-vision list vision only.
How do you transcribe speech locally?
whisper.cpp runs OpenAI's Whisper models as a plain C/C++ implementation without dependencies, on the CPU or on Metal, CUDA, and Vulkan GPUs. Whisper's code and weights are MIT-licensed, per OpenAI's repository, and whisper.cpp is MIT as well. Its bundled server turns it into a local HTTP service.
git clone https://github.com/ggml-org/whisper.cpp.git
cd whisper.cpp
sh ./models/download-ggml-model.sh base.en
cmake -B build
cmake --build build -j --config Release
# whisper.cpp expects 16-bit WAV; convert recordings to 16 kHz mono first
ffmpeg -i question.m4a -ar 16000 -ac 1 -c:a pcm_s16le question.wav
# Port 8081, because llama-server also defaults to 8080
./build/bin/whisper-server -m models/ggml-base.en.bin --port 8081
# From a second terminal: returns {"text": "..."}
curl 127.0.0.1:8081/inference -F [email protected] -F response_format=json
Three details from the README save time. Input must be 16-bit WAV; the server can convert uploads itself with --convert if ffmpeg is on the host. Memory scales with model size: tiny needs about 273 MB, base about 388 MB, small about 852 MB, medium about 2.1 GB, and large about 3.9 GB, so start with base and move up only if your recordings demand it. And the server README warns against running it with administrative privileges, because it accepts uploads and hands them to ffmpeg. Keep it on loopback and sandboxed.
One-process alternatives exist. llama.cpp's server accepts audio input for models such as Voxtral Mini 3B and Qwen3-ASR, plus image input, per its multimodal docs. Keep Whisper as the baseline and compare both on the same recordings before you switch.
How do you generate speech locally?
Piper describes itself as a fast, local neural text-to-speech engine. Its voices are ONNX files, and it runs on the CPU unless you enable GPU acceleration. Development moved from rhasspy/piper, now archived under MIT, to the Open Home Foundation's piper1-gpl, and piper-tts 1.8.0 on PyPI declares GPL-3.0-or-later. Check that against how you distribute your software.
python3 -m pip install "piper-tts[http]"
python3 -m piper.download_voices en_US-lessac-medium
# Bind to loopback; without --host the server listens on all interfaces
python3 -m piper.http_server -m en_US-lessac-medium --host 127.0.0.1 --port 5000
# From a second terminal: POST text, get a WAV file back
curl -X POST -H 'Content-Type: application/json' -d '{"text": "The build passed."}' -o reply.wav 127.0.0.1:5000/synthesize
Use the server for anything interactive, because the CLI loads the voice model on every run. The HTTP API takes JSON with a required text field plus optional speed and speaker controls, and returns WAV audio. GPU acceleration needs --cuda and the onnxruntime-gpu package.
Check the voice license separately from the engine license. The model card for en_US-lessac-medium, the example voice in Piper's CLI and HTTP docs, points to the Blizzard 2013 Lessac dataset license, a research license that excludes commercial use of the data, including voice synthesis products. Read the card for every voice you ship, and let counsel decide what a dataset license means for a model trained on it. Piper's README also says the project is looking for maintainers, which matters for anything long-lived.
How do you generate images locally?
Two engines suit different needs. ComfyUI (GPL-3.0) is a node-graph application for image, video, audio, and 3D workflows with a local API, and it supports models from Stable Diffusion 1.5 through FLUX.2 and Qwen Image. stable-diffusion.cpp (MIT) is a C/C++ program built on ggml that works the same way as llama.cpp, per its README, with CPU, CUDA, Vulkan, Metal, OpenCL, and SYCL backends and GGUF support. One command, sd-cli -m model.safetensors -p "a lovely cat", produces an image.
git clone https://github.com/comfyanonymous/ComfyUI
cd ComfyUI
# Install the PyTorch build for your GPU first, as the README describes
pip install -r requirements.txt
# Checkpoints go in models/checkpoints. The UI serves on 127.0.0.1:8188.
python main.py --disable-api-nodes
ComfyUI's README says the core runs fully offline and downloads nothing unless you ask, but it also ships optional paid API nodes that call hosted closed models. --disable-api-nodes turns those off and, per the command-line help, stops the frontend from talking to the internet. The README also claims the largest open models run on as little as 4 GB of VRAM plus 8 GB of RAM by streaming weights; time it on your own card. A --cpu mode exists, and ComfyUI's help text calls it slow.
Image model licenses differ within a single family, so check each checkpoint:
| Model | License | Note |
|---|---|---|
| FLUX.2 [klein] 4B | Apache 2.0 | About 13 GB of VRAM, RTX 3090 or 4070 and up, per its card |
| FLUX.2 [klein] 9B | FLUX Non-Commercial License | Same family, different terms, per the 4B card |
| FLUX.1 [schnell] | Apache 2.0 | About 12B parameters, three times klein 4B |
| FLUX.1 [dev] | FLUX.1 [dev] Non-Commercial License | Commercial use needs separate terms from the publisher |
| SDXL base 1.0 | CreativeML Open RAIL++-M | Use-based restrictions you must pass on to your users |
| Stable Diffusion 3.5 | Stability AI Community License | Custom terms; read before shipping |
How do you fit several models on one machine?
Add up the resident pieces before you buy hardware. With the starter models above:
| Component | Footprint | Where it can run |
|---|---|---|
| qwen3:8b, chat and tools | 5.2 GB download | GPU, or CPU at lower speed |
| qwen3-vl:4b, vision | 3.3 GB download | GPU, or CPU at lower speed |
| Whisper base.en | About 388 MB in memory | CPU or GPU; --no-gpu forces CPU |
| Piper en_US-lessac-medium | 63 MB voice file | CPU by default; --cuda for GPU |
| FLUX.2 [klein] 4B | About 13 GB of VRAM | GPU |
Chat, vision, and both speech models come to roughly 9 GB of weights. Treat download size as a floor, because context adds memory on top; the self-hosted LLM serving guide explains why context length is a memory budget. That set leaves headroom on a 16 GB GPU. Add the image model and it no longer fits, so something has to move.
Ollama manages part of this, per its FAQ. It keeps each model loaded for five minutes after its last request and allows up to three loaded models per GPU by default (OLLAMA_MAX_LOADED_MODELS), but on a GPU a new model loads alongside the others only if it fits entirely in VRAM. Otherwise requests queue while idle models unload, so the symptom is a long pause after a switch rather than an error. Nothing coordinates across engines, and ComfyUI cannot see what Ollama holds. Unload explicitly before a heavy job:
# Unload the chat model now instead of waiting out the 5-minute keep-alive
curl http://127.0.0.1:11434/api/generate -d '{"model": "qwen3:8b", "keep_alive": 0}'
# Show loaded models and whether each sits on GPU, CPU, or both
ollama ps
Move the small models to the CPU so the GPU is free for the two that need it. Whisper base.en needs about 388 MB, the Piper voice file is 63 MB, and both run on the CPU. On Apple Silicon, whisper.cpp runs inference on the GPU through Metal, and unified memory means the whole budget above comes out of one pool.
How do you wire the pieces into a local voice agent?
The script below chains the engines. whisper-server transcribes a WAV file, a Qwen3 model on Ollama answers and may call a web search tool, and Piper speaks the reply. It uses only the Python standard library, and everything runs on your machine except the optional search call. Qwen3 lists tool calling on its Ollama library page.
import json
import os
import sys
import urllib.request
import uuid
WHISPER = "http://127.0.0.1:8081/inference" # whisper-server started with --port 8081
OLLAMA = "http://127.0.0.1:11434/api/chat"
PIPER = "http://127.0.0.1:5000/synthesize" # Piper started with --host 127.0.0.1
SEARCH = "https://ydc-index.io/v1/search" # the only call that leaves the machine
MODEL = "qwen3:8b"
SEARCH_TOOL = {
"type": "function",
"function": {
"name": "web_search",
"description": "Search the web for facts that may have changed recently.",
"parameters": {
"type": "object",
"required": ["query"],
"properties": {"query": {"type": "string", "description": "Search query"}},
},
},
}
def post(url, body, headers, timeout=120):
req = urllib.request.Request(url, data=body, headers=headers, method="POST")
with urllib.request.urlopen(req, timeout=timeout) as resp:
return resp.read()
def transcribe(wav_path):
boundary = uuid.uuid4().hex
with open(wav_path, "rb") as f:
audio = f.read()
head = (
f"--{boundary}\r\n"
'Content-Disposition: form-data; name="response_format"\r\n\r\njson\r\n'
f"--{boundary}\r\n"
'Content-Disposition: form-data; name="temperature"\r\n\r\n0.0\r\n'
f"--{boundary}\r\n"
'Content-Disposition: form-data; name="file"; filename="input.wav"\r\n'
"Content-Type: audio/wav\r\n\r\n"
).encode()
body = head + audio + f"\r\n--{boundary}--\r\n".encode()
raw = post(WHISPER, body, {"Content-Type": f"multipart/form-data; boundary={boundary}"})
text = json.loads(raw)["text"].strip()
if not text:
raise RuntimeError("no speech found; is the file 16 kHz mono 16-bit WAV?")
return text
def web_search(query, count=5):
body = json.dumps({"query": query, "count": count}).encode()
headers = {"Content-Type": "application/json", "X-API-Key": os.environ["YDC_API_KEY"]}
data = json.loads(post(SEARCH, body, headers, timeout=30))
hits = data.get("results", {}).get("web", [])[:count]
return json.dumps([
{"title": h.get("title"), "url": h.get("url"),
"text": (h.get("snippets") or [h.get("description", "")])[0]}
for h in hits
])
def chat(messages, tools):
payload = {"model": MODEL, "messages": messages, "stream": False,
"think": False, "keep_alive": "10m"}
if tools:
payload["tools"] = tools
raw = post(OLLAMA, json.dumps(payload).encode(), {"Content-Type": "application/json"})
return json.loads(raw)["message"]
def answer(question, max_rounds=3):
tools = [SEARCH_TOOL] if os.environ.get("YDC_API_KEY") else []
messages = [
{"role": "system", "content": "Your reply will be read aloud. Use at most three "
"short sentences. Call web_search only when the answer may have changed recently."},
{"role": "user", "content": question},
]
for _ in range(max_rounds):
msg = chat(messages, tools)
messages.append(msg)
calls = msg.get("tool_calls") or []
if not calls:
return msg.get("content", "").strip()
for call in calls:
fn = call.get("function", {})
args = fn.get("arguments") or {}
if isinstance(args, str): # tolerate servers that send a JSON string
args = json.loads(args)
if fn.get("name") == "web_search" and args.get("query"):
result = web_search(args["query"])
else:
result = json.dumps({"error": "unknown tool or missing query"})
messages.append({"role": "tool", "tool_name": fn.get("name", ""), "content": result})
raise RuntimeError(f"no final answer after {max_rounds} tool rounds")
def speak(text, out_path):
wav = post(PIPER, json.dumps({"text": text}).encode(), {"Content-Type": "application/json"})
if wav[:4] != b"RIFF":
raise RuntimeError("Piper did not return a WAV file")
with open(out_path, "wb") as f:
f.write(wav)
if __name__ == "__main__":
heard = transcribe(sys.argv[1])
print("heard:", heard)
reply = answer(heard)
print("reply:", reply)
speak(reply, "reply.wav")
Start the three servers from the earlier sections, then run python3 voice_agent.py question.wav. It prints what it heard and what it will say, and writes reply.wav.
Four choices are deliberate. Tool calls follow Ollama's native format from its tool-calling docs: arguments arrive as a JSON object, and results go back as a tool message with tool_name; the client also accepts string-encoded arguments in case you swap servers. "think": false requests no thinking output, per the chat API reference, so no reasoning trace reaches the speaker. The loop stops after three tool rounds, because a model that keeps requesting tools would otherwise never answer. And the search tool is offered only when a key is set, so the agent runs fully offline without one.
Search is there because local weights stop at their training cutoff, so "what changed recently" is the question a local agent cannot answer alone. The tool sends only the query to the You.com Web Search API: a POST to https://ydc-index.io/v1/search with an X-API-Key header, which returns results under results.web with titles, URLs, and snippets. As of September 2026 it costs $5.00 per 1,000 calls, per the billing docs. New accounts get $100 in credits, and self-serve keys default to 10 requests per second, per the rate limits page.
That query does leave your network. Zero Data Retention, available on enterprise agreements, limits what is retained, not what is sent, so keep private details out of queries. If your agent lives in an MCP client instead, the keyless You.com MCP server at https://api.you.com/mcp?profile=free exposes you-search and you-discover for up to 100 queries per day, and the MCP explainer covers how clients connect. For a full answer pipeline with citation checks, see the self-hosted AI search engine guide; for coding agents, the local coding LLM guide.
What should you test before relying on a local AI stack?
Run these checks once per machine, and again after every engine or model update.
- Offline run. Pull every model and voice first, then disconnect the network and run the voice agent with
YDC_API_KEYunset. Anything that fails has a hidden download or cloud dependency. - Cloud-feature audit. Local tools now ship hosted options behind the same interface. The Ollama library lists cloud tags such as
gemma4:31b-cloudbeside local ones, and the Ollama FAQ documentsOLLAMA_NO_CLOUD=1to turn cloud features off. Keep--disable-api-nodesin ComfyUI's start command. - Bind audit. List listening sockets and confirm every engine is on 127.0.0.1, especially Piper. For outbound traffic, run the boundary test in the on-premise AI guide, which also covers the cost of owning hardware.
- License audit. Record the engine, weights, and, for voices, training-data license for each workload. The models in this guide alone span MIT, Apache 2.0, GPL-3.0, the Gemma Terms of Use, the Llama 3.2 Community License, non-commercial FLUX terms, and a research-only dataset license.
- Accuracy on your inputs. Run 20 real recordings through Whisper base.en and small, and 20 real images through the vision model, and compare outputs by hand. Generic benchmarks will not show how a model handles your accents, jargon, or layouts.
- Swap latency. Time the first image after a chat turn and the first reply after an image. If the pause is too long, move speech to the CPU, use a smaller vision model, or give images their own GPU.
Cost follows utilization. Local inference has no per-request fee, but hardware, power, and upkeep are fixed costs that pay off only when the machine stays busy. In this stack, the only metered call is search.
LI Test
LI Test
Share Article:
Related resources.

What Is Jev? TypeSafe AI's System One Model, Explained for Developers
September 20, 2026
Blog

Company Data Enrichment API: Providers, Pricing, and How to Test Them
September 18, 2026
Blog

Best Local LLM for Coding: A Developer's Guide to AI-Powered Programming
August 20, 2026
Blog

Local LLM: Running Large Language Models on Your Own Infrastructure
August 19, 2026
Blog

Lead Enrichment API: Automated Contact and Company Data Enhancement
August 18, 2026
Blog
