Running LLMs Locally on Linux with Ollama
Running a language model on your own hardware stopped being difficult around the time Ollama appeared. It handles model downloading, quantization formats, GPU offloading, and serving behind an API, and the setup is genuinely two commands.
What has not changed is the physics. Models need memory proportional to their size, and the honest version of this guide spends as much time on what you cannot run as on what you can.
Installing
curl -fsSL https://ollama.com/install.sh | sh
Piping a script from the internet into a shell is a habit worth resisting. Read it first:
curl -fsSL https://ollama.com/install.sh -o install.sh
less install.sh
sh install.sh
The script installs a binary and a systemd service. Check it came up:
systemctl status ollama
Most distributions also package it, which is preferable if yours does.
GPU support
NVIDIA needs the proprietary driver and CUDA libraries. Our NVIDIA drivers guide covers installation.
nvidia-smi # driver working, VRAM visible
ollama run llama3.1 --verbose # check it reports GPU layers
AMD works through ROCm on supported cards, which is a shorter list than AMD’s marketing implies. Check your specific card before assuming.
If the GPU is not detected, Ollama silently falls back to CPU. The symptom is inference that works and is inexplicably slow, so verify rather than assume.
Running a model
ollama run llama3.1
First run downloads several gigabytes. After that it drops you into a prompt.
ollama list # what you have
ollama pull qwen2.5-coder:7b # download without running
ollama rm llama3.1 # free the disk
ollama ps # what is currently loaded in memory
Models live in /usr/share/ollama/.ollama/models by default and they are large. Point it elsewhere if your root filesystem is small:
sudo systemctl edit ollama
[Service]
Environment="OLLAMA_MODELS=/mnt/storage/ollama"
How much memory you need
This is the whole question, and the arithmetic is simple.
At 4-bit quantization, budget roughly 1GB of VRAM per billion parameters, plus 1 to 2GB for context and overhead.
| Model size | VRAM at Q4 | Realistic on |
|---|---|---|
| 3B | ~3GB | Almost anything, including integrated graphics |
| 7-8B | ~6GB | RTX 3060 12GB, RTX 4060, most modern cards |
| 14B | ~10GB | RTX 3080, 4070, 12GB cards |
| 32B | ~20GB | RTX 3090, 4090, 24GB cards |
| 70B | ~42GB | Two 24GB cards, or an A6000 |
If a model does not fit in VRAM, Ollama splits it between GPU and system RAM. It runs, and the layers on the CPU are dramatically slower, so a model that is 80 percent offloaded performs closer to CPU-only than to GPU speed.
ollama run llama3.1 --verbose
# reports layers offloaded to GPU; you want all of them
Quantization
Models are trained at 16-bit precision. Quantization reduces that to store weights in fewer bits.
The naming looks cryptic and follows a pattern: Q4_K_M is 4-bit, K-quant method, medium size.
Q8 is nearly indistinguishable from full precision and needs roughly twice the memory of Q4.
Q5_K_M is slightly better than Q4 for a modest memory increase.
Q4_K_M is the default for good reason. The quality loss against 8-bit is small enough that most people cannot identify it blind, and it halves the memory requirement.
Q3 and below degrade visibly: more repetition, weaker instruction following, more confident errors.
ollama pull llama3.1:8b-instruct-q4_K_M
ollama pull llama3.1:8b-instruct-q8_0
The practical rule: a larger model at Q4 beats a smaller model at Q8 at the same memory budget. A 14B at Q4 will outperform an 8B at Q8.
Which models
The landscape moves monthly, so treat specific names as examples rather than recommendations.
General purpose: the Llama, Qwen, Mistral, and Gemma families all have instruction-tuned releases in the 7B to 14B range that are genuinely useful.
Code: Qwen Coder and the DeepSeek Coder models are strong for their size. A 7B coding model that runs locally and completes functions is a different and more useful thing than a general model of the same size.
Small: 3B models run on almost anything and are adequate for classification, extraction, and simple summarisation. They are not adequate for reasoning.
Check the licence. “Open weights” is not the same as open source, and several popular models carry usage restrictions.
The API
Ollama serves HTTP on port 11434, which is what makes it useful as infrastructure rather than a toy.
curl http://localhost:11434/api/generate -d '{
"model": "llama3.1",
"prompt": "Explain what a systemd unit file does in two sentences.",
"stream": false
}'
It also exposes an OpenAI-compatible endpoint, so most tooling written against the OpenAI API works by changing the base URL.
curl http://localhost:11434/v1/chat/completions -d '{
"model": "llama3.1",
"messages": [{"role": "user", "content": "Hello"}]
}'
By default Ollama binds to localhost. Leave it that way. There is no authentication whatsoever, so binding it to 0.0.0.0 on a machine reachable from the internet hands anyone free use of your GPU. If you need remote access, reach it over WireGuard or put it behind an authenticating reverse proxy.
A web interface
Ollama has none. Open WebUI is the usual pairing:
services:
open-webui:
image: ghcr.io/open-webui/open-webui:main
ports:
- "127.0.0.1:3000:8080"
environment:
- OLLAMA_BASE_URL=http://host.docker.internal:11434
volumes:
- open-webui:/app/backend/data
extra_hosts:
- "host.docker.internal:host-gateway"
restart: unless-stopped
volumes:
open-webui:
Note the 127.0.0.1: prefix on the port binding. Without it Docker publishes the port on every interface and bypasses your firewall rules, which surprises people regularly.
Our AI and ML self-hosted apps directory covers the wider set of options.
Custom models
A Modelfile sets a system prompt and parameters without retraining anything:
FROM llama3.1
PARAMETER temperature 0.3
PARAMETER num_ctx 8192
SYSTEM """
You are a Linux systems administration assistant. Prefer concise,
correct answers. When a command is destructive, say so explicitly.
"""
ollama create sysadmin -f ./Modelfile
ollama run sysadmin
num_ctx is worth knowing about: context length costs memory, and raising it on a model that barely fits will push layers onto the CPU.
An honest comparison
Local models at consumer sizes are not equivalent to the large hosted ones. Anyone telling you a 7B model replaces Claude or GPT is either not using it for much or has not compared carefully.
What local models are good at: summarising text you provide, drafting and rewriting, straightforward code completion, extraction and classification, answering questions about material in the prompt.
Where they are weaker: multi-step reasoning, long context, accuracy on obscure facts, and knowing when they do not know. Small models hallucinate confidently and frequently.
The reasons to run locally are not capability:
Privacy. The prompt never leaves the machine. For anything involving client data, medical information, or unpublished work, this is not a preference but a requirement.
Cost at volume. Processing ten thousand documents through an API has a bill. Locally it has an electricity cost.
Offline. Works on a plane, in a datacentre with no egress, on an air-gapped network.
No rate limits, no deprecation. The model on your disk works the same next year.
Learning. Running inference teaches you what these systems actually are in a way using an API does not.
If you want the best available answer to a hard question, use a frontier model. If you want a capable tool that runs on your hardware and tells nobody what you asked, this is that.
Frequently Asked Questions
How much VRAM do I need to run a local LLM?
As a rough guide, a model needs about one gigabyte of VRAM per billion parameters at 4-bit quantization, plus one to two gigabytes of overhead for context. An 8 billion parameter model fits comfortably in 8GB, a 14 billion parameter model needs around 12GB, and a 70 billion parameter model needs roughly 48GB or a pair of large cards.
Can I run a local LLM without a GPU?
Yes, Ollama falls back to CPU inference and it works. It is slow, typically two to five tokens per second for a 7 or 8 billion parameter model against fifty or more on a modern GPU. Enough for occasional queries or batch processing, frustrating for conversation.
What does quantization do to model quality?
Quantization reduces the numeric precision of the model weights, shrinking memory use at some cost to accuracy. The drop from 8-bit to 4-bit is usually barely noticeable in practice, while 3-bit and below degrade output quality visibly. Q4_K_M is the common default because it sits at a good point on that curve.
Is a local model as good as ChatGPT or Claude?
No, not at the sizes most people can run. A good 8 to 14 billion parameter model is genuinely useful for summarising, drafting, simple code, and answering questions, but it will be weaker at complex reasoning, long context, and accuracy on obscure facts. The reasons to run locally are privacy, cost, offline capability, and control rather than raw capability.
Does Ollama send my data anywhere?
Inference runs entirely on your machine and prompts are not transmitted. Ollama does contact its registry to download models and to check for updates, which is the only outbound traffic in normal use. You can verify this with a packet capture or block it at the firewall after downloading the models you want.
How do I give Ollama a web interface?
Open WebUI is the common choice and runs as a container pointed at the Ollama API on port 11434. It provides a chat interface, conversation history, model switching, and document upload. Ollama itself is a command-line tool and an HTTP API with no interface of its own.