Ollama: Run Language Models Locally
Ollama runs large language models locally. ollama run llama3 downloads the model and gives you a prompt, which is most of why it became the default entry point.
Why run models yourself
Prompts sent to a hosted API are prompts you have handed to a third party, which rules out cloud models for confidential documents, client work, and anything under regulatory constraint. Local inference removes that entirely, works offline, and costs nothing per token after the hardware.
Hardware is the real constraint
Model size determines what you need. Roughly, a 7B parameter model quantized to 4-bit wants around 5GB of memory, a 13B around 9GB, and a 70B considerably more than most people have. A GPU with enough VRAM is dramatically faster than CPU inference; Apple Silicon performs well because memory is unified. Running a large model on CPU works and can be slow enough to be unpleasant for interactive use.
The API is what makes it useful
Ollama exposes an HTTP API, including an OpenAI-compatible endpoint, so existing tools point at it with a URL change. Open WebUI is the usual front-end, giving a ChatGPT-style interface over local models, and many editors and applications support Ollama directly.
Practical notes
Models are large downloads and accumulate quickly; check disk before pulling several. Quantized variants trade a little quality for substantially lower memory use and are usually the right choice on consumer hardware.
Alternatives
llama.cpp is the engine underneath much of this ecosystem and can be used directly. LocalAI targets drop-in OpenAI API compatibility more broadly.
License
Ollama is released under the MIT License.