Run AI locally with Ollama: a plain-English guide
How to run AI models on your own computer with Ollama: what hardware you need, which models to try, privacy and cost benefits, and the honest limitations.
By Outline Digital · · 6 min read
What is Ollama and why would you run AI locally?
Ollama is a free, open-source tool that downloads AI language models and runs them on your own computer. Once a model is downloaded, it runs without sending anything to the internet: your questions, your documents and the answers stay on your machine.
There are three good reasons to do it:
- Privacy: client lists, contracts or business plans never leave your computer. Useful for accountants, solicitors, clinics and anyone handling sensitive data.
- Cost: no subscription and no charge per message. Once you have the computer, running the model costs only electricity.
- Control: the model doesn’t change overnight, doesn’t get withdrawn, and works the same way every time.
The trade-off is power. The very largest models from OpenAI, Anthropic or Google run on data-centre hardware that no laptop can match. Local models are smaller: very good at many tasks, weaker at long, complex reasoning.
What hardware do you need to run Ollama?
The most important number is memory: the model has to fit in RAM (or in the graphics card’s memory). Model sizes are measured in billions of parameters (written 3B, 8B, 14B), and Ollama uses compressed (“quantised”) versions that need roughly 0.6–0.7 GB per billion parameters, plus some headroom.
- 8 GB RAM: small models of 1–4B. Fine for simple writing, summaries and classification.
- 16 GB RAM: 7–8B models comfortably. The sweet spot for many people.
- 32 GB RAM: 14B models and some larger ones, noticeably smarter.
- 64 GB or more: 30B-class models and beyond, approaching cloud quality on many tasks.
Apple Silicon Macs (M1 and later) are particularly good, because the processor and graphics share the same fast memory. A MacBook Air with 16 GB runs an 8B model well. On Windows and Linux, an NVIDIA graphics card with 8–12 GB of video memory makes a big difference; without one, models run on the processor, more slowly. Leave plenty of disk space too: models range from 1 to 20 GB or more each.
How do you install Ollama and run your first model?
- Download Ollama from its official website for macOS, Windows or Linux and install it like any app.
- Open a terminal (Terminal on Mac, PowerShell on Windows).
- Run a model by typing “ollama run” followed by the model name, for example a small Llama, Qwen or Gemma model. The first time, it downloads the model; afterwards it starts in seconds.
- Chat in the terminal, or connect a friendlier app on top. Many apps can talk to Ollama, because it exposes a local interface on your computer.
- Try a second model and compare. Different models are good at different things.
If the model is too big for your memory, everything slows to a crawl as the computer swaps data to disk. If that happens, pick a smaller model: a fast small model is far more useful than a slow large one.
Which models should you try first?
The model library changes quickly, so take this as a starting point rather than a ranking. Good families to try include Llama (Meta), Qwen (Alibaba, strong at coding and many languages), Gemma (Google), Mistral and DeepSeek. Most come in several sizes: start with the size that fits your memory, then move up if your computer copes.
For writing code, look for models with “coder” in the name. For general business writing, a general instruction-tuned model of 7–8B is a good start. Always check the model’s licence if you plan to use it commercially; most popular ones allow it, with some conditions.
Do you need Ollama, or can AI run in the browser?
There is a lighter option: some web apps can run small models directly in the browser, using your computer’s graphics chip through a technology called WebGPU. Nothing to install, nothing to configure: the model downloads once and is cached.
Browser models are convenient for trying things out and for small tasks, but they’re limited to smaller models and depend on a recent browser and a reasonably modern computer. Ollama is the better choice when you want bigger models, faster responses, or one model shared by several apps on the same machine. Many people start in the browser and install Ollama once they know they’ll use local AI regularly.
What are local models good for, and where do they struggle?
- Good: drafting emails and product descriptions, summarising documents, rewriting text, extracting data from notes, answering questions about your own files, simple coding tasks.
- Mixed: longer documents (local models often have smaller context windows by default), tasks in less common languages.
- Weaker: long multi-step reasoning, large software projects, the latest facts (they don’t browse the web by themselves).
Can you build apps and software with a local model?
Yes. Plannify is designed to work with the AI you choose, including free models running on your own computer with Ollama, or directly in the browser with nothing to install. You describe what you want (business software, an app, a website, an online shop) and a team of AI agents builds it on your computer, with the model doing the thinking.
To make small models reliable, the team starts from cases it has already studied, uses proven building blocks and tests every piece: if the model gets something wrong, the mistake goes back until it works. You can switch to a bigger model, or to your own API key from OpenRouter, Anthropic or OpenAI, whenever a project needs more power.
See it in practice on build an app or build business software.
Frequently asked questions
Yes. Ollama is free and open source, and most models it runs are free to download. Your only costs are the computer and the electricity.
Yes. A laptop with 16 GB of memory runs popular 7–8B models comfortably, and Apple Silicon MacBooks are especially good at it. With 8 GB, stick to small models.
Once the model is downloaded, your prompts and documents stay on your computer and aren’t sent to a cloud service. Apps you connect on top of Ollama have their own privacy terms, so check them.
Not the biggest ones. Local models are very capable for everyday tasks, but the largest cloud models are still better at complex reasoning and large projects.
Once a model is downloaded, Ollama itself runs offline. Apps that use it may still need a connection for their own features.
Start your project
A team of AI agents builds your app or website. The project stays on your computer. It’s free.