How to run your first local LLM on a Mac with Ollama
Short answer
Install Ollama from ollama.com or with Homebrew, open Terminal, and run "ollama run gemma4". Ollama downloads the model the first time and then runs it entirely on your Mac. Pick a model whose size fits comfortably in your Mac's memory.
Ollama is the quickest way to get a language model running on a Mac. It handles downloading models, picks sensible settings for Apple Silicon, and gives you a chat in the terminal plus a local API other apps can use.
What you need
- A Mac with an Apple Silicon chip (M1 or later). Intel Macs work but are much slower.
- At least 8 GB of memory, ideally 16 GB or more. See how much memory you need.
- A few gigabytes of free disk space per model.
Step 1: Install Ollama
Download the Mac app from ollama.com, unzip it and drag it to Applications, then open it once. Or, if you use Homebrew:
brew install ollama
Check that it works:
ollama --version
Step 2: Download and run a model
Open Terminal and run:
ollama run gemma4
The first time, Ollama downloads the model (several gigabytes), then drops you into a chat. Type a question and press Return. Type /bye to leave.
Good first choices by Mac memory:
| Your Mac's memory | Try | Command |
|---|---|---|
| 8 GB | Qwen 3.5 4B (3.4 GB download) | ollama run qwen3.5:4b |
| 16 GB | Qwen 3.5 9B (6.6 GB) or Gemma 4 | ollama run qwen3.5:9b |
| 32 to 36 GB | Gemma 4 26B (about 17 GB) or Qwen 3.5 27B | ollama run gemma4:26b |
| 64 GB or more | Qwen 3.6 35B (about 24 GB) and larger | ollama run qwen3.6:35b |
On Apple Silicon, some models also come with an -mlx tag, such as qwen3.6:35b-mlx. These run on Apple's MLX engine and are often faster on a Mac, so try both.
Model names change as new versions come out. The Ollama library lists what is current and shows each version's download size, which is close to how much memory it needs.
Step 3: Manage your models
ollama list # models you have downloaded
ollama ps # models currently loaded in memory
ollama rm gemma4 # delete a model
Step 4: Use it from other apps
While Ollama is running, it serves an API at http://localhost:11434. Many apps can use it directly, including chat front ends like Open WebUI and coding assistants in editors such as VS Code. A quick test from Terminal:
curl http://localhost:11434/api/generate -d '{"model": "gemma4", "prompt": "Why is the sky blue?", "stream": false}'
If it is slow
- The model is too big. If your Mac's memory pressure turns yellow or red in Activity Monitor, try a smaller model or a smaller quantization.
- Other apps are using memory. Close browsers with many tabs before running large models.
- The context is very long. Long documents use more memory and slow down replies.
Next, compare Ollama with the alternatives in Ollama vs LM Studio vs llama.cpp vs MLX.
Frequently asked questions
Is Ollama free?
Yes. Ollama is free, open-source software, and the models it downloads are free to use under their own licenses.
Does Ollama need an internet connection?
Only to download models. Once a model is on your Mac, chatting with it works offline and your prompts never leave the machine.
Where does Ollama store models on a Mac?
In the .ollama/models folder in your home directory. Run "ollama rm" followed by the model name to delete one and free the space.