How to run your first local LLM on a Mac with Ollama

By Tried AI Tools · Published

Short answer

Install Ollama from ollama.com or with Homebrew, open Terminal, and run "ollama run gemma4". Ollama downloads the model the first time and then runs it entirely on your Mac. Pick a model whose size fits comfortably in your Mac's memory.

Ollama is the quickest way to get a language model running on a Mac. It handles downloading models, picks sensible settings for Apple Silicon, and gives you a chat in the terminal plus a local API other apps can use.

What you need

Step 1: Install Ollama

Download the Mac app from ollama.com, unzip it and drag it to Applications, then open it once. Or, if you use Homebrew:

brew install ollama

Check that it works:

ollama --version

Step 2: Download and run a model

Open Terminal and run:

ollama run gemma4

The first time, Ollama downloads the model (several gigabytes), then drops you into a chat. Type a question and press Return. Type /bye to leave.

Good first choices by Mac memory:

Your Mac's memory Try Command
8 GB Qwen 3.5 4B (3.4 GB download) ollama run qwen3.5:4b
16 GB Qwen 3.5 9B (6.6 GB) or Gemma 4 ollama run qwen3.5:9b
32 to 36 GB Gemma 4 26B (about 17 GB) or Qwen 3.5 27B ollama run gemma4:26b
64 GB or more Qwen 3.6 35B (about 24 GB) and larger ollama run qwen3.6:35b

On Apple Silicon, some models also come with an -mlx tag, such as qwen3.6:35b-mlx. These run on Apple's MLX engine and are often faster on a Mac, so try both.

Model names change as new versions come out. The Ollama library lists what is current and shows each version's download size, which is close to how much memory it needs.

Step 3: Manage your models

ollama list          # models you have downloaded
ollama ps            # models currently loaded in memory
ollama rm gemma4     # delete a model

Step 4: Use it from other apps

While Ollama is running, it serves an API at http://localhost:11434. Many apps can use it directly, including chat front ends like Open WebUI and coding assistants in editors such as VS Code. A quick test from Terminal:

curl http://localhost:11434/api/generate -d '{"model": "gemma4", "prompt": "Why is the sky blue?", "stream": false}'

If it is slow

Next, compare Ollama with the alternatives in Ollama vs LM Studio vs llama.cpp vs MLX.

Frequently asked questions

Is Ollama free?

Yes. Ollama is free, open-source software, and the models it downloads are free to use under their own licenses.

Does Ollama need an internet connection?

Only to download models. Once a model is on your Mac, chatting with it works offline and your prompts never leave the machine.

Where does Ollama store models on a Mac?

In the .ollama/models folder in your home directory. Run "ollama rm" followed by the model name to delete one and free the space.