Best local AI models for a 64 GB and 128 GB Mac
Short answer
On a 64 GB Mac, Qwen 3.6 35B, Qwen 3.5 27B and Gemma 4 31B are the best all-rounders, and Gemma 4 26B is the fastest. On a 128 GB Mac, add gpt-oss 120B and Qwen 3.5 122B, two large mixture-of-experts models that run surprisingly fast, or Llama 3.3 70B. Most people get the best results keeping a fast 27B to 35B model for daily use and a larger one for hard problems.
New open models arrive constantly, so this list sticks to models available in Ollama's library as of October 2026, with the download sizes Ollama lists. Sizes are close to the memory each model needs before context.
For a 64 GB Mac
By default macOS lets the GPU use about 48 GB of a 64 GB Mac, so aim for models under about 35 GB.
| Model | Download | Good at | Command |
|---|---|---|---|
| Qwen 3.6 35B | ~24 GB | Coding, agents, general use; reads images | ollama run qwen3.6:35b |
| Qwen 3.5 27B | ~17 GB | Strong all-rounder, many languages | ollama run qwen3.5:27b |
| Gemma 4 31B | ~20 GB | Writing, general knowledge; reads images | ollama run gemma4:31b |
| Gemma 4 26B (mixture of experts, 4B active) | ~17 GB | Very fast everyday chat | ollama run gemma4:26b |
| gpt-oss 20B | ~14 GB | Reasoning and tool use, fast | ollama run gpt-oss:20b |
| DeepSeek R1 32B (distilled) | ~20 GB | Step-by-step reasoning, math | ollama run deepseek-r1:32b |
Our pick: Qwen 3.6 35B as the main model, with Gemma 4 26B when you want speed.
For a 128 GB Mac
The GPU can use about 96 GB by default (more if you raise the limit). Everything in the 64 GB list runs here, with room to keep two loaded at once, plus:
| Model | Download | Good at | Command |
|---|---|---|---|
| gpt-oss 120B (mixture of experts, ~5B active) | ~65 GB | Reasoning, tool use; fast for its size | ollama run gpt-oss:120b |
| Qwen 3.5 122B (mixture of experts) | ~81 GB | Long, complex tasks and planning | ollama run qwen3.5:122b |
| Llama 3.3 70B | ~43 GB | Dependable general use, widely supported | ollama run llama3.3:70b |
| DeepSeek R1 70B (distilled) | ~43 GB | Reasoning | ollama run deepseek-r1:70b |
Our pick: gpt-oss 120B for hard problems, alongside Qwen 3.6 35B for everyday work. Qwen 3.5 122B fits but leaves little room for long contexts unless you raise the GPU memory limit.
How to judge a new model yourself
- Check its download size on the Ollama library or Hugging Face.
- Add about 30 percent for context.
- If that fits inside the GPU share of your memory, it will run. If it is a mixture-of-experts model, expect it to be faster than its size suggests.
Speed expectations
Dense models slow down as they grow: on an M5 Max, a 70B model generates roughly 10 to 15 tokens per second, while a 27B model is several times faster. See M5 Max vs M5 Ultra for how to estimate speed on any chip.
Frequently asked questions
Should I run the biggest model that fits?
Usually not. A model that fills your memory leaves no room for long conversations or other apps, and dense models get slow as they grow. A 27B to 35B model is the daily sweet spot on both 64 GB and 128 GB Macs.
How often does this list change?
Often. New open models come out every few weeks. This guide shows the date it was last updated; the method for judging whether a new model fits stays the same.