Best local AI models for a 64 GB and 128 GB Mac

By Tried AI Tools · Published

Short answer

On a 64 GB Mac, Qwen 3.6 35B, Qwen 3.5 27B and Gemma 4 31B are the best all-rounders, and Gemma 4 26B is the fastest. On a 128 GB Mac, add gpt-oss 120B and Qwen 3.5 122B, two large mixture-of-experts models that run surprisingly fast, or Llama 3.3 70B. Most people get the best results keeping a fast 27B to 35B model for daily use and a larger one for hard problems.

New open models arrive constantly, so this list sticks to models available in Ollama's library as of October 2026, with the download sizes Ollama lists. Sizes are close to the memory each model needs before context.

For a 64 GB Mac

By default macOS lets the GPU use about 48 GB of a 64 GB Mac, so aim for models under about 35 GB.

Model Download Good at Command
Qwen 3.6 35B ~24 GB Coding, agents, general use; reads images ollama run qwen3.6:35b
Qwen 3.5 27B ~17 GB Strong all-rounder, many languages ollama run qwen3.5:27b
Gemma 4 31B ~20 GB Writing, general knowledge; reads images ollama run gemma4:31b
Gemma 4 26B (mixture of experts, 4B active) ~17 GB Very fast everyday chat ollama run gemma4:26b
gpt-oss 20B ~14 GB Reasoning and tool use, fast ollama run gpt-oss:20b
DeepSeek R1 32B (distilled) ~20 GB Step-by-step reasoning, math ollama run deepseek-r1:32b

Our pick: Qwen 3.6 35B as the main model, with Gemma 4 26B when you want speed.

For a 128 GB Mac

The GPU can use about 96 GB by default (more if you raise the limit). Everything in the 64 GB list runs here, with room to keep two loaded at once, plus:

Model Download Good at Command
gpt-oss 120B (mixture of experts, ~5B active) ~65 GB Reasoning, tool use; fast for its size ollama run gpt-oss:120b
Qwen 3.5 122B (mixture of experts) ~81 GB Long, complex tasks and planning ollama run qwen3.5:122b
Llama 3.3 70B ~43 GB Dependable general use, widely supported ollama run llama3.3:70b
DeepSeek R1 70B (distilled) ~43 GB Reasoning ollama run deepseek-r1:70b

Our pick: gpt-oss 120B for hard problems, alongside Qwen 3.6 35B for everyday work. Qwen 3.5 122B fits but leaves little room for long contexts unless you raise the GPU memory limit.

How to judge a new model yourself

  1. Check its download size on the Ollama library or Hugging Face.
  2. Add about 30 percent for context.
  3. If that fits inside the GPU share of your memory, it will run. If it is a mixture-of-experts model, expect it to be faster than its size suggests.

Speed expectations

Dense models slow down as they grow: on an M5 Max, a 70B model generates roughly 10 to 15 tokens per second, while a 27B model is several times faster. See M5 Max vs M5 Ultra for how to estimate speed on any chip.

Frequently asked questions

Should I run the biggest model that fits?

Usually not. A model that fills your memory leaves no room for long conversations or other apps, and dense models get slow as they grow. A 27B to 35B model is the daily sweet spot on both 64 GB and 128 GB Macs.

How often does this list change?

Often. New open models come out every few weeks. This guide shows the date it was last updated; the method for judging whether a new model fits stays the same.