How much memory do you need to run an LLM on a Mac?
Short answer
Multiply the model's parameter count in billions by about 0.6 to get gigabytes at 4-bit quantization, then add 20 to 30 percent for context and keep a few gigabytes free for macOS. An 8B model fits comfortably in 16 GB, a 32B model wants 32 to 48 GB, and a 70B model wants 64 GB or more.
On an Apple Silicon Mac, the CPU and GPU share one pool of unified memory. That is what makes Macs good at running AI locally: a model has to fit in memory to run at a usable speed, and Macs offer far more GPU-accessible memory per dollar than most PC graphics cards. It also means memory is the first thing to get right.
The formula
A model's memory footprint is mostly its weights. Each parameter takes a certain number of bits depending on how the model is quantized (compressed):
| Format | Bits per parameter (approx.) | GB per billion parameters |
|---|---|---|
| FP16 / BF16 (full size) | 16 | 2.0 |
| 8-bit (Q8_0, MLX 8-bit) | 8.5 | 1.06 |
| 4-bit (Q4_K_M, MLX 4-bit) | 4.5 to 5 | 0.6 |
So the weights take roughly parameters (in billions) × GB per billion. On top of that you need:
- Context (the KV cache). The longer the conversation or document, the more memory it takes. Budget 20 to 30 percent extra for normal chat, more for very long contexts.
- Headroom for macOS and your apps. By default macOS only lets the GPU use part of unified memory, roughly 65 to 75 percent depending on the total. Plan on the model plus context fitting inside that share.
Worked examples at 4-bit
| Model size | Weights | With context | Comfortable Mac memory |
|---|---|---|---|
| 3B | ~2 GB | ~2.5 GB | 8 GB |
| 8B | ~5 GB | ~6.5 GB | 16 GB |
| 14B | ~8.5 GB | ~11 GB | 16 to 24 GB |
| 32B | ~19 GB | ~25 GB | 36 to 48 GB |
| 70B | ~42 GB | ~52 GB | 64 to 96 GB |
| 120B+ (mixture of experts) | ~65 GB+ | ~80 GB+ | 128 GB |
Mixture-of-experts models only activate part of their weights for each word, so they run faster than their size suggests, but all of the weights still have to fit in memory.
Can you raise the GPU memory limit?
Yes. On recent versions of macOS you can raise the cap with a terminal command, for example to let the GPU use 56 GB on a 64 GB Mac:
sudo sysctl iogpu.wired_limit_mb=57344
The setting resets when you restart. Leave at least 8 GB for macOS, or the whole system will slow down as it swaps to disk.
Why speed depends on memory bandwidth
Generating each word means reading the whole active model from memory once. A rough ceiling on speed is therefore memory bandwidth ÷ model size. A 5 GB model on a chip with 120 GB/s of bandwidth tops out around 24 tokens (word pieces) per second in theory, and real results land somewhat lower. Higher chip tiers (Pro, Max, Ultra) have more bandwidth, which is why they answer faster with the same model.
How to choose
- Decide the largest model you want to run every day.
- Work out its 4-bit size with the formula above and add 30 percent.
- Make sure that number fits within about 70 percent of the Mac's memory.
If you are choosing between more memory and a faster chip at the same price, more memory usually wins, because a model that does not fit will not run at all. See which Mac to buy for local AI for how this plays out across the lineup.
Frequently asked questions
Can a Mac with 8 GB of memory run an LLM?
Yes, but only small models of roughly 1 to 4 billion parameters at 4-bit quantization. They are useful for quick drafting and summaries but noticeably weaker than larger models.
Does a Mac use RAM or GPU memory for AI models?
Apple Silicon Macs use unified memory, so the same pool serves the CPU and the GPU. By default macOS lets the GPU use roughly two thirds to three quarters of it, which is why a 64 GB Mac cannot quite load a 64 GB model.
Is more memory or a faster chip more important?
Memory decides which models you can load at all. Memory bandwidth, which rises with the chip tier, decides how fast they answer. Buy enough memory first.