Guide · Local AI
How to run LLMs locally on a Mac
Install Ollama, LM Studio, llama.cpp or MLX and pick a 4-bit model that fits. Apple Silicon’s GPU shares unified memory, but by default it may use only about two-thirds on Macs with up to 32 GB and three-quarters above: 10.7 GB on a 16 GB Mac, enough for 8B to 12B models, and 48 GB on a 64 GB Mac, enough for a 70B model.
By Tahir NazirUpdated 12 min read
On this page
How much of a Mac’s memory can the GPU use?
About two-thirds of unified memory on Macs with 32 GB or less, and three-quarters on larger ones. Apple Silicon has no separate graphics memory: the CPU and GPU share one pool, which is why a Mac can run models that won’t fit on a 24 GB graphics card. But macOS limits how much of it the GPU may use at once, a figure Metal calls the recommended working set size.
Apple gave two data points in its Metal compute tech talk: on a 32 GB M1 Pro or M1 Max the GPU can access 21 GB, and on a 64 GB M1 Max 48 GB. Our VRAM calculator applies that rule to every size (21.3 GB of 32 GB, 48 GB of 64 GB), and so does this guide. Your Mac’s exact figure can differ a little; llama.cpp prints it as recommendedMaxWorkingSetSize when it loads a model.
Unified memory · Q4_K_M · 8K context · GB
- 16 GB10.7 GB for the GPU10.7Qwen3 14B
- 24 GB16.0 GB for the GPU10.9Phi-4 14B
- 32 GB21.3 GB for the GPU21.1Gemma 4 31B
- 36 GB27.0 GB for the GPU22.7Qwen3 32B
- 48 GB36.0 GB for the GPU22.7Qwen3 32B
- 64 GB48.0 GB for the GPU46.9Llama 3.3 70B
- 96 GB72.0 GB for the GPU69.7Llama 4 Scout
- 128 GB96.0 GB for the GPU69.7Llama 4 Scout
- 192 GB144 GB for the GPU69.7Llama 4 Scout
- 256 GB192 GB for the GPU149Qwen3 235B-A22B
- 512 GB384 GB for the GPU149Qwen3 235B-A22B
Can you raise the limit?
Yes. MLX’s documentation (MLX is Apple’s own machine-learning framework) describes raising it with a system setting, choosing a value larger than the model in megabytes but smaller than the Mac’s memory. On a 32 GB Mac, this gives the GPU 24 GB and leaves 8 GB for macOS:
sudo sysctl iogpu.wired_limit_mb=24576What can you run on a 16, 24, 32, 64 or 128 GB Mac?
Here are the two largest models from our calculator that fit each memory size at Q4_K_M, the usual 4-bit choice. Each name links to that model’s full VRAM table.
| Memory | GPU limit | 8K context | 32K context | Limit raised (8 GB kept free), 8K |
|---|---|---|---|---|
| 16 GB | 10.7 GB | Qwen3 14B (10.7 GB), Mistral Nemo 12B (9.2 GB) | Gemma 3 12B (10.3 GB), Qwen3 8B (10.2 GB) | No gain over the default |
| 24 GB | 16.0 GB | Phi-4 14B (10.9 GB), Qwen3 14B (10.7 GB) | Qwen3 14B (14.7 GB), Mistral Nemo 12B (13.2 GB) | No gain over the default |
| 32 GB | 21.3 GB | Gemma 4 31B (21.1 GB), Qwen3 30B-A3B (19.9 GB) | Mistral Small 3.2 24B (20.5 GB), Gemma 3 27B (20.4 GB) | Qwen3 32B (22.7 GB), Gemma 4 31B (21.1 GB) |
| 36 GB | 27.0 GB | Qwen3 32B (22.7 GB), Gemma 4 31B (21.1 GB) | Gemma 4 31B (23.2 GB), Qwen3 30B-A3B (22.4 GB) | Qwen3 32B (22.7 GB), Gemma 4 31B (21.1 GB) |
| 48 GB | 36.0 GB | Qwen3 32B (22.7 GB), Gemma 4 31B (21.1 GB) | Qwen3 32B (29.3 GB), Gemma 4 31B (23.2 GB) | Qwen3 32B (22.7 GB), Gemma 4 31B (21.1 GB) |
| 64 GB | 48.0 GB | Llama 3.3 70B (46.9 GB), Qwen3 32B (22.7 GB) | Qwen3 32B (29.3 GB), Gemma 4 31B (23.2 GB) | Qwen2.5 72B (48.3 GB), Llama 3.3 70B (46.9 GB) |
| 96 GB | 72.0 GB | Llama 4 Scout (69.7 GB), Qwen2.5 72B (48.3 GB) | Llama 4 Scout (70.9 GB), Qwen2.5 72B (56.5 GB) | Llama 4 Scout (69.7 GB), Qwen2.5 72B (48.3 GB) |
| 128 GB | 96.0 GB | Llama 4 Scout (69.7 GB), Qwen2.5 72B (48.3 GB) | Llama 4 Scout (70.9 GB), Qwen2.5 72B (56.5 GB) | Llama 4 Scout (69.7 GB), Qwen2.5 72B (48.3 GB) |
Q4_K_M weights + FP16 KV cache + overhead, one conversation, same estimates as the VRAM calculator. Models whose maximum context is under 32K are left out of that column, and gpt-oss is covered below.
OpenAI’s gpt-oss doesn’t follow these quantisation levels: it is released in MXFP4, a 4-bit format, and its GGUF files are about the same size whichever label they carry. Going by the official files, gpt-oss 20B needs about 13.2 GB with a 32K context, so it runs on a 24 GB Mac but not a 16 GB one, and gpt-oss 120B about 65.3 GB with 8K, which fits the 72 GB limit of a 96 GB Mac.
The GPU limit is a ceiling, not a promise that the memory is free: macOS, your browser and other apps share the same pool. On a 16 GB Mac, Qwen3 14B (10.7 GB against a 10.7 GB limit) fits on paper but leaves nothing spare, so stick to 12B and below. MLX’s 4-bit files run slightly smaller than Q4_K_M: mlx-community’s Qwen3 14B 4-bit weighs 7.74 GB against 8.38 GB for Qwen’s Q4_K_M file, so the table is a little conservative for MLX.
Which tool should you use: Ollama, LM Studio, llama.cpp or MLX?
Start with Ollama if you’re happy in a terminal, or LM Studio if you want a desktop app. Both run models on the Mac’s GPU and give you a local API. Move to llama.cpp or MLX when you want to control every setting.
| Tool | Best for | Model formats | Needs |
|---|---|---|---|
| Ollama | The quickest start: one command per model, plus a local API | Ollama’s model library, or GGUF and Safetensors files you import | macOS 14 Sonoma or newer; Apple Silicon for GPU use |
| LM Studio | A desktop app for browsing, chatting and serving models | GGUF (llama.cpp) and MLX | macOS 14 or newer, Apple Silicon; 16 GB or more recommended |
| llama.cpp | Full control over any GGUF file on Hugging Face | GGUF | Homebrew or a release download; Metal is built in |
| MLX LM | Python scripting and fine-tuning with Apple’s MLX | MLX (thousands in mlx-community) | Python; macOS 15 for its large-model memory wiring |
Requirements from each project’s documentation, checked 2026-10-08.
Ollama
# Or download the app from ollama.com/download
curl -fsSL https://ollama.com/install.sh | sh
ollama run qwen3:14b-q4_K_M
# In another terminal: the PROCESSOR column should say 100% GPU
ollama psLM Studio
Download the app from lmstudio.ai and search for a model; on Apple Silicon it offers both GGUF and MLX builds. The lms command-line tool ships with the app (open the app once first):
lms get llama-3.1-8b@q4_k_m
lms chatllama.cpp
brew install llama.cpp
# Downloads Qwen's official Q4_K_M file; chat at http://localhost:8080
llama-server -hf Qwen/Qwen3-14B-GGUF:Q4_K_M -c 16384MLX LM
python3 -m venv ~/mlx && source ~/mlx/bin/activate
pip install mlx-lm
mlx_lm.chat --model mlx-community/Qwen3-14B-4bitHow fast does an LLM run on a Mac?
Generation speed is set mainly by memory bandwidth, because every new token reads all the active weights from memory once. Reading your prompt (prompt processing) depends more on GPU compute. llama.cpp’s community benchmark runs the same small model on every Apple chip:
| Chip (GPU cores) | Memory bandwidth | Generation | Share of the bandwidth ceiling | Prompt processing |
|---|---|---|---|---|
| M1 (8) | 68 GB/s | 14.2 | 79% | 118 |
| M4 (10) | 120 GB/s | 24.1 | 76% | 221 |
| M4 Pro (20) | 273 GB/s | 50.7 | 70% | 440 |
| M4 Max (40) | 546 GB/s | 83.1 | 58% | 886 |
| M3 Ultra (80) | 800 GB/s | 92.1 | 44% | 1,471 |
| M5 (10) | 154 GB/s | 31.9 | 78% | 723 |
| M5 Pro (20) | 307 GB/s | 66.3 | 82% | 1,621 |
| M5 Max (40) | 614 GB/s | 119.9 | 74% | 3,220 |
| M5 Ultra (80) | 1,228 GB/s | 179.1 | 55% | 4,945 |
From llama.cpp discussion 4167. The ceiling is bandwidth ÷ the model’s 3.79 GB file, the most a chip could generate if reading the weights were the only cost.
That gives a quick way to estimate any model: divide your chip’s bandwidth by the model’s size, then expect 44 to 82% of the result, the range these runs achieved (the Ultra chips sit at the low end). Qwen3 14B at Q4_K_M (9.0 GB of weights) has a ceiling of about 17 tokens a second on an M5, 34 on an M5 Pro and 68 on a 40-core M5 Max. Llama 3.3 70B (43 GB) tops out near 14 on an M5 Max and 28 on an M5 Ultra.
Two more patterns stand out. Mixture-of-experts models read only the experts they use: Qwen3 30B-A3B runs 8 of its 128 experts per token, so it generates far faster than a dense 30B model while needing the same memory. And the M5 generation, which Apple builds with Neural Accelerators in its GPU, reads prompts much faster at similar bandwidth: the M5 Pro processed 1,621 prompt tokens a second against 440 for the M4 Pro. Those runs used different llama.cpp builds, so part of the gap may be software, but it matters for long documents and coding agents.
Worked examples: a 16 GB MacBook Air, a 36 GB MacBook Pro and a 128 GB Mac Studio
A 16 GB MacBook Air (M5)
The GPU may use 10.7 GB. Llama 3.1 8B fits even with a 32K context (9.6 GB), and Gemma 3 12B needs 8.8 GB at 8K or 10.3 GB at 32K, close to the limit. gpt-oss 20B doesn’t fit: its official MXFP4 file alone is 11.28 GB. With 153 GB/s of bandwidth, Gemma 3 12B’s ceiling is about 21 tokens a second; at the 78% the M5 reached in the benchmark, that’s roughly 16.
A 36 GB MacBook Pro (M5 Max, 32-core GPU)
The limit is 27 GB, enough for Qwen3 32B at Q4_K_M (22.7 GB with an 8K context). At 32K it needs 29 GB, but an 8-bit KV cache brings that to 25.2 GB, which fits. At 460 GB/s the bandwidth ceiling for Qwen3 32B is about 23 tokens a second; Qwen3 30B-A3B (22 GB at 32K) will feel much quicker.
A 128 GB Mac Studio or MacBook Pro (M5 Max, 40-core GPU)
The GPU may use 96 GB. That runs gpt-oss 120B (about 65 GB with its official MXFP4 file), Llama 3.3 70B at Q8_0 (80 GB), or Llama 3.3 70B at Q4_K_M with its full 128K context (88 GB). Qwen3 235B-A22B at Q4_K_M (149 GB) needs a 256 GB Mac Studio. Speed is the limit here, not memory: Llama 3.3 70B at Q4_K_M has a ceiling of about 14 tokens a second at 614 GB/s.
Free toolLLM VRAM calculatorChoose a model, quantisation and context length to see which Mac memory sizes can run it, with and without raising the GPU limit.Context length is where memory surprises happen. Count a typical prompt with the token counter, and use the context window checker to see whether a long document fits a model at all. For graphics cards and PCs, see how much VRAM you need to run an LLM locally.
FAQ
Questions people ask
Can a MacBook Air run an LLM locally?
Yes. Any Apple Silicon MacBook Air can run a local model on its GPU. With 16 GB, the GPU may use about 11 GB by default, enough for 8B to 12B models at 4-bit such as Llama 3.1 8B or Gemma 3 12B. A 24 GB Air raises that to 16 GB, enough for 14B models with a long context.
How much RAM does a Mac need to run a 70B model?
64 GB. Llama 3.3 70B at Q4_K_M with an 8K context needs about 47 GB, just under the 48 GB the GPU may use on a 64 GB Mac by default. For longer contexts or 8-bit weights, 96 or 128 GB gives real headroom. On 48 GB or less, a 70B model doesn’t fit at 4-bit.
Is 8 GB of RAM enough to run an LLM on a Mac?
Only for small models. The GPU may use about 5.3 GB of 8 GB, which fits a 3B model such as Llama 3.2 3B at Q4_K_M (3.7 GB with an 8K context). LM Studio’s requirements say 8 GB Macs can work if you stick to smaller models and modest context sizes.
Is it safe to raise iogpu.wired_limit_mb?
Within limits. MLX’s README suggests it for large models, and a value set with sysctl lasts only until you restart, so a mistake is easy to undo. Set it too high and macOS runs short of memory, so keep several gigabytes free; our calculator assumes 8 GB. It only helps on Macs with 32 GB or more.
Should I use MLX or GGUF models on a Mac?
Both run on the GPU, so try both with your model. MLX is Apple’s own framework and its 4-bit files are slightly smaller than Q4_K_M; GGUF, used by llama.cpp and Ollama, offers more quantisation levels and works across Mac, Windows and Linux. LM Studio runs both, which makes a side-by-side test easy.
Does Ollama use the Mac’s GPU?
Yes, on Apple Silicon. Ollama’s macOS documentation lists Apple M-series chips with CPU and GPU support, and Intel Macs as CPU only. Run ollama ps while a model is loaded: the PROCESSOR column shows 100% GPU when it fits, or a CPU/GPU split when it doesn’t.
Try it
Tools from this guide
Keep reading