Skip to content
AI Dev Toolkit.
Esc
  • AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.Tool
  • LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.Tool
  • AI Model ComparisonCompare prices, context windows and features across models.Tool
  • AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.Tool
  • Context Window CheckerSee whether your text fits each model's context window.Tool
  • Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.Tool
  • GPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.Tool
  • Claude Code Error DatabaseExact Claude Code error messages with tested fixes.Tool

Guide · Local AI

How to run LLMs locally on a Mac

Install Ollama, LM Studio, llama.cpp or MLX and pick a 4-bit model that fits. Apple Silicon’s GPU shares unified memory, but by default it may use only about two-thirds on Macs with up to 32 GB and three-quarters above: 10.7 GB on a 16 GB Mac, enough for 8B to 12B models, and 48 GB on a 64 GB Mac, enough for a 70B model.

By Tahir NazirUpdated 12 min read

On this page
  1. How much of a Mac’s memory can the GPU use?
  2. What can you run on a 16, 24, 32, 64 or 128 GB Mac?
  3. Which tool should you use: Ollama, LM Studio, llama.cpp or MLX?
  4. How fast does an LLM run on a Mac?
  5. Worked examples: a 16 GB MacBook Air, a 36 GB MacBook Pro and a 128 GB Mac Studio
  6. Questions people ask

How much of a Mac’s memory can the GPU use?

About two-thirds of unified memory on Macs with 32 GB or less, and three-quarters on larger ones. Apple Silicon has no separate graphics memory: the CPU and GPU share one pool, which is why a Mac can run models that won’t fit on a 24 GB graphics card. But macOS limits how much of it the GPU may use at once, a figure Metal calls the recommended working set size.

Apple gave two data points in its Metal compute tech talk: on a 32 GB M1 Pro or M1 Max the GPU can access 21 GB, and on a 64 GB M1 Max 48 GB. Our VRAM calculator applies that rule to every size (21.3 GB of 32 GB, 48 GB of 64 GB), and so does this guide. Your Mac’s exact figure can differ a little; llama.cpp prints it as recommendedMaxWorkingSetSize when it loads a model.

Each bar is the Mac’s whole memory. The lime part is what the GPU may use by default; the violet bar is the largest model in our calculator that fits inside it.

Can you raise the limit?

Yes. MLX’s documentation (MLX is Apple’s own machine-learning framework) describes raising it with a system setting, choosing a value larger than the model in megabytes but smaller than the Mac’s memory. On a 32 GB Mac, this gives the GPU 24 GB and leaves 8 GB for macOS:

Raise the GPU memory limit (until the next restart)
sudo sysctl iogpu.wired_limit_mb=24576

What can you run on a 16, 24, 32, 64 or 128 GB Mac?

Here are the two largest models from our calculator that fit each memory size at Q4_K_M, the usual 4-bit choice. Each name links to that model’s full VRAM table.

Largest models that fit, by unified memory
MemoryGPU limit8K context32K contextLimit raised (8 GB kept free), 8K
16 GB10.7 GBQwen3 14B (10.7 GB), Mistral Nemo 12B (9.2 GB)Gemma 3 12B (10.3 GB), Qwen3 8B (10.2 GB)No gain over the default
24 GB16.0 GBPhi-4 14B (10.9 GB), Qwen3 14B (10.7 GB)Qwen3 14B (14.7 GB), Mistral Nemo 12B (13.2 GB)No gain over the default
32 GB21.3 GBGemma 4 31B (21.1 GB), Qwen3 30B-A3B (19.9 GB)Mistral Small 3.2 24B (20.5 GB), Gemma 3 27B (20.4 GB)Qwen3 32B (22.7 GB), Gemma 4 31B (21.1 GB)
36 GB27.0 GBQwen3 32B (22.7 GB), Gemma 4 31B (21.1 GB)Gemma 4 31B (23.2 GB), Qwen3 30B-A3B (22.4 GB)Qwen3 32B (22.7 GB), Gemma 4 31B (21.1 GB)
48 GB36.0 GBQwen3 32B (22.7 GB), Gemma 4 31B (21.1 GB)Qwen3 32B (29.3 GB), Gemma 4 31B (23.2 GB)Qwen3 32B (22.7 GB), Gemma 4 31B (21.1 GB)
64 GB48.0 GBLlama 3.3 70B (46.9 GB), Qwen3 32B (22.7 GB)Qwen3 32B (29.3 GB), Gemma 4 31B (23.2 GB)Qwen2.5 72B (48.3 GB), Llama 3.3 70B (46.9 GB)
96 GB72.0 GBLlama 4 Scout (69.7 GB), Qwen2.5 72B (48.3 GB)Llama 4 Scout (70.9 GB), Qwen2.5 72B (56.5 GB)Llama 4 Scout (69.7 GB), Qwen2.5 72B (48.3 GB)
128 GB96.0 GBLlama 4 Scout (69.7 GB), Qwen2.5 72B (48.3 GB)Llama 4 Scout (70.9 GB), Qwen2.5 72B (56.5 GB)Llama 4 Scout (69.7 GB), Qwen2.5 72B (48.3 GB)

Q4_K_M weights + FP16 KV cache + overhead, one conversation, same estimates as the VRAM calculator. Models whose maximum context is under 32K are left out of that column, and gpt-oss is covered below.

OpenAI’s gpt-oss doesn’t follow these quantisation levels: it is released in MXFP4, a 4-bit format, and its GGUF files are about the same size whichever label they carry. Going by the official files, gpt-oss 20B needs about 13.2 GB with a 32K context, so it runs on a 24 GB Mac but not a 16 GB one, and gpt-oss 120B about 65.3 GB with 8K, which fits the 72 GB limit of a 96 GB Mac.

The GPU limit is a ceiling, not a promise that the memory is free: macOS, your browser and other apps share the same pool. On a 16 GB Mac, Qwen3 14B (10.7 GB against a 10.7 GB limit) fits on paper but leaves nothing spare, so stick to 12B and below. MLX’s 4-bit files run slightly smaller than Q4_K_M: mlx-community’s Qwen3 14B 4-bit weighs 7.74 GB against 8.38 GB for Qwen’s Q4_K_M file, so the table is a little conservative for MLX.

Which tool should you use: Ollama, LM Studio, llama.cpp or MLX?

Start with Ollama if you’re happy in a terminal, or LM Studio if you want a desktop app. Both run models on the Mac’s GPU and give you a local API. Move to llama.cpp or MLX when you want to control every setting.

Four ways to run a model on Apple Silicon
ToolBest forModel formatsNeeds
OllamaThe quickest start: one command per model, plus a local APIOllama’s model library, or GGUF and Safetensors files you importmacOS 14 Sonoma or newer; Apple Silicon for GPU use
LM StudioA desktop app for browsing, chatting and serving modelsGGUF (llama.cpp) and MLXmacOS 14 or newer, Apple Silicon; 16 GB or more recommended
llama.cppFull control over any GGUF file on Hugging FaceGGUFHomebrew or a release download; Metal is built in
MLX LMPython scripting and fine-tuning with Apple’s MLXMLX (thousands in mlx-community)Python; macOS 15 for its large-model memory wiring

Requirements from each project’s documentation, checked 2026-10-08.

Ollama

Install Ollama, run Qwen3 14B at Q4_K_M, check where it runs
# Or download the app from ollama.com/download
curl -fsSL https://ollama.com/install.sh | sh

ollama run qwen3:14b-q4_K_M

# In another terminal: the PROCESSOR column should say 100% GPU
ollama ps

LM Studio

Download the app from lmstudio.ai and search for a model; on Apple Silicon it offers both GGUF and MLX builds. The lms command-line tool ships with the app (open the app once first):

Download a model and chat from the terminal
lms get llama-3.1-8b@q4_k_m
lms chat

llama.cpp

Install with Homebrew and serve a model from Hugging Face
brew install llama.cpp

# Downloads Qwen's official Q4_K_M file; chat at http://localhost:8080
llama-server -hf Qwen/Qwen3-14B-GGUF:Q4_K_M -c 16384

MLX LM

Install MLX LM in a virtual environment and chat
python3 -m venv ~/mlx && source ~/mlx/bin/activate
pip install mlx-lm

mlx_lm.chat --model mlx-community/Qwen3-14B-4bit

How fast does an LLM run on a Mac?

Generation speed is set mainly by memory bandwidth, because every new token reads all the active weights from memory once. Reading your prompt (prompt processing) depends more on GPU compute. llama.cpp’s community benchmark runs the same small model on every Apple chip:

Llama 2 7B at Q4_0 in llama.cpp, tokens per second
Chip (GPU cores)Memory bandwidthGenerationShare of the bandwidth ceilingPrompt processing
M1 (8)68 GB/s14.279%118
M4 (10)120 GB/s24.176%221
M4 Pro (20)273 GB/s50.770%440
M4 Max (40)546 GB/s83.158%886
M3 Ultra (80)800 GB/s92.144%1,471
M5 (10)154 GB/s31.978%723
M5 Pro (20)307 GB/s66.382%1,621
M5 Max (40)614 GB/s119.974%3,220
M5 Ultra (80)1,228 GB/s179.155%4,945

From llama.cpp discussion 4167. The ceiling is bandwidth ÷ the model’s 3.79 GB file, the most a chip could generate if reading the weights were the only cost.

That gives a quick way to estimate any model: divide your chip’s bandwidth by the model’s size, then expect 44 to 82% of the result, the range these runs achieved (the Ultra chips sit at the low end). Qwen3 14B at Q4_K_M (9.0 GB of weights) has a ceiling of about 17 tokens a second on an M5, 34 on an M5 Pro and 68 on a 40-core M5 Max. Llama 3.3 70B (43 GB) tops out near 14 on an M5 Max and 28 on an M5 Ultra.

Two more patterns stand out. Mixture-of-experts models read only the experts they use: Qwen3 30B-A3B runs 8 of its 128 experts per token, so it generates far faster than a dense 30B model while needing the same memory. And the M5 generation, which Apple builds with Neural Accelerators in its GPU, reads prompts much faster at similar bandwidth: the M5 Pro processed 1,621 prompt tokens a second against 440 for the M4 Pro. Those runs used different llama.cpp builds, so part of the gap may be software, but it matters for long documents and coding agents.

Worked examples: a 16 GB MacBook Air, a 36 GB MacBook Pro and a 128 GB Mac Studio

A 16 GB MacBook Air (M5)

The GPU may use 10.7 GB. Llama 3.1 8B fits even with a 32K context (9.6 GB), and Gemma 3 12B needs 8.8 GB at 8K or 10.3 GB at 32K, close to the limit. gpt-oss 20B doesn’t fit: its official MXFP4 file alone is 11.28 GB. With 153 GB/s of bandwidth, Gemma 3 12B’s ceiling is about 21 tokens a second; at the 78% the M5 reached in the benchmark, that’s roughly 16.

A 36 GB MacBook Pro (M5 Max, 32-core GPU)

The limit is 27 GB, enough for Qwen3 32B at Q4_K_M (22.7 GB with an 8K context). At 32K it needs 29 GB, but an 8-bit KV cache brings that to 25.2 GB, which fits. At 460 GB/s the bandwidth ceiling for Qwen3 32B is about 23 tokens a second; Qwen3 30B-A3B (22 GB at 32K) will feel much quicker.

A 128 GB Mac Studio or MacBook Pro (M5 Max, 40-core GPU)

The GPU may use 96 GB. That runs gpt-oss 120B (about 65 GB with its official MXFP4 file), Llama 3.3 70B at Q8_0 (80 GB), or Llama 3.3 70B at Q4_K_M with its full 128K context (88 GB). Qwen3 235B-A22B at Q4_K_M (149 GB) needs a 256 GB Mac Studio. Speed is the limit here, not memory: Llama 3.3 70B at Q4_K_M has a ceiling of about 14 tokens a second at 614 GB/s.

Free toolLLM VRAM calculatorChoose a model, quantisation and context length to see which Mac memory sizes can run it, with and without raising the GPU limit.

Context length is where memory surprises happen. Count a typical prompt with the token counter, and use the context window checker to see whether a long document fits a model at all. For graphics cards and PCs, see how much VRAM you need to run an LLM locally.

FAQ

Questions people ask

Can a MacBook Air run an LLM locally?

Yes. Any Apple Silicon MacBook Air can run a local model on its GPU. With 16 GB, the GPU may use about 11 GB by default, enough for 8B to 12B models at 4-bit such as Llama 3.1 8B or Gemma 3 12B. A 24 GB Air raises that to 16 GB, enough for 14B models with a long context.

How much RAM does a Mac need to run a 70B model?

64 GB. Llama 3.3 70B at Q4_K_M with an 8K context needs about 47 GB, just under the 48 GB the GPU may use on a 64 GB Mac by default. For longer contexts or 8-bit weights, 96 or 128 GB gives real headroom. On 48 GB or less, a 70B model doesn’t fit at 4-bit.

Is 8 GB of RAM enough to run an LLM on a Mac?

Only for small models. The GPU may use about 5.3 GB of 8 GB, which fits a 3B model such as Llama 3.2 3B at Q4_K_M (3.7 GB with an 8K context). LM Studio’s requirements say 8 GB Macs can work if you stick to smaller models and modest context sizes.

Is it safe to raise iogpu.wired_limit_mb?

Within limits. MLX’s README suggests it for large models, and a value set with sysctl lasts only until you restart, so a mistake is easy to undo. Set it too high and macOS runs short of memory, so keep several gigabytes free; our calculator assumes 8 GB. It only helps on Macs with 32 GB or more.

Should I use MLX or GGUF models on a Mac?

Both run on the GPU, so try both with your model. MLX is Apple’s own framework and its 4-bit files are slightly smaller than Q4_K_M; GGUF, used by llama.cpp and Ollama, offers more quantisation levels and works across Mac, Windows and Linux. LM Studio runs both, which makes a side-by-side test easy.

Does Ollama use the Mac’s GPU?

Yes, on Apple Silicon. Ollama’s macOS documentation lists Apple M-series chips with CPU and GPU support, and Intel Macs as CPU only. Run ollama ps while a model is loaded: the PROCESSOR column shows 100% GPU when it fits, or a CPU/GPU split when it doesn’t.

Try it

Tools from this guide

Keep reading