Guide · Local AI
Ollama vs LM Studio vs llama.cpp: which local LLM tool?
Pick LM Studio if you want a desktop app, Ollama if you want one command and a local API, and llama.cpp if you want every setting. They are layers more than rivals: llama.cpp is the engine, Ollama and LM Studio both run it, and on Apple Silicon both can also run Apple’s MLX.
By Tahir NazirUpdated 10 min read
On this page
- Ollama vs LM Studio vs llama.cpp: what’s the difference?
- Does Ollama use llama.cpp?
- How do you install each one?
- GGUF or MLX: which model formats does each support?
- Which has the best local API? Endpoints and default ports
- GPU support on Windows, Linux and Mac
- What context length does each use by default?
- Is each one open source, and does it keep your data local?
- Which should you choose?
- Questions people ask
Ollama vs LM Studio vs llama.cpp: what’s the difference?
They differ in how you use them, not in what fundamentally runs the model. LM Studio is a desktop app for browsing, chatting with and serving models. Ollama is a command-line tool and background service with a desktop app and its own model library. llama.cpp is the open-source C/C++ engine itself, with a server, a command line and now its own small app.
| Ollama | LM Studio | llama.cpp | |
|---|---|---|---|
| What it is | CLI, background server and desktop app | Desktop app with lms CLI and the llmster daemon | Engine, llama-server, llama-cli and the Llama app |
| Install | Script, Mac/Windows installer or Docker image | Installer (Linux AppImage) or llmster script | Script, Homebrew, Winget, conda, Nix, Docker or binaries |
| Model formats | Its library, or GGUF and Safetensors you import | GGUF; MLX on Apple Silicon | GGUF |
| Find models | ollama.com/library | Hugging Face search in the app | Any GGUF repo on Hugging Face (-hf) |
| Engines | llama.cpp; MLX on Apple Silicon | llama.cpp; MLX on Apple Silicon | llama.cpp (on ggml) |
| Local API | 127.0.0.1:11434 | 127.0.0.1:1234 | 127.0.0.1:9931 |
| Default context | 4K, 32K or 256K by VRAM | Set when you load a model | The model’s own, fitted to memory |
| Licence | MIT | Proprietary, free at home and work | MIT |
From each project’s documentation and repository, 2026-10-11.
Does Ollama use llama.cpp?
Yes. Ollama’s README lists llama.cpp as its supported backend, and LM Studio’s docs say it runs models on Mac, Windows and Linux with llama.cpp. On Apple Silicon both add a second engine, Apple’s MLX: LM Studio runs MLX models there, and since Ollama 0.40.0 “model architectures supported by the MLX runtime will automatically run on MLX”.
What you install · default API port
Ollama
:11434CLI, desktop app and local API
- llama.cpp
- MLXApple Silicon, default since v0.40
LM Studio
:1234Desktop app, lms CLI, llmster daemon
- llama.cpp
- MLXApple Silicon Macs
llama.cpp’s own tools
:9931llama-server, llama-cli, Llama app
- llama.cpp
MLX LM
:8080Python package, mlx_lm.server
- MLXApple silicon
Engines that run the model
llama.cpp
ggml-org · built on the ggml library
FormatGGUF files
MLX
Apple · machine-learning framework
FormatMLX weights (safetensors)
llama.cpp backends shown are a selection from its README; MLX is shown for Apple silicon, where these tools use it.
So the same GGUF file runs on the same engine in all three. What differs is the version of llama.cpp each one bundles, the defaults it picks (context length, GPU offload, sampling) and the extras around it. That’s why we don’t quote speed rankings: a result depends on those versions and settings, and we found no like-for-like comparison published by the projects themselves. Test your own model on your own machine.
How do you install each one?
Each project offers a one-line installer as well as the usual packages. These are the commands from their own READMEs and docs:
# Ollama (macOS or Linux; Windows: irm https://ollama.com/install.ps1 | iex)
curl -fsSL https://ollama.com/install.sh | sh
ollama run gemma4
# LM Studio: the desktop app from lmstudio.ai/download, or the headless daemon
curl -fsSL https://lmstudio.ai/install.sh | bash
lms daemon up
lms get llama-3.1-8b@q4_k_m
lms server start
# llama.cpp (or: brew install llama.cpp, winget install llama.cpp)
curl -LsSf https://llama.app/install.sh | sh
llama serve -hf ggml-org/Qwen3.5-0.8B-GGUFOllama needs macOS 14 or newer and uses the GPU only on Apple M-series Macs; Intel Macs run on the CPU. LM Studio needs macOS 14 on Apple Silicon (Intel Macs aren’t supported), AVX2 on x64 Windows, and Ubuntu 20.04 or newer on Linux, and recommends 16 GB of RAM. llama.cpp’s team, with Hugging Face, also ships Llama, a free, open-source app (menu bar on Mac, system tray on Windows) that serves the same API on port 9931.
GGUF or MLX: which model formats does each support?
GGUF is llama.cpp’s single-file format, and all three read it. MLX models are safetensors weights prepared for Apple’s MLX, used only on Apple Silicon here.
- Ollama pulls from its own library (
ollama pull gemma4) and imports GGUF or Safetensors through a Modelfile withFROM /path/to/file.ggufandollama create. It doesn’t quantise GGUF on import: its docs send you to llama.cpp’sllama-quantize. - LM Studio searches Hugging Face from its Discover tab and offers GGUF and, on Apple Silicon, MLX downloads (
lms get --mlxor--gguf). Its docs suggest a 4-bit option or higher if your machine can run it. - llama.cpp downloads any GGUF repository with
-hf user/model:quant, choosing Q4_K_M when you don’t name a quantisation. The Llama app keeps models in the Hugging Face cache, shared with llama.cpp. - MLX LM runs MLX models from Hugging Face, thousands of them in the mlx-community organisation, and can quantise and fine-tune them.
Whatever the format, the file has to fit in memory with room for the context. The VRAM calculator shows how much a model needs at each quantisation, and how much VRAM you need to run an LLM locally explains the sums.
Which has the best local API? Endpoints and default ports
All of them speak OpenAI’s API, so any OpenAI client can switch between them by changing the base URL. Each listens only on 127.0.0.1 by default:
| Server | Default address | OpenAI-compatible | Also offers |
|---|---|---|---|
| Ollama | http://localhost:11434 | /v1/chat/completions, /v1/responses, /v1/completions, /v1/embeddings, /v1/models | Native /api/* API, Anthropic-compatible /v1/messages |
| LM Studio | http://localhost:1234 | /v1/chat/completions, /v1/responses, /v1/completions, /v1/embeddings, /v1/models | Its REST API, Python and TypeScript SDKs, Anthropic-compatible /v1/messages |
| llama-server | http://localhost:9931 | /v1/chat/completions, /v1/responses, /v1/completions, /v1/embeddings, /v1/models | Anthropic-compatible /v1/messages, reranking, web UI |
| mlx_lm.server | http://localhost:8080 | /v1/chat/completions, /v1/models | MLX LM warns it isn’t meant for production |
LM Studio reuses the last port you chose if you don’t pass `--port`. llama-server’s README sets 8080 in its Docker examples, so check which port your build uses.
from openai import OpenAI
# Pick the server you run. Each one needs a model already downloaded.
client = OpenAI(base_url="http://localhost:11434/v1/", api_key="ollama") # Ollama: key required, ignored
# client = OpenAI(base_url="http://localhost:1234/v1", api_key="lm-studio") # LM Studio
# client = OpenAI(base_url="http://localhost:9931/v1", api_key="none") # llama-server
reply = client.chat.completions.create(
model="gemma4", # Ollama's name; use LM Studio's model identifier there
messages=[{"role": "user", "content": "Say hello in five words."}],
)
print(reply.choices[0].message.content)Ollama’s OpenAI layer can’t set the context size, so for a longer context you create a model with PARAMETER num_ctx in a Modelfile. If you move between local servers and hosted APIs, our guide to OpenAI, Anthropic and Gemini message formats covers the request shapes these endpoints copy.
GPU support on Windows, Linux and Mac
All three use NVIDIA and Apple GPUs. AMD and Intel support depends on the backend:
| Hardware | Ollama | LM Studio | llama.cpp |
|---|---|---|---|
| NVIDIA | Compute capability 5.0 or newer, driver 550+ (570+ for 5.0 to 6.2) | 4 GB or more of dedicated VRAM recommended (Windows) | CUDA |
| AMD | ROCm v7 on listed Radeon and Instinct cards; Vulkan for others | Not listed in its docs | HIP and Vulkan |
| Intel GPUs | Vulkan on Windows and Linux | Not listed in its docs | SYCL and Vulkan |
| Apple Silicon | Metal, plus MLX for supported models | Metal (llama.cpp) and MLX | Metal |
| Intel Macs | CPU only | Not supported | CPU builds |
LM Studio’s docs don’t list GPU vendors; it downloads the llama.cpp runtime for your system. llama.cpp also builds for CANN, MUSA, OpenCL, WebGPU and more, and can split a model between CPU and GPU.
On a Mac the GPU can use only part of unified memory by default. Running LLMs locally on a Mac shows what fits at each memory size.
What context length does each use by default?
This is the setting most likely to surprise you. Ollama sizes the context by GPU memory: 4K tokens below 24 GiB of VRAM, 32K from 24 to 48 GiB, and 256K at 48 GiB or more. llama-server starts from the context the model was trained with and, with its default --fit option, shrinks unset values to fit your memory (to no less than 4,096 tokens). LM Studio sets the context when you load a model, in the app or with lms load --context-length.
Ollama recommends at least 64,000 tokens for agents, web search and coding tools, and LM Studio’s Claude Code guide asks for more than about 25,000. Raising it costs memory. By our estimator, Llama 3.1 8B at Q4_K_M needs about 6.1 GB with Ollama’s 4K default but 13.6 GB at 64,000 tokens; Gemma 3 12B, whose sliding-window layers keep its cache small, goes from 8.5 GB to 12.3 GB. Both still fit a 16 GB card.
# Ollama: for every model the server loads
OLLAMA_CONTEXT_LENGTH=64000 ollama serve
ollama ps # the CONTEXT column shows what each loaded model got
# LM Studio
lms load <model_key> --context-length 64000
# llama.cpp: llama-server comes with brew, winget and source builds
# (the llama.app script installs only the "llama" command, not llama-server)
llama-server -hf ggml-org/Qwen3.5-0.8B-GGUF -c 64000To see whether a document or a coding session fits, paste it into the context window checker.
Is each one open source, and does it keep your data local?
llama.cpp, Ollama, MLX LM and the Llama app are MIT-licensed. LM Studio is proprietary: its terms grant use “for Your personal and / or internal business purposes”, and since July 2025 it has been free “both at home and at work”. Its lms CLI, SDKs and MLX engine are open source under MIT.
- Ollama: “We don’t see your prompts or data when you run locally.” Its optional cloud models do process prompts (Ollama says it doesn’t store or train on them); set
OLLAMA_NO_CLOUD=1to switch cloud features off. - LM Studio: once a model is downloaded, “Nothing you enter into LM Studio when chatting with LLMs leaves your device”. Searching, downloading and update checks need the internet.
- llama.cpp: a local program with no account;
-hfdownloads model files from Hugging Face.
Which should you choose?
- Your first local model: LM Studio if you’d rather click than type: you search, download and chat inside one app. Ollama if you’re happy in a terminal: one command downloads and runs a model.
- A local API for your app or coding agent: Ollama or llama-server. Both are scriptable and OpenAI-compatible; Ollama manages models for you, llama-server lets you set every option. Raise the context first.
- A Mac: LM Studio or Ollama, both of which use MLX where it’s supported. Choose MLX LM if you work in Python or want to fine-tune.
- A headless server: Ollama (Linux installer or Docker image), llama-server (Docker images and a router mode for several models), or LM Studio’s llmster daemon.
- Fine control: llama.cpp. You pick the exact GGUF file, GPU layers, cache type and backend, and you can build it for the hardware you have.
FAQ
Questions people ask
Is LM Studio open source?
No. The LM Studio app is proprietary software, licensed for personal and internal business use, and free at home and at work since July 2025. Some parts are open source under the MIT licence: the lms command-line tool, the Python and TypeScript SDKs and its MLX engine. Ollama and llama.cpp are fully open source under MIT.
Is Ollama faster than LM Studio?
There’s no reliable general answer. Both run GGUF models with llama.cpp and, on Apple Silicon, supported models with MLX, so speed depends mostly on which engine version each bundles and on settings such as context length and GPU offload. We found no like-for-like benchmark from either project. Run the same model file in both on your own machine.
Can LM Studio run on a server without the GUI?
Yes. LM Studio’s docs recommend llmster, a standalone daemon packaged from the core of the desktop app, which runs on Linux servers, cloud machines and GPU rigs without a GUI. Install it with the script from lmstudio.ai, start it with lms daemon up, then serve models with lms server start.
Which ports do Ollama, LM Studio and llama.cpp use?
Ollama listens on 127.0.0.1:11434, LM Studio on 127.0.0.1:1234 and current llama-server builds on 127.0.0.1:9931. MLX LM’s server uses port 8080. Each serves OpenAI-compatible endpoints under /v1, so an OpenAI client only needs the base URL changed, plus a model name the server knows.
Is LM Studio free for commercial use?
LM Studio says it is free to use both at home and at work, with no form or separate licence needed since July 2025. Its terms licence the app for personal and internal business purposes. LM Studio also offers Teams and Enterprise plans, with features such as single sign-on and private sharing. Check the terms if you plan to build it into a product.
Do Ollama, LM Studio or llama.cpp send my prompts anywhere?
Not when you run models locally. Ollama says it doesn’t see prompts from local runs, LM Studio says nothing you type leaves your device, and llama.cpp is a local program with no account. Prompts leave your machine only if you choose Ollama’s cloud models or expose a server to a network.
Try it
Tools from this guide
Keep reading