Skip to content
AI Dev Toolkit.
Esc
  • AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.Tool
  • LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.Tool
  • AI Model ComparisonCompare prices, context windows and features across models.Tool
  • AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.Tool
  • Context Window CheckerSee whether your text fits each model’s context window.Tool
  • Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.Tool
  • GPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.Tool
  • Prompt Caching CalculatorEstimate savings from prompt caching.Tool

Guide · Local AI

Ollama vs LM Studio vs llama.cpp: which local LLM tool?

Pick LM Studio if you want a desktop app, Ollama if you want one command and a local API, and llama.cpp if you want every setting. They are layers more than rivals: llama.cpp is the engine, Ollama and LM Studio both run it, and on Apple Silicon both can also run Apple’s MLX.

By Tahir NazirUpdated 10 min read

On this page
  1. Ollama vs LM Studio vs llama.cpp: what’s the difference?
  2. Does Ollama use llama.cpp?
  3. How do you install each one?
  4. GGUF or MLX: which model formats does each support?
  5. Which has the best local API? Endpoints and default ports
  6. GPU support on Windows, Linux and Mac
  7. What context length does each use by default?
  8. Is each one open source, and does it keep your data local?
  9. Which should you choose?
  10. Questions people ask

Ollama vs LM Studio vs llama.cpp: what’s the difference?

They differ in how you use them, not in what fundamentally runs the model. LM Studio is a desktop app for browsing, chatting with and serving models. Ollama is a command-line tool and background service with a desktop app and its own model library. llama.cpp is the open-source C/C++ engine itself, with a server, a command line and now its own small app.

The three tools side by side
OllamaLM Studiollama.cpp
What it isCLI, background server and desktop appDesktop app with lms CLI and the llmster daemonEngine, llama-server, llama-cli and the Llama app
InstallScript, Mac/Windows installer or Docker imageInstaller (Linux AppImage) or llmster scriptScript, Homebrew, Winget, conda, Nix, Docker or binaries
Model formatsIts library, or GGUF and Safetensors you importGGUF; MLX on Apple SiliconGGUF
Find modelsollama.com/libraryHugging Face search in the appAny GGUF repo on Hugging Face (-hf)
Enginesllama.cpp; MLX on Apple Siliconllama.cpp; MLX on Apple Siliconllama.cpp (on ggml)
Local API127.0.0.1:11434127.0.0.1:1234127.0.0.1:9931
Default context4K, 32K or 256K by VRAMSet when you load a modelThe model’s own, fitted to memory
LicenceMITProprietary, free at home and workMIT

From each project’s documentation and repository, 2026-10-11.

Does Ollama use llama.cpp?

Yes. Ollama’s README lists llama.cpp as its supported backend, and LM Studio’s docs say it runs models on Mac, Windows and Linux with llama.cpp. On Apple Silicon both add a second engine, Apple’s MLX: LM Studio runs MLX models there, and since Ollama 0.40.0 “model architectures supported by the MLX runtime will automatically run on MLX”.

What you install sits on top of one of two engines. Lime marks llama.cpp, violet marks MLX; the numbers are each tool’s default API port.

So the same GGUF file runs on the same engine in all three. What differs is the version of llama.cpp each one bundles, the defaults it picks (context length, GPU offload, sampling) and the extras around it. That’s why we don’t quote speed rankings: a result depends on those versions and settings, and we found no like-for-like comparison published by the projects themselves. Test your own model on your own machine.

How do you install each one?

Each project offers a one-line installer as well as the usual packages. These are the commands from their own READMEs and docs:

Install and run a first model
# Ollama (macOS or Linux; Windows: irm https://ollama.com/install.ps1 | iex)
curl -fsSL https://ollama.com/install.sh | sh
ollama run gemma4

# LM Studio: the desktop app from lmstudio.ai/download, or the headless daemon
curl -fsSL https://lmstudio.ai/install.sh | bash
lms daemon up
lms get llama-3.1-8b@q4_k_m
lms server start

# llama.cpp (or: brew install llama.cpp, winget install llama.cpp)
curl -LsSf https://llama.app/install.sh | sh
llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF

Ollama needs macOS 14 or newer and uses the GPU only on Apple M-series Macs; Intel Macs run on the CPU. LM Studio needs macOS 14 on Apple Silicon (Intel Macs aren’t supported), AVX2 on x64 Windows, and Ubuntu 20.04 or newer on Linux, and recommends 16 GB of RAM. llama.cpp’s team, with Hugging Face, also ships Llama, a free, open-source app (menu bar on Mac, system tray on Windows) that serves the same API on port 9931.

GGUF or MLX: which model formats does each support?

GGUF is llama.cpp’s single-file format, and all three read it. MLX models are safetensors weights prepared for Apple’s MLX, used only on Apple Silicon here.

  • Ollama pulls from its own library (ollama pull gemma4) and imports GGUF or Safetensors through a Modelfile with FROM /path/to/file.gguf and ollama create. It doesn’t quantise GGUF on import: its docs send you to llama.cpp’s llama-quantize.
  • LM Studio searches Hugging Face from its Discover tab and offers GGUF and, on Apple Silicon, MLX downloads (lms get --mlx or --gguf). Its docs suggest a 4-bit option or higher if your machine can run it.
  • llama.cpp downloads any GGUF repository with -hf user/model:quant, choosing Q4_K_M when you don’t name a quantisation. The Llama app keeps models in the Hugging Face cache, shared with llama.cpp.
  • MLX LM runs MLX models from Hugging Face, thousands of them in the mlx-community organisation, and can quantise and fine-tune them.

Whatever the format, the file has to fit in memory with room for the context. The VRAM calculator shows how much a model needs at each quantisation, and how much VRAM you need to run an LLM locally explains the sums.

Which has the best local API? Endpoints and default ports

All of them speak OpenAI’s API, so any OpenAI client can switch between them by changing the base URL. Each listens only on 127.0.0.1 by default:

Local API servers
ServerDefault addressOpenAI-compatibleAlso offers
Ollamahttp://localhost:11434/v1/chat/completions, /v1/responses, /v1/completions, /v1/embeddings, /v1/modelsNative /api/* API, Anthropic-compatible /v1/messages
LM Studiohttp://localhost:1234/v1/chat/completions, /v1/responses, /v1/completions, /v1/embeddings, /v1/modelsIts REST API, Python and TypeScript SDKs, Anthropic-compatible /v1/messages
llama-serverhttp://localhost:9931/v1/chat/completions, /v1/responses, /v1/completions, /v1/embeddings, /v1/modelsAnthropic-compatible /v1/messages, reranking, web UI
mlx_lm.serverhttp://localhost:8080/v1/chat/completions, /v1/modelsMLX LM warns it isn’t meant for production

LM Studio reuses the last port you chose if you don’t pass `--port`. llama-server’s README sets 8080 in its Docker examples, so check which port your build uses.

One OpenAI client, three local servers (Python)
from openai import OpenAI

# Pick the server you run. Each one needs a model already downloaded.
client = OpenAI(base_url="http://localhost:11434/v1/", api_key="ollama")  # Ollama: key required, ignored
# client = OpenAI(base_url="http://localhost:1234/v1", api_key="lm-studio")  # LM Studio
# client = OpenAI(base_url="http://localhost:9931/v1", api_key="none")  # llama-server

reply = client.chat.completions.create(
    model="gemma4",  # Ollama's name; use LM Studio's model identifier there
    messages=[{"role": "user", "content": "Say hello in five words."}],
)
print(reply.choices[0].message.content)

Ollama’s OpenAI layer can’t set the context size, so for a longer context you create a model with PARAMETER num_ctx in a Modelfile. If you move between local servers and hosted APIs, our guide to OpenAI, Anthropic and Gemini message formats covers the request shapes these endpoints copy.

GPU support on Windows, Linux and Mac

All three use NVIDIA and Apple GPUs. AMD and Intel support depends on the backend:

GPU support as each project documents it
HardwareOllamaLM Studiollama.cpp
NVIDIACompute capability 5.0 or newer, driver 550+ (570+ for 5.0 to 6.2)4 GB or more of dedicated VRAM recommended (Windows)CUDA
AMDROCm v7 on listed Radeon and Instinct cards; Vulkan for othersNot listed in its docsHIP and Vulkan
Intel GPUsVulkan on Windows and LinuxNot listed in its docsSYCL and Vulkan
Apple SiliconMetal, plus MLX for supported modelsMetal (llama.cpp) and MLXMetal
Intel MacsCPU onlyNot supportedCPU builds

LM Studio’s docs don’t list GPU vendors; it downloads the llama.cpp runtime for your system. llama.cpp also builds for CANN, MUSA, OpenCL, WebGPU and more, and can split a model between CPU and GPU.

On a Mac the GPU can use only part of unified memory by default. Running LLMs locally on a Mac shows what fits at each memory size.

What context length does each use by default?

This is the setting most likely to surprise you. Ollama sizes the context by GPU memory: 4K tokens below 24 GiB of VRAM, 32K from 24 to 48 GiB, and 256K at 48 GiB or more. llama-server starts from the context the model was trained with and, with its default --fit option, shrinks unset values to fit your memory (to no less than 4,096 tokens). LM Studio sets the context when you load a model, in the app or with lms load --context-length.

Ollama recommends at least 64,000 tokens for agents, web search and coding tools, and LM Studio’s Claude Code guide asks for more than about 25,000. Raising it costs memory. By our estimator, Llama 3.1 8B at Q4_K_M needs about 6.1 GB with Ollama’s 4K default but 13.6 GB at 64,000 tokens; Gemma 3 12B, whose sliding-window layers keep its cache small, goes from 8.5 GB to 12.3 GB. Both still fit a 16 GB card.

Raise the context, then check it
# Ollama: for every model the server loads
OLLAMA_CONTEXT_LENGTH=64000 ollama serve
ollama ps   # the CONTEXT column shows what each loaded model got

# LM Studio
lms load <model_key> --context-length 64000

# llama.cpp: llama-server comes with brew, winget and source builds
# (the llama.app script installs only the "llama" command, not llama-server)
llama-server -hf ggml-org/Qwen3.5-0.8B-GGUF -c 64000

To see whether a document or a coding session fits, paste it into the context window checker.

Is each one open source, and does it keep your data local?

llama.cpp, Ollama, MLX LM and the Llama app are MIT-licensed. LM Studio is proprietary: its terms grant use “for Your personal and / or internal business purposes”, and since July 2025 it has been free “both at home and at work”. Its lms CLI, SDKs and MLX engine are open source under MIT.

  • Ollama: “We don’t see your prompts or data when you run locally.” Its optional cloud models do process prompts (Ollama says it doesn’t store or train on them); set OLLAMA_NO_CLOUD=1 to switch cloud features off.
  • LM Studio: once a model is downloaded, “Nothing you enter into LM Studio when chatting with LLMs leaves your device”. Searching, downloading and update checks need the internet.
  • llama.cpp: a local program with no account; -hf downloads model files from Hugging Face.

Which should you choose?

  • Your first local model: LM Studio if you’d rather click than type: you search, download and chat inside one app. Ollama if you’re happy in a terminal: one command downloads and runs a model.
  • A local API for your app or coding agent: Ollama or llama-server. Both are scriptable and OpenAI-compatible; Ollama manages models for you, llama-server lets you set every option. Raise the context first.
  • A Mac: LM Studio or Ollama, both of which use MLX where it’s supported. Choose MLX LM if you work in Python or want to fine-tune.
  • A headless server: Ollama (Linux installer or Docker image), llama-server (Docker images and a router mode for several models), or LM Studio’s llmster daemon.
  • Fine control: llama.cpp. You pick the exact GGUF file, GPU layers, cache type and backend, and you can build it for the hardware you have.
Free toolGPU and VRAM calculatorCheck whether a model fits your GPU or Mac at each quantisation and context length before you download it.

FAQ

Questions people ask

Is LM Studio open source?

No. The LM Studio app is proprietary software, licensed for personal and internal business use, and free at home and at work since July 2025. Some parts are open source under the MIT licence: the lms command-line tool, the Python and TypeScript SDKs and its MLX engine. Ollama and llama.cpp are fully open source under MIT.

Is Ollama faster than LM Studio?

There’s no reliable general answer. Both run GGUF models with llama.cpp and, on Apple Silicon, supported models with MLX, so speed depends mostly on which engine version each bundles and on settings such as context length and GPU offload. We found no like-for-like benchmark from either project. Run the same model file in both on your own machine.

Can LM Studio run on a server without the GUI?

Yes. LM Studio’s docs recommend llmster, a standalone daemon packaged from the core of the desktop app, which runs on Linux servers, cloud machines and GPU rigs without a GUI. Install it with the script from lmstudio.ai, start it with lms daemon up, then serve models with lms server start.

Which ports do Ollama, LM Studio and llama.cpp use?

Ollama listens on 127.0.0.1:11434, LM Studio on 127.0.0.1:1234 and current llama-server builds on 127.0.0.1:9931. MLX LM’s server uses port 8080. Each serves OpenAI-compatible endpoints under /v1, so an OpenAI client only needs the base URL changed, plus a model name the server knows.

Is LM Studio free for commercial use?

LM Studio says it is free to use both at home and at work, with no form or separate licence needed since July 2025. Its terms licence the app for personal and internal business purposes. LM Studio also offers Teams and Enterprise plans, with features such as single sign-on and private sharing. Check the terms if you plan to build it into a product.

Do Ollama, LM Studio or llama.cpp send my prompts anywhere?

Not when you run models locally. Ollama says it doesn’t see prompts from local runs, LM Studio says nothing you type leaves your device, and llama.cpp is a local program with no account. Prompts leave your machine only if you choose Ollama’s cloud models or expose a server to a network.

Try it

Tools from this guide

Keep reading