AI glossary · Agents and tools
What is an OpenAI-compatible API?
Also called: OpenAI compatibility, OpenAI-compatible endpoint, base_url
Definition
An OpenAI-compatible API is a model API that accepts OpenAI’s Chat Completions request format, so you can call it with the official OpenAI SDK by changing only the base URL, the API key and the model name.
Explained
How it works
OpenAI’s Chat Completions shape, a messages list in and a choices array out, became the format other providers copy. Groq, OpenRouter, Google’s Gemini API, Anthropic and local servers such as Ollama all accept it, and the OpenAI SDK sends the same request to whatever base_url you set.
Compatible means the core works, not everything. Groq returns a 400 error for logprobs, logit_bias, top_logprobs and messages[].name, and turns temperature: 0 into 1e-8. Ollama doesn’t support tool_choice. Anthropic’s layer ignores response_format and the strict flag on tools, has no prompt caching, and is meant for testing and comparing models rather than long-term production use in most cases. Gemini’s is in beta and takes Gemini-only options through extra_body.
Example
Five base URLs, checked today
We sent the same Chat Completions request through the OpenAI Python SDK (3.28.0) to the four hosted URLs below with a deliberately invalid key. Each answered with an authentication error, which confirms the URL and path. Groq, OpenRouter and Anthropic replied 401 with an error object; Gemini replied 400 with its error object inside a list, so even error handling can differ.
Switching providers is then a three-line change: the base URL, the key and the model name, as in the OpenRouter example below.
| Provider | Base URL | Worth knowing |
|---|---|---|
| Groq | https://api.groq.com/openai/v1 | 400 error on logprobs, logit_bias, top_logprobs |
| OpenRouter | https://openrouter.ai/api/v1 | Model IDs look like provider/model |
| Google Gemini | https://generativelanguage.googleapis.com/v1beta/openai/ | Beta; Gemini options via extra_body |
| Anthropic | https://api.anthropic.com/v1/ | For testing; response_format ignored |
| Ollama (local) | http://localhost:11434/v1/ | Key required but ignored |
Base URLs from each provider’s compatibility docs, 2026-10-11. Ollama runs on your machine and was not probed.
import os
from openai import OpenAI
client = OpenAI(
base_url="https://openrouter.ai/api/v1",
api_key=os.environ["OPENROUTER_API_KEY"],
)
response = client.chat.completions.create(
model="google/gemini-3.8-flash",
messages=[
{"role": "system", "content": "Answer in one sentence."},
{"role": "user", "content": "What is an OpenAI-compatible API?"},
],
)
print(response.choices[0].message.content)Cost and quality
Why it matters
One client library lets you compare models and prices across providers, or move to a cheaper one, without rewriting your code. Keep each provider’s API key in its own environment variable, never in the code.
Test the features you rely on, such as function calling, JSON output, streaming and usage fields, on each provider. Unsupported fields are often ignored silently rather than rejected, and provider-only features, such as Anthropic’s prompt caching, need the native API.
Don’t mix up
Common confusions
- Chat Completions vs the Responses API
- “OpenAI-compatible” almost always means the Chat Completions endpoint. Support for OpenAI’s newer Responses API varies: Ollama, for example, offers a non-stateful
/v1/responses. - Same format vs same model
- A compatible API copies the request format, not the model. Prompts tuned for one model may need rework on another, as Anthropic’s compatibility docs point out.
Go deeper
Try it and read more
- Free toolGet an AI API key: free tiers comparedWhich AI APIs you can use for free, which need a card, and step-by-step guides for Gemini, Groq, OpenRouter, Mistral, OpenAI and Claude keys.
- Free toolAI Model ComparisonCompare prices, context windows and features across models.
- Free toolLLM API Cost CalculatorEstimate per-request, daily and monthly API costs.
- Guide · 13 min readOpenAI vs Anthropic vs Gemini API message formats comparedOpenAI, Anthropic and Gemini message formats side by side: system prompts, roles, images, tool calls, streaming, usage fields and a tested converter.
- Guide · 9 min readThe cheapest LLM APIs right now, ranked from daily price dataThe cheapest LLM APIs ranked by blended price from daily data, the cheapest for tools, vision and long context, free tiers, and how to cut cost per task.
Related
Related terms
- Function callingFunction calling, also called tool use, is an LLM API feature in which the model replies with a structured request to run a function you described, with JSON arguments, which your code executes before sending the result back.
- Structured outputsStructured outputs is an LLM API feature that constrains the model’s reply to a JSON Schema you supply, so the response parses and has the fields and types you asked for, unlike JSON mode, which only promises valid JSON.
- API keyAn API key is a secret string that identifies your account to a service such as the OpenAI, Claude or Gemini API, so every request made with it is authorised, rate-limited and billed to you.
- BYOKBYOK (bring your own key) is a model where an AI app or tool runs on your own API key from a provider such as OpenAI, Anthropic or Google, so usage is billed to your account instead of the app’s.
- Rate limitA rate limit is a cap on how many requests or tokens an account may send to an API per minute or per day, and going over it makes the API reject requests with HTTP 429 until the allowance refills.