Skip to content
AI Dev Toolkit.
Esc
  • AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.Tool
  • LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.Tool
  • AI Model ComparisonCompare prices, context windows and features across models.Tool
  • AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.Tool
  • Context Window CheckerSee whether your text fits each model’s context window.Tool
  • Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.Tool
  • GPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.Tool
  • Prompt Caching CalculatorEstimate savings from prompt caching.Tool

AI glossary · Models and context

What is a system prompt?

Also called: system message, system instructions, developer message

Definition

A system prompt is the set of standing instructions a developer sends with every request to set a model’s role, rules and output format, kept separate from what the user types.

Explained

How it works

It typically says who the assistant is, what it may and may not do, which tools to call and how answers should look. Keep secrets such as API keys out of it: it is ordinary text that the model may repeat if a user asks the right way.

Each API takes it in a different place. OpenAI’s Chat Completions uses a developer message (or the older system role), and its Responses API has a top-level instructions field. Anthropic’s Messages API takes a top-level system, and Gemini’s generateContent takes systemInstruction, which accepts text only. Our message format guide shows all four side by side.

Models keep no state between calls, so the system prompt is sent, and billed as input tokens, on every request and every turn of a conversation.

Example

A 78-token system prompt sent 1,000,000 times

The prompt below is 68 words and 78 tokens with OpenAI’s o200k_base tokenizer, counted as plain text. On GPT-6.1 Sol at $2 per million input tokens, sending it with 1,000,000 requests costs $156.00 before a single user message is counted.

It is too short to cache: OpenAI’s minimum cacheable prefix is 1,024 tokens on GPT-5.6 and later. Once a system prompt grows past that with tool definitions and examples, prompt caching makes the repeated part much cheaper, as long as it stays identical at the top of every request.

System prompt (78 tokens, o200k_base)
You are the support assistant for an online bike shop.
Answer questions about orders, deliveries, returns and bike parts.
Keep replies to three sentences or fewer unless the customer asks for more detail.
Call the check_order tool before you say anything about an order, and never guess a delivery date.
If you do not know the answer, say so and offer to pass the question to a person.

Cost and quality

Why it matters

It is the cheapest lever on quality you have: role, rules and format in one place, applied to every turn without retraining anything. Anthropic’s prompting guide notes that even a single sentence setting a role makes a difference.

It is also a fixed cost on every call, and a long list of rarely needed rules competes for the model’s attention. Keep it focused, put reference material in retrieval or tools, and measure it with the token counter when it grows.

Don’t mix up

Common confusions

System prompt vs user prompt
The user prompt comes from the person using the app and changes every turn. The system prompt is yours and stays the same across the conversation.
System message vs developer message
OpenAI’s newer models use the developer role for the same job: its reasoning best practices guide says reasoning models support developer messages rather than system messages from o1-2024-12-17 on. Anthropic and Gemini have no developer role; use their top-level fields.

Go deeper

Try it and read more

Related

All 40 terms in the AI glossary