AI glossary · Models and context
What is a system prompt?
Also called: system message, system instructions, developer message
Definition
A system prompt is the set of standing instructions a developer sends with every request to set a model’s role, rules and output format, kept separate from what the user types.
Explained
How it works
It typically says who the assistant is, what it may and may not do, which tools to call and how answers should look. Keep secrets such as API keys out of it: it is ordinary text that the model may repeat if a user asks the right way.
Each API takes it in a different place. OpenAI’s Chat Completions uses a developer message (or the older system role), and its Responses API has a top-level instructions field. Anthropic’s Messages API takes a top-level system, and Gemini’s generateContent takes systemInstruction, which accepts text only. Our message format guide shows all four side by side.
Models keep no state between calls, so the system prompt is sent, and billed as input tokens, on every request and every turn of a conversation.
Example
A 78-token system prompt sent 1,000,000 times
The prompt below is 68 words and 78 tokens with OpenAI’s o200k_base tokenizer, counted as plain text. On GPT-6.1 Sol at $2 per million input tokens, sending it with 1,000,000 requests costs $156.00 before a single user message is counted.
It is too short to cache: OpenAI’s minimum cacheable prefix is 1,024 tokens on GPT-5.6 and later. Once a system prompt grows past that with tool definitions and examples, prompt caching makes the repeated part much cheaper, as long as it stays identical at the top of every request.
You are the support assistant for an online bike shop.
Answer questions about orders, deliveries, returns and bike parts.
Keep replies to three sentences or fewer unless the customer asks for more detail.
Call the check_order tool before you say anything about an order, and never guess a delivery date.
If you do not know the answer, say so and offer to pass the question to a person.Cost and quality
Why it matters
It is the cheapest lever on quality you have: role, rules and format in one place, applied to every turn without retraining anything. Anthropic’s prompting guide notes that even a single sentence setting a role makes a difference.
It is also a fixed cost on every call, and a long list of rarely needed rules competes for the model’s attention. Keep it focused, put reference material in retrieval or tools, and measure it with the token counter when it grows.
Don’t mix up
Common confusions
- System prompt vs user prompt
- The user prompt comes from the person using the app and changes every turn. The system prompt is yours and stays the same across the conversation.
- System message vs developer message
- OpenAI’s newer models use the
developerrole for the same job: its reasoning best practices guide says reasoning models support developer messages rather than system messages fromo1-2024-12-17on. Anthropic and Gemini have nodeveloperrole; use their top-level fields.
Go deeper
Try it and read more
- Free toolAI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.
- Free toolPrompt Caching CalculatorEstimate savings from prompt caching.
- Free toolGemini API PlaygroundTry the Gemini API free with your own key.
- Guide · 13 min readOpenAI vs Anthropic vs Gemini API message formats comparedOpenAI, Anthropic and Gemini message formats side by side: system prompts, roles, images, tool calls, streaming, usage fields and a tested converter.
Related
Related terms
- Prompt cachingPrompt caching is an API feature that stores the processed start of a prompt, so later requests that begin with exactly the same tokens are billed at a much lower cached-input price and start answering sooner.
- Few-shot promptingFew-shot prompting is giving a model a handful of worked input-and-output examples inside the prompt, so it copies the pattern for a new input without any retraining.
- TokenA token in AI is the unit of text a language model reads and writes, usually a whole word, part of a word or a punctuation mark, which the model sees only as a number from its vocabulary.
- CLAUDE.mdCLAUDE.md is a Markdown file of project instructions that Claude Code loads into context at the start of every session, and AGENTS.md is the open, tool-neutral equivalent read by many coding agents.