Playground
Gemini API playground: try Gemini free
Chat with Google’s free Gemini models using your own free API key. Replies stream into your browser straight from Google, with tokens, speed and cost shown for each one.
All free on Google’s free tier, within your daily limits.
Settings
Thinking counts towards this limit, so very low values can cut answers short.
Messages go straight from your browser to Google. On the free tier, Google may use them to improve its products, so don’t send anything confidential.
Chat
Add your free Gemini API key to start chatting.
Steps
How to use the gemini API playground
- Get a free key from Google AI Studio (see the guide), then paste it in. It’s kept for this visit unless you tick “Remember on this device”.
- Pick a model. Gemini 3.8 Flash is the newest; the Flash-Lite models are faster and lighter.
- Type a message and press Enter. Press Stop to cut a reply short.
- Check the readout under each reply to see its token use, speed and what it would cost on the paid tier.
Method
How it works
The playground is a direct line between your browser and the Gemini API. When you send a message, your browser calls Google’s streamGenerateContent endpoint with your key in a request header, and the reply streams back word by word. This site’s server never sees your key or your messages, and the security policy of this page only allows connections to Google’s API.
A real conversation, as the API sees it
The API has no memory: each message you send includes the whole conversation so far, which is why the “in” count grows as you chat. That is exactly how an app you build would work, so the readout is a realistic preview of your own costs. Use the subscription vs API calculator to turn it into a monthly figure.
Thinking tokens
Current Gemini models reason before replying. That reasoning isn’t shown, but it is counted: the readout lists it separately, and on the paid tier it’s billed at the output price. A short answer can use far more thinking tokens than visible ones, so it’s worth watching when you estimate costs.
Free tier and privacy
Google’s free tier needs no billing account. Limits apply per project and reset daily at midnight Pacific time; your current limits are in AI Studio. On the free tier, Google says content may be used to improve its products, so keep confidential data out of it, or use a paid-tier key, where it isn’t.
Examples
Worked examples
What a typical reply would cost on the paid tier
| Model | Input / output per 1M | Typical reply |
|---|---|---|
| Gemini 3.8 Flash | $0.75 / $3.75 | $0.00638 |
| Gemini 3.7 Flash | $0.75 / $3.75 | $0.00638 |
| Gemini 3.6 Flash | $0.75 / $3.75 | $0.00638 |
| Gemini 3.5 Flash | $1.50 / $9 | $0.015 |
| Gemini 3.5 Flash-Lite | $0.30 / $2.50 | $0.00405 |
| Gemini 3.1 Flash-Lite | $0.25 / $1.50 | $0.0025 |
A typical reply here is 1,000 tokens in, 500 visible tokens out and 1,000 thinking tokens. All of these models cost nothing on the free tier.
FAQ
Frequently asked questions
Is the Gemini API really free?
Yes, within limits. Google’s free tier offers Gemini 3.8 Flash, Gemini 3.7 Flash, Gemini 3.6 Flash, Gemini 3.5 Flash, Gemini 3.5 Flash-Lite, Gemini 3.1 Flash-Lite free of charge (checked 2026-10-08), with no billing account needed. Limits are set per Google Cloud project and shown in AI Studio, and the daily quota resets at midnight Pacific time.
Is my conversation private?
Your messages and key go straight from your browser to Google; this site never receives them, and the conversation is only kept in this tab. But on Google’s free tier, Google says your content may be used to improve its products, so don’t paste anything confidential. On the paid tier it isn’t used that way.
Why does the first word take a few seconds?
Current Gemini models think before they answer. The playground shows “Thinking…” with a timer until the first words arrive, and the readout under each reply shows how many thinking tokens were used and the time to the first token.
What do the numbers under each reply mean?
“In” is the tokens sent (your message plus the conversation so far and any system instruction), “out” is the visible reply, and “thinking” is the model’s hidden reasoning. On the paid tier, thinking is billed as output. The cost shown is what the reply would cost at paid-tier prices; on the free tier it costs nothing.
Why do I get a “busy” or “high demand” error?
Google sometimes returns a “high demand” error for a popular model when it’s overloaded. It’s temporary: wait a moment or pick another model, which may answer straight away. A “limit” error means you’ve used your free-tier quota for now.
How do I get a key?
Create one free in Google AI Studio; it takes a few minutes and no card. The step-by-step guide walks you through it and lets you test the key.
Related
Related tools
- AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.
- LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.
- AI Model ComparisonCompare prices, context windows and features across models.
- AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.
- Context Window CheckerSee whether your text fits each model's context window.
- Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.
Read