Skip to content
AI Dev Toolkit.
Esc
  • AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.Tool
  • LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.Tool
  • AI Model ComparisonCompare prices, context windows and features across models.Tool
  • AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.Tool
  • Context Window CheckerSee whether your text fits each model’s context window.Tool
  • Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.Tool
  • GPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.Tool
  • Prompt Caching CalculatorEstimate savings from prompt caching.Tool

Guide · Developer tools

Fine-tuning JSONL format for OpenAI, Gemini and Mistral

A fine-tuning file is JSONL: one complete JSON object per line, UTF-8, no blank lines. OpenAI and Mistral expect a messages array with system, user and assistant roles; Gemini on Google Cloud expects contents with user and model roles and text in parts. Check availability first: OpenAI closed fine-tuning to new organisations in May 2026.

By Tahir NazirUpdated 13 min read

On this page
  1. What is the JSONL format for fine-tuning?
  2. Where can you still fine-tune? (October 2026)
  3. OpenAI fine-tuning JSONL example
  4. Gemini tuning JSONL format
  5. Mistral fine-tuning JSONL format
  6. How to validate a JSONL file before you upload it
  7. Common fine-tuning file errors and how to fix them
  8. How to convert CSV to JSONL (Python and JavaScript)
  9. How many training tokens, and what will it cost?
  10. Pre-upload checklist
  11. Questions people ask

What is the JSONL format for fine-tuning?

JSONL (JSON Lines) is a text file where every line is a complete JSON value. For fine-tuning, each line is one training example: a short conversation that ends with the answer you want the model to learn. The JSON Lines specification has three rules: the file is UTF-8 without a byte order mark, every line is valid JSON (a blank line is not), and lines end with \n.

The assistant message is the only part the model is trained to write. Here it is 1 token, while the instructions repeated on every line are 20, which matters when you pay per training token.

The most common mistake is pretty-printing. An object spread over five lines is valid JSON but five broken JSONL lines. Write each example with a JSON serialiser (json.dumps in Python, JSON.stringify in JavaScript), which escapes line breaks inside strings as \n and keeps the object on one line.

Where can you still fine-tune? (October 2026)

Check this before you prepare data, because all three providers have narrowed self-serve tuning:

Self-serve fine-tuning status, checked 2026-10-11
ProviderStatusModels
OpenAIWinding down. Since 7 May 2026, organisations that had never fine-tuned can’t start; since 2 July 2026, nor can those with no fine-tuned model inference in the past 60 days. Active customers can create jobs until 6 January 2027. Existing fine-tuned models keep working until their base model is deprecated.Supervised and DPO: gpt-4.1, gpt-4.1-mini, gpt-4.1-nano (2025-04-14). Vision: gpt-4o-2024-08-06. Reinforcement: o4-mini-2025-04-16.
Google GeminiNot in the Gemini API or AI Studio: no model there has supported tuning since Gemini 1.5 Flash-001 was deprecated in May 2025. Supervised tuning is offered on Google Cloud’s Gemini Enterprise Agent Platform (the old Vertex AI tuning docs now redirect there).Gemini 3.5 Flash, 3.1 Flash-Lite, 2.5 Pro, 2.5 Flash, 2.5 Flash-Lite
MistralIts fine-tuning docs now open with “This feature is deprecated and is no longer actively supported.” Its fine-tuning overview lists a minimum fee of $4 per job and $2 a month to store each model.open-mistral-7b, mistral-small, codestral, open-mistral-nemo, mistral-large, ministral-8b, ministral-3b; pixtral-12b for vision

From OpenAI’s deprecations page and fine-tuning guides, Google’s Gemini API model-tuning page and Google Cloud’s supervised tuning docs, and Mistral’s fine-tuning docs.

If you can’t fine-tune, the usual alternative is a well-tested prompt with examples, and prompt caching makes the repeated part cheaper. The JSONL you build is still useful as an evaluation set.

OpenAI fine-tuning JSONL example

OpenAI supervised fine-tuning uses the chat format: a messages array of system, user and assistant messages, one example per line, uploaded with the purpose fine-tune. Multi-turn examples simply contain several user and assistant turns, and the optional weight key (0 or 1) on an assistant message turns training on that message off or on:

train.jsonl (OpenAI, supervised)
{"messages":[{"role":"system","content":"Classify the support ticket. Reply with one label: billing, bug, feature_request or other."},{"role":"user","content":"I was charged twice for my March invoice."},{"role":"assistant","content":"billing"}]}
{"messages":[{"role":"system","content":"Marv is a factual chatbot that is also sarcastic."},{"role":"user","content":"What's the capital of France?"},{"role":"assistant","content":"Paris","weight":0},{"role":"user","content":"Can you be more sarcastic?"},{"role":"assistant","content":"Paris, as if everyone doesn't know that already.","weight":1}]}

To teach function calling, the assistant message carries tool_calls instead of content, and the line includes the tools definitions (and optionally parallel_tool_calls) next to messages. Note that arguments is a JSON string, not an object:

train.jsonl (OpenAI, function calling)
{"messages":[{"role":"user","content":"What is the weather in San Francisco?"},{"role":"assistant","tool_calls":[{"id":"call_id","type":"function","function":{"name":"get_current_weather","arguments":"{\"location\": \"San Francisco, USA\", \"format\": \"celsius\"}"}}]}],"parallel_tool_calls":false,"tools":[{"type":"function","function":{"name":"get_current_weather","description":"Get the current weather","parameters":{"type":"object","properties":{"location":{"type":"string"},"format":{"type":"string","enum":["celsius","fahrenheit"]}},"required":["location","format"]}}}]}

Direct preference optimisation (DPO) uses a different shape: an input with the prompt, a preferred_output and a non_preferred_output. OpenAI only trains DPO on one-turn conversations, and both outputs must be the final assistant message:

train.jsonl (OpenAI, DPO)
{"input":{"messages":[{"role":"user","content":"Hello, can you tell me how cold San Francisco is today?"}],"tools":[],"parallel_tool_calls":true},"preferred_output":[{"role":"assistant","content":"Today in San Francisco, it is not quite cold as expected. Morning clouds will give away to sunshine, with a high near 68°F (20°C) and a low around 57°F (14°C)."}],"non_preferred_output":[{"role":"assistant","content":"It is not particularly cold in San Francisco today."}]}

Vision fine-tuning puts image_url parts in user messages (at most 10 images per example, each up to 10 MB, JPEG, PNG or WebP). Reinforcement fine-tuning lines hold messages plus whatever reference fields your grader reads, for example a compliant label.

Gemini tuning JSONL format

Gemini supervised tuning on Google Cloud takes one JSON object per line with a required contents list and an optional systemInstruction. Roles are user and model (not assistant), and text goes inside parts. Google’s docs say the role inside systemInstruction is ignored:

train.jsonl (Gemini)
{"systemInstruction":{"role":"system","parts":[{"text":"Classify the support ticket. Reply with one label: billing, bug, feature_request or other."}]},"contents":[{"role":"user","parts":[{"text":"I was charged twice for my March invoice."}]},{"role":"model","parts":[{"text":"billing"}]}]}
  • Other parts: fileData (with mimeType and fileUri) for images, audio, video and documents, and functionCall / functionResponse parts plus a tools field for function calling.
  • Limits: up to 131,072 input and output tokens per example, a 1 GB JSONL training file, and up to 10 million text examples, on all five tunable models.
  • How many: Google recommends starting with 100 examples, and its job docs suggest 100 to 500 for best results.
  • Where: upload the file to a Cloud Storage bucket; the tuning job reads it from there.
  • Thinking models: Google advises a thinking budget of 0 (Gemini 2.5) or thinking level MINIMAL (Gemini 3 and later) for tuned tasks, because tuning trains on answers without the thinking.

Mistral fine-tuning JSONL format

Mistral’s format follows the same chat shape as OpenAI’s, and its docs say it accepts OpenAI-format files. Messages go under messages, roles are system, user, assistant or tool, and loss is computed only on assistant tokens. Its docs add explicit rules for function-calling data and file sizes:

  • An assistant message with tool_calls can’t also have content, and must be followed by a tool message and then another assistant message.
  • Each tool_call_id must match the id of an earlier tool call, and both are random strings of exactly 9 characters.
  • The line’s tools list must define every tool the conversation uses.
  • Training files can be up to 512 MB each; validation data is capped at 1 MB (Mistral’s rule of thumb is the smaller of 1 MB and 5% of the training data).

How to validate a JSONL file before you upload it

Validate locally: a provider only reports a bad file after you upload it, and some problems (an over-long example, a missing answer) don’t fail at all but quietly waste training. This script checks the rules above for OpenAI and Mistral files and flags examples over OpenAI’s limit of 65,536 tokens per example for the GPT-4.1 family:

validate_jsonl.py
# pip install tiktoken
import json
import sys

import tiktoken

ROLES = {"system", "user", "assistant", "tool"}  # Gemini files use "user" and "model" instead
MIN_EXAMPLES = 10        # OpenAI's documented minimum
MAX_TOKENS = 65_536      # OpenAI's per-example limit for the GPT-4.1 family
enc = tiktoken.get_encoding("o200k_base")

def problems(example):
    if not isinstance(example, dict):
        return ["line is not a JSON object"]
    messages = example.get("messages")
    if not isinstance(messages, list) or not messages:
        return ["missing a non-empty \"messages\" list"]
    found = []
    for i, m in enumerate(messages):
        role = m.get("role") if isinstance(m, dict) else None
        if role not in ROLES:
            found.append(f"message {i}: unknown role {role!r}")
        elif not isinstance(m.get("content"), (str, list)) and not m.get("tool_calls"):
            found.append(f"message {i}: content must be a string")
        if isinstance(m, dict) and m.get("weight", 1) not in (0, 1):
            found.append(f"message {i}: weight must be 0 or 1")
    if not any(isinstance(m, dict) and m.get("role") == "assistant" for m in messages):
        found.append("no assistant message to learn from")
    return found

path = sys.argv[1]
raw = open(path, "rb").read()
if raw.startswith(b"\xef\xbb\xbf"):
    print("file starts with a UTF-8 byte order mark: save as UTF-8 without BOM")
text = raw.decode("utf-8-sig")  # raises on bytes that aren't UTF-8

count, errors = 0, 0
for n, line in enumerate(text.split("\n"), start=1):
    if not line.strip():
        if n < text.count("\n") + 1:
            print(f"line {n}: blank line"); errors += 1
        continue
    try:
        example = json.loads(line)
    except json.JSONDecodeError as e:
        print(f"line {n}: invalid JSON ({e.msg} at column {e.colno})"); errors += 1
        continue
    count += 1
    for p in problems(example):
        print(f"line {n}: {p}"); errors += 1
    tokens = len(enc.encode(line))  # rough: counts the JSON syntax too
    if tokens > MAX_TOKENS:
        print(f"line {n}: about {tokens:,} tokens, over {MAX_TOKENS:,} (will be truncated)")

if count < MIN_EXAMPLES:
    print(f"only {count} examples: OpenAI needs at least {MIN_EXAMPLES}")
print(f"{count} examples, {errors} errors")

We ran it on a nine-line file with one mistake per line (the error explanations are in the next section):

python validate_jsonl.py broken.jsonl
file starts with a UTF-8 byte order mark: save as UTF-8 without BOM
line 2: blank line
line 3: invalid JSON (Expecting value at column 97)
line 4: message 1: unknown role 'model'
line 4: no assistant message to learn from
line 5: invalid JSON (Expecting property name enclosed in double quotes at column 2)
line 6: invalid JSON (Extra data at column 13)
line 7: invalid JSON (Expecting value at column 1)
line 8: no assistant message to learn from
line 9: invalid JSON (Expecting property name enclosed in double quotes at column 2)
only 3 examples: OpenAI needs at least 10
3 examples, 9 errors

To repair one broken line by hand, paste it into our JSON repair tool, which fixes trailing commas, single quotes and Python literals in your browser. For exact per-example token counts, use the token counter or the code in counting tokens in Python and JavaScript.

Common fine-tuning file errors and how to fix them

What the errors mean
Error or symptomCauseFix
“Expecting property name enclosed in double quotes” at column 2A pretty-printed object split over lines ({ alone on a line), or single-quoted keysWrite one object per line with json.dumps or JSON.stringify
“Expecting value” near the end of a lineA trailing comma, often in data a model generatedSerialise again from real objects; see fixing invalid JSON from LLMs
“Extra data”Two JSON values on one line, or the leftovers of a split objectOne example per line
Blank lineAn empty line between examples or a double newlineDelete it (a single newline at the very end is fine)
Byte order markSaved as “UTF-8 with BOM”, an option in many Windows editors and spreadsheet exportsSave as UTF-8 without BOM, or read with utf-8-sig
Unknown role modelA Gemini-format line in an OpenAI or Mistral fileUse assistant for OpenAI and Mistral, model only for Gemini
No assistant messageThe example has a prompt but no answerAdd the target answer as the last assistant message
Example over the token limitOpenAI truncates long examples by removing tokens from the endShorten or split it: the end is where the answer is
Wrong columns after converting a CSVCommas or line breaks inside cells, split with split(",")Use a real CSV parser, as below

Python’s json module wording (Python 3.12 and older; from 3.13 a trailing comma reports “Illegal trailing comma before end of array” or “… of object” instead). JavaScript’s JSON.parse words the same problems differently.

How to convert CSV to JSONL (Python and JavaScript)

Read the CSV with a proper parser, build one example per row, and write one JSON object per line. Both scripts handle a quoted comma, a line break inside a cell, accented text and Excel’s byte order mark. Run on the same CSV, they wrote byte-identical files:

csv_to_jsonl.py
import csv
import json

SYSTEM = "Classify the support ticket. Reply with one label: billing, bug, feature_request or other."

# newline="" lets the csv module read line breaks inside quoted cells;
# utf-8-sig also accepts files that Excel saved with a byte order mark.
with open("tickets.csv", newline="", encoding="utf-8-sig") as src, \
     open("train.jsonl", "w", encoding="utf-8", newline="\n") as out:
    for row in csv.DictReader(src):
        ticket = row["ticket"].replace("\r\n", "\n").strip()
        example = {"messages": [
            {"role": "system", "content": SYSTEM},
            {"role": "user", "content": ticket},
            {"role": "assistant", "content": row["label"].strip()},
        ]}
        # one object per line; ensure_ascii=False writes "é" as UTF-8, not an escape
        out.write(json.dumps(example, ensure_ascii=False, separators=(",", ":")) + "\n")
csv-to-jsonl.mjs
// npm install csv-parse
import { readFileSync, writeFileSync } from "node:fs";
import { parse } from "csv-parse/sync";

const SYSTEM = "Classify the support ticket. Reply with one label: billing, bug, feature_request or other.";

// columns: true gives one object per row; bom: true drops an Excel byte order mark
const rows = parse(readFileSync("tickets.csv", "utf8"), { columns: true, bom: true, skip_empty_lines: true });

const lines = rows.map((row) =>
  JSON.stringify({
    messages: [
      { role: "system", content: SYSTEM },
      { role: "user", content: row.ticket.replace(/\r\n/g, "\n").trim() },
      { role: "assistant", content: row.label.trim() },
    ],
  }),
);
writeFileSync("train.jsonl", lines.join("\n") + "\n", "utf8");

The row with a line break inside a quoted cell came out as a single line, with the break escaped:

train.jsonl, line 3
{"messages":[{"role":"system","content":"Classify the support ticket. Reply with one label: billing, bug, feature_request or other."},{"role":"user","content":"Could you add dark mode?\nMy eyes would thank you."},{"role":"assistant","content":"feature_request"}]}

For Gemini, build {"systemInstruction": {"role": "system", "parts": [{"text": SYSTEM}]}, "contents": [...]} per row instead, with user and model roles.

How many training tokens, and what will it cost?

Training tokens are the tokens in your training file multiplied by the number of epochs (full passes over the data). Google’s tuning docs state this formula; OpenAI’s data-preparation cookbook (now archived) estimates billed tokens the same way. Both providers pick the epoch count automatically from your dataset size unless you set it, so read it from the job and multiply.

The ticket example above is 30 tokens of message text and about 45 tokens as a chat request. With 3 epochs, here is what two datasets would cost to train:

Training cost at 3 epochs
ModelTraining price per 1M tokens1,000 tickets × 45 tokens (135,000 training tokens)2,000 chats × 1,500 tokens (9,000,000 training tokens)
OpenAI gpt-4.1$25$3.38$225.00
OpenAI gpt-4.1-mini$5$0.68$45.00
OpenAI gpt-4.1-nano$1.50$0.20$13.50
Gemini 2.5 Flash-Lite$1.50$0.20$13.50
Gemini 2.5 Flash$5$0.68$45.00
Gemini 3.5 Flash$10$1.35$90.00
Gemini 2.5 Pro$25$3.38$225.00

Training prices from OpenAI’s and Google Cloud’s pricing pages, checked 2026-10-11 (Google lists them per 1,000 tokens). Mistral adds a $4 minimum per job. Inference on the tuned model is billed separately.

For datasets this size, training is cheap; the ongoing cost is using the tuned model. OpenAI charges $0.80 input and $3.20 output per million tokens for a fine-tuned gpt-4.1-mini, against $0.40 and $1.60 for the base model in our data (2026-10-11), twice as much. Google prices tuned endpoints from Gemini 3 onwards at 1.5 times the base model. Model the running cost in the LLM cost calculator before you commit.

Free toolAI token counterPaste a few training examples to see their exact token counts on OpenAI’s tokenizers, then multiply by examples and epochs.

Pre-upload checklist

  1. The provider still lets your account create tuning jobs (see the status table).
  2. The file is UTF-8 without a byte order mark, ends in .jsonl, and has exactly one JSON object per line with no blank lines.
  3. Every line uses the right shape: messages with assistant for OpenAI and Mistral, contents with model for Gemini, input plus preferred_output and non_preferred_output for OpenAI DPO.
  4. Every example ends with the answer you want, and the answers are consistent with each other.
  5. The instructions in each example match what you’ll send in production.
  6. You have at least 10 examples for OpenAI (start with 50) or 100 for Gemini, plus a separate validation or test set.
  7. No example is over the per-example token limit, and you’ve estimated tokens × epochs × price.
  8. The data contains nothing you aren’t allowed to send to the provider: strip API keys, passwords and personal data.

FAQ

Questions people ask

What is the difference between JSON and JSONL for fine-tuning?

A JSON file holds one value, often an array of every example. JSONL holds one complete JSON object per line, so each training example is a line and the file can be read, split or appended one line at a time. Fine-tuning APIs expect JSONL: an array in a .json file, or pretty-printed objects spread over several lines, will fail validation.

How many examples do I need to fine-tune?

OpenAI’s minimum is 10 examples, and it suggests starting with 50 well-made ones, noting that 50 to 100 often show improvement. Google recommends starting Gemini tuning with 100 examples, and 100 to 500 for best results. Quality and consistency matter more than volume, so check results on a held-out set before adding more.

Can I use the same JSONL file for OpenAI and Gemini?

Not as is. OpenAI and Mistral use messages with role and content, and assistant for the model’s turn. Gemini uses contents with parts, the role model, and puts the system prompt in systemInstruction. Converting is a short script: map each message, rename assistant to model, and wrap each text in a parts list.

Is OpenAI fine-tuning still available?

Only to existing customers, and not for long. OpenAI stopped new organisations from fine-tuning on 7 May 2026, closed it on 2 July 2026 to organisations with no fine-tuned model inference in the previous 60 days, and active customers can create jobs until 6 January 2027. Fine-tuned models already trained keep working until their base model is deprecated.

Can I fine-tune Gemini in Google AI Studio?

No. Google’s Gemini API docs say no model available in the Gemini API or AI Studio supports fine-tuning since Gemini 1.5 Flash-001 was deprecated in May 2025, with no immediate plans to bring it back. Supervised tuning of Gemini 2.5 and later models is available on Google Cloud’s Gemini Enterprise Agent Platform, with data in Cloud Storage.

Do system prompts count towards training tokens?

Yes. Everything in the training file counts, and it is multiplied by the number of epochs. In our ticket example the instructions are 20 tokens and the answer is 1, so the repeated instructions dominate the bill. OpenAI still advises keeping your best instructions in every example, especially with fewer than 100 examples, so weigh cost against quality.

Try it

Tools from this guide

Keep reading