Skip to content
AI Dev Toolkit.
Esc
  • AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.Tool
  • LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.Tool
  • AI Model ComparisonCompare prices, context windows and features across models.Tool
  • AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.Tool
  • Context Window CheckerSee whether your text fits each model's context window.Tool
  • Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.Tool
  • GPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.Tool
  • Claude Code Error DatabaseExact Claude Code error messages with tested fixes.Tool

Guide · Developer tools

How to fix invalid JSON from LLMs (and stop getting it)

Stop invalid JSON at the source: turn on your provider’s structured outputs so the model is constrained to your JSON Schema, then validate every reply with Zod or Pydantic anyway. Structured outputs can still fail when the reply hits the token limit or the model refuses. Repair broken JSON only as a last resort, because a repair can hide missing data.

By Tahir NazirUpdated 10 min read

On this page
  1. Why do LLMs return invalid JSON?
  2. How do you make an LLM return valid JSON?
  3. How to validate LLM JSON with a schema
  4. What to do when the JSON is cut off
  5. How to retry when validation fails
  6. When should you repair broken JSON?
  7. How to parse streaming partial JSON
  8. Questions people ask

Why do LLMs return invalid JSON?

Because, unless you constrain it, a model writes JSON the way it writes prose: one token at a time, with nothing checking the syntax. It learned from JavaScript, Python and chat transcripts, where single quotes, trailing commas, comments and True are normal, and chat models like to wrap code in Markdown and add a friendly sentence. The common failures:

  • Wrapping: a sentence before the JSON, or a Markdown code fence around it.
  • Loose syntax: single quotes, unquoted keys, trailing commas, comments, Python’s True, False and None.
  • Bad strings: unescaped double quotes or raw line breaks inside a value.
  • Truncation: the reply hits max_tokens and stops mid-value, leaving brackets open.
  • Placeholders: ... in an array where the model shortened a list instead of writing every item.
One reply, five problems. The first four are safe syntax fixes. The fifth isn’t: the repaired JSON parses, but the title stops at “Bernoulli num” because the rest was never generated.

A single one of these makes JSON.parse or json.loads reject the whole reply. The fix is mostly upstream.

How do you make an LLM return valid JSON?

Turn on structured outputs. All three major APIs can constrain the model to a JSON Schema you supply, which is far more reliable than asking for JSON in the prompt. The OpenAI and Anthropic SDKs also accept a Pydantic model (Python) or a Zod schema (TypeScript) and convert it for you; for Gemini you pass the JSON Schema that Pydantic or Zod generates, as in the Gemini example below:

Structured output options, as each provider documents them
ProviderHow to turn it onWatch out for
OpenAItext.format with type: "json_schema" and strict: true (Responses API), or responses.parse with a Pydantic or Zod schemaEvery field must be required and objects need additionalProperties: false. Check for refusal content and status: "incomplete".
Anthropicoutput_config.format with type: "json_schema", or messages.parse; strict: true on tools for tool inputsAdds a system prompt (more input tokens). First use of a schema is slower while its grammar compiles. Check stop_reason for refusal and max_tokens.
Google Geminiresponse_format with mime_type: "application/json" and a schema (responseFormat in JavaScript)Supports a subset of JSON Schema and ignores unsupported keywords. responseSchema is deprecated in the API reference.

From the OpenAI, Anthropic and Gemini structured output docs and API references, checked 2026-10-08.

OpenAI (Python)
from openai import OpenAI
from pydantic import BaseModel

client = OpenAI()


class ContactInfo(BaseModel):
    name: str
    email: str
    plan_interest: str
    demo_requested: bool


response = client.responses.parse(
    model="gpt-6-astra",
    input=[
        {"role": "system", "content": "Extract the contact information."},
        {"role": "user", "content": "John Smith (john@example.com) is interested in our Enterprise plan and wants to schedule a demo for next Tuesday at 2pm."},
    ],
    text_format=ContactInfo,
)

contact = response.output_parsed
Claude (Python)
from anthropic import Anthropic
from pydantic import BaseModel


class ContactInfo(BaseModel):
    name: str
    email: str
    plan_interest: str
    demo_requested: bool


client = Anthropic()

response = client.messages.parse(
    model="claude-opus-5-5",
    max_tokens=1024,
    messages=[
        {
            "role": "user",
            "content": "Extract the key information from this email: John Smith (john@example.com) is interested in our Enterprise plan and wants to schedule a demo for next Tuesday at 2pm.",
        }
    ],
    output_format=ContactInfo,
)
print(response.parsed_output)
Gemini (Python)
from google import genai
from pydantic import BaseModel


class ContactInfo(BaseModel):
    name: str
    email: str
    plan_interest: str
    demo_requested: bool


client = genai.Client()

response = client.models.generate_content(
    model="gemini-3.8-flash",
    contents="Extract the contact information: John Smith (john@example.com) is interested in our Enterprise plan and wants to schedule a demo for next Tuesday at 2pm.",
    config={
        "response_format": {"text": {"mime_type": "application/json", "schema": ContactInfo.model_json_schema()}},
    },
)

contact = ContactInfo.model_validate_json(response.text)

Is JSON mode enough?

No. OpenAI’s older JSON mode (type: "json_object") makes the output parse but doesn’t enforce any schema, and you must still say “JSON” in a message or the API returns an error. OpenAI recommends Structured Outputs instead wherever the model supports it. The main reason to stay on JSON mode is an older model without schema support, and then you have to validate everything yourself.

How to validate LLM JSON with a schema

Parse, then validate every reply against a schema, even with structured outputs on. Google’s docs say it plainly: structured output guarantees syntactically correct JSON, not correct values. Anthropic adds that string enum values may come back with different capitalisation. A schema catches both:

validate.ts (Zod 4)
import { z } from "zod";

const Contact = z.object({
  name: z.string(),
  plan: z.enum(["free", "pro", "enterprise"]),
  seats: z.number().int().positive(),
});

const raw = '{"name": "Ada", "plan": "Enterprise", "seats": "5"}';
const result = Contact.safeParse(JSON.parse(raw));
if (!result.success) console.log(z.prettifyError(result.error));

// ✖ Invalid option: expected one of "free"|"pro"|"enterprise"
//   → at plan
// ✖ Invalid input: expected number, received string
//   → at seats
validate.py (Pydantic 2)
from typing import Literal
from pydantic import BaseModel, Field, ValidationError

class Contact(BaseModel):
    name: str
    plan: Literal["free", "pro", "enterprise"]
    seats: int = Field(gt=0)

raw = '{"name": "Ada", "plan": "Enterprise", "seats": "5"}'
try:
    Contact.model_validate_json(raw)
except ValidationError as e:
    print(e)

# 1 validation error for Contact
# plan
#   Input should be 'free', 'pro' or 'enterprise' [type=literal_error, input_value='Enterprise', input_type=str]
#     For further information visit https://errors.pydantic.dev/2.14/v/literal_error

Notice the difference. Zod rejected the string "5"; Pydantic quietly converted it to the number 5, because its default “lax” mode coerces numeric strings. That’s often what you want from a model, but if it isn’t, set model_config = ConfigDict(strict=True): the same input then fails with “Input should be a valid integer”. model_validate_json also reports syntax errors with a position, for example “Invalid JSON: trailing comma at line 1 column 13”, so one call covers both checks.

What to do when the JSON is cut off

Check why generation stopped before you parse anything. A reply that ran out of tokens is incomplete, and no repair can bring back what was never written. Structured outputs don’t protect you here: OpenAI and Anthropic both document that a reply cut off by the token limit may not match your schema.

Where each API tells you the reply was cut off
APIField to checkValue meaning “ran out of tokens”
OpenAI Responsesstatus and incomplete_details.reason"incomplete" and "max_output_tokens"
Anthropic Messagesstop_reason"max_tokens"
GeminifinishReason on the candidateMAX_TOKENS

When it happens, retry with a higher output limit, ask for less (fewer items, shorter strings), or split the job into smaller requests. To size the limit, count a typical good reply with the token counter and leave generous headroom: reasoning models also spend output tokens before they write the JSON. Our guide to counting tokens in Python and JavaScript shows how to do it in code.

How to retry when validation fails

Send the error back. A model that sees “Input should be ‘free’, ‘pro’ or ‘enterprise’” usually fixes its next reply, which is cheaper and safer than guessing what it meant. Cap the attempts so a stubborn failure doesn’t loop forever:

retry.py
from typing import Literal
from pydantic import BaseModel, Field, ValidationError

class Contact(BaseModel):
    name: str
    plan: Literal["free", "pro", "enterprise"]
    seats: int = Field(gt=0)

def get_contact(prompt, call_model, max_attempts=3):
    """call_model(messages) -> str is your API call. Retries with the error."""
    messages = [{"role": "user", "content": prompt}]
    for _ in range(max_attempts):
        text = call_model(messages)
        try:
            return Contact.model_validate_json(text)
        except ValidationError as e:
            messages += [
                {"role": "assistant", "content": text},
                {"role": "user", "content": f"That JSON failed validation:\n{e}\nReply with the corrected JSON only."},
            ]
    raise RuntimeError(f"No valid JSON after {max_attempts} attempts")

# Stand-in for a real model: wrong the first time, right the second.
replies = iter(['{"name": "Ada", "plan": "Enterprise", "seats": 5}',
                '{"name": "Ada", "plan": "enterprise", "seats": 5}'])
print(get_contact("Extract the contact.", lambda messages: next(replies)))
# name='Ada' plan='enterprise' seats=5
  • Don’t retry blindly. A refusal or a token-limit stop needs a different fix (a higher limit, a smaller task or a changed prompt), not the same request again.
  • Keep two or three attempts. Each retry resends the conversation, so it costs input tokens again; the LLM cost calculator shows what a retry rate does to a monthly bill, and estimating LLM API costs explains the sums.
  • Log failures. Repeated errors on the same field usually mean the schema or prompt is unclear. Describe the field better instead of retrying harder.
  • Fall back deliberately. After the last attempt, fail loudly or queue the item for a person. Don’t silently store half-valid data.

When should you repair broken JSON?

Repair when you can’t change the request: old models without structured outputs, logs you’re cleaning up, or replies pasted by a person. A repair library fixes syntax well, but it has to guess when the text is ambiguous, and some guesses change your data without an error:

What a repair does with ambiguous input
Model outputAfter repairThe problem
{"total": 12{"total": 12}Was it 12, 125 or 12.50? The number may be cut short.
{"done": tru{"done": "tru"}A cut-off boolean becomes the string “tru”.
{"items": [1, 2, 3, ...]}{"items": [1, 2, 3 ]}The model’s placeholder is dropped, so the list looks complete.
{"msg": "Hello{"msg": "Hello"}The rest of the message is missing, but the JSON looks fine.

Run with jsonrepair 3.15.0, the library our JSON repair tool uses. The tool flags the three cut-off cases as truncated; the dropped ellipsis gets no warning.

So treat repair as a way to see the data, not a way to trust it. Validate the repaired value with your schema, record that it was repaired, and if the reply was truncated, retry instead of keeping the patched version. In Python, the json_repair package does the same job as jsonrepair; for a one-off, our JSON repair tool shows every change it makes.

Free toolJSON repairPaste broken JSON or a whole model reply. It fixes the syntax in your browser, shows every change, and warns when the output was cut off.

How to parse streaming partial JSON

Accumulate the chunks and parse the full text once the stream ends; parse partial text only to preview it. Gemini’s docs say streamed structured output arrives as partial JSON strings that concatenate into the final object, and Anthropic streams tool inputs as input_json_delta events whose partial_json fragments you join and parse when the block stops. OpenAI recommends using its SDK helpers to stream structured outputs.

To show progress before the end, use a parser that accepts incomplete input. Pydantic’s from_json has an allow_partial option, which Anthropic’s streaming docs point to:

partial.py
from pydantic_core import from_json

chunk = '{"name": "Ada Lovelace", "skills": ["maths", "poe'
print(from_json(chunk, allow_partial=True))
print(from_json(chunk, allow_partial="trailing-strings"))

# {'name': 'Ada Lovelace', 'skills': ['maths']}
# {'name': 'Ada Lovelace', 'skills': ['maths', 'poe']}

By default it drops the unfinished string; "trailing-strings" keeps it. In JavaScript, running jsonrepair on the text so far gives a similar preview. Either way, render partial values but don’t act on them: run your schema check and the stop-reason check on the final text before saving anything or calling a tool.

FAQ

Questions people ask

How do I force ChatGPT or the OpenAI API to return JSON?

Use Structured Outputs: pass a JSON Schema in text.format with strict: true, or pass a Pydantic or Zod schema to the SDK’s responses.parse. The model is then constrained to your schema. Still check status for an incomplete reply and look for a refusal, then validate the parsed result.

What’s the difference between JSON mode and structured outputs?

JSON mode only makes the reply parse as JSON; it can still miss fields or use the wrong types. Structured outputs constrain the reply to a JSON Schema you supply. OpenAI recommends Structured Outputs over JSON mode wherever the model supports it. Either way, a reply cut off by the token limit can still be incomplete.

Do structured outputs guarantee valid JSON?

Almost always, but not in every case. OpenAI and Anthropic both document exceptions: a refusal, or a reply that stops at the token limit, may not match your schema. Anthropic also says string enum values may differ in capitalisation. Check the stop reason and validate every reply with a schema.

How do I get JSON from Claude?

Use Anthropic’s structured outputs: pass a JSON Schema in output_config.format, or a Pydantic model or Zod schema to client.messages.parse, which validates the reply and returns it in parsed_output. For tool inputs, set strict: true on the tool. Check stop_reason for refusal and max_tokens.

How do I parse JSON from an LLM response in Python?

With structured outputs on, validate the text directly with YourModel.model_validate_json(text), which checks syntax and schema in one step. Without them, strip any Markdown fence and surrounding sentences first, then validate. Use a repair package such as json_repair only as a fallback, and validate its output too.

Can you fix JSON that was cut off?

You can make it parse by closing the open strings and brackets, which is what repair tools do, but the missing data is gone. A cut-off number or boolean can even turn into a wrong value. The real fix is to retry with a higher output token limit or a smaller request.

Try it

Tools from this guide

Keep reading