AI glossary · Models and context
What is an AI hallucination?
Also called: AI hallucination, LLM hallucination
Definition
A hallucination is a confident, plausible-sounding statement from a language model that is false or unsupported by its input, such as an invented citation, date, function or quote.
Explained
How it works
A model writes a likely continuation of its context; nothing in it checks that a sentence is true. Hallucinations can contradict the input you supplied, such as misreading a document, or the world, such as citing a paper that doesn’t exist. In one example from the research below, an open-weights model asked for a researcher’s birthday “if you know” gave three different wrong dates in three tries.
A 2025 paper by OpenAI and Georgia Tech researchers, “Why Language Models Hallucinate”, gives two causes. In pretraining, errors arise from ordinary statistical pressure whenever true and false statements can’t be told apart, even with error-free data. Afterwards, most benchmarks grade answers right or wrong and give nothing for “I don’t know”, so guessing scores better and models are tuned to be good test-takers.
Example
Why guessing wins under right-or-wrong grading
The paper proposes stating a confidence threshold t in the task: a right answer scores 1, “I don’t know” scores 0, and a wrong answer loses t ÷ (1 − t) points. Take a model that is 30% sure of an answer.
Under plain right-or-wrong grading, guessing is worth 0.30 on average and abstaining 0, so the model should always guess. With t = 0.5 or 0.9, the same guess is worth −0.40 or −6.00, and “I don’t know” wins. The paper argues leaderboards mostly use the first rule, so models learn to bluff.
| Grading | Points lost for a wrong answer | Guess | Say “I don’t know” |
|---|---|---|---|
| Right or wrong (t = 0) | 0 | 0.30 | 0.00 |
| Confidence target t = 0.5 | 1 | −0.40 | 0.00 |
| Confidence target t = 0.9 | 9 | −6.00 | 0.00 |
Computed from the scoring rule in the paper; the 30% confidence is an illustration.
Cost and quality
Why it matters
Hallucinations are the main reason to check model output before it reaches users or runs as code. The fixes that work are about evidence, not settings: ground answers in documents you supply (RAG), ask for direct quotes or citations and drop claims without one, and say explicitly that “I don’t know” is acceptable. Anthropic’s guide to reducing hallucinations recommends all three.
We don’t quote hallucination rates here: they depend heavily on the task and benchmark, and change with every model release. Test on your own questions.
Don’t mix up
Common confusions
- Hallucination vs invalid output
- Broken JSON or a reply cut off at the token limit is a format failure, fixed with structured outputs or a higher cap. A hallucination can be perfectly formatted and still false; a schema guarantees the shape, not the facts.
- Hallucination vs temperature
- Lowering temperature makes output more repeatable, not more true: a model can give the same wrong answer every time.
Go deeper
Try it and read more
- Free toolContext Window CheckerSee whether your text fits each model’s context window.
- Free toolJSON RepairFix broken JSON from LLM output.
- Guide · 12 min readRAG chunk size and overlap: how to choose the right settingsHow to choose a RAG chunk size and overlap: token vs character splitters, embedding limits, what research shows, overlap cost maths and a test you can run.
- Guide · 10 min readHow to fix invalid JSON from LLMs (and stop getting it)Why ChatGPT, Claude and Gemini return broken JSON, and how to fix it: structured outputs, schema validation, retries, safe repair and streaming.
Related
Related terms
- RAGRAG (retrieval-augmented generation) is a technique where an application searches your own documents for passages relevant to a question and adds them to the prompt, so the model answers from that text rather than from memory alone.
- Structured outputsStructured outputs is an LLM API feature that constrains the model’s reply to a JSON Schema you supply, so the response parses and has the fields and types you asked for, unlike JSON mode, which only promises valid JSON.
- TemperatureTemperature is a sampling setting that divides a model’s next-token scores before they become probabilities, so low values make the likeliest token dominate and high values spread the choice across more tokens.
- Chain of thoughtChain-of-thought prompting is asking a language model to write out intermediate reasoning steps before its final answer, usually by showing worked examples or telling it to think step by step.