Hallucinations: Why AI Makes Things Up & Lies
On this page5 sections
Hallucinations: Why AI Makes Things Up & Lies
The phenomenon where a large language model generates fluent, confident text that is factually false — invented citations, fabricated quotes, non-existent legal cases. Hallucinations arise from the model’s training as a next-token predictor, not a truth-teller, and are the central reliability problem for LLM deployment.
What Is a Hallucination?
In the context of AI, a hallucination is when a language model generates text that is false, fabricated, or unsupported by evidence, but that sounds plausible. The model does not know that the text is false — it has no concept of truth or falsity. It is simply generating the text that, based on its training data, is most likely to come next in the sequence. Training data scraped from the internet contains rumours, speculation, errors, and misinformation — which the model sometimes reproduces. Sometimes that text is accurate. Sometimes it is not. And the model cannot tell the difference.
The term “hallucination” is borrowed from psychology, where it refers to a perception of something that is not there. The analogy is apt: the model “perceives” patterns that do not correspond to reality, and it reports them as if they were real. But the analogy is also misleading. A human who is hallucinating usually knows, at some level, that something is wrong. A language model has no such awareness. It generates text with the same confidence whether the text is accurate or fabricated.
Hallucinations take many forms. The model might invent facts — dates, names, numbers, events — that are not real. It might invent citations — case names, paper titles, author names — that do not exist. It might invent quotes — attributing statements to people who never said them. It might invent biographical details — claiming that a real person held a job they never held or attended a school they never attended. Or it might simply get facts wrong — stating that the Eiffel Tower is in London, or that Abraham Lincoln was the sixteenth president (he was the sixteenth, but the model might say seventeenth).
The common thread is that the hallucination sounds plausible. The model does not generate random nonsense. It generates text that follows the patterns of its training data — patterns of grammar, logic, and style. The text reads like it could be true, even when it is not. This is what makes hallucinations so dangerous: they are hard to detect, especially for users who do not have the expertise to evaluate the claims.
Why Do Language Models Hallucinate?
To understand why language models hallucinate, you have to understand what they are and how they work.
A large language model is, at its core, a statistical system that predicts what word comes next in a sequence of text. It has been trained on enormous amounts of text — billions of words from books, articles, websites, and other sources — and it has learned the statistical patterns (the regularities a model learns from training data) in that text. When you give it a prompt, it generates a response by predicting, one word at a time, what should come next, based on the patterns it has learned.
The model does not have a database of facts. It does not look up information. It does not verify its claims. It simply generates the text that is most statistically likely to follow the prompt. Most of the time, this produces accurate or at least plausible text, because the training data contains accurate information, and the statistical patterns reflect that accuracy. But sometimes the statistical patterns lead the model astray.
There are several specific reasons why hallucinations occur:
Statistical noise. The model’s training data contains both accurate and inaccurate information. If the model has seen more inaccurate information about a topic than accurate information — or if the inaccurate information is more statistically prevalent in the training data — the model may generate the inaccurate version.
Pattern completion. The model is very good at completing patterns. If you ask it about a topic it knows little about, it will generate text that follows the pattern of what it has seen about similar topics — even if the specific details are wrong. This is why the model can generate plausible-sounding but fictional case citations: it has seen many real case citations, and it can generate text that follows the same pattern.
Overconfidence. The model generates text with the same level of confidence regardless of whether it knows the answer. It does not say “I don’t know” — it generates a response, and the response sounds confident even when it is wrong. This is partly a design choice: models that frequently say “I don’t know” are less useful to users, so the models are trained to be helpful and confident.
Training on unreliable data. The model’s training data includes text from the internet, which contains a lot of inaccurate information — rumours, speculation, errors, and deliberate misinformation. The model learns from this information, and it sometimes reproduces it.
Lack of grounding. The model has no connection to the real world. It cannot check its claims against reality. It cannot look up facts in a database. It cannot verify its citations. It generates text based solely on the patterns in its training data, and those patterns are sometimes wrong. Grounding — connecting a model’s responses to verified external information — is what RAG (retrieval-augmented generation) provides.
The Consequences
The consequences of hallucination range from minor to severe. At the minor end, a hallucination might produce a mildly inaccurate answer to a casual question — the model says the wrong date for a historical event, or misattributes a quote. The user notices the error, corrects it, and moves on.
At the severe end, hallucinations can cause real harm. The Steven Schwartz case — the lawyer who submitted fictional case citations — is an example. In that case, the hallucination was caught, and the consequence was professional sanction. But in other cases, hallucinations might not be caught. A doctor using an AI system to research a medical condition might receive inaccurate information. A journalist using an AI system to research a story might publish false claims. A student using an AI system to write an essay might submit work that contains fabricated facts.
The most dangerous hallucinations are those that are hard to detect. If the model generates a fact that is obviously wrong — “the Eiffel Tower is in London” — the user can catch it. But if the model generates a fact that is plausible but wrong — “the population of Bolivia is 12.5 million” (it is actually about 12 million) — the user may not catch it, especially if they are not familiar with the topic.
The hallucination problem is compounded by the fact that language models are becoming more capable. As the models get better at generating plausible text, their hallucinations become more convincing. A model that generates obviously wrong facts is less dangerous than a model that generates subtly wrong facts, because the subtle errors are harder to detect.
What Can Be Done?
The AI industry has been working on the hallucination problem since the release of ChatGPT, and several approaches have been developed.
Retrieval-Augmented Generation (RAG). RAG is a technique that connects the language model to an external database. Instead of generating answers solely from its training data, the model first retrieves relevant documents from the database, then generates its answer based on those documents. This “grounds” the model’s response in verified information, reducing the likelihood of hallucination. RAG is widely used in enterprise AI systems, where the database contains the company’s own documents. But RAG is not a complete solution: the model can still hallucinate when interpreting the retrieved documents, and the quality of the answer depends on the quality of the retrieval.
Attribution and citation. Some AI systems are being designed to provide citations for their claims — to link each factual statement to a source. This allows the user to verify the claim, and it makes hallucinations easier to detect. But citation is not a complete solution either: the model can cite a real source that does not actually support the claim, or it can misrepresent the source’s content.
Calibration. Researchers are working on “calibrating” language models — training them to express appropriate levels of confidence in their claims. A well-calibrated model would say “I’m not sure” when it is uncertain, rather than generating a confident but potentially false answer. But calibration is difficult, and even well-calibrated models can hallucinate.
Fact-checking and verification. Some systems use a second model — or a human reviewer — to fact-check the first model’s outputs. This can catch hallucinations, but it adds cost and latency (the delay added by fact-checking — one reason the technique is not always used), and the fact-checker itself can make mistakes.
Training improvements. Researchers are working on training techniques that reduce hallucination — for example, training the model to prefer “I don’t know” over a fabricated answer, or training it to recognise when a question is outside its area of competence. RLHF, the technique behind InstructGPT and ChatGPT, was partly designed to address this — by training models on human feedback about which responses were helpful and accurate. These techniques can reduce the frequency of hallucination, but they cannot eliminate it entirely.
The fundamental problem is that hallucination is not a bug that can be fixed. It is a feature of the architecture. Language models generate text by predicting the next word, based on statistical patterns. They do not have a concept of truth. They do not verify their claims. They generate plausible text, and sometimes plausible text is false. No amount of training or engineering can change this fundamental characteristic — at least not without fundamentally changing the architecture of the models.
What It All Means
The hallucination problem is a reminder that language models are not search engines, databases, or encyclopedias. They are text generators. They generate text that follows the statistical patterns of their training data, and that text is usually — but not always — accurate. The user must always verify the model’s claims, especially when the claims are about facts that matter.
The problem is also a reminder that the gap between what AI can do and what users expect it to do is still large. Users often treat language models as if they were omniscient oracles — asking them questions and accepting their answers without verification. This is understandable: the models are very good at generating plausible text, and the vast majority of their answers are accurate. But the minority of answers that are not accurate can cause real harm, especially in high-stakes contexts like law, medicine, and journalism.
The hallucination problem will not be solved by a single technique or a single breakthrough. It will be addressed through a combination of approaches — RAG, attribution, calibration, fact-checking, and training improvements — each of which reduces the problem but does not eliminate it. The user will always need to be the final check. The AI can generate the text, but the human must verify the facts. This is not a limitation that will go away. It is a permanent feature of the technology, and the sooner users understand it, the better.
- “Survey of Hallucination in Natural Language Generation” — Ziwei Ji et al., ACM Computing Surveys 55(12), 2023. The most comprehensive academic survey of the hallucination problem. dl.acm.org/doi/10.1145/3571733
- Mata v. Avianca — Court order sanctioning Steven Schwartz, June 2023. The primary source for the most famous hallucination case. courtlistener.com
- “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks” — Lewis et al., NeurIPS 2020. The foundational paper on the RAG technique. arxiv.org/abs/2005.11401
- “What AI Cannot Do: The Limits of the Possible” — Minds & Machines, A20. The main series’ exploration of AI’s limitations, including hallucination. series link
- OpenAI: “New AI classifier for indicating AI-written text” — OpenAI’s (now-withdrawn) attempt at a hallucination-detection tool, illustrating the difficulty of the problem. openai.com
This piece is part of Minds & Machines: Beyond the Series. The companion pieces B88 — Prompt Engineering and the New Art of Talking to Machines (the technique users developed to work around hallucinations and other LLM limitations), the main-series A20 — What AI Cannot Do: The Limits of the Possible (the broader limitations context), and the main-series A16 — The Attention Economy: How the Transformer Changed Everything (the architecture that makes hallucination structurally inevitable) cover the related milestones.
You interact with the technologies described here every day. The story of ai hallucinations is, in a real sense, the story of how your daily life came to be shaped by machines that think — and understanding it is the first step toward having a say in what comes next.
Subscribe
Get new articles delivered to your inbox. No spam — just the story behind the screen.