Spot the AI Hallucination: Can You Tell When It's Making It Up?
An interactive quiz. Judge 5 AI answers as grounded or hallucinated, learn why you often can't tell from the text, and what actually catches it.
Contents
An AI hallucination never announces itself. It reads exactly like a correct answer: fluent, specific, and delivered without a flicker of doubt. That is what makes it dangerous, and it is why trust, not raw capability, is usually the real blocker to AI adoption.
Here is the uncomfortable part, and the whole point of this quiz. You often cannot tell a hallucination from the text alone. Neither can your team. Below are five answers a business AI assistant might give. Each one tells you what the assistant actually had in front of it. Judge whether the answer is grounded in that, or made up.
Scenario 1 of 5
You uploaded invoice #4521. You ask when it was paid. The assistant answers: "Invoice #4521 was marked paid on March 3rd for $12,400."
Grounded, because a real source, the invoice you provided, actually exists for this claim. Note what makes it grounded is not that it sounds specific, but that you can open #4521 and check it. Grounded means checkable, not automatically correct, so open it anyway.
Scenario 2 of 5
The assistant has no company documents in its context. You ask about your history. It replies: "Your company was founded in 2014 by two Stanford graduates."
Hallucinated. There was nothing in its context to support this, so the confident detail is invented to fill the gap. The tell is not the wording, which is flawless, it is that no real source was available for it to draw on.
Scenario 3 of 5
You ask whether AI copilots pay off quickly. It answers: "Studies show 73% of companies using AI copilots see gains within the first month."
Hallucinated. An authoritative-sounding statistic with no named, traceable study behind it. If you cannot follow a number to a real source, treat it as invented, no matter how quotable it sounds.
Scenario 4 of 5
You never uploaded a board deck. The assistant answers: "Per your Q3 board deck, slide 12, revenue rose 18% quarter over quarter."
Hallucinated, and this is the dangerous one. It cites a specific source, which feels trustworthy, but that source was never in front of it. Models fabricate citations as fluently as they fabricate facts. A named document is not proof of grounding, a source you can actually open is.
Scenario 5 of 5
You uploaded last month's sales CSV. You ask for the top region. It answers: "Based on your CSV, the Northeast led at $84,200."
Grounded, because the data is genuinely there. But grounded does not mean the math is right, models are notoriously weak at summing rows. The source is real, so this is grounding, yet you should still verify the total. Grounded and correct are two different checks.
Why you probably missed at least one
If a few of those fooled you, that is the lesson, not a failure. The fabricated board deck cites a specific slide and a specific number. The invented founding story is fluent and plausible. Nothing in the wording separates them from the real answers. That is exactly why “just read it carefully” is not a strategy, for you or for the people using your product all day.
Grounded vs hallucinated: what actually separates them
The real distinction is not how an answer sounds. It is whether the claim is tied to real evidence the system actually retrieved, and whether you can trace it back to that evidence yourself.
Two ideas worth keeping separate here. Faithfulness is whether the answer is consistent with the source the model was given. Factuality is whether it is true in the world. Grounding buys you faithfulness, the answer sticks to real retrieved context, which is a huge win. It does not automatically buy you factuality: the retrieved document can be outdated, or the model can miscount over it. So grounding is the foundation, not the finish line.
”Your Returns Policy page (retrieved, linked) states 30 days, if unused and in original packaging.”
Links to the exact passage it pulled, so you can click through and confirm. Still worth checking the page is current, but now you can.
”Per your returns policy, items can be returned within 60 days.”
Names a source, sounds authoritative, and you cannot open it or check it. This is how a fabricated citation looks, right up until it costs you.
What actually catches hallucinations
Since you cannot eyeball them reliably, the answer is not a smarter reader, it is a system that checks itself. The levers that actually move the needle:
- Retrieval quality. Most hallucinations are really retrieval failures. If the right passage never reaches the model, it fills the gap. Better retrieval prevents more hallucinations than a bigger model does.
- Citations you can click. Every claim should link to the exact source passage, so a person can verify in one click instead of trusting the tone.
- A verifier step. A second model or a set of rules can check whether the answer is actually supported by the retrieved context before it ever reaches a user.
- A human-in-the-loop on anything irreversible, and evals against a known-good set so quality is a number you track, not a feeling you hope for.
Verification, not vibes
The grounding check
Is every claim actually supported by what was retrieved, before a user sees it?
An answer you can verify, not just believe
Why this matters more than it seems
One hallucination caught early is a shrug. One that slips into a customer email, a financial summary, or a compliance answer is how trust in an entire AI product collapses in an afternoon. Teams rarely abandon AI because the model was incapable. They abandon it because one confident, wrong, well-cited answer taught everyone to stop believing the rest.
The fix is not hoping the model never makes something up. It is building the checks that catch it before anyone has to, so your team does not have to run this quiz on every answer they see.
Worried your AI product might be confidently, plausibly wrong somewhere you can’t see? That is exactly what grounding, citations, and evals are built to catch. Book a free consult and we will look at where yours could be hallucinating and how to make every answer traceable.
Frequently asked questions
What is an AI hallucination?
An AI hallucination is when a model produces a confident, fluent claim that is not supported by any real source, whether or not it happens to be true. The danger is that a hallucination reads exactly like a correct answer. Fluency and even a cited source are not proof of accuracy, which is why you cannot reliably spot one by eye.
How do you detect AI hallucinations?
Honestly, not by reading the answer, that is the trap. You detect them by tracing every claim back to a real, retrieved source you can open, by using citations that link to the exact passage, and by running evals against known-good examples. A confident tone or a named document is not evidence, because models fabricate citations too. Verification, not vibes, is the only reliable check.
Can AI hallucinations be fixed?
You cannot eliminate them, but you can sharply reduce how often they happen and how much damage they do. Grounding answers in your real documents through RAG, showing click-through citations, using a second model or rules to verify claims, keeping a human-in-the-loop on risky actions, and measuring with evals all shrink the risk. The goal is not a perfect model, it is a system that catches its own mistakes.
Why do AI answers sound so confident even when they are wrong?
A language model is trained to produce fluent, plausible text, and its tone is not tied to whether the claim is true. Newer models can surface some uncertainty, through confidence signals or by abstaining, but those signals are imperfect and easy to over-trust. The safe assumption is that confidence tells you nothing about correctness.
Found this useful?
Share this with your network on LinkedIn, it helps more than you think.
Enjoyed this read? Get the next one in your inbox.
When we publish something worth your time, you will be first to know. No spam, unsubscribe anytime.
Keep reading
What You Should Refuse to Automate
The hype says automate everything. The discipline is knowing where the line goes. A task belongs to a person, not a model, when it is hard to undo, needs real judgment, or puts money, health, a job, or safety at stake. Grade any task on those three axes here, and see where the boundary actually falls.
Read articleOnboard It Like a Hire: The 30-60-90 Plan for Your First AI Agent
Nobody gives a new hire production access on day one, yet teams hand it to a week-old AI agent and get burned, or grant nothing and get nothing. The fix is the oldest tool in management: a scoped job, a named manager, probation, and promotions the agent has to earn.
Read articleAnswer Engine Optimization: How Customers Find You When AI Answers First
Your next customer is asking ChatGPT, not scrolling Google. The answer cites two or three sources, and either you are one of them or you are invisible. A scorecard shows whether an answer engine would cite you today, and the fixes are more honest than any SEO trick.
Read articleHave software that should be smarter?
Let’s map a free AI-transformation roadmap for your product.