
Start here
Asking a chatbot to back up a claim with sources feels like the responsible move, and it is a better habit than not asking. It is not, on its own, a guarantee that what comes back is real, and the vendor's own help page and a documented case both say so directly rather than leaving it implied.
What the documents say
OpenAI's help article on what ChatGPT is, retrieved 16 September 2026, states plainly that outputs may be inaccurate, untruthful and otherwise misleading, and separately that ChatGPT will occasionally make up facts, a failure the page names as hallucinating, recommending a reader check whether a response is accurate rather than trust it by default. That is the vendor's own name for the failure, not a critic's label. A documented case shows what that looks like with citations specifically: Artificial Hallucinations in ChatGPT: Implications in Scientific Writing, published in the journal Cureus on 19 February 2023 by Hussam Alkaissi and Samy McFarlane, asked the model for references supporting a specific medical claim about homocysteine and bone metabolism. The authors report that none of the provided paper titles existed, and every PubMed identifier belonged to a different, unrelated paper; checking one by hand turned up an unrelated paper on laparoscopic surgery. Asking again for more recent sources did not fix it: the model returned the same fabricated list with only the dates changed. This is one documented case with named authors, not a measured rate across many chatbots or questions, and should be read at that scale.
Check this
The one-minute check the case study performs by hand is the same one any reader can run: take one citation a model provides, its title, author or identifier, and search for it independently in a source the model did not generate, rather than asking the same model to confirm itself. If the title, the author, or the identifier does not resolve to a real, matching document, the citation was invented regardless of how confident or specific it sounded.
What holds and what fails
Asking for a source remains useful because it gives a reader something checkable that a bare assertion does not; that much holds. What fails is treating the presence of a citation, a name, a year, an identifier, as evidence on its own, since the case study shows a fabricated reference can carry every surface feature of a real one. Whether this happens at a high or low rate across current tools is not something either document measures, and neither should be stretched into a claim about every assistant's current citation accuracy.
- Look up any citation a model provides in an independent source before repeating it.
- Treat a specific-sounding identifier, a PMID, a DOI, a case number, as a claim to verify, not proof by itself.
- If a checked citation turns out fabricated, assume neighbouring citations from the same answer need checking too.
The vendor names the failure and the case study shows its shape; neither replaces the minute it takes to look a source up directly.
Sources & reading trail
States that outputs may be inaccurate and that ChatGPT will occasionally hallucinate facts, recommending readers check accuracy.
Source published: Not established · Retrieved: 16 September 2026
Documents a case where every requested scientific reference from ChatGPT was fabricated, including PubMed IDs pointing to unrelated papers.
Source published: 19 February 2023 · Retrieved: 16 September 2026
Documentation, regulator guidance and studies establish the record; the checks and the boundary are AI Use Field Guide editorial analysis. This retrospective draft does not imply the site published on the event date.