On Faithfulness and Factuality in Abstractive Summarization
- Document
- 2 May 2020
- Event
- no single event
- Retrieved
- 16 September 2026
Start here
Ask an assistant to summarise a report and you get something shorter, readable, and confident. The question worth asking before you rely on it is what that shortening cost. A summary is a compression of a document, and compression means choices about what to keep; a qualifier such as 'roughly' or 'in one region' is exactly the kind of detail a shorter version tends to drop first, because it does not change the headline of a sentence. None of that makes summarising unsafe to use — it means the summary is a draft of the original's meaning, not a replacement for reading the part that matters.
What the documents say
The research on this has moved. A widely cited 2020 study, On Faithfulness and Factuality in Abstractive Summarization, ran a large human evaluation of neural summarization systems of that era and reported that annotators 'found substantial amounts of hallucinated content' in the output of every system tested, meaning the summaries stated things the source document did not. That study predates today's instruction-tuned chat assistants. A 2023 benchmark, Benchmarking Large Language Models for News Summarization, ran human evaluation across ten LLMs and found instruction tuning, not model size, was the main driver of summary quality, with outputs judged on par with high-quality human-written references once weak reference summaries were corrected for. A second 2023 paper, Summarization is (Almost) Dead, found evaluators preferred LLM summaries to both human-written and fine-tuned model summaries and reported better factual consistency and fewer extrinsic hallucinations — but the authors still called for higher-quality datasets and more reliable evaluation methods, which is itself a limit on how far 'preferred by evaluators' should be stretched.
Check this
Pick the two or three numbers, dates, or qualifying words that actually matter in the source and search for them in the summary. If a number survived but changed — a percentage rounded the wrong way, a range collapsed to one figure — that is the compression mechanism at work, not a one-off glitch. Ask the assistant to summarise again 'without adding any number not in the original,' and compare the two versions; a difference tells you where it was filling a gap rather than reporting one.
What holds and what fails
Overall preference ratings, of the kind both 2023 studies report, describe how evaluators judged fluency and general coverage — not whether every fact-level detail survived. That distinction holds regardless of how favourably a study rates a model. The approach fails on documents where the qualifier is the point: a study's confidence interval, a contract's exception clause, a report's stated limitation. This is an editorial reading of the evidence, not a claim any of the three papers makes directly.
- Identify the two or three load-bearing numbers before you read the summary.
- Ask for a second summary that adds nothing, and compare the two.
- Open the original section whenever a qualifier or exception is missing.
A summary rated highly by evaluators and a summary that preserved every number you need are two different claims, and only the source document can confirm the second one.
Sources & reading trail
Human evaluation of pre-LLM neural summarizers found substantial hallucinated content in outputs from every system tested.
Source published: 2 May 2020 · Retrieved: 16 September 2026
Human evaluation across ten LLMs found instruction tuning, not model size, drives zero-shot summary quality, rated on par with corrected human references.
Source published: 31 January 2023 · Retrieved: 16 September 2026
Human evaluators preferred LLM summaries over human-written and fine-tuned model summaries and found fewer extrinsic hallucinations, while calling for better evaluation datasets.
Source published: 18 September 2023 · Retrieved: 16 September 2026
Documentation, regulator guidance and studies establish the record; the checks and the boundary are AI Use Field Guide editorial analysis. This retrospective draft does not imply the site published on the event date.