RETROSPECTIVE RECORD · PREPARED 16 SEPTEMBER 2026Start here · 100 retrospective records ↗
AI Use Field Guide

Start here / Start here

Start here / Start-here guide · Start-here guide · prepared 16 September 2026

A good summary can still drop the number that mattered

Two 2023 evaluations rate AI summaries highly overall, but neither measured whether every number and qualifier survived.

arxiv.orgprimary record

On Faithfulness and Factuality in Abstractive Summarization

Document
2 May 2020
Event
no single event
Retrieved
16 September 2026
No visual was published with this record, so its primary document stands in its place.

Start here

Ask an assistant to summarise a report and you get something shorter, readable, and confident. The question worth asking before you rely on it is what that shortening cost. A summary is a compression of a document, and compression means choices about what to keep; a qualifier such as 'roughly' or 'in one region' is exactly the kind of detail a shorter version tends to drop first, because it does not change the headline of a sentence. None of that makes summarising unsafe to use — it means the summary is a draft of the original's meaning, not a replacement for reading the part that matters.

What the documents say

The research on this has moved. A widely cited 2020 study, On Faithfulness and Factuality in Abstractive Summarization, ran a large human evaluation of neural summarization systems of that era and reported that annotators 'found substantial amounts of hallucinated content' in the output of every system tested, meaning the summaries stated things the source document did not. That study predates today's instruction-tuned chat assistants. A 2023 benchmark, Benchmarking Large Language Models for News Summarization, ran human evaluation across ten LLMs and found instruction tuning, not model size, was the main driver of summary quality, with outputs judged on par with high-quality human-written references once weak reference summaries were corrected for. A second 2023 paper, Summarization is (Almost) Dead, found evaluators preferred LLM summaries to both human-written and fine-tuned model summaries and reported better factual consistency and fewer extrinsic hallucinations — but the authors still called for higher-quality datasets and more reliable evaluation methods, which is itself a limit on how far 'preferred by evaluators' should be stretched.

Check this

Pick the two or three numbers, dates, or qualifying words that actually matter in the source and search for them in the summary. If a number survived but changed — a percentage rounded the wrong way, a range collapsed to one figure — that is the compression mechanism at work, not a one-off glitch. Ask the assistant to summarise again 'without adding any number not in the original,' and compare the two versions; a difference tells you where it was filling a gap rather than reporting one.

What holds and what fails

Overall preference ratings, of the kind both 2023 studies report, describe how evaluators judged fluency and general coverage — not whether every fact-level detail survived. That distinction holds regardless of how favourably a study rates a model. The approach fails on documents where the qualifier is the point: a study's confidence interval, a contract's exception clause, a report's stated limitation. This is an editorial reading of the evidence, not a claim any of the three papers makes directly.

  • Identify the two or three load-bearing numbers before you read the summary.
  • Ask for a second summary that adds nothing, and compare the two.
  • Open the original section whenever a qualifier or exception is missing.

A summary rated highly by evaluators and a summary that preserved every number you need are two different claims, and only the source document can confirm the second one.

Sources & reading trail

On Faithfulness and Factuality in Abstractive Summarization ↗

Human evaluation of pre-LLM neural summarizers found substantial hallucinated content in outputs from every system tested.

Source published: 2 May 2020 · Retrieved: 16 September 2026

Benchmarking Large Language Models for News Summarization ↗

Human evaluation across ten LLMs found instruction tuning, not model size, drives zero-shot summary quality, rated on par with corrected human references.

Source published: 31 January 2023 · Retrieved: 16 September 2026

Summarization is (Almost) Dead ↗

Human evaluators preferred LLM summaries over human-written and fine-tuned model summaries and found fewer extrinsic hallucinations, while calling for better evaluation datasets.

Source published: 18 September 2023 · Retrieved: 16 September 2026

Documentation, regulator guidance and studies establish the record; the checks and the boundary are AI Use Field Guide editorial analysis. This retrospective draft does not imply the site published on the event date.