RETROSPECTIVE RECORD · PREPARED 16 SEPTEMBER 2026Start here · 100 retrospective records ↗
AI Use Field Guide

Start here / History

History / From the archive · 14 March 2023 event · prepared 16 September 2026

GPT-4's own system card names overreliance as a real risk

OpenAI's GPT-4 materials pair a bar-exam score with an explicit warning about hallucination and overreliance.

Visual for this record: GPT-4's own system card names overreliance as a real risk
Visual published by images.ctfassets.net, shown for identification of the record. Credit: images.ctfassets.net · source page ↗ Rights: owner-review-pending.

Start here

One of the most useful habits in this guide is the least glamorous: read the limitations section before trusting a tool with something that matters. OpenAI made that easy to practice on 14 March 2023, publishing GPT-4 alongside a technical report and system card naming the model's failure modes in the vendor's own words. The page states GPT-4 'passes a simulated bar exam with a score around the top 10% of test takers,' and in the same breath admits the model 'still is not fully reliable (it "hallucinates" facts and makes reasoning errors).' Both sentences belong to the same launch; a reader who absorbs only the first is working from half the document.

What the documents say

The GPT-4 Technical Report describes a 'large-scale, multimodal model' that is nonetheless 'less capable than humans in many real-world scenarios' despite beating human performance on specific academic benchmarks. OpenAI's own page adds a number to the reliability claim: 'GPT-4 scores 40% higher than our latest GPT-3.5 on our internal adversarial factuality evaluations' -- an internal test the company designed and ran itself, not an independently verified figure. The GPT-4 System Card goes further, naming a risk category, 'Overreliance,' stating the model 'maintains a tendency to make up facts, to double-down on incorrect information, and to perform tasks incorrectly,' and warning this grows more dangerous, not less, 'as models become more truthful,' because users learn to trust them.

Check this

The check is simple and repeatable for any model release: search the technical report or system card for 'limitations,' 'hallucinat,' or 'overreliance' before adopting a tool for something consequential, rather than stopping at the headline capability. Then hold any benchmark claim, like a bar-exam percentile, against your actual task: a standardized test score describes performance on that test, under conditions the vendor chose, and says nothing directly about the real-world question you plan to ask.

What holds and what fails

This is an editorial reading of what a vendor's limitations section is worth: it holds up as a genuine, checkable source, because naming a product's own weaknesses in writing is a real admission, not a formality, and 'Overreliance' as a named section shows the maker anticipated the exact mistake of trusting the tool too much. It fails as a ceiling on caution, since these sections describe pre-release testing, not your specific document or stakes -- treat the stated limitations as a floor for how careful to be, not proof the risk is fully covered.

  • Before a high-stakes use, search the model's own documentation for its stated limitations rather than relying on the headline.
  • Treat a benchmark score as a description of that test, not a guarantee for your task.
  • When a document names 'overreliance' as a risk, take that as the vendor's own warning that convincing answers are not the same as accurate ones.

A system card that names its own product's habit of making things up sound plausible is a rare kind of source: a vendor's admission working against its own sales pitch, which is exactly what makes it worth reading closely.

Sources & reading trail

GPT-4 ↗

States GPT-4's benchmark performance alongside its own admission of hallucination and reasoning errors.

Source published: 14 March 2023 · Retrieved: 16 September 2026

GPT-4 Technical Report ↗

Describes GPT-4 as less capable than humans in many real-world scenarios despite benchmark performance.

Source published: 15 March 2023 · Retrieved: 16 September 2026

GPT-4 System Card ↗

Names 'Overreliance' as an explicit risk category and describes the hallucination-trust dynamic.

Source published: 14 March 2023 · Retrieved: 16 September 2026

Documentation, regulator guidance and studies establish the record; the checks and the boundary are AI Use Field Guide editorial analysis. This retrospective draft does not imply the site published on the event date.