
Start here
One of the most useful habits in this guide is the least glamorous: read the limitations section before trusting a tool with something that matters. OpenAI made that easy to practice on 14 March 2023, publishing GPT-4 alongside a technical report and system card naming the model's failure modes in the vendor's own words. The page states GPT-4 'passes a simulated bar exam with a score around the top 10% of test takers,' and in the same breath admits the model 'still is not fully reliable (it "hallucinates" facts and makes reasoning errors).' Both sentences belong to the same launch; a reader who absorbs only the first is working from half the document.
What the documents say
The GPT-4 Technical Report describes a 'large-scale, multimodal model' that is nonetheless 'less capable than humans in many real-world scenarios' despite beating human performance on specific academic benchmarks. OpenAI's own page adds a number to the reliability claim: 'GPT-4 scores 40% higher than our latest GPT-3.5 on our internal adversarial factuality evaluations' -- an internal test the company designed and ran itself, not an independently verified figure. The GPT-4 System Card goes further, naming a risk category, 'Overreliance,' stating the model 'maintains a tendency to make up facts, to double-down on incorrect information, and to perform tasks incorrectly,' and warning this grows more dangerous, not less, 'as models become more truthful,' because users learn to trust them.
Check this
The check is simple and repeatable for any model release: search the technical report or system card for 'limitations,' 'hallucinat,' or 'overreliance' before adopting a tool for something consequential, rather than stopping at the headline capability. Then hold any benchmark claim, like a bar-exam percentile, against your actual task: a standardized test score describes performance on that test, under conditions the vendor chose, and says nothing directly about the real-world question you plan to ask.
What holds and what fails
This is an editorial reading of what a vendor's limitations section is worth: it holds up as a genuine, checkable source, because naming a product's own weaknesses in writing is a real admission, not a formality, and 'Overreliance' as a named section shows the maker anticipated the exact mistake of trusting the tool too much. It fails as a ceiling on caution, since these sections describe pre-release testing, not your specific document or stakes -- treat the stated limitations as a floor for how careful to be, not proof the risk is fully covered.
- Before a high-stakes use, search the model's own documentation for its stated limitations rather than relying on the headline.
- Treat a benchmark score as a description of that test, not a guarantee for your task.
- When a document names 'overreliance' as a risk, take that as the vendor's own warning that convincing answers are not the same as accurate ones.
A system card that names its own product's habit of making things up sound plausible is a rare kind of source: a vendor's admission working against its own sales pitch, which is exactly what makes it worth reading closely.
Sources & reading trail
States GPT-4's benchmark performance alongside its own admission of hallucination and reasoning errors.
Source published: 14 March 2023 · Retrieved: 16 September 2026
Describes GPT-4 as less capable than humans in many real-world scenarios despite benchmark performance.
Source published: 15 March 2023 · Retrieved: 16 September 2026
Names 'Overreliance' as an explicit risk category and describes the hallucination-trust dynamic.
Source published: 14 March 2023 · Retrieved: 16 September 2026
Documentation, regulator guidance and studies establish the record; the checks and the boundary are AI Use Field Guide editorial analysis. This retrospective draft does not imply the site published on the event date.