
Start here
Asking an assistant to sketch a trip — flights, a hotel near the right neighbourhood, a few restaurant options within budget — produces something that reads like a finished plan well before it should be treated as one. The task is not generating the itinerary, which is quick; it is knowing how much of that confident detail is actually current and how much is a plausible-sounding guess filling gaps in what the assistant can verify.
What the documents say
A 2024 benchmark, TravelPlanner, built a sandbox with roughly four million real travel records and 1,225 planning scenarios that combine multiple constraints, such as budget, dates, and group size, the way an actual trip request does. It reports that even GPT-4, given tools to query the sandbox, achieved a 0.6% success rate on the full task, with failures traced to poor tool selection and difficulty juggling several constraints at once rather than any single missing fact. That figure describes an agent working through a structured, tool-using benchmark, not a casual conversation, but it is a direct measurement of how far confident-sounding planning can be from a plan that actually satisfies every stated requirement. Separately, OpenAI's ChatGPT search announcement, published 31 October 2024, describes a search feature that returns links to sources alongside an answer and states a plan to keep improving results 'particularly in areas like shopping and travel,' an acknowledgment, from the vendor, that this use case was not yet solved at launch. Google's Gemini Apps privacy notice states outputs are 'for informational purposes only' and instructs users to 'supervise Gemini's web browsing and tasks closely,' a direct instruction rather than a footnote.
Check this
Before booking anything, open the actual airline, hotel, or venue page named in the plan and confirm the price and availability yourself; the assistant's citation mechanism, where present, points at a source but does not guarantee the source still says what it said when retrieved. Ask specifically whether a detail came from a live search or from the model's general knowledge, since the two are handled differently and only one reflects the current page.
What holds and what fails
Assistants are useful for narrowing options and drafting a starting itinerary, which is a lower bar than the full-constraint planning task the benchmark measured. They fail, by the benchmark's own numbers, at reliably satisfying every constraint of a real request at once, and they fail on freshness whenever an answer draws on general knowledge instead of a live source, which is not always obvious from the reply itself.
- Verify price, date, and availability directly with the airline, hotel, or venue.
- Ask whether a specific detail came from a live search or general knowledge.
- Treat a full itinerary as a draft, not a booking-ready plan.
A confident plan and a current, verified plan are not the same output, and only checking the underlying source closes that gap.
Sources & reading trail
Built a sandbox benchmark of 1,225 travel-planning scenarios with about four million records and found GPT-4 succeeded on 0.6% of them, citing tool selection and constraint-juggling as the main failure modes.
Source published: 2 February 2024 · Retrieved: 16 September 2026
Describes how ChatGPT's web search surfaces sources and links, and states a plan to keep improving search 'particularly in areas like shopping and travel.'
Source published: 31 October 2024 · Retrieved: 16 September 2026
States Gemini Apps outputs are for informational purposes only, may be inaccurate, and that users should supervise Gemini's web browsing and tasks closely.
Source published: Not established · Retrieved: 16 September 2026
Documentation, regulator guidance and studies establish the record; the checks and the boundary are AI Use Field Guide editorial analysis. This retrospective draft does not imply the site published on the event date.