RETROSPECTIVE RECORD · PREPARED 16 SEPTEMBER 2026Start here · 100 retrospective records ↗
AI Use Field Guide

Start here / Start here

Start here / Start-here guide · Start-here guide · prepared 16 September 2026

Translation quality tracks how much data a language has

A 2023 GPT evaluation and Google's own 2024 language expansion both show translation quality is uneven, not uniformly good.

Visual for this record: Translation quality tracks how much data a language has
Visual published by storage.googleapis.com, shown for identification of the record. Credit: storage.googleapis.com · source page ↗ Rights: owner-review-pending.

Start here

Ask an assistant to translate a paragraph of French or Spanish and the result usually reads naturally. Ask it to translate the same paragraph into a language with fewer digitised books, news archives, and parallel texts online, and the result can be fluent-sounding but wrong in ways a non-speaker cannot catch. The practical task is knowing which side of that line your language pair falls on before you send the translation anywhere that matters — a job application, a medical form, a message to a relative.

What the documents say

A 2023 evaluation, How Good Are GPT Models at Machine Translation?, tested three GPT variants across eighteen translation directions using both automatic metrics and human judgment, including high-resource, low-resource, and non-English-centric pairs. It found the models achieve 'very competitive translation quality for high resource languages, while having limited capabilities for low resource languages' — a direct statement of the asymmetry, from a study that also tested document-level and domain robustness. The WMT24 general translation task page shows how the machine-translation research community structures a shared evaluation: eleven language pairs including low-resource and morphologically rich languages, judged with a human protocol that reads document context rather than isolated sentences, a design choice that matters because word order and reference resolution both change across a paragraph. Separately, Google's own June 2024 announcement of 110 new Translate languages, built on its PaLM 2 model and called its 'largest expansion ever,' notes in passing that closely related languages are 'tricky to find data and train models' for — a vendor's own acknowledgment that resource scarcity, not just model size, shapes quality.

Check this

Before trusting a translation into an unfamiliar language, run it back the other way: paste the output back in and ask for a translation to your own language, and compare the round trip to your original meaning. It is not proof of accuracy, since the same blind spots can survive a round trip, but a meaning that drifts noticeably is a signal to find a fluent speaker or a dedicated translation service instead.

What holds and what fails

The high-resource case holds up consistently across the cited evaluation and general use: major world languages with large training data are where assistants perform closest to dedicated translation tools. It fails, by the evaluation's own description, on low-resource pairs, and it fails quietly — a fluent wrong sentence gives no warning sign the way a garbled one would. This is an editorial boundary, drawn from a study of one model family and a shared research benchmark, not a ranking of every assistant on every language today.

  • Check whether your target language is high- or low-resource before trusting nuance.
  • Round-trip an unfamiliar-language translation back to your own language.
  • Use a dedicated, human-reviewed translation service for anything official.

Fluency is not the same signal as accuracy, and the gap between them is largest exactly where the training data is thinnest.

Sources & reading trail

How Good Are GPT Models at Machine Translation? A Comprehensive Evaluation ↗

Evaluated GPT models across 18 translation directions and found competitive quality for high-resource languages but limited capability for low-resource ones.

Source published: 18 February 2023 · Retrieved: 16 September 2026

Shared Task: General Machine Translation ↗

Describes the shared evaluation's language pairs, including low-resource ones, and its document-level human evaluation protocol.

Source published: Not established · Retrieved: 16 September 2026

110 new languages are coming to Google Translate ↗

States Google's largest-ever Translate expansion used the PaLM 2 model and that closely related, lower-resource languages are harder to source training data for.

Source published: 27 June 2024 · Retrieved: 16 September 2026

Documentation, regulator guidance and studies establish the record; the checks and the boundary are AI Use Field Guide editorial analysis. This retrospective draft does not imply the site published on the event date.