
Start here
Someone deciding whether a chatbot can help them learn a subject, rather than just finish a task, is really asking about two different tools both called AI. Two randomised trials published in 2025 tested that difference directly, and did not find the same answer.
What the documents say
A study in PNAS, ‘Generative AI without guardrails can harm learning: Evidence from high school mathematics,’ ran a field randomised controlled trial with close to 1,000 high school students in Turkey across four 90-minute sessions in the fall semester of the 2023–2024 academic year, funded in part by the Wharton AI & Analytics Initiative. It reports that plain access to GPT and a version built to tutor rather than answer directly would ‘increase performance on the assisted practice sessions by 48% and 127%, respectively,’ relative to a control group. But on a later exam taken without AI access, GPT Base ‘diminished the average control student's performance on the unassisted exam by 17%,’ a statistically significant decline, while the tutor-styled group showed no significant difference from control. A separate study in Scientific Reports, ‘AI tutoring outperforms in-class active learning,’ ran a randomised crossover trial with 194 Harvard introductory-physics students comparing an AI tutor built around research-based pedagogy against in-class active learning on the same material. It reports a median post-test score of 4.5 for the AI-tutored group against 3.5 for the in-class group, with an effect size of 0.73 to 1.3 standard deviations, achieved in less time on task.
Check this
Before trusting either result for your own study habits, check what the AI was built to do: give answers, or withhold them and ask questions back. The PNAS study's own contrast — the same model, worse when it answered directly, roughly neutral when structured as a tutor — is the mechanism, not the brand of chatbot.
What holds and what fails
What holds across both trials is that a plain answer-generating assistant used during practice can substitute for the practice itself, showing up only later as a gap once the assistant is removed. What differs is that a tool designed to withhold answers and prompt reasoning avoided that gap and, in the Harvard trial's narrower physics context, outperformed a standard classroom method. Editorially: these are two trials in two subjects and age groups, not a verdict on all AI-assisted learning, and the Harvard authors state they do not presume their approach would outperform in-class learning on material requiring complex synthesis and higher-order thinking.
- Before using an assistant to study, ask whether it is answering your question or making you produce the answer.
- Test yourself without the assistant afterward, the way the PNAS exam condition did, to check what actually stuck.
- Do not extend either trial's finding past its own subject, age group and sample size.
Both studies measured a defined outcome — a practice score, an exam score, a post-test — in a specific classroom, and neither describes every use of AI in learning. The distinction they support is narrower than ‘AI helps’ or ‘AI hurts’: whether the tool is built to answer or built to teach.
Sources & reading trail
Field RCT with about 1,000 Turkish high school students: GPT access raised practice scores (up to 127% for a tutor-styled version) but plain GPT access lowered follow-up exam scores by 17%.
Source published: 25 June 2025 · Retrieved: 16 September 2026
Crossover RCT with 194 Harvard physics students: AI tutor group had roughly double the learning gains of in-class active learning, effect size 0.73-1.3 SD.
Source published: 3 June 2025 · Retrieved: 16 September 2026
Documentation, regulator guidance and studies establish the record; the checks and the boundary are AI Use Field Guide editorial analysis. This retrospective draft does not imply the site published on the event date.